Findings
What forecasting a whole World Cup three ways actually showed. Every figure below is recomputed at build time from the frozen ledger — the same numbers the Divergence Log renders, read from forecasts that were locked before each kickoff and never revised.
The model lost.
Over 104 matches, the statistical model finished last of the three sources on aggregate calibration — Brier 0.525, against the sportsbook's 0.498 and Polymarket's 0.498, where lower is better — and last on the simplest measure there is: its favourite won 61% of the time, against 63% for the sportsbook and 63% for Polymarket.
The one measure the model leads is the narrowest. It was the distinctly closest source on 26 of 104 matches, against 5 for Polymarket and 0 for the sportsbook — and that measure flatters it. A distinct win requires one source to separate from the others by more than two percentage points on the outcome that actually happened, and the two markets track each other so closely that they almost never separate; when they jointly beat the model, neither is credited. The model wins that count largely by being the only source that is ever alone.
So: the markets were better priced, the model was more often boldly right, and those two sentences are not in conflict. They measure different things, which is the whole reason this board showed three numbers instead of one.
Frozen on 19 July 2026.
The three-source scoreboard
Brier and log loss score every source on every match it priced, independent of who "won" — lower is better. Favourite hit rate is the blunt version: how often that source's most likely outcome was the outcome. Both are graded at 90 minutes, so extra time and penalties score as draws.
| Source | Brier | Log loss | Favourite hit rate | Distinct verdict wins |
|---|---|---|---|---|
| Model | 0.525 n=104 | 0.869 | 61% n=104 | 26 |
| Sportsbook | 0.498 n=104 | 0.844 | 63% n=104 | 0 |
| Polymarket | 0.498 n=104 | 0.842 | 63% n=104 | 5 |
Of 104 scored matches, 6 went to the two markets jointly (credited to neither), 32 were agreements too close to separate, and in 35 every single source's favourite lost.
That last figure is worth sitting with: on roughly a third of the tournament, no source — model or market — had the eventual result as its most likely outcome. International football at this level is genuinely close to a coin flip in a large share of matches, which is exactly what the calibration page was built to show.
Where the model diverged hardest
The premise of this site was that the gap between the model and the market is itself a signal. Splitting every match by the size of that gap shows the model's record is not flat across it.
| Model–market gap | Matches | Model verdict wins | Model Brier | Sportsbook Brier | Polymarket Brier |
|---|---|---|---|---|---|
| Under 5pp | 17 | 0 | 0.645 n=17 | 0.626 n=17 | 0.632 n=17 |
| 5–15pp | 53 | 19 | 0.483 n=53 | 0.488 n=53 | 0.485 n=53 |
| 15pp+ | 34 | 7 | 0.530 n=34 | 0.450 n=34 | 0.451 n=34 |
The model out-scores both markets in exactly one band — 5–15pp — and is beaten in the others. Its calibration is worst precisely where it disagrees with the market most.
Of the 34 matches where model and sportsbook differed by 15 percentage points or more — the threshold this site flagged as a divergence all along — the model was the distinctly closest source 7 times, and its own favourite lost 15 times. When the model backed itself hardest against the market, it was wrong more than twice as often as it was distinctly right. That asymmetry, not the headline Brier gap, is the finding.
The biggest calls, both ways
The largest model-vs-market gaps, split by result. Percentages are each source's frozen pre-match probability on the outcome that actually happened.
Called it (showing top 3 of 7)
- Croatia v Ghana 2–1 Home win Model 78%Sportsbook 54%Polymarket 54%
- Mexico v Czechia 3–0 Home win Model 71%Sportsbook 50%Polymarket 51%
- Colombia v Ghana 1–0 Home win Model 86%Sportsbook 67%Polymarket 68%
Got it wrong (showing top 3 of 15)
- Ghana v Panama 1–0 Home win Model 11%Sportsbook 44%Polymarket 44%
- DR Congo v Uzbekistan 3–1 Home win Model 24%Sportsbook 56%Polymarket 56%
- Belgium v Iran 0–0 Draw Model 32%Sportsbook 21%Polymarket 21%
Both lists are drawn from the same rows as the Divergence Log archive, which carries all 104 matches and a CSV export.
The AI-commentary verdict
The site's second layer — Claude-written match previews and divergence commentary — was declared an experiment from day one: would AI commentary add genuine value, or would it be decorative? After 104 matches the honest answer is that it was decorative. The previews were fluent and never wrong, but they restated what the bars above them already showed. The prompt was deliberately tight — three paragraphs, numerical inputs only, deliberately non-definitive language, and no AI-issued scoreline (the forecasts were never AI-generated). Whether a looser brief would have earned its place is untested; we never ran the variant. What can be said is that under these constraints it added length, not information.
What this does not show
- One tournament, 104 matches. The separations here are small — hundredths of a Brier point, a couple of percentage points of hit rate. Over 104 matches that is suggestive, not decisive. A different World Cup could reorder the table without anything about the three sources having changed.
- These are normal numbers, not a failure. Good international-football forecasts hit roughly 55–60% match accuracy, and all three sources landed near that band (61% / 63% / 63%). "The model lost" means it was beaten by a market that aggregates injuries, form, team news and money — none of which the model can see. It does not mean the forecasts were bad.
- The reading is post-hoc even though the forecasts are not. Every probability was frozen before its kickoff and the thresholds — the 15pp divergence flag, the gap bands, the 2pp verdict margin — were all set before the tournament finished. But which comparisons to feature on this page was chosen with the results in hand. Treat the framing with more suspicion than the figures.
- The model's known limits are structural. Pure Elo plus Dixon-Coles, with no injury, form, or squad data, is more top-heavy than the markets by construction — which is exactly the shape of its worst band above. The methodology page lists these in full.
Where to look next: the Divergence Log for every match scored source by source, the calibration page for the reliability picture behind these scores, and the methodology for how the model, the markets and the AI layer each worked.