Calibration

How well do the probabilities match what actually happened?

Live WC predictions (n = 104)

0.525 Brier score — lower is better · random ≈ 0.667
60.6% Match accuracy · random ≈ 37%

Full backtest (Full backtest (2024, all competitions), n = 290): Brier 0.495, accuracy 63.8% — inflated by friendlies and qualifiers.

Reliability diagram — pooled outcomes

0% 100% 0% 100% Predicted probability Observed frequency 10% predicted, 12% actual (n=84) 28% predicted, 24% actual (n=138) 48% predicted, 61% actual (n=38) 70% predicted, 78% actual (n=40) 86% predicted, 58% actual (n=12)

Each dot is a probability bin pooling home-win, draw, and away-win predictions. Dots on the dashed diagonal = perfect calibration. Dot size scales with sample count.

Per-outcome summary

Outcome Predicted (avg) Observed Gap
Home win 45.1% 43.3% -1.8pp
Draw 25.4% 28.8% +3.5pp
Away win 29.5% 27.9% -1.6pp

Statistics are from live World Cup matches. The reliability diagram and table will fill in as results come in.

Known limitations: on Euro-like fields with clustered Elo ratings (the closest analog to World Cup group play) the model hits roughly 49% match accuracy and a Brier score of around 1.05 per outcome — only marginally better than random. Draws are systematically under-predicted by ~4 percentage points, with away wins over-predicted by a matching amount. These biases are corrected in the per-match probability displays; see the methodology for details.