- Our call
- Argentina to advance
- Result
- 1–0
Spain through (won 1–0 after extra time) — our pre-match lean missed.
Our running record for the 2026 tournament: every call locked before kickoff, then graded in public against the real result. The number moves as matches finish. Below, a retrospective backtest across past World Cups for credibility.
This World Cup — 2026 (so far)
How often our pre-match call was right so far. In the group stage that's the win / draw / loss result; in a knockout it's whether the side we favoured advanced — because a knockout call is read as winning the tie, extra time and penalties included. We also keep a separate 90-minute ledger (below). A moving stat on a smallish sample — a snapshot of this tournament, not a fixed accuracy claim.
68%
71 of 104 results right
graded so far · this tournament
90-minute reads: 66/104
Most recent first — the ones we got wrong are right here alongside the ones we got right.
Spain through (won 1–0 after extra time) — our pre-match lean missed.
England through (won 6–4) — our pre-match lean missed.
Right call — Argentina through (won 2–1).
Right call — Spain through (won 2–0).
Right call — Argentina through (won 3–1 after extra time).
Right call — England through (won 2–1 after extra time).
Right call — Spain through (won 2–1).
Right call — France through (won 2–0).
Switzerland through on penalties — our pre-match lean missed. Shootouts are near coin flips; we grade the call, we don't claim the shootout.
Right call — Argentina through (won 3–2).
Right call — Belgium through (won 4–1).
Right call — Spain through (won 1–0).
Right call — England through (won 3–2).
Norway through (won 2–1) — our pre-match lean missed.
Right call — France through (won 1–0).
Right call — Morocco through (won 3–0).
Right call — Colombia through (won 1–0).
Right call — Argentina through (won 3–2 after extra time).
Egypt through on penalties — our pre-match lean missed. Shootouts are near coin flips; we grade the call, we don't claim the shootout.
Right call — Switzerland through (won 2–0).
Right call — Portugal through (won 2–1).
Right call — Spain through (won 3–0).
Right call — USA through (won 2–0).
Right call — Belgium through (won 3–2 after extra time).
Right call — England through (won 2–1).
Mexico through (won 2–0) — our pre-match lean missed.
Right call — France through (won 3–0).
Right call — Norway through (won 2–1).
Morocco through on penalties — our pre-match lean missed. Shootouts are near coin flips; we grade the call, we don't claim the shootout.
Paraguay through on penalties — our pre-match lean missed. Shootouts are near coin flips; we grade the call, we don't claim the shootout.
Right call — Brazil through (won 2–1).
Right call — Canada through (won 1–0).
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
The result went against the pre-match lean — a reminder that any single match can defy the model. That uncertainty is expected, not a bug.
Called the right result. One match never validates a model — results are noisy by nature.
Called the right result. One match never validates a model — results are noisy by nature.
For credibility beyond this tournament, we re-ran the model over the 2018 and 2022 World Cups — all 128 matches, about 55% of results called correctly — and graded every call against what really happened. The misses are shown right alongside the hits, no cherry-picking.
One note: this historical backtest uses the core Elo + Dixon–Coles model. The small set-piece and xG-form adjustments in our live 2026 model aren't applied to it.
This is a retrospective test. For each match, the model was given only the data available before that match — no future results — then it predicted the outcome, and only afterwards did we check it against the real result. These were never locked in ahead of time for real, and never money. It is evidence the method is reasonable, not a promise about the future.
For the live, going-forward record — predictions we lock before kickoff and grade in public as matches finish — see the Results on today's page. That is the real-time test; this page is the historical one.
Outcome accuracy
55%
Outcome accuracy
55%
Outcome accuracy
55%
Showing 128 matches — 70 correct, 58 missed
Retrospective backtest, not live picks -- these were never locked in for real money. The model parameters (K, home advantage, supremacy, rho) were tuned on >=2023 data, so 2018 and 2022 are out-of-sample for the parameters as well as for the ratings. Metrics: outcomeAccuracy is top-pick (1X2) hit rate; RPS is the ranked probability score for ordered Home>Draw>Away (lower is better, 0 perfect); Brier is multiclass (lower is better); exactScoreRate is how often the rounded predicted scoreline equalled the real scoreline. Football is high-variance and a single tournament is only 64 matches -- treat these as credibility evidence, not a guarantee.
A note on scope: the model can genuinely call a draw, and those are graded like any other outcome. A knockout that goes to extra time or penalties was level after 90′, so we grade it on two ledgers: the headline is whether the side we favoured advanced, and a separate 90-minute ledger grades the regulation read (level at 90′ is a hit only if we called a draw) — the same 90-minute basis our historical backtest uses. The scoreboard shows the final result (e.g. 3–2 a.e.t., or the level score with the shootout tally); the 90′ read is noted alongside. Only an awarded (forfeit) result is shown but never graded, since it was never played.
Outcome accuracy = how often the model's top win/draw/loss call was right. This is the number we lead with — calling the result is the model's real strength. RPS and Brier score the full probabilities — lower is better, 0 is perfect.
For context: over these same 128 matches, always picking the home side would have been right about 41% of the time, and roughly one in five results was a draw — so a single percentage is only meaningful against a baseline like that, never on its own.
These predictions are statistical estimates for context and fun — not advice, and not a betting product. A single tournament is high-variance; treat the backtest as credibility evidence, never a guarantee.