Three matches, two draws, and a masterclass in how reputation-weighted models struggle when big sides fail to convert. Monday was a day of partial credit and collective embarrassment.
A single match on Sunday's slate, and the models got it right — Canada beat South Africa 0-1. The more interesting story is who was most confident, and why.
Six matches across Friday and the early hours of Saturday tested the models on everything from heavy favourites to coin-flip three-way splits — and draws, as ever, remained the system's blind spot.
Thursday's six-match slate delivered two significant upsets against heavily-favoured opponents, exposing a shared blind spot across all models for underdog home wins in high-stakes group games. The easy calls landed cleanly; the hard ones hurt.
Wednesday's six-game slate produced four comfortable home/favourite wins and one collective embarrassment as South Africa stunned every model by beating South Korea. Grok led on the day; nobody saw Bafana Bafana coming.
Tuesday's four matches gave the models a straightforward Portuguese hammering, a Croatia win they all expected, a tight Colombian victory they mostly called correctly — and then England vs Ghana arrived and made everyone look foolish. Failed predictions are findings, and today produced a big one.
Four matches on Monday, four home wins — but the probabilities told very different stories. Grok's betting-market lean paid handsomely against Argentina and France; Norway's nervy 3-2 over Senegal left everyone underconfident.
Four matches on Sunday produced one comfortable Spain demolition and two stubborn draws that punished every model. The day's scorecard: two wins, two costly misses, and a reminder that collective confidence is not collective wisdom.
A five-goal Dutch demolition, three exact-score hits in Toronto, and a 0-0 nobody saw coming made Saturday a day of contrasts — Grok led on the numbers, but Ecuador vs Curacao humbled the whole field.
Three confident calls and one collective faceplant made for a revealing Friday. Paraguay's win over Turkey exposed every model's blind spot in one clean 0-1.
Thursday's four-match slate gave the models two comfortable wins, one collective miss on a draw, and a scoreline that made every exact-score pick look timid. Canada 6–0 Qatar will be talked about for a while.
Tuesday's four Group I and J matches all went to the pre-match favourite, giving every model a clean results sheet — yet the Brier scores reveal a wide spread in how confidently, and correctly, they called it.
DR Congo held Portugal to a draw that no model really wanted to believe in, while England's 4-2 demolition of Croatia papered over some shaky probability estimates. Two findings from a busy Wednesday slate.
Before the opening whistle at the Azteca, all 360 group-stage predictions are logged and timestamped. England edges Argentina as favourite, Gemini backs Portugal, and all five models are cool on Brazil.
The 25-question Tactical Knowledge panel puts GPT-5.5 at the top of the field at 8.67/10 — but only after fixing a 500-token cap that had been silently hiding its responses.
League data isn't World Cup data. But the patterns we found point to which models will be reliable in June, where the corrections will matter most, and where the system might break.
Claude crushes it in La Liga, struggles in Bundesliga. Grok is steady but unspectacular. GPT-5.4 surprises in MLS. The full breakdown by model and league.
A pattern emerged across 18 leagues: the further a league sits from the AI training-data centre of gravity, the more bias correction helps. The map matters.
A 13% Brier improvement, every model lifting, no model degrading. La Liga gave us our cleanest validation of the methodology — here's what it tells us about Spain at the World Cup.
Of every league we tested, the Premier League moved the least when we corrected for bias. We think we know why — and what it means for English clubs at the World Cup.
Our research reveals systematic bias toward Big 5 league players. When given identical stats, models prefer the player from the more prestigious league 58-71% of the time.
An introduction to GPT-5.4, GPT-5.5, Claude, Grok, and Gemini — their personalities, strengths, and blind spots. Understanding why they disagree is key to our methodology.
We tested our prediction methodology on international friendlies from March-April 2026. The ensemble beat market odds, and models showed distinct behavioral patterns.