Four Matches, One Shock: Friday's World Cup AI Report Card
Friday's four Group C and D fixtures gave the models a mixed evening: two comfortable calls, one tidy away win they largely saw coming, and one result that left every single model holding a wrong scoreline pick. That last one is worth dwelling on — but so is the good news, because there was plenty of it too.
USA 2–0 Australia — Group D, Gillette Stadium
A home win for the USA in front of their own crowd was the consensus call, and it landed correctly. The spread of home-win probabilities ran from Claude's cautious 52% up to Grok's 58%, with everyone else clustered in between. All models also picked 1-0 as their exact scoreline — reasonable enough, but the USA scored twice, so no exact hits here. That's not a failure; 1-0 is always the modal low-scoring home win, and missing the extra goal is entirely normal.
| Model | P(USA) | P(Draw) | P(Australia) | Brier ↓ |
|---|---|---|---|---|
| Grok | 0.58 | 0.23 | 0.19 | 0.265 |
| Gemini | 0.55 | 0.26 | 0.19 | 0.306 |
| Ensemble | 0.54 | 0.24 | 0.21 | 0.313 |
| Naive Avg | 0.54 | 0.25 | 0.22 | 0.321 |
| Claude | 0.52 | 0.23 | 0.25 | 0.346 |
| GPT-5.4 | 0.50 | 0.26 | 0.24 | 0.375 |
Grok wins this one comfortably on Brier score, helped by its higher home-win probability. Ironically, Claude — whose known fingerprint is over-pricing home advantage — was actually the least bullish on the USA here, at just 52%, and paid for it with the worst score of the group. Claude also gave Australia a 25% chance, the highest of any model, which looks generous in hindsight. GPT-5.4 was the joint-softest on the USA and ends up last. A clean result for the models overall, with Grok doing best by leaning into what the betting markets were likely already pricing.
Scotland 0–1 Morocco — Group C, Lumen Field
This was the models' tidiest piece of collective forecasting all day. Morocco were installed as clear favourites by every model, with away-win probabilities ranging from 52% (Grok and GPT-5.4) up to 59% (Claude). The result duly arrived: Scotland 0-1 Morocco. What makes this match particularly striking is that every model with a recorded scoreline pick called 0-1 exactly — and every one of those exact picks landed. That's a genuine rarity: the scoreline was the single most likely result and it happened.
| Model | P(Scotland) | P(Draw) | P(Morocco) | Brier ↓ |
|---|---|---|---|---|
| Claude | 0.17 | 0.24 | 0.59 | 0.255 |
| Ensemble | 0.19 | 0.27 | 0.54 | 0.317 |
| Naive Avg | 0.19 | 0.27 | 0.54 | 0.317 |
| Gemini | 0.18 | 0.28 | 0.54 | 0.322 |
| Grok | 0.21 | 0.27 | 0.52 | 0.347 |
| GPT-5.4 | 0.21 | 0.27 | 0.52 | 0.347 |
Claude takes top honours here, rewarded for its most confident Morocco probability. This is a case where Claude's home-advantage fingerprint actually helped: it was less willing to credit Scotland with any home-ish edge, driving a higher Morocco number. Grok, which leans on betting markets, was slightly more hedged — those markets may have been giving Scotland a touch more credit than Claude was. A good evening for the Sonnet.
Brazil 3–0 Haiti — Group C, Lincoln Financial Field
Nobody was going to distinguish themselves with careful probabilistic reasoning here — this was always going to be a case of who stacked the most chips on Brazil and whether the mismatch rewarded them. Grok went hardest at 92%, Gemini and the ensemble sat around 88%, while Claude and GPT-5.4 were the most conservative at 83-84%. Brazil duly won 3-0, and all models called 3-0 as their exact scoreline. Given Haiti's status, 3-0 is a perfectly defensible modal pick and it landed — but this is a match where even a wrong scoreline pick would have been forgivable.
| Model | P(Brazil) | P(Draw) | P(Haiti) | Brier ↓ |
|---|---|---|---|---|
| Grok | 0.92 | 0.07 | 0.01 | 0.011 |
| Gemini | 0.88 | 0.09 | 0.03 | 0.023 |
| Ensemble | 0.88 | 0.09 | 0.03 | 0.024 |
| Naive Avg | 0.87 | 0.10 | 0.04 | 0.028 |
| GPT-5.4 | 0.84 | 0.11 | 0.05 | 0.040 |
| Claude | 0.83 | 0.11 | 0.06 | 0.045 |
Grok's betting-market lean pays off handsomely: a Brier score of just 0.011 is about as good as it gets in a three-outcome framework. There's a mild note of caution worth logging though: the reputation-versus-form bias that we've identified across all models is essentially being validated by results like this — Brazil were the right pick, and they happened to be the famous name. The models won't always be this lucky when reputation and current form diverge.
Türkiye 0–1 Paraguay — Group D, Levi's Stadium
And here's the one that stings. Paraguay won in Santa Clara, and not a single model came close to calling it. Home-win probabilities ranged from a low of 41% (GPT-5.4) to a high of 51% (Grok). Paraguay's chances were rated at just 21-28% across the board. Every model's scoreline pick was 1-0 to Türkiye. Every single one was wrong on both counts.
| Model | P(Türkiye) | P(Draw) | P(Paraguay) | Brier ↓ |
|---|---|---|---|---|
| GPT-5.4 | 0.41 | 0.31 | 0.28 | 0.783 |
| Claude | 0.44 | 0.28 | 0.28 | 0.790 |
| Naive Avg | 0.45 | 0.30 | 0.26 | 0.843 |
| Gemini | 0.42 | 0.33 | 0.25 | 0.848 |
| Ensemble | 0.46 | 0.30 | 0.25 | 0.863 |
| Grok | 0.51 | 0.28 | 0.21 | 0.963 |
GPT-5.4 and Claude emerge as the 'least wrong', for what that's worth — both gave Paraguay 28% and had the lowest Brier scores on the day. Grok had its worst result here by some distance: a Brier score of 0.963 reflects just how hard it leaned on Türkiye. That 51% home-win figure, almost certainly pulled from betting-market odds, looks like it mispriced the match badly. Gemini's league-prestige bias probably contributed too — Türkiye compete in a higher-profile European context than Paraguay's CONMEBOL environment, and that may have nudged the model's numbers upward. The collective failure here is a proper finding: the models broadly cannot separate Türkiye's reputation from their actual current-tournament form, and Paraguay — a team that doesn't carry the same name recognition — went underrated across the board.
The Türkiye–Paraguay result is Friday's sharpest lesson. Every model picked Türkiye to win; Paraguay won instead, and Grok — usually steadied by market signals — posted its worst Brier score of the day at 0.963 by going most aggressively the wrong way. When betting markets and model priors agree on a European side over a South American underdog, that consensus may be exactly the gap where value hides.
Day Summary
- 3 of 4 results called correctly across all models — the collective hit rate is solid.
- Claude was best on Scotland–Morocco (Brier 0.255) and joint-best on Türkiye–Paraguay (0.790); a quietly good evening.
- Grok was best on both USA–Australia and Brazil–Haiti, but worst of all on Türkiye–Paraguay — a high-variance day that mirrors its high-conviction style.
- The Scotland–Morocco 0-1 exact-score landslide — every model nailing the scoreline — is a genuine statistical curiosity worth tracking.
- All five models called Türkiye's scoreline wrong, which is a collective bias signal, not bad luck.
- Brazil–Haiti was essentially a free Brier-score harvest; the interesting matches are the ones where reputations are closer.
What Saturday Will Test
Tomorrow's slate will be worth watching for how the models handle teams whose early-tournament form is now on record. After one matchday each, a few sides have already overperformed or underperformed their pre-tournament ratings — the question is whether the models update meaningfully or whether reputation continues to anchor their numbers. Any match involving a team that lost on matchday one will be a direct test of whether form is getting through. Keep an eye on whether Grok's market-following strategy recovers, and whether Claude's home-advantage tendency creates any mispricing when a strong away side visits a weaker host.
Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.