MBModelBall
June 20, 2026

Four Matches, One Shock: Friday's World Cup AI Report Card

resultsgroup-stage

Friday's four Group C and D fixtures gave the models a mixed evening: two comfortable calls, one tidy away win they largely saw coming, and one result that left every single model holding a wrong scoreline pick. That last one is worth dwelling on — but so is the good news, because there was plenty of it too.

USA 2–0 Australia — Group D, Gillette Stadium

A home win for the USA in front of their own crowd was the consensus call, and it landed correctly. The spread of home-win probabilities ran from Claude's cautious 52% up to Grok's 58%, with everyone else clustered in between. All models also picked 1-0 as their exact scoreline — reasonable enough, but the USA scored twice, so no exact hits here. That's not a failure; 1-0 is always the modal low-scoring home win, and missing the extra goal is entirely normal.

ModelP(USA)P(Draw)P(Australia)Brier ↓
Grok0.580.230.190.265
Gemini0.550.260.190.306
Ensemble0.540.240.210.313
Naive Avg0.540.250.220.321
Claude0.520.230.250.346
GPT-5.40.500.260.240.375

Grok wins this one comfortably on Brier score, helped by its higher home-win probability. Ironically, Claude — whose known fingerprint is over-pricing home advantage — was actually the least bullish on the USA here, at just 52%, and paid for it with the worst score of the group. Claude also gave Australia a 25% chance, the highest of any model, which looks generous in hindsight. GPT-5.4 was the joint-softest on the USA and ends up last. A clean result for the models overall, with Grok doing best by leaning into what the betting markets were likely already pricing.

Scotland 0–1 Morocco — Group C, Lumen Field

This was the models' tidiest piece of collective forecasting all day. Morocco were installed as clear favourites by every model, with away-win probabilities ranging from 52% (Grok and GPT-5.4) up to 59% (Claude). The result duly arrived: Scotland 0-1 Morocco. What makes this match particularly striking is that every model with a recorded scoreline pick called 0-1 exactly — and every one of those exact picks landed. That's a genuine rarity: the scoreline was the single most likely result and it happened.

ModelP(Scotland)P(Draw)P(Morocco)Brier ↓
Claude0.170.240.590.255
Ensemble0.190.270.540.317
Naive Avg0.190.270.540.317
Gemini0.180.280.540.322
Grok0.210.270.520.347
GPT-5.40.210.270.520.347

Claude takes top honours here, rewarded for its most confident Morocco probability. This is a case where Claude's home-advantage fingerprint actually helped: it was less willing to credit Scotland with any home-ish edge, driving a higher Morocco number. Grok, which leans on betting markets, was slightly more hedged — those markets may have been giving Scotland a touch more credit than Claude was. A good evening for the Sonnet.

Brazil 3–0 Haiti — Group C, Lincoln Financial Field

Nobody was going to distinguish themselves with careful probabilistic reasoning here — this was always going to be a case of who stacked the most chips on Brazil and whether the mismatch rewarded them. Grok went hardest at 92%, Gemini and the ensemble sat around 88%, while Claude and GPT-5.4 were the most conservative at 83-84%. Brazil duly won 3-0, and all models called 3-0 as their exact scoreline. Given Haiti's status, 3-0 is a perfectly defensible modal pick and it landed — but this is a match where even a wrong scoreline pick would have been forgivable.

ModelP(Brazil)P(Draw)P(Haiti)Brier ↓
Grok0.920.070.010.011
Gemini0.880.090.030.023
Ensemble0.880.090.030.024
Naive Avg0.870.100.040.028
GPT-5.40.840.110.050.040
Claude0.830.110.060.045

Grok's betting-market lean pays off handsomely: a Brier score of just 0.011 is about as good as it gets in a three-outcome framework. There's a mild note of caution worth logging though: the reputation-versus-form bias that we've identified across all models is essentially being validated by results like this — Brazil were the right pick, and they happened to be the famous name. The models won't always be this lucky when reputation and current form diverge.

Türkiye 0–1 Paraguay — Group D, Levi's Stadium

And here's the one that stings. Paraguay won in Santa Clara, and not a single model came close to calling it. Home-win probabilities ranged from a low of 41% (GPT-5.4) to a high of 51% (Grok). Paraguay's chances were rated at just 21-28% across the board. Every model's scoreline pick was 1-0 to Türkiye. Every single one was wrong on both counts.

ModelP(Türkiye)P(Draw)P(Paraguay)Brier ↓
GPT-5.40.410.310.280.783
Claude0.440.280.280.790
Naive Avg0.450.300.260.843
Gemini0.420.330.250.848
Ensemble0.460.300.250.863
Grok0.510.280.210.963

GPT-5.4 and Claude emerge as the 'least wrong', for what that's worth — both gave Paraguay 28% and had the lowest Brier scores on the day. Grok had its worst result here by some distance: a Brier score of 0.963 reflects just how hard it leaned on Türkiye. That 51% home-win figure, almost certainly pulled from betting-market odds, looks like it mispriced the match badly. Gemini's league-prestige bias probably contributed too — Türkiye compete in a higher-profile European context than Paraguay's CONMEBOL environment, and that may have nudged the model's numbers upward. The collective failure here is a proper finding: the models broadly cannot separate Türkiye's reputation from their actual current-tournament form, and Paraguay — a team that doesn't carry the same name recognition — went underrated across the board.

The Türkiye–Paraguay result is Friday's sharpest lesson. Every model picked Türkiye to win; Paraguay won instead, and Grok — usually steadied by market signals — posted its worst Brier score of the day at 0.963 by going most aggressively the wrong way. When betting markets and model priors agree on a European side over a South American underdog, that consensus may be exactly the gap where value hides.

Day Summary

  • 3 of 4 results called correctly across all models — the collective hit rate is solid.
  • Claude was best on Scotland–Morocco (Brier 0.255) and joint-best on Türkiye–Paraguay (0.790); a quietly good evening.
  • Grok was best on both USA–Australia and Brazil–Haiti, but worst of all on Türkiye–Paraguay — a high-variance day that mirrors its high-conviction style.
  • The Scotland–Morocco 0-1 exact-score landslide — every model nailing the scoreline — is a genuine statistical curiosity worth tracking.
  • All five models called Türkiye's scoreline wrong, which is a collective bias signal, not bad luck.
  • Brazil–Haiti was essentially a free Brier-score harvest; the interesting matches are the ones where reputations are closer.

What Saturday Will Test

Tomorrow's slate will be worth watching for how the models handle teams whose early-tournament form is now on record. After one matchday each, a few sides have already overperformed or underperformed their pre-tournament ratings — the question is whether the models update meaningfully or whether reputation continues to anchor their numbers. Any match involving a team that lost on matchday one will be a direct test of whether form is getting through. Keep an eye on whether Grok's market-following strategy recovers, and whether Claude's home-advantage tendency creates any mispricing when a strong away side visits a weaker host.

Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.

Discussion