MBModelBall
June 27, 2026

Friday 26 June: France Rout, a 5-0 Demolition, and the Draw Problem

resultsgroup-stage

Friday gave us six Group Stage matches spanning three time zones and about eighteen hours of football, from Foxborough to Vancouver to Seattle. The models went 4-for-6 on outcomes — a respectable return on paper, but the two misses were both draws, and the aggregate scoreline data revealed the usual gap between knowing who wins and knowing by how much. Let's take them one by one.

Norway 1–4 France — Gillette Stadium

A comfortable call, and the models delivered. All five live models had France as clear favourites — probabilities of 0.57 to 0.60 for an away win — and France duly won, albeit rather more emphatically than the expected scoreline of 1–2 that every single model plumped for. Getting the outcome right but underestimating the margin is a familiar pattern with strong favourites; the models see France as good, not quite as *that* good.

ModelP(Norway)P(Draw)P(France)Result correct?Brier
GPT-5.40.180.220.600.241
Gemini0.180.220.600.241
Claude0.190.220.590.253
Grok0.190.240.570.279
Ensemble0.190.230.590.253

GPT-5.4 and Gemini were joint-best here, sharing a Brier of 0.241. Grok was the softest on France at 0.57 — in a match with no meaningful home-side advantage to bias towards, that difference is harder to attribute to a specific fingerprint. Claude's known home-advantage inflation didn't really show here either; Norway in a neutral-ish US venue presumably didn't trigger that particular habit strongly enough to matter. All scoreline picks landed on 1–2; outcome-hit, exact-miss — which is the normal state of affairs when a team wins by four.

Senegal 5–0 Iraq — BMO Field, Toronto

The models knew this one. The question was only how confident to be, and Grok answered that most boldly: 0.81 for a Senegal win, against Gemini and Claude at 0.72 and GPT-5.4 at 0.69. Grok's confidence was rewarded handsomely — a Brier of 0.057, the best single-match score of the day by some distance.

ModelP(Senegal)P(Draw)P(Iraq)Result correct?Brier
GPT-5.40.690.190.120.147
Gemini0.720.200.080.125
Claude0.720.170.110.119
Grok0.810.130.060.057
Ensemble0.750.160.090.098

Every model picked 2–0 as their scoreline. They got the outcome right and the rough shape right (clean sheet for Senegal), but five goals is five goals. Grok's market-anchored aggression paid off here: in a match where the betting markets would have been similarly bullish, leaning into that signal was the correct play. Gemini's league-prestige bias is worth flagging in passing — Senegal are an established African top side, which Gemini would likely weight heavily — but with all models agreeing roughly, it's hard to call that a fingerprint moment rather than just a correct read.

Cape Verde 0–0 Saudi Arabia — NRG Stadium, Houston

Here the models ran into the draw problem. With three-way probabilities hovering around 0.31–0.38 for each outcome, this was as close to a genuine coin-flip as you get, and the models treated it accordingly. The draw — the outcome with the narrowest probability band across all models — duly arrived and punished everyone equally.

ModelP(Cape Verde)P(Draw)P(Saudi Arabia)Result correct?Brier
GPT-5.40.340.310.350.714
Gemini0.380.310.310.717
Claude0.340.310.350.714
Grok0.380.310.310.717
Ensemble0.360.310.330.715

The Brier scores cluster tightly around 0.714–0.717 because all models gave roughly the same probability (0.31) to the outcome that occurred. There's no winner or loser among the models here; they were equally wrong, equally honestly. Scoreline picks tell a mild story: GPT-5.4 and Claude picked 0–1 (Saudi Arabia), Gemini, Grok, and the Ensemble picked 1–0 (Cape Verde) — meaning the field was split on *which* non-draw outcome to prefer, which is a decent signal that nobody had strong conviction. The actual 0–0 was the one result everyone rated least likely to be the *scoreline*, even if the draw outcome itself was on the table. This is a finding, not an embarrassment: three-way near-uniform matches are genuinely hard, and models that confidently predicted a winner here would be the ones to distrust.

Uruguay 0–1 Spain — Estadio Akron, Zapopan

A good day for Claude, a bad day for Grok. Spain were favourites across the board, but the spread of convictions was notable: Claude at 0.58 for a Spanish win, Grok at only 0.48. That gap is significant — Grok was effectively refusing to commit to either side, treating Uruguay as roughly co-favourites, which Spain's actual win proved overly generous to the hosts.

ModelP(Uruguay)P(Draw)P(Spain)Result correct?Brier
GPT-5.40.180.260.560.294
Gemini0.200.250.550.305
Claude0.180.240.580.266
Grok0.230.290.480.407
Ensemble0.200.260.540.316

Claude's best match of the day by a clear margin: Brier of 0.266, compared to Grok's 0.407. Grok's known tendency to anchor heavily on betting-market odds appears to have backfired here — the markets presumably respected Uruguay's World Cup pedigree, keeping their implied probability high and dragging Grok's Spain number down. Claude, meanwhile, was the most willing to back Spain at 0.58. Worth noting: four models plus the ensemble all picked 0–1 as their exact scoreline, and that is precisely what happened. A clean sweep of exact hits — which is genuinely uncommon and worth a moment's appreciation, even if the individual probability of any exact score is always low.

Egypt 1–1 Iran — Lumen Field, Seattle

The second draw of the day, and another collective miss — though a more forgivable one than Cape Verde vs Saudi Arabia. The models gave Egypt a meaningful edge (0.36–0.41 for a home win) against Iran (0.23–0.30), with draws in the 0.34–0.36 range. The draw was the second-most-likely outcome, and it happened. The log-loss scores are heavy (~1.02–1.08) because no model had the draw as their top pick.

ModelP(Egypt)P(Draw)P(Iran)Result correct?Brier
GPT-5.40.360.340.300.655
Gemini0.380.360.260.622
Claude0.380.350.270.640
Grok0.410.360.230.631
Ensemble0.390.350.260.636

Gemini was closest, Brier 0.622, by dint of having the highest draw probability (0.36) while keeping home probability at 0.38. Grok was most convinced by Egypt (0.41) — perhaps the home side's AFCON profile and African football profile inflated things slightly, though in this case all models were directionally similar. Every model picked Egypt 1–0 as their scoreline; the actual 1–1 is an outcome-miss but a scoreline near-miss, which reflects the genuine uncertainty here. The two draws on Friday are a reminder that models broadly assigned ~30–36% to draws in both matches and still missed — which is statistically expected over a small sample, not evidence of a broken draw model.

New Zealand 1–5 Belgium — BC Place, Vancouver

The last match of the day's slate, and probably the most comfortable call of all six. Belgium were massive favourites, with Claude most bullish at 0.81 and GPT-5.4 and Grok both at 0.77. The actual result — a 1–5 Belgium win — vindicated the high confidence fully.

ModelP(New Zealand)P(Draw)P(Belgium)Result correct?Brier
GPT-5.40.080.150.770.082
Gemini0.070.130.800.062
Claude0.060.130.810.057
Grok0.080.150.770.082
Ensemble0.070.140.790.070

Claude's best Brier of the match at 0.057, sharing the day's top score with Grok's Senegal call. Interestingly, Claude — normally the model most inclined to over-inflate home advantage — here gave New Zealand only 0.06, the lowest home probability of any model. That is either Claude correctly reading that New Zealand at a World Cup are not a meaningful home-advantage side (they're playing in Canada, not New Zealand), or the fact that Belgium's pedigree simply overwhelmed any home-field nudge. Either way, the result is Claude's clearest win of the day. All scoreline picks underestimated the final margin — 0–2 or 0–3 versus an actual 1–5 — but the outcome was never in serious doubt.

Day Summary

  • Outcome accuracy: 4 from 6 correct. Both misses were draws — Cape Verde vs Saudi Arabia and Egypt vs Iran.
  • Best model on the day: Claude, by a narrow margin, led on Uruguay vs Spain and New Zealand vs Belgium and avoided the worst Brier scores elsewhere.
  • Worst model on the day: Grok, whose soft 0.48 for Spain in Zapopan proved most costly. Its market-anchoring appeared to keep Uruguay's odds inflated beyond what form merited.
  • Scoreline exact hits: Four models (Claude, Gemini, GPT-5.4, Grok, Ensemble) landed Uruguay 0–1 Spain as an exact score. Given how hard exact scores are, that is a notable collective win.
  • Draw blindness: Two draws, two collective misses. Both had draw probabilities of ~0.31–0.36, meaning the models weren't ignoring the possibility — they just didn't have it as their mode. That is the correct epistemic stance, but it still burns in the Brier scoring.
  • Margin calibration: France 1–4, Belgium 1–5, Senegal 5–0 all exceeded the scoreline picks. Models remain conservative about the scale of one-sided matches.

Uruguay 0–1 Spain produced a mass exact-score hit: Claude, Gemini, GPT-5.4, Grok, and the Ensemble all locked in 0–1 before kickoff. It happened. Exact scores are low-probability targets even when you're directionally right, so a clean sweep across all models on one line is unusual enough to flag — and it coincided with Claude's best Brier performance of the day, suggesting that when the models agree strongly on both outcome and margin, they're occasionally doing something right beyond luck.

What Tomorrow Tests

Saturday's slate brings further Group Stage conclusions, likely including matches where qualification stakes sharpen and teams may play conservatively — which historically produces more draws than open football. If that pattern holds, the models' structural preference for decisive outcomes will be tested again. Watch also for any matches involving hosts USA or Mexico, where Claude's home-advantage inflation could become material in a way it wasn't for Norway in Massachusetts. And if Grok continues anchoring to markets in matches where form and reputation diverge from the odds, we'll learn whether Friday's Spain result was signal or noise.

Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.

Discussion