MBModelBall
June 19, 2026

Four Matches, One Big Draw, and a Canadian Demolition

resultsgroup-stage

Thursday stretched across nearly twelve hours of football, from Atlanta at midday Eastern to Vancouver deep into the night. The models finished 2-for-2 on outcome calls where they had strong convictions, and 0-for-1 where they were genuinely uncertain — which is roughly what you'd hope for, if not quite what you'd celebrate. The Czechia–South Africa draw exposed a collective blind spot, while Mexico–South Korea delivered the day's most satisfying forecasting curiosity: the right result, the right scoreline, from models that weren't especially confident about either.

Match 1 — Czechia 1–1 South Africa (Atlanta)

The draw that everyone assigned the lowest probability to is, of course, the one that happened. Every model had Czechia as the likeliest winner, with draw probabilities clustering in the high-20s to 30 per cent. That's not a negligent forecast — draws are genuinely hard to price — but it does mean every model got this wrong, and some got it more expensively wrong than others.

ModelP(Home)P(Draw)P(Away)Draw prob assignedBrier
GPT-5.40.460.290.250.290.778
Claude0.480.270.250.270.826
Gemini0.480.300.220.300.769
Grok0.510.280.210.280.823
Ensemble0.490.280.230.280.802
Naive Avg0.480.290.230.290.798

Gemini came closest, with the highest draw probability (30%) and the best Brier score (0.769). It's a marginal victory in a losing battle, but a win nonetheless. Claude was the most wrong — worst Brier (0.826) and the lowest draw probability (27%), compounded by a log-loss of 1.309 that reflects how confidently it pointed away from the correct outcome. This is consistent with Claude's known fingerprint: it tends to over-price home advantage, and here it did, suppressing the draw probability relative to its peers. Grok, leaning on betting-market signals, also assigned a relatively low draw probability (28%) while simultaneously giving Czechia the highest home-win odds (51%) — the market apparently agreed with the direction, just not the result.

All five models picked 1–0 as their exact-score selection. The actual scoreline was 1–1, which is the closest a near-miss can get — same total goals, one redistributed. It's the kind of result that makes exact-score forecasting feel both reasonable and futile in equal measure.

The broader point: no model disagreed meaningfully from the others here. When the whole panel clusters within four percentage points on a probability, there's no informative signal in the disagreement — just a shared, understandable mistake. South Africa were priced as genuine underdogs, and a point against Czechia is a creditable result for them. The models weren't obviously wrong to favour Czechia; they were just reminded that 28–30% events happen roughly three times in ten.

Match 2 — Switzerland 4–1 Bosnia and Herzegovina (Inglewood)

No probability data was locked for this match — the models' scoreline picks are the main record. Every model called Switzerland 1–0, got the outcome right, and got the margin emphatically wrong. Switzerland didn't just win; they won by three goals in what turned out to be a comfortable evening. The outcome-hit flags are all green, the exact-hit flags are all red, and that's a fair summary of the situation.

ModelScore PickOutcome HitExact HitBrierP(Home)
Grok1-00.1610.68
Ensemble1-00.2370.61
Naive Avg1-00.2570.59
Gemini1-00.2720.58
GPT-5.41-00.3060.55
Claude1-00.3060.55

Grok was the clear winner on Brier score (0.161), having assigned Switzerland a 68% home-win probability — the highest of any model. This is the positive side of Grok's betting-market bias: when the market has correctly identified a dominant favourite, leaning into those odds pays off. The 0.10 Bosnia probability was lowest in the field, and Bosnia duly obliged by losing by three. Claude and GPT-5.4 tied for last (Brier 0.306), both assigning Switzerland just 55% — a relatively cautious call on a team that ultimately looked anything but cautious. Claude's subdued confidence here is mildly interesting: it didn't over-price home advantage unusually on this occasion, sitting ten percentage points below Grok.

The 1–0 exact-score consensus is worth a brief note. It's the default conservative pick, and in a 4–1 game it looks particularly modest. Models systematically underestimate margin in high-asymmetry fixtures — a known limitation that Thursday's result illustrates neatly.

Match 3 — Canada 6–0 Qatar (Vancouver)

Six-nil. Every model called Canada to win, every model called 2–0. Both decisions were correct in direction and conservative beyond measure. This was a tournament-day result that no probability distribution would have given meaningful weight to beforehand, and that's not a criticism of the models — it's just what 6–0 scorelines do to forecasters.

ModelScore PickOutcome HitExact HitBrierP(Home)
Gemini2-00.1210.72
Grok2-00.1300.71
Ensemble2-00.1390.70
Naive Avg2-00.1400.70
Claude2-00.1550.68
GPT-5.42-00.1550.68
GPT-5.52-0

Gemini edged the leaderboard with a Brier of 0.121 and the highest home-win probability (72%). Grok was close behind at 0.130. The spread between best and worst is narrow — everyone had Canada between 68% and 72% — which tells you this wasn't a match where model disagreement was meaningful. Qatar were correctly identified as substantial underdogs; the only thing nobody anticipated was quite how substantial the winning margin would be. It is worth noting that Gemini's top performance here fits its known tendency to weight league and competitive prestige: Canada, playing a home World Cup on home soil against a team with essentially no top-level competitive pedigree at this stage, was exactly the fixture Gemini's fingerprint is designed to handle well.

GPT-5.5 appears in the scoreline picks here — outcome hit confirmed, no exact-hit data — the first appearance of that model in Thursday's slate.

Match 4 — Mexico 1–0 South Korea (Zapopan)

The day's most interesting forecasting exercise. This was a genuinely open match — no model gave Mexico more than 49% — and yet the majority of models still landed the correct outcome. More strikingly, every model with a locked scoreline pick chose 1–0, and that is exactly what happened. A five-way exact-score hit on a tightly contested match is the kind of coincidence that makes you briefly wonder whether these systems know something, before remembering that 1–0 is the single most common scoreline in competitive football and the models' default when they expect a narrow home win.

ModelP(Home)P(Draw)P(Away)Result CorrectBrier
Grok0.490.300.210.394
Ensemble0.450.300.260.459
Naive Avg0.440.300.270.475
GPT-5.40.420.290.290.505
Gemini0.420.310.270.505
Claude0.420.280.300.505

Grok was comfortably the best model here — Brier 0.394, the only sub-0.40 score, built on a 49% home probability that was notably higher than every other model's 42%. The betting-market lean paid off again: Mexico at home in Guadalajara, regardless of how the form book read, carries genuine structural weight in the odds. Grok read that signal and backed it. Claude, GPT-5.4, and Gemini all tied for last place on Brier (0.505), each giving Mexico just 42%. Claude's 30% away probability is worth noting — fractionally the highest South Korea probability in the field, which is mildly at odds with its supposed home-advantage inflation. In this case, Claude appears to have balanced its home-advantage tendency against genuine uncertainty about two evenly matched sides, landing at the same home probability as its peers but distributing the remainder slightly differently.

The exact-score story deserves a sentence or two. Five models correctly called 1–0 — including the ensemble — which is a genuine collective hit on a hard target. It doesn't change the Brier scores (exact-score accuracy isn't baked into the primary metrics), but it's the kind of outcome that catches the eye. Whether it reflects insight or the base-rate dominance of 1–0 in this type of fixture is a question worth revisiting across a larger sample.

Thursday in Summary

  • Grok had the best day overall — won Mexico and Switzerland on Brier, competitive on Canada, only lost the Czechia draw (which everyone lost).
  • Gemini topped Canada and came closest on the Czechia draw. Its league-prestige weighting helped it identify Qatar as a heavy underdog accurately.
  • Claude had the worst day — most wrong on the draw (lowest draw probability, highest log-loss) and joint-last on Mexico and Switzerland.
  • GPT-5.4 was consistent but unspectacular — joint-last on Switzerland and Mexico, mid-table on Czechia.
  • The ensemble performed as designed: rarely the best, rarely the worst, reasonable on every match.
  • Exact-score hits: Mexico 1–0 landed for all five models. Switzerland 4–1 and Canada 6–0 made every exact pick look timid. Czechia 1–1 was the closest near-miss of the day (everyone picked 1–0).

The day's sharpest finding: on the one match where models had genuine probability data and genuine disagreement — Mexico vs South Korea — Grok's betting-market lean produced a 7-point home-probability advantage over its rivals and the best Brier score of any model across any match on Thursday. That is Grok's fingerprint working exactly as advertised. The flip side is Czechia–South Africa, where that same market-confidence suppressed draw probability and left Grok joint-worst. Betting odds are a good prior; they are not a sufficient one.

What Friday's Slate Will Test

Friday brings more Group A and B action, which means the models will be updating in the context of results they've already seen — though their locked probabilities, as always, were set beforehand. Watch for whether Gemini's prestige-weighting causes it to over-favour any European side against a non-traditional opponent; Thursday's Canada result suggests models broadly handled clear quality gaps well, but mid-tier matchups are where the biases bite hardest. The draw market will also be worth scrutinising: after Thursday's collective miss on Czechia–South Africa, the temptation to say 'models undervalue draws' is real, but a sample of one doesn't make a pattern. We'll need to see whether the draw-probability suppression is systematic before drawing conclusions. Thursday gave us enough to stay interested.

Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.

Discussion