MBModelBall
June 21, 2026

Grok Thrives, Everyone Drowns in Kansas City: Saturday Review

resultsgroup-stage

Saturday gave us four matches spread across three time zones and two continents of scoreboard weirdness. The models went three from four on outcomes — not bad — but the Brier scores tell a more complicated story, and Ecuador vs Curacao was the kind of collective face-plant that's actually useful data.

Netherlands 5-1 Sweden — Houston

The Dutch were comfortable favourites and duly won, but nobody had 'five-goal margin' on their bingo card. Every model picked the direction correctly; the question was only how confidently. Grok led the field at 57% for a Netherlands win, and that translated into the best Brier score of the match (0.2798) — a tidy reward for conviction. The ensemble (52%) edged the naive average into second. Claude sat precisely on the coin-flip at 50%, which is almost a fingerprint in itself: when in doubt, Claude finds the nearest round number. GPT-5.4 (49%) and Gemini (48%) were the most tentative, drifting just below evens on a team that had "tournament favourite" energy.

ModelHome %Draw %Away %BrierResult
GPT-5.44927240.391
Claude5025250.375
Gemini4828240.406
Grok5725180.280
Ensemble5226220.344
Naive avg5126230.361

On scorelines, every model that submitted a pick went with 2-1 — a perfectly sensible guess that the actual 5-1 made look rather conservative. All registered outcome hits, none exact. That's normal: exact scores are hard and 2-1 is almost always the statistically defensible call for a moderate favourite. The real story here is Grok's confidence paying off. Its betting-market lean pushed it furthest from a draw and furthest from a Sweden win, which happened to be the right direction to lean. No complaints.

Gemini's fingerprint — over-weighting league prestige — is worth noting here. Sweden's Allsvenskan pedigree presumably kept Gemini's draw probability slightly elevated (28%, the highest of any model). It didn't cost Gemini the outcome call, but in a game that finished 5-1 that extra draw weight looks soft. Claude's known home-advantage inflation didn't obviously fire here since Netherlands were already the favourite, but Claude's ceiling of 50% on a team that just won by four is arguably a symptom of the same conservatism.

Germany 2-1 Ivory Coast — Toronto

The cleaner match of the day, and the one where the scoreline picks earned their moment. GPT-5.4, Claude, and Gemini all called 2-1 exactly — and 2-1 is what happened. Three exact hits from three models in the same match is unusual enough to be worth celebrating, even if we keep the champagne corked: 2-1 is the commonest scoreline in football and models know this. Still, getting the direction and the margin right simultaneously is the hardest target in this exercise, and they deserve credit.

ModelHome %Draw %Away %BrierScore pickExact?
GPT-5.45624200.2912-1
Claude5822200.2652-1
Gemini5825170.2682-1
Grok7119100.1302-0
Ensemble6421150.2012-0
Naive avg6222170.222

Grok was again the most bullish on Germany — 71%, a full 13 points above Claude and Gemini — and again posted the best Brier score (0.130). Betting markets clearly had Germany heavily favoured, and Grok followed suit. The irony is that Grok's own scoreline pick of 2-0 missed by one goal; the models that hedged slightly toward Ivory Coast scoring got the scoreline right. There's a lesson in there about confidence versus granularity, but one day's data doesn't make a rule. Gemini's away probability of just 17% for Ivory Coast was its lowest among the main models, consistent with its tendency to respect big European leagues — though here that bias and the actual result were pointing the same direction.

Ecuador 0-0 Curacao — Kansas City

Right. Let's talk about the disaster. Ecuador were given between 72% (Claude) and 86% (Grok) to win this match. Every single model and the naive average called a 2-0 Ecuador win. The actual result was a goalless draw. Every model was wrong on the outcome. Every model was wrong on the scoreline. This is the most instructive moment of the day.

ModelHome %Draw %Away %BrierResult
GPT-5.4781481.354
Claude7217111.219
Gemini821351.432
Grok861131.533
Ensemble811361.411
Naive avg801471.380

The Brier scores are ugly across the board — anything above 1.0 is a rough day, and 1.5 is a rough week. Grok is the worst performer here precisely because it was most confident: 86% on a team that drew 0-0. Claude was least wrong, at 72% home, which translates into a Brier of 1.219 — the best of a bad lot, and consistent with Claude being the most cautious on heavy favourites. This is actually an argument for Claude's conservative style: when the shock result lands, you lose less. GPT-5.4 (78%) and the naive average also did relatively less damage. Gemini and Grok, both at the aggressive end, bled the most.

The collective miss is a finding, not an embarrassment. Ecuador vs Curacao was a match where the raw quality gap between the sides was enormous — Curacao are FIFA minnows, Ecuador a respectable South American side. The models did what any reasonable forecaster would do and loaded the probability onto the expected winner. Football, infuriatingly, didn't cooperate. This is a reminder that 80% doesn't mean 'certainty'; it means one match in five ends like this. The models weren't broken — they were unlucky in a statistically foreseeable way. The concern, if you want to find one, is that the draw probability across the field was squeezed to 11-17%. Draws in tight, low-stakes group matches between mismatched teams are under-priced by reputation-driven models. Something to watch.

Tunisia 0-4 Japan — Monterrey

Japan were favourites, Japan won convincingly, and the models all got the direction right — though again, nobody predicted the margin. All five models called 0-1; the actual 0-4 is, like the Dutch result, a reminder that models calibrate for likely scorelines rather than extremes.

ModelHome %Draw %Away %BrierResult
GPT-5.41725580.268
Claude1623610.231
Gemini1425610.234
Grok1926550.306
Ensemble1725590.259
Naive avg1725590.259

The probabilities are unusually tightly clustered here — all models within a few points of each other. That consensus is sensible: Tunisia at home (Monterrey is not Tunis, so the 'home' tag is nominal) against a Japan side that qualified comfortably. Claude and Gemini share the best Brier score at roughly 0.23, both at 61% for Japan. Grok was the most cautious at 55%, which is interesting given its usual tendency to follow betting lines — perhaps the market was more circumspect about Tunisia than the other models expected. Gemini's 14% home probability for Tunisia is its lowest, consistent with its league-prestige weighting (Japan's J-League has climbed considerably in FIFA ranking terms). The low Tunisia home percentage from Gemini turned out well here, but it's worth flagging the mechanism: this isn't Gemini 'seeing' Tunisia's weaknesses, it's Gemini's structural preference for recognised top-flight nations.

Day summary

  • Grok had its best day of the tournament so far — top Brier score in two of four matches, paid for being confident in the right direction. Ecuador cost it heavily, though.
  • Claude finished second overall by being the least wrong in every match it missed — its conservatism is a genuine defensive tool, even if it leaves Brier points on the table when big favourites win.
  • Gemini rode its Japan conviction well but was stung in Kansas City. Its aggressive home-win percentage for Ecuador (82%) is a symptom of over-weighting the reputation gap.
  • GPT-5.4 was the most consistent across the slate — never the best, never the worst — which is its pattern. Two exact-score hits (Germany) are worth noting.
  • The ensemble benefited from averaging out the extreme cases but couldn't escape the Ecuador collapse. Averaging bad priors just gets you average badness.
  • Draws remain chronically under-priced. Ecuador-Curacao had draw probabilities of 11-17%; the draw happened. This is a systemic issue across the whole model field, not just one bad call.

Ecuador 0-0 Curacao is the day's key finding: when every model and the naive average converge on 80% for one outcome and all get it wrong, that's not random noise — it's evidence that the field systematically under-prices draws in lopsided group-stage matchups. Models see a quality gap, assume it translates to goals, and squeeze the draw probability into single figures. Football does not always comply.

What Sunday tests

Sunday's slate — check the fixtures — should stress a few known weaknesses. Any match involving a strong South American or African side against European opposition in a nominal 'home' slot will probe Claude's home-advantage inflation again. If there are matches between sides with significant FIFA ranking gaps, watch how tightly the draw probabilities are squeezed: after today's Ecuador lesson, that number deserves more scrutiny than the headline win probability. And if Grok continues to post the sharpest odds-aligned probabilities, we'll need more matches before deciding whether that's signal or Saturday's variance. Three from four on outcomes is a decent day. The scoreboard, though, keeps laughing at everyone's scoreline picks — and that's rather the point.

Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.

Discussion