Grok Thrives, Everyone Drowns in Kansas City: Saturday Review
Saturday gave us four matches spread across three time zones and two continents of scoreboard weirdness. The models went three from four on outcomes — not bad — but the Brier scores tell a more complicated story, and Ecuador vs Curacao was the kind of collective face-plant that's actually useful data.
Netherlands 5-1 Sweden — Houston
The Dutch were comfortable favourites and duly won, but nobody had 'five-goal margin' on their bingo card. Every model picked the direction correctly; the question was only how confidently. Grok led the field at 57% for a Netherlands win, and that translated into the best Brier score of the match (0.2798) — a tidy reward for conviction. The ensemble (52%) edged the naive average into second. Claude sat precisely on the coin-flip at 50%, which is almost a fingerprint in itself: when in doubt, Claude finds the nearest round number. GPT-5.4 (49%) and Gemini (48%) were the most tentative, drifting just below evens on a team that had "tournament favourite" energy.
| Model | Home % | Draw % | Away % | Brier | Result |
|---|---|---|---|---|---|
| GPT-5.4 | 49 | 27 | 24 | 0.391 | ✓ |
| Claude | 50 | 25 | 25 | 0.375 | ✓ |
| Gemini | 48 | 28 | 24 | 0.406 | ✓ |
| Grok | 57 | 25 | 18 | 0.280 | ✓ |
| Ensemble | 52 | 26 | 22 | 0.344 | ✓ |
| Naive avg | 51 | 26 | 23 | 0.361 | ✓ |
On scorelines, every model that submitted a pick went with 2-1 — a perfectly sensible guess that the actual 5-1 made look rather conservative. All registered outcome hits, none exact. That's normal: exact scores are hard and 2-1 is almost always the statistically defensible call for a moderate favourite. The real story here is Grok's confidence paying off. Its betting-market lean pushed it furthest from a draw and furthest from a Sweden win, which happened to be the right direction to lean. No complaints.
Gemini's fingerprint — over-weighting league prestige — is worth noting here. Sweden's Allsvenskan pedigree presumably kept Gemini's draw probability slightly elevated (28%, the highest of any model). It didn't cost Gemini the outcome call, but in a game that finished 5-1 that extra draw weight looks soft. Claude's known home-advantage inflation didn't obviously fire here since Netherlands were already the favourite, but Claude's ceiling of 50% on a team that just won by four is arguably a symptom of the same conservatism.
Germany 2-1 Ivory Coast — Toronto
The cleaner match of the day, and the one where the scoreline picks earned their moment. GPT-5.4, Claude, and Gemini all called 2-1 exactly — and 2-1 is what happened. Three exact hits from three models in the same match is unusual enough to be worth celebrating, even if we keep the champagne corked: 2-1 is the commonest scoreline in football and models know this. Still, getting the direction and the margin right simultaneously is the hardest target in this exercise, and they deserve credit.
| Model | Home % | Draw % | Away % | Brier | Score pick | Exact? |
|---|---|---|---|---|---|---|
| GPT-5.4 | 56 | 24 | 20 | 0.291 | 2-1 | ✓ |
| Claude | 58 | 22 | 20 | 0.265 | 2-1 | ✓ |
| Gemini | 58 | 25 | 17 | 0.268 | 2-1 | ✓ |
| Grok | 71 | 19 | 10 | 0.130 | 2-0 | ✗ |
| Ensemble | 64 | 21 | 15 | 0.201 | 2-0 | ✗ |
| Naive avg | 62 | 22 | 17 | 0.222 | — | — |
Grok was again the most bullish on Germany — 71%, a full 13 points above Claude and Gemini — and again posted the best Brier score (0.130). Betting markets clearly had Germany heavily favoured, and Grok followed suit. The irony is that Grok's own scoreline pick of 2-0 missed by one goal; the models that hedged slightly toward Ivory Coast scoring got the scoreline right. There's a lesson in there about confidence versus granularity, but one day's data doesn't make a rule. Gemini's away probability of just 17% for Ivory Coast was its lowest among the main models, consistent with its tendency to respect big European leagues — though here that bias and the actual result were pointing the same direction.
Ecuador 0-0 Curacao — Kansas City
Right. Let's talk about the disaster. Ecuador were given between 72% (Claude) and 86% (Grok) to win this match. Every single model and the naive average called a 2-0 Ecuador win. The actual result was a goalless draw. Every model was wrong on the outcome. Every model was wrong on the scoreline. This is the most instructive moment of the day.
| Model | Home % | Draw % | Away % | Brier | Result |
|---|---|---|---|---|---|
| GPT-5.4 | 78 | 14 | 8 | 1.354 | ✗ |
| Claude | 72 | 17 | 11 | 1.219 | ✗ |
| Gemini | 82 | 13 | 5 | 1.432 | ✗ |
| Grok | 86 | 11 | 3 | 1.533 | ✗ |
| Ensemble | 81 | 13 | 6 | 1.411 | ✗ |
| Naive avg | 80 | 14 | 7 | 1.380 | ✗ |
The Brier scores are ugly across the board — anything above 1.0 is a rough day, and 1.5 is a rough week. Grok is the worst performer here precisely because it was most confident: 86% on a team that drew 0-0. Claude was least wrong, at 72% home, which translates into a Brier of 1.219 — the best of a bad lot, and consistent with Claude being the most cautious on heavy favourites. This is actually an argument for Claude's conservative style: when the shock result lands, you lose less. GPT-5.4 (78%) and the naive average also did relatively less damage. Gemini and Grok, both at the aggressive end, bled the most.
The collective miss is a finding, not an embarrassment. Ecuador vs Curacao was a match where the raw quality gap between the sides was enormous — Curacao are FIFA minnows, Ecuador a respectable South American side. The models did what any reasonable forecaster would do and loaded the probability onto the expected winner. Football, infuriatingly, didn't cooperate. This is a reminder that 80% doesn't mean 'certainty'; it means one match in five ends like this. The models weren't broken — they were unlucky in a statistically foreseeable way. The concern, if you want to find one, is that the draw probability across the field was squeezed to 11-17%. Draws in tight, low-stakes group matches between mismatched teams are under-priced by reputation-driven models. Something to watch.
Tunisia 0-4 Japan — Monterrey
Japan were favourites, Japan won convincingly, and the models all got the direction right — though again, nobody predicted the margin. All five models called 0-1; the actual 0-4 is, like the Dutch result, a reminder that models calibrate for likely scorelines rather than extremes.
| Model | Home % | Draw % | Away % | Brier | Result |
|---|---|---|---|---|---|
| GPT-5.4 | 17 | 25 | 58 | 0.268 | ✓ |
| Claude | 16 | 23 | 61 | 0.231 | ✓ |
| Gemini | 14 | 25 | 61 | 0.234 | ✓ |
| Grok | 19 | 26 | 55 | 0.306 | ✓ |
| Ensemble | 17 | 25 | 59 | 0.259 | ✓ |
| Naive avg | 17 | 25 | 59 | 0.259 | ✓ |
The probabilities are unusually tightly clustered here — all models within a few points of each other. That consensus is sensible: Tunisia at home (Monterrey is not Tunis, so the 'home' tag is nominal) against a Japan side that qualified comfortably. Claude and Gemini share the best Brier score at roughly 0.23, both at 61% for Japan. Grok was the most cautious at 55%, which is interesting given its usual tendency to follow betting lines — perhaps the market was more circumspect about Tunisia than the other models expected. Gemini's 14% home probability for Tunisia is its lowest, consistent with its league-prestige weighting (Japan's J-League has climbed considerably in FIFA ranking terms). The low Tunisia home percentage from Gemini turned out well here, but it's worth flagging the mechanism: this isn't Gemini 'seeing' Tunisia's weaknesses, it's Gemini's structural preference for recognised top-flight nations.
Day summary
- Grok had its best day of the tournament so far — top Brier score in two of four matches, paid for being confident in the right direction. Ecuador cost it heavily, though.
- Claude finished second overall by being the least wrong in every match it missed — its conservatism is a genuine defensive tool, even if it leaves Brier points on the table when big favourites win.
- Gemini rode its Japan conviction well but was stung in Kansas City. Its aggressive home-win percentage for Ecuador (82%) is a symptom of over-weighting the reputation gap.
- GPT-5.4 was the most consistent across the slate — never the best, never the worst — which is its pattern. Two exact-score hits (Germany) are worth noting.
- The ensemble benefited from averaging out the extreme cases but couldn't escape the Ecuador collapse. Averaging bad priors just gets you average badness.
- Draws remain chronically under-priced. Ecuador-Curacao had draw probabilities of 11-17%; the draw happened. This is a systemic issue across the whole model field, not just one bad call.
Ecuador 0-0 Curacao is the day's key finding: when every model and the naive average converge on 80% for one outcome and all get it wrong, that's not random noise — it's evidence that the field systematically under-prices draws in lopsided group-stage matchups. Models see a quality gap, assume it translates to goals, and squeeze the draw probability into single figures. Football does not always comply.
What Sunday tests
Sunday's slate — check the fixtures — should stress a few known weaknesses. Any match involving a strong South American or African side against European opposition in a nominal 'home' slot will probe Claude's home-advantage inflation again. If there are matches between sides with significant FIFA ranking gaps, watch how tightly the draw probabilities are squeezed: after today's Ecuador lesson, that number deserves more scrutiny than the headline win probability. And if Grok continues to post the sharpest odds-aligned probabilities, we'll need more matches before deciding whether that's signal or Saturday's variance. Three from four on outcomes is a decent day. The scoreboard, though, keeps laughing at everyone's scoreline picks — and that's rather the point.
Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.