Six Matches, One Shocker: Models Navigate a Wild Saturday
Saturday's six-match slate ran the full gamut: a routine Croatia win, a dominant England stroll, an Argentina procession, a DR Congo upset, a six-goal thriller that ended level, and the evening's headline act — Colombia holding Portugal to a goalless draw that every single model got wrong. A productive day for collecting findings, if not always for scoring points.
Croatia 2–1 Ghana (Group L, Philadelphia)
Croatia won, which all models expected, but the spread of confidence told its own story. Grok led the field at 56% for a Croatia win, tracking the betting market (52.6%) most closely — consistent with its known tendency to anchor on odds. Claude and GPT-5.5 were the most cautious at 48%, giving Ghana a non-trivial 22–23% shot. Claude's relative reluctance to back Croatia is slightly at odds with its fingerprint of over-pricing home advantage; perhaps that bias is most pronounced when home sides are clear favourites rather than moderate ones. Grok's Brier score of 0.300 was the best among the AI models, with Gemini (0.335) second. The market edged both with 0.344, but this was a close enough game that no one should feel smug.
| Model | Croatia % | Draw % | Ghana % | Brier |
|---|---|---|---|---|
| Market | 52.6 | 29.6 | 17.9 | 0.344 |
| Grok | 56.0 | 29.0 | 15.0 | 0.300 |
| Gemini | 53.0 | 28.0 | 19.0 | 0.335 |
| Ensemble | 51.7 | 29.0 | 19.4 | 0.355 |
| Naive Avg | 50.8 | 29.0 | 20.2 | 0.367 |
| GPT-5.4 | 49.0 | 29.0 | 22.0 | 0.393 |
| GPT-5.5 | 48.0 | 30.0 | 22.0 | 0.409 |
| Claude | 48.0 | 29.0 | 23.0 | 0.407 |
On scoreline picks, the unanimity was almost comical: every model filed 1–0 to Croatia. The actual scoreline was 2–1, so all took an outcome hit but not an exact hit. Backing a low-scoring Croatia win was defensible; the game simply delivered more goals than anticipated. No disgrace there.
Panama 0–2 England (Group L, East Rutherford)
The most straightforward result of the day. England were priced between 78% and 85% to win across all models, and they duly obliged. Gemini was the most bullish at 85%, followed by Claude at 84% — both outperforming the market (82.5%) and earning the best Brier scores (0.035 and 0.040 respectively). Grok, leaning on betting lines, came in at 78% — the lowest of the group and consequently its worst relative performance here, though a Brier of 0.076 is still excellent in absolute terms.
| Model | Panama % | Draw % | England % | Brier |
|---|---|---|---|---|
| Gemini | 5.0 | 10.0 | 85.0 | 0.035 |
| Claude | 5.0 | 11.0 | 84.0 | 0.040 |
| Market | 5.9 | 11.6 | 82.5 | 0.048 |
| GPT-5.5 | 6.0 | 12.0 | 82.0 | 0.050 |
| Ensemble | 6.0 | 12.2 | 81.8 | 0.052 |
| Naive Avg | 6.0 | 12.2 | 81.8 | 0.052 |
| GPT-5.4 | 7.0 | 13.0 | 80.0 | 0.062 |
| Grok | 7.0 | 15.0 | 78.0 | 0.076 |
A notable scoreline moment: every model that filed a pick chose 0–2, and every one of them was right. Six exact hits in a single match is the kind of number that looks impressive until you remember this was among the most predictable scorelines of the tournament so far — England beating a weak side 2–0 is close to a default answer. Still, hits are hits.
Colombia 0–0 Portugal (Group K, Miami Gardens)
Here is where the day gets interesting. Portugal were the clear favourites — Claude had them at 50%, the market at 47%, with most models clustering in the 42–47% range for a Portugal win. Colombia were rated a modest home side at 27–31%. The draw probability sat between 24% and 28% across the board. The actual result: a goalless draw. Every model got the outcome wrong, and the Brier scores reflect it — all sitting between 0.787 and 0.895, with the market (0.859) ironically performing worse than some of the AI models despite carrying the most confident Portugal probability.
| Model | Colombia % | Draw % | Portugal % | Brier |
|---|---|---|---|---|
| Gemini | 29.0 | 28.0 | 43.0 | 0.787 |
| Grok | 31.0 | 27.0 | 42.0 | 0.805 |
| Naive Avg | 28.0 | 26.2 | 45.8 | 0.833 |
| Ensemble | 28.0 | 26.2 | 45.8 | 0.833 |
| GPT-5.4 | 27.0 | 26.0 | 47.0 | 0.841 |
| GPT-5.5 | 27.0 | 26.0 | 47.0 | 0.841 |
| Market | 27.7 | 25.1 | 47.1 | 0.859 |
| Claude | 26.0 | 24.0 | 50.0 | 0.895 |
Claude was the worst performer here, partly because its 50% Portugal figure left the least room for a draw or home win. This is the reputation-over-form bias in plain view: Portugal carry the weight of Ronaldo and a historic brand, and the models collectively loaded up on them. Gemini and Grok were least wrong — both gave the draw around 27–28% and kept Colombia above 29%, giving them slightly more slack. All six scoreline picks were 1–2 to Portugal. All six were wrong on outcome. A collective failure, and a useful one: it confirms that when a prestige nation plays a decent-but-unheralded side, the models are probably undervaluing the draw.
DR Congo 3–1 Uzbekistan (Group K, Atlanta)
This match had no full probability predictions filed by the AI models — only scoreline picks and market data are available. What we do have is telling. Four models (GPT-5.4, GPT-5.5, Grok, Claude, and the Ensemble) all picked 0–1 to Uzbekistan. Gemini picked 1–1. DR Congo won 3–1. The market, meanwhile, had DR Congo at 59.1% — comfortably the favourite. Grok, which is supposed to lean on betting odds, nonetheless picked a Uzbekistan win, which suggests its scoreline logic doesn't always follow its probability instincts. The market's Brier score of 0.253 was the best of the day for this fixture; Claude and Grok (both at 0.346–0.348 on their outright probabilities) were the closest AI models, having at least given DR Congo above 50%.
| Model | DR Congo % | Draw % | Uzbekistan % | Brier |
|---|---|---|---|---|
| Market | 59.1 | 23.7 | 17.2 | 0.253 |
| Grok | 52.0 | 28.0 | 20.0 | 0.349 |
| Claude | 52.0 | 26.0 | 22.0 | 0.346 |
| Gemini | 50.0 | 28.0 | 22.0 | 0.377 |
| Ensemble | 48.4 | 27.9 | 23.7 | 0.400 |
| Naive Avg | 48.0 | 27.8 | 24.2 | 0.406 |
| GPT-5.4 | 43.0 | 29.0 | 28.0 | 0.487 |
| GPT-5.5 | 43.0 | 28.0 | 29.0 | 0.487 |
GPT-5.4 and GPT-5.5 were notably the weakest performers, giving Uzbekistan close to equal billing with DR Congo — a miscalibration the market did not share. If the models are broadly overrating reputation, Uzbekistan may have benefited from a mild novelty premium: an unfamiliar side from a non-traditional footballing region that the models struggled to place on the quality spectrum.
Algeria 3–3 Austria (Group J, Kansas City)
A spectacular match that ended in a draw, which most models — barely — gave a chance. The probability distributions were: draw roughly 32–38% across AI models, Austria favourites at 34–43%, Algeria the underdogs at 25–29%. The market was the outlier, assigning the draw 45.6% — the single highest draw probability of any model. It was also the only model to correctly call the draw outright, earning a Brier score of 0.449 against scores of 0.582–0.710 for the AI models.
| Model | Algeria % | Draw % | Austria % | Brier | Correct? |
|---|---|---|---|---|---|
| Market | 22.5 | 45.6 | 31.9 | 0.449 | ✓ |
| Claude | 26.0 | 38.0 | 36.0 | 0.582 | ✓ |
| Grok | 29.0 | 37.0 | 34.0 | 0.597 | ✓ |
| Ensemble | 26.6 | 35.0 | 38.4 | 0.641 | ✗ |
| Naive Avg | 26.6 | 35.0 | 38.4 | 0.641 | ✗ |
| GPT-5.4 | 27.0 | 34.0 | 39.0 | 0.661 | ✗ |
| GPT-5.5 | 26.0 | 34.0 | 40.0 | 0.663 | ✗ |
| Gemini | 25.0 | 32.0 | 43.0 | 0.710 | ✗ |
Claude and Grok were the AI models that called this correctly — Claude with 38% draw, Grok with 37%. Gemini was worst, anchoring hard on Austria's presumably superior league pedigree (43% Austria) and assigning the lowest draw probability in the field. That's the league-prestige fingerprint in action. Both Claude and Grok filed 0–0 as their scoreline pick, so they called the outcome right while missing the scoreline by a considerable distance — the actual result was a six-goal draw, which no model would have reasonably anticipated. Outcome hit, exact miss; that's a normal outcome when football decides to be ridiculous.
Jordan 1–3 Argentina (Group J, Arlington)
Argentina were the most nailed-on favourites of the day, and the result was straightforward. Probabilities ranged from 80% (Grok) to 86% (Gemini) for an Argentina win. Gemini was the most confident and earned the best Brier score (0.031), closely followed by Claude at 85% and Brier 0.035. The market at 83.9% sat comfortably in the middle of the pack. Grok, again the least bullish of the AI models on the clear favourite, had a Brier of 0.062 — still fine, but a recurring pattern worth noting: Grok's betting-market anchor appears to pull it slightly *away* from the implied probability when a team is a very heavy favourite, perhaps because bookmakers already shade those prices.
| Model | Jordan % | Draw % | Argentina % | Brier |
|---|---|---|---|---|
| Gemini | 4.0 | 10.0 | 86.0 | 0.031 |
| Claude | 5.0 | 10.0 | 85.0 | 0.035 |
| Market | 5.6 | 10.5 | 83.9 | 0.040 |
| Naive Avg | 5.2 | 11.4 | 83.4 | 0.043 |
| Ensemble | 5.2 | 11.4 | 83.4 | 0.043 |
| GPT-5.4 | 5.0 | 12.0 | 83.0 | 0.046 |
| GPT-5.5 | 5.0 | 12.0 | 83.0 | 0.046 |
| Grok | 7.0 | 13.0 | 80.0 | 0.062 |
All models picked 0–2 to Argentina. The actual score was 1–3. Jordan got a goal, Argentina got three — outcome correct, scoreline wrong in both directions. Again, entirely normal at this target difficulty.
Day Summary
Across the six matches: the models went 4-from-6 on outcomes collectively (every model got Croatia, England, and Argentina right; every model got Colombia–Portugal wrong; DR Congo and Algeria split the field). The Colombia–Portugal draw was the collective failure of the day, and the Algeria–Austria result showed the market doing something the AI models couldn't quite manage — properly pricing in a high draw probability for an evenly matched fixture. Gemini's league-prestige bias cost it on both Algeria–Austria and the DR Congo match. Grok's betting-market anchor served it well on Croatia but left it trailing on heavy-favourite games.
- Best model, Saturday: Gemini on the two blowout wins (England, Argentina); Claude across the balanced fixtures.
- Worst model, Saturday: Claude on Colombia–Portugal (50% Portugal, lowest draw allocation, highest Brier of the group); GPT-5.4 and GPT-5.5 on DR Congo.
- Market vs AI: The market won the day's most interesting match — Algeria–Austria — by a wide margin, thanks to its elevated draw probability. A reminder that implied draw probabilities from odds markets deserve more respect.
- Bias confirmed: Reputation-over-form across all models on Portugal; Gemini's prestige weighting on Austria; Grok's moderated confidence on heavy favourites.
- Exact score highlights: Six models nailed Panama 0–2 England. Zero models called the DR Congo, Algeria, or Colombia scorelines correctly.
Colombia 0–0 Portugal: every model, every scoreline pick, wrong. The draw probability sat at 24–28% across all AI models — never the top outcome, never even close. The market gave it 25.1%. This is the clearest single-match illustration yet of the shared reputation bias: when a glamour side meets a respectable underdog, AI models collectively underweight the draw to a degree that will, over a large enough sample, cost them. File this one.
Looking Ahead: Sunday's Tests
Tomorrow's slate will be worth watching for a few specific reasons. Any match featuring a European side from a prestigious league against an African or Asian opponent will test whether Gemini has learned anything from today's Algeria and DR Congo performances — probably not, given the weights are locked, but the pattern will be visible. If there are any matches where betting markets assign an unusually high draw probability relative to the AI models, today's Algeria–Austria result gives us reason to trust the market's read. And if another 'obvious' favourite game delivers an upset, the reputation-over-form question will move from a known bias to an active liability.
Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.