Sunday Sweep: Spain Deliver, Two Upsets Bite Everyone
Sunday's slate ran from the very comfortable to the quietly humiliating. Spain obliged the forecasters in Atlanta. Belgium and Uruguay, however, decided that form, reputation, and betting markets could all take the afternoon off. Four matches, four different stories — here is how the models fared.
Spain 4-0 Saudi Arabia — Group H, Atlanta
The easiest call of the day, and broadly the models delivered. Every model gave Spain north of 70% to win, and win they did — convincingly enough that nobody is taking a bow for merely pointing at the obvious.
| Model | P(Spain) | P(Draw) | P(Saudi) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 0.84 | 0.11 | 0.05 | ✓ | 0.040 |
| Gemini | 0.85 | 0.11 | 0.04 | ✓ | 0.036 |
| Grok | 0.87 | 0.10 | 0.03 | ✓ | 0.028 |
| Ensemble | 0.83 | 0.12 | 0.05 | ✓ | 0.044 |
| Naive Avg | 0.82 | 0.12 | 0.06 | ✓ | 0.051 |
| Claude | 0.72 | 0.17 | 0.11 | ✓ | 0.119 |
Grok takes the best Brier score here (0.028), its market-anchored confidence paying off when the favourite strolls home. Claude is the clear laggard — 0.72 against a field sitting at 0.84–0.87 — which is Claude's home-advantage fingerprint running in reverse: Spain were technically the home side in name only (a neutral American venue), but Claude's general tendency to temper large favourite probabilities left it notably softer. Still a correct call, just the least rewarded one. All five models picked 3-0 as their exact-score tip; the actual 4-0 means nobody got it precisely, but all had the right outcome — that's a normal miss on an exact target, not a failure worth labouring.
Belgium 0-0 Iran — Group G, Inglewood
Here is where Sunday starts to sting. Belgium were given between 58% and 71% to win, Iran between 10% and 20% to win, and a draw sat in the 19–27% range across the panel. The match finished goalless — the second-most-likely outcome for most models, but still clearly not what anyone expected.
| Model | P(Belgium) | P(Draw) | P(Iran) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 0.61 | 0.23 | 0.16 | ✗ Draw | 0.991 |
| Gemini | 0.58 | 0.27 | 0.15 | ✗ Draw | 0.892 |
| Claude | 0.58 | 0.22 | 0.20 | ✗ Draw | 0.985 |
| Ensemble | 0.65 | 0.21 | 0.14 | ✗ Draw | 1.070 |
| Naive Avg | 0.63 | 0.21 | 0.15 | ✗ Draw | 1.043 |
| Grok | 0.71 | 0.19 | 0.10 | ✗ Draw | 1.170 |
Gemini absorbs the smallest loss here, having assigned the most probability to the draw (27%). That is a rare case of Gemini's reputation-weighting arguably doing some inadvertent good: it was less bullish on Belgium than the others, which softened the blow. Grok is the biggest casualty — its 71% Belgium figure is the highest on the panel and reflects its betting-market lean; the markets were wrong, and Grok followed them off the cliff. Brier scores above 1.0 for four of the six outputs make this a genuinely bad day. Every model picked 2-0 Belgium as their exact score; every model got neither the scoreline nor the outcome. A collective miss this uniform is useful data: it tells us the models are jointly over-confident on established European sides against lower-ranked opponents.
Uruguay 2-2 Cape Verde — Group H, Miami Gardens
The other simultaneous kick-off in the Americas produced an equally awkward result. Cape Verde — given at most 18% to win by any model, and as low as 5% by Grok — came away with a point. Uruguay's draw probability sat between 17% and 25%; it was the outcome nobody really budgeted for.
| Model | P(Uruguay) | P(Draw) | P(Cape Verde) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 0.61 | 0.24 | 0.15 | ✗ Draw | 0.972 |
| Gemini | 0.62 | 0.25 | 0.13 | ✗ Draw | 0.964 |
| Claude | 0.58 | 0.24 | 0.18 | ✗ Draw | 0.946 |
| Ensemble | 0.67 | 0.21 | 0.11 | ✗ Draw | 1.084 |
| Naive Avg | 0.65 | 0.23 | 0.13 | ✗ Draw | 1.036 |
| Grok | 0.78 | 0.17 | 0.05 | ✗ Draw | 1.300 |
Claude edges the best Brier (0.946) and Gemini is close behind (0.964) — both having been marginally more cautious about Uruguay and slightly more generous to the draw. Grok is again the worst performer, its 78% Uruguay figure the most exposed when the final whistle blew at 2-2. That 5% probability for Cape Verde — Grok's lowest for any outcome today — looks particularly ungenerous in hindsight. The overarching fingerprint that 'models broadly overrate reputation versus form' is being written in capital letters across both the Belgium and Uruguay games on the same afternoon. All models except Grok picked 1-0 Uruguay; Grok went 2-0. Every single pick had the wrong outcome.
New Zealand 1-3 Egypt — Group G, Vancouver
Sunday ends on a more satisfying note for the panel. Egypt were clear favourites, the models said so, and Egypt delivered. This was the tidiest collective performance of the day after Spain, with Egypt probabilities clustering between 60% and 63%.
| Model | P(NZ) | P(Draw) | P(Egypt) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 0.16 | 0.24 | 0.60 | ✓ | 0.243 |
| Gemini | 0.12 | 0.25 | 0.63 | ✓ | 0.214 |
| Claude | 0.14 | 0.22 | 0.63 | ✓ | 0.206 |
| Grok | 0.16 | 0.24 | 0.60 | ✓ | 0.243 |
| Ensemble | 0.15 | 0.24 | 0.62 | ✓ | 0.226 |
| Naive Avg | 0.15 | 0.24 | 0.62 | ✓ | 0.226 |
Claude just edges the best Brier here (0.206), marginally ahead of Gemini (0.214). Both assigned 63% to Egypt and kept New Zealand below 15%, which proved apt. GPT-5.4 and Grok were identical at 60/16/24 and score identically as a result. All models picked 0-1 Egypt as their scoreline; the actual 1-3 means no exact hits, but every model had the right outcome — a perfectly acceptable near-miss. Notably this is one of the matches where Gemini's league-prestige weighting may have helped rather than hurt, given Egypt's stronger continental pedigree relative to New Zealand.
Day Summary: The Numbers
- Results correct today — GPT-5.4: 2/4, Gemini: 2/4, Claude: 2/4, Grok: 2/4, Ensemble: 2/4. Everyone equal, but the quality of the misses is not.
- Grok's cumulative Brier pain from Belgium and Uruguay (1.170 + 1.300 = 2.470) is the heaviest single-model two-match toll of the day, driven by its market-anchored overconfidence on both favourites.
- Gemini's relative caution on Belgium (0.58, lowest on the panel) paid off with the best draw-game Brier — a rare day where being the least certain about a favourite was the right call.
- Claude's home-advantage fingerprint was largely irrelevant today (no clear neutral-advantage situation), but its general tendency to temper extremes left it closest on both draw games.
- Every single model picked a home/away win as their scoreline in the two draw games. The collective blind spot for draws in mismatched fixtures is structural, not random.
Two different group-stage favourites (Belgium vs Iran, Uruguay vs Cape Verde) both drew on the same afternoon — and every model predicted a home/away win in both. The models' combined draw probability across those two matches averaged roughly 21–25%, which is not absurd, but the fact that both games simultaneously landed in that minority outcome is a sharp reminder: when models uniformly assign 60–70% to a favourite, the 20–25% draw bucket is not a rounding error. It happens. Reputation continues to outweigh form in how these systems think, and today was the invoice.
What Monday Might Test
Tomorrow's slate brings more group-stage football, and with it fresh opportunities to probe the same fault lines. Watch for any match where the models show wide disagreement on draw probability — today's data suggests that when the panel collectively buries the draw below 25% against a lower-ranked side, the market may be underpricing it. Grok in particular will want to see some favourites actually win; its two worst Brier scores of the entire project so far likely sit in today's ledger. If tomorrow features a true home-soil advantage match (an actual host-nation crowd, not just a hosted venue), Claude's systematic home-premium fingerprint will be back under scrutiny. Form over reputation remains the standing challenge for the entire panel.
Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.