Ecuador Stun Germany, Türkiye Beat USA: A Costly Day for Model Confidence
Six matches, two continents, and a thorough reminder that probability is not destiny. Thursday's slate split neatly in two: a pair of comfortable away wins that the models called correctly, then four matches where the football had other ideas. Ecuador beating Germany 2-1 and Türkiye defeating the USA 3-2 were the headline shocks — and every single model got both wrong. Japan vs Sweden drew when nearly everyone expected a Japan home win. Paraguay vs Australia served up a goalless draw that, to be fair, most models at least considered plausible. A good day for the football, a mixed-to-poor day for the machines.
Curacao 0-2 Ivory Coast — Group E, MetLife Stadium
Start with the easy one. Ivory Coast were overwhelming favourites and duly won. Every model had the Elephants at 79–83% and every model was right. Gemini led the pack with a Brier score of 0.046 on the back of its 83% away probability — the highest of any model here — and GPT-5.4, Grok, Claude, and the ensemble all scored respectably too. Claude's well-documented tendency to over-price home advantage was, for once, conspicuous in the right direction: giving Curacao only 6% at home was the most aggressive suppression of home advantage on the card, and it paid off with the joint-best log-loss of the individual models.
| Model | P(Home) | P(Draw) | P(Away) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 7% | 14% | 79% | ✅ Away | 0.069 |
| Claude | 6% | 14% | 80% | ✅ Away | 0.063 |
| Gemini | 5% | 12% | 83% | ✅ Away | 0.046 |
| Grok | 7% | 14% | 79% | ✅ Away | 0.069 |
| Ensemble | 6% | 14% | 80% | ✅ Away | 0.061 |
A rare scoreline sweep as well: all five models with locked picks called 0-2, and all five hit it exactly. Treat that as a lucky bonus rather than evidence of supernatural precision — 0-2 was simply the most probable scoreline for a dominant away win — but it's a clean sheet nonetheless. Gemini's league-prestige bias aligned perfectly with reality here: Ivory Coast are meaningfully the better side by any measure, and the model's higher away probability reflected that.
Ecuador 2-1 Germany — Group E, Lincoln Financial Field
Here is where the day gets interesting and uncomfortable. Germany were 58–64% favourites across the panel. Ecuador were given between 16% and 19%. The actual result — a 2-1 Ecuador home win — was the lowest-probability outcome according to every model, and every model paid the price with Brier scores above 1.0. In probability terms, backing the 18% shot is not a scandal; upsets happen. But the spread of model opinions is worth examining.
| Model | P(Home) | P(Draw) | P(Away) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 18% | 22% | 60% | ❌ Home | 1.081 |
| Claude | 16% | 20% | 64% | ❌ Home | 1.155 |
| Gemini | 18% | 24% | 58% | ❌ Home | 1.066 |
| Grok | 19% | 23% | 58% | ❌ Home | 1.045 |
| Ensemble | 18% | 22% | 60% | ❌ Home | 1.086 |
Claude was the furthest wrong, handing Germany 64% and Ecuador only 16% — the most emphatic endorsement of the German side. This is a known pattern: Claude has a documented tendency to over-price home advantage, yet here it gave Ecuador the *lowest* home probability of all models. That apparent contradiction resolves when you consider form: Germany arrived in better recent shape and Claude leans hard on reputational form signals. Grok was the closest to correct by Brier score (1.045), having given Germany only 58% — consistent with its habit of shadowing betting-market odds, which presumably reflected some Ecuador market money. Gemini, despite its league-prestige fingerprint which should have inflated Germany further, was second-best, having also stayed at 58%. None of this is a vindication; it is a question of which model lost by the least.
Every scoreline pick had Germany winning — GPT-5.4, Claude, and the ensemble all called 0-2, while Gemini and Grok picked 1-2. All missed both the outcome and the scoreline. This is not a failure of the picking mechanism so much as a reflection of the underlying probabilities: when your model gives the winner only 18%, your scoreline pick will not feature a home win.
Japan 1-1 Sweden — Group F, AT&T Stadium
Full probability breakdowns are absent from the data for this match (the predictions array is empty), so we work from the accuracy scores and scoreline picks. What we can say: the ensemble had Japan at 46%, the draw at 29%, and Sweden at 25%. Grok went furthest for Japan at 51%. GPT-5.4 was the most cautious on Japan at 39% and consequently the least wrong when the match finished 1-1, posting the best Brier of 0.718.
| Model | P(Japan) | P(Draw) | P(Sweden) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 39% | 31% | 30% | ❌ Draw | 0.718 |
| Claude | 46% | 27% | 27% | ❌ Draw | 0.817 |
| Gemini | 45% | 28% | 27% | ❌ Draw | 0.794 |
| Grok | 51% | 29% | 20% | ❌ Draw | 0.804 |
| Ensemble | 46% | 29% | 25% | ❌ Draw | 0.779 |
All five models with scoreline picks chose a Japan home win — GPT-5.4, GPT-5.5, Claude, Grok, and the ensemble all picked 1-0; Gemini went for 2-0. Not a single model entertained a draw scoreline as its headline pick. Grok's 51% Japan probability is the starkest overreach, consistent with its betting-market reliance: Japan were presumably favoured in the markets, and Grok followed them into a result it was least prepared for. The draw probability ranged from 27% to 31% — models knew it was plausible, but none led with it.
Tunisia 1-3 Netherlands — Group F, Arrowhead Stadium
Back to comfortable territory. The Netherlands were heavy favourites and won 3-1. GPT-5.4 was the standout performer here, giving the Dutch 84% and posting a Brier of just 0.040 — the best single-match score across the entire Thursday card. Gemini at 80% and Claude at 78% also scored well. Grok, interestingly, was the most hesitant at 67% for a Dutch win — notably lower than the rest of the field and its lowest away probability of the three 'easy' matches today. Whether that reflects genuine model uncertainty or simply less aggressive betting-market odds for this fixture isn't clear from the data, but it cost Grok relative to its peers.
| Model | P(Home) | P(Draw) | P(Away) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 5% | 11% | 84% | ✅ Away | 0.040 |
| Claude | 7% | 15% | 78% | ✅ Away | 0.076 |
| Gemini | 5% | 15% | 80% | ✅ Away | 0.065 |
| Grok | 11% | 22% | 67% | ✅ Away | 0.169 |
| Ensemble | 7% | 16% | 77% | ✅ Away | 0.081 |
No exact-score hits here — GPT-5.5 and Claude both picked 1-2, GPT-5.4, Gemini, Grok, and the ensemble all picked 0-2, but the actual was 1-3. All models got the correct outcome, which is the right level of expectation for scoreline picks.
Paraguay 0-0 Australia — Group D, Levi's Stadium
A genuinely competitive fixture, and the models reflected that. Draw probabilities ranged from a low of 34% (GPT-5.4) to a high of 43% (Gemini), with Grok at 40%. Paraguay at home was given 33–37% by most models. The actual result was a goalless draw, and all models except GPT-5.4 called the correct outcome — though no model had the draw as a commanding favourite.
| Model | P(Paraguay) | P(Draw) | P(Australia) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 37% | 34% | 29% | ❌ Draw | 0.657 |
| Claude | 36% | 38% | 26% | ✅ Draw | 0.582 |
| Gemini | 33% | 43% | 24% | ✅ Draw | 0.491 |
| Grok | 37% | 40% | 23% | ✅ Draw | 0.550 |
| Ensemble | 36% | 39% | 25% | ✅ Draw | 0.571 |
Gemini was the clear winner here — its 43% draw probability was the highest of any model and produced the best Brier of 0.491. This is an interesting inversion of the usual pattern: Gemini's tendency to over-weight league prestige should theoretically have pushed it towards Paraguay or away from a draw, but the relatively even quality of the sides dampened that bias. GPT-5.4 was the only model to miss the outcome, having given Paraguay the highest win probability at 37% and the draw the lowest at 34%. Scoreline picks were more mixed: Claude, Grok, Gemini, and the ensemble all called 0-0 exactly — all four hit. GPT-5.4 picked 1-0 and missed. A nice result for the models that leaned into defensive expectations for this tie.
Türkiye 3-2 USA — Group D, SoFi Stadium
The most narratively loaded match of the day, and a comprehensive miss by every model. The USA were favourites across the board — Claude gave them 51%, Grok 47%, GPT-5.4 45%, the ensemble 46%, and Gemini 40%. Türkiye won 3-2. Gemini was least wrong with a Brier of 0.701, having given Türkiye the highest home probability at 32% and the USA the lowest away probability at 40%. Claude was furthest off at 0.841, having handed the USA 51% and Türkiye only 27%.
| Model | P(Türkiye) | P(Draw) | P(USA) | Result | Brier |
|---|---|---|---|---|---|
| GPT-5.4 | 29% | 26% | 45% | ❌ Home | 0.774 |
| Claude | 27% | 22% | 51% | ❌ Home | 0.841 |
| Gemini | 32% | 28% | 40% | ❌ Home | 0.701 |
| Grok | 29% | 24% | 47% | ❌ Home | 0.783 |
| Ensemble | 29% | 25% | 46% | ❌ Home | 0.772 |
Claude's fingerprint is right on the surface here: despite a stated tendency to over-price home advantage, it consistently gave Türkiye the *lowest* home probability of any model. That is the form-versus-reputation tension at work — the USA's record and profile apparently overrode whatever home-advantage signal Claude was processing. Every scoreline pick across all five models was 1-2 to the USA. Not a single model's headline pick featured a Türkiye lead. This was a collective failure, not a model-specific one.
Day in Review: What We Learned
- The two clear mismatches — Curacao vs Ivory Coast and Tunisia vs Netherlands — were handled well. These are the matches models are built for: large quality gaps, predictable outcomes.
- Ecuador and Türkiye each won as home underdogs against higher-ranked opposition. All models preferred the stronger side in both cases. This is not irrational — both were low-probability outcomes — but two such results in the same session stings cumulatively.
- The 'reputation over form' bias is the common thread. Germany and the USA arrived with better tournament pedigrees; the models loaded probability onto them accordingly. Whether Ecuador or Türkiye's recent form was available and weighted appropriately is a question worth re-examining.
- Grok's betting-market proximity showed both sides of its coin: marginally closest on Ecuador (Germany at only 58%), but most overconfident on Japan (51%), and cautious to a fault on Netherlands (only 67%).
- Gemini had the best single-match Brier of the day (0.046 vs Ivory Coast) and was also the most valuable on Paraguay-Australia (43% draw). Its league-prestige fingerprint was unhelpful on Germany-Ecuador — but it was the least wrong of all models on Türkiye-USA.
- GPT-5.4 was the standout performer on Tunisia-Netherlands (Brier 0.040) and the sharpest on Japan-Sweden (Brier 0.718). A quietly solid day on the games it got right.
- Exact-score hits: a remarkable five-from-five on Ivory Coast 0-2, and four-from-four on Paraguay 0-0. In tighter matches, the models appropriately called the favoured outcome's most likely scoreline. When the outcome was wrong, there were no scoreline hits — as expected.
Two home underdogs won in the same evening session — Ecuador over Germany and Türkiye over the USA — and every model preferred the favourite in both. Combined Brier scores above 1.0 for the Ecuador match represent the worst per-match performance of the tournament so far. This is not random noise; it is a measurable signal that the models are systematically undervaluing tournament-context home wins against prestigious opponents, likely because reputation and recent form both point the same way and leave insufficient mass on the home side.
Looking Ahead: What Tomorrow Tests
Friday's card brings more group-stage conclusions, which means results matter more: teams with something to play for versus those already through or already out can behave very differently to their 90-minute averages. If there are matches where one side needs a win and the other would happily take a draw, the models will need to process that strategic dimension — something they have historically priced inconsistently. After today's twin home-underdog shocks, it is worth watching whether Grok's market-following tendency recalibrates, and whether Claude continues to suppress home probabilities in matches where form and reputation conflict. Any fixture pairing a well-ranked nation against a motivated lower-ranked home side should be approached with extra scepticism.
Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.