Models Back Canada, Claude Leads the Scoring — Sunday Review
One match on Sunday, and a clean result: Canada won 0-1 in what was, from a forecasting perspective, a satisfyingly unambiguous afternoon. Every model had Canada as favourites, the market agreed, and the scoreline itself was the consensus pick across five of the six AI models. When everything aligns like this, the post writes itself — but there are still meaningful differences in how confident each model was, and those differences are where the learning lives.
South Africa 0-1 Canada — Match 73, Round of 32
The result was about as close to a 'predicted' outcome as you get in football. Canada were favourites across the board, won by a single goal, and the canonical scoreline — 0-1 — was the locked pick for every model that submitted one. What separated the models was the degree of conviction assigned to a Canada win.
| Model | P(South Africa) | P(Draw) | P(Canada) | Brier Score | Result Correct |
|---|---|---|---|---|---|
| Claude | 0.14 | 0.24 | **0.62** | 0.222 | ✓ |
| GPT-5.5 | 0.18 | 0.25 | 0.57 | 0.280 | ✓ |
| Market | 0.16 | 0.26 | 0.58 | 0.272 | ✓ |
| Ensemble | 0.185 | 0.263 | 0.553 | 0.303 | ✓ |
| Naive Avg | 0.185 | 0.263 | 0.553 | 0.303 | ✓ |
| GPT-5.4 | 0.19 | 0.27 | 0.54 | 0.321 | ✓ |
| Gemini | 0.18 | 0.28 | 0.54 | 0.322 | ✓ |
| Grok | 0.23 | 0.29 | **0.48** | 0.407 | ✓ |
Claude was the standout performer today, posting the best Brier score (0.222) and the lowest log-loss (0.478). Its 62% probability for Canada was the most decisive read of the fixture from any of the AI models, and it was rewarded. This is a mildly ironic result given Claude's known fingerprint — a tendency to over-price home advantage. Here, that bias arguably worked in reverse: Claude was least seduced by South Africa's home support and gave Canada the most credit. Whether that's a corrected bias or a happy accident is worth monitoring, but the number stands.
GPT-5.5 came in second among the AI models (Brier 0.280), with a 57% Canada call — nudging it close to the market's 57.8%. The market itself, as is often the case, sat near the top of the leaderboard (Brier 0.272), which is a useful reminder that it remains a strong benchmark. GPT-5.4 and Gemini were essentially tied (Brier 0.321 and 0.322), both sitting at 54% for Canada — reasonable, if a touch cautious.
Grok was the clear laggard today, and the data is worth sitting with. It gave Canada only a 48% chance of winning — making South Africa and the draw collectively more likely than a Canada victory. That's a notably different read from the rest of the field, and Grok's Brier score of 0.407 reflects the cost of that diffidence. Grok's known fingerprint is that it leans hardest on betting-market odds, which makes its underperformance relative to the market today genuinely curious. The market had Canada at 57.8%; Grok had them at 48%. Something in Grok's weighting pushed against the grain of the very data source it typically mirrors. It still picked the right outcome and the right scoreline, but at a calibration cost.
On Gemini: its 54% Canada call is fine, but its 28% draw probability is the highest in the field. The known Gemini fingerprint — over-weighting league prestige — is a plausible explanation. South Africa's domestic football infrastructure doesn't carry the same prestige signals as a CONCACAF or European opponent, which may have nudged Gemini toward treating the fixture as more open than it was.
Scoreline Picks: A Rare Clean Sweep
Every model that submitted a scoreline pick — Grok, Gemini, GPT-5.4, GPT-5.5, Claude, and the Ensemble — called 0-1 exactly. That's a genuine clean sweep on a hard target, and it's worth a moment's acknowledgement. Exact-score prediction is partly luck even when the logic is sound: a 0-1 scoreline in a low-expectation fixture between a modest tournament entrant and a solid but unspectacular CONCACAF side is, frankly, the statistically sensible call. The models found it. We'll flag it, but we won't oversell it — the real test is whether that read was principled or whether they'd all pick 0-1 reflexively for any slight-favourite away win.
Claude's known bias is over-pricing home advantage — yet today it gave South Africa the lowest home win probability (14%) and scored best across all AI models. One match proves nothing, but it's a datapoint worth tracking: is the bias situational rather than universal, or is Claude gradually recalibrating? We'll be watching.
What Tomorrow's Slate Will Test
Monday should give us more to chew on. As the Round of 32 continues, we'll be looking for fixtures that stress-test the models' known weak points: matches where form and reputation diverge sharply (a classic bias trap), any tie involving a prestigious European league nation as slight underdogs (Gemini's pressure test), and any fixture with a strong home-crowd narrative that might pull Claude's probability toward the home side. The broader over-reliance on reputation versus form is something the whole field has shown across the tournament so far — a match where a lower-ranked side arrives in better shape than their name suggests will be the real examination.
Written by claude-sonnet-4-6 from locked pre-match predictions and final results — part of the Modelball study.