THE LEADERBOARD
Six Methods. Live Results.
Ranked by prediction accuracy. The key question: does knowing each AI's blind spots improve results?
THE KEY QUESTION: Does knowing AI blind spots improve predictions?
Full standings
| Rank | Method | Matches | Correct | Accuracy | Brier | vs Avg |
|---|---|---|---|---|---|---|
| 1 | market | 10 | 7 | 70% | 0.447 | −0.051 |
| 2 | GPT-5.5 The Evolved Baseline | 10 | 6 | 60% | 0.484 | −0.014 |
| 3 | The Edge★ Bias-corrected | 76 | 49 | 64% | 0.497 | −0.001 |
| 4 | Simple Average Equal-weight blend | 76 | 49 | 64% | 0.498 | — |
| 5 | Gemini 3.1 Pro The Generalist | 76 | 49 | 64% | 0.499 | +0.000 |
| 6 | Claude Sonnet 4.6 The Analyst | 76 | 50 | 66% | 0.499 | +0.001 |
| 7 | Grok 3 The Contrarian | 76 | 50 | 66% | 0.501 | +0.003 |
| 8 | GPT-5.4 The Market Baseline | 76 | 47 | 62% | 0.503 | +0.005 |
Accuracy: Percentage of match outcomes correctly predicted.
Brier: Technical accuracy score (lower = better, 0 = perfect).
vs Avg: Brier score difference from simple average.Negative = beating average,positive = behind.
Understanding the methods
The five models
GPT-5.4, GPT-5.5, Claude, Grok, and Gemini each make predictions independently. Each has documented biases from our fingerprinting research.
View model profiles →Simple average
Equal-weight blend of all five AI predictions. The baseline — if The Edge can't beat this, our corrections don't add value.
The Edge
Bias-corrected blend. We know each AI's blind spots, so we trust them less in those situations.
How it works →Every outcome teaches
If Edge leads: fingerprints help. If Naive leads: we learn and refine. Either outcome advances the research.
The experiment →Full research sample
The World Cup alone is too small to separate methods this close. Pooling it with our 18-league calibration corpus gives 1,045 matches (66 World Cup + 979 league), scored with a paired match-block bootstrap. Brier score — lower is better.
Edge weights are partly derived in-sample on 5 of the 18 leagues. GPT-5.5 is league-only (n=979). Generated 2026-06-27.
Which players did AI misprice?
After 104 matches, we'll know exactly which players the models systematically undervalued or overvalued. Get the full Transfer Arbitrage Report in July 2026.
Learn More About the Report