MBModelBall

THE LEADERBOARD

Six Methods. Live Results.

Ranked by prediction accuracy. The key question: does knowing each AI's blind spots improve results?

Matches scored: 10
Edge ahead by 0.001

THE KEY QUESTION: Does knowing AI blind spots improve predictions?

The Edge
64%
49/76 correct
-0%
Edge more accurate
Simple Average
64%
49/76 correct
Bias corrections are helping — The Edge is beating the simple average.

Full standings

RankMethodMatchesCorrectAccuracyBriervs Avg
1
market
10770%0.447−0.051
2
GPT-5.5
The Evolved Baseline
10660%0.484−0.014
3
The Edge
Bias-corrected
764964%0.497−0.001
4
Simple Average
Equal-weight blend
764964%0.498
5
Gemini 3.1 Pro
The Generalist
764964%0.499+0.000
6
Claude Sonnet 4.6
The Analyst
765066%0.499+0.001
7
Grok 3
The Contrarian
765066%0.501+0.003
8
GPT-5.4
The Market Baseline
764762%0.503+0.005

Accuracy: Percentage of match outcomes correctly predicted.

Brier: Technical accuracy score (lower = better, 0 = perfect).

vs Avg: Brier score difference from simple average.Negative = beating average,positive = behind.

Understanding the methods

The five models

GPT-5.4, GPT-5.5, Claude, Grok, and Gemini each make predictions independently. Each has documented biases from our fingerprinting research.

View model profiles →

Simple average

Equal-weight blend of all five AI predictions. The baseline — if The Edge can't beat this, our corrections don't add value.

The Edge

Bias-corrected blend. We know each AI's blind spots, so we trust them less in those situations.

How it works →

Every outcome teaches

If Edge leads: fingerprints help. If Naive leads: we learn and refine. Either outcome advances the research.

The experiment →

Full research sample

The World Cup alone is too small to separate methods this close. Pooling it with our 18-league calibration corpus gives 1,045 matches (66 World Cup + 979 league), scored with a paired match-block bootstrap. Brier score — lower is better.

The Edge
0.599n=1045
Naive average
0.599n=1045
GPT-5.4
0.600n=1045
Grok 3
0.600n=1045
Gemini 3.1 Pro
0.604n=1045
GPT-5.5
0.604n=979
Claude Sonnet 4.6
0.610n=1045
A note on calibration. We tested whether bias-correcting the models before combining them improves accuracy. A cross-validated re-fit lands essentially on the identity (no shrinkage) — out-of-sample it does not beat the raw ensemble on either the leagues or the World Cup. In other words the raw ensemble is already well-calibrated, so the Edge's value is the ensembling itself, not a correction layer. Calibration is not applied to the live predictions; the figures above are raw.
82%
probability the Edge weighting beats a naive average (paired bootstrap over 1,045 matches). Directional, not yet conclusive — the interval still includes zero.
The Edge significantly beats
Claude Sonnet 4.6, Gemini 3.1 Pro (95% CI excludes zero). It also leads every other model, with GPT-5.4 the closest single model.

Edge weights are partly derived in-sample on 5 of the 18 leagues. GPT-5.5 is league-only (n=979). Generated 2026-06-27.

Which players did AI misprice?

After 104 matches, we'll know exactly which players the models systematically undervalued or overvalued. Get the full Transfer Arbitrage Report in July 2026.

Learn More About the Report