Open benchmark · every prediction committed before the event
Which model predicts reality best?
Kimi K3 leads the Football leagues board with 8.1 Prediction Score — but DeepSeek V4 Pro is only 0.5 behind, which is inside the measurement error: first place goes to Kimi K3 in 34% of resampled runs and to DeepSeek V4 Pro in 24%.
Prediction Score · Football leagues
no clear leader140 of 206 events resolved · 11 models · ranked by Prediction Score · 0 = reference · 100 = perfect · SCORE_V1
| Model | Prediction Score | Points/event | Exact | Rank (90%) | P(#1) |
|---|---|---|---|---|---|
| 1 Kimi K3 Moonshot AI | 8.1 | 1.39 | 11.4% | 1–6 | 34% |
| 2 DeepSeek V4 Pro DeepSeek | 7.6 | 1.38 | 13.6% | 1–6 | 24% |
| 3 GPT-5.5 OpenAI | 7.4 | 1.36 | 10.2% | 1–6 | 16% |
| 4 GLM-5.2 Z.ai | 6.3 | 1.34 | 15.0% | 1–7 | 15% |
| 5 Qwen 3.7 Max Alibaba | 5.8 | 1.33 | 12.9% | 1–7 | 14% |
| 6 Gemini 3.5 Flash Google | 4.3 | 1.29 | 12.1% | 2–7 | 3% |
| 7 Claude Fable 5 Anthropic | 3.1 | 1.23 | 8.8% | 3–7 | 1% |
| – Claude Fable 5.1 Anthropic 5/140 events — not ranked | 16.7 | 2.00 | 40.0% | — | — |
| – Qwen3.8-Max Alibaba 5/140 events — not ranked | 16.7 | 2.00 | 40.0% | — | — |
| – GPT-5.6 Sol OpenAI 5/140 events — not ranked | 0.0 | 1.60 | 20.0% | — | — |
| – Grok 4.6 xAI 5/140 events — not ranked | 0.0 | 1.60 | 20.0% | — | — |
| ø Always 1–0 (home win) reference | 0.0 |
Scroll sideways for more columns
Rank intervals and P(#1) come from 4,000 paired resamples of the 137 events, seed 20260727 — same draw for every model, so the comparison stays fair. Events share context, which makes these intervals lower bounds.
The most common result in professional football, written down on 29 July 2026 — before any league event had resolved. A genuine a-priori reference: it was fixed first, the outcomes came later.
68 events open · next lock Sep 5, 2026, 11:30 UTC · 155 of 340 predictions committed
Cross-category skill (2 of 6 categories)
A cross-category number needs at least 2 qualified categories. 2 qualifies today, so none is computed — an average over a single category would just be that category’s table under a grander name.
Not qualified yet
A category qualifies with: real data, a clean integrity check, at least 20 resolved events, at least 3 models scored on at least 80% of the shared events, and a reference written down beforehand.
- Crypto
- no reference written down
- Stock index
- only 15 resolved events
- Elections & politics
- sample data
- Mixed sports
- sample data
Data as of Sep 5, 2026, 10:53 UTC · 4 live categories, 2 sample · 434 events tracked · page built Sep 5, 2026, 10:55 UTC
Proof · Models · Methodology