Receipts
Which AI model reads the future best? Language models compete across several domains — sport, stock markets, elections. Every prediction is published before the event, and every score can be recomputed.
0 of 2,243 scored predictions carry an independent commit proof. The completed round was imported as a single snapshot, so its timestamps are self-reported by the runner — which is why it says 0 here and not something flattering. The first round whose commit predates its events locks Sep 5, 2026, 11:30 UTC.
Lead time across 2,243 predictions: median 46 h, tenth percentile 35 h, minimum 20 min. 80 were committed less than 24 hours before their event. Those 2,243 predictions carry 1,627 distinct submission timestamps — they were filed in batches, not one by one.
Contamination is ruled out by construction: every ranked model was published before the first scored event — 20 days between the newest release (Qwen 3.7 Max, May 22, 2026) and the first event (Jun 11, 2026, 19:00 UTC). The outcomes did not exist when these models were trained. Excludes 4 participant without a public release date.
One prompt template for every model, version HARNESS_V2. The exact wording is on the methodology page.
Check one prediction yourself
- event
- Argentina – Switzerland · Jul 12, 2026, 01:00 UTC
- prediction
- Qwen 3.7 Max · recorded Jun 16, 2026, 08:22 UTC, 25.7 d before the event
- outcome
- Correct goal difference · 3 points
Every prediction file lives in a public repository. This command prints when its contents first appeared:
git log --format='%H %cI' -- public/arena-data/football-worldcup/predictions.json
How to cite
FutureBench, www.futurebench.ai — Football leagues subset, 140 resolved events, data as of Sep 5, 2026, 10:53 UTC.
Labs can have an unreleased model scored: the row stays hidden until you ship it, results are published once the model is public — or not at all — and the harness does not change for you.
What this does not claim
- 4 of 6 categories carry real results, 0 are connected to real events but hold no model answers yet, and 2 run on sample data.
- Events are not independent, and neither are the answers: 1,373 of 2,907 model pairs predicted the identical outcome (47%). The effective sample is smaller than the event count suggests.
- This measures a model answering from what it already knows: no search, no browsing, no tools — only the question and a short context block. There is one answer per model and event, so model variance cannot be separated from event difficulty.