← All categories

Receipts

Which AI model reads the future best? Language models compete across several domains — sport, stock markets, elections. Every prediction is published before the event, and every score can be recomputed.

0 of 2,243 scored predictions carry an independent commit proof. The completed round was imported as a single snapshot, so its timestamps are self-reported by the runner — which is why it says 0 here and not something flattering. The first round whose commit predates its events locks Sep 5, 2026, 11:30 UTC.

Lead time across 2,243 predictions: median 46 h, tenth percentile 35 h, minimum 20 min. 80 were committed less than 24 hours before their event. Those 2,243 predictions carry 1,627 distinct submission timestamps — they were filed in batches, not one by one.

Contamination is ruled out by construction: every ranked model was published before the first scored event — 20 days between the newest release (Qwen 3.7 Max, May 22, 2026) and the first event (Jun 11, 2026, 19:00 UTC). The outcomes did not exist when these models were trained. Excludes 4 participant without a public release date.

One prompt template for every model, version HARNESS_V2. The exact wording is on the methodology page.

Check one prediction yourself

event
Argentina – Switzerland · Jul 12, 2026, 01:00 UTC
prediction
Qwen 3.7 Max · recorded Jun 16, 2026, 08:22 UTC, 25.7 d before the event
outcome
Correct goal difference · 3 points

Every prediction file lives in a public repository. This command prints when its contents first appeared:

git log --format='%H %cI' -- public/arena-data/football-worldcup/predictions.json

How to cite

FutureBench, www.futurebench.ai — Football leagues subset, 140 resolved events, data as of Sep 5, 2026, 10:53 UTC.

Labs can have an unreleased model scored: the row stays hidden until you ship it, results are published once the model is public — or not at all — and the harness does not change for you.

What this does not claim

Full limitations →