Math · Reasoning
IMO-ProofBench
IMO-Proof Bench — the proof-writing benchmark of the IMO-Bench suite
Sixty olympiad problems marked on the written proof, not the final answer.
- Released
- 2025
- Built by
- Google DeepMind
- Size
- 60 tasks
- Status
- Near ceiling
- Reported by
- Google DeepMind, DeepSeek
- Judge
- 58
- Mixed grading, so part of the score inherits the weakest judge.
- Headroom
- 40
- The top models crowd the ceiling; gaps are shrinking.
- Resolution
- 36
- 60 instances, so one item moves the score by 1.67 points.
- Fine print
- 45
- 5 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 48
- Reported by 2 labs; 21 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
One prompt, no tools — on paper
The benchmark itself asks for a single natural-language proof with no tools and no environment, and DeepMind's own scaling study ran each model exactly once per problem at each compute scale with tool use disabled. What actually sits behind a published number is usually something far larger. Gemini Deep Think explores many candidate proofs in parallel and its compute can be dialled across orders of magnitude. Huang and Yang's open-source harness wraps Gemini 2.5 Pro in a self-verification loop that repeats up to 5 times, restarts on failure, exits after 10 consecutive failed verifications, and runs 100 of those pipelines in parallel. DeepSeekMath-V2's Heavy configuration starts from 64 proof samples with 64 verification analyses each and refines for up to 16 iterations. Aletheia is a generator-verifier-reviser agent whose total compute, DeepMind says, cannot be precisely controlled.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single run per problem in the paper's own protocol; unbounded in the agentic systems on the leaderboard
Where the tasks came from
60 problems in two subsets of 30, each problem worth 7 points. One fully solved problem is 3.33 points of a subset score, so gaps under about 5 points are one problem wide. 33 of the 60 have no short answer at all.
Written and vetted by a panel of IMO medallists and mathematicians at Google DeepMind. Basic problems are largely rephrased existing olympiad problems; the advanced set is 18 novel problems plus 6 robustified IMO 2024 problems and 6 USAMO 2025 problems. Released as a CSV under CC-BY carrying the problem, a reference solution, per-problem grading guidelines, category, difficulty level and source.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The advanced breakdown by source is an overfitting detector, and it fires. Grok 4 (heavy) scores 76.2% on the six USAMO 2025 problems and 11.1% on the eighteen novel ones. o3 is 52.4% against 15.1%, and Gemini 2.5 Pro under the Huang-Yang harness is 52.4% against 17.5%. Gemini Deep Think (IMO Gold) is the only system in the September 2025 table without a large gap, at 69.0% USAMO against 61.1% novel. Beyond the recycled problems, the entire benchmark — statements, reference solutions and grading guidelines — sits in a public CC-BY CSV on GitHub, and the paper names future contamination as one of its two headline limitations.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelAletheiaGoogle DeepMind | Score95.1% | ConfigurationAletheia paper, graded by human experts, no internet access. Conditional accuracy was 98.3% on the 29 of 30 problems where the agent returned a solution — it declines to answer rather than bluff, which the overall score absorbs as a zero. | SourceLab-reported2026-02-10 |
| ModelAletheiaGoogle DeepMind | Score91.9% | ConfigurationIMO-Bench leaderboard, query date 2026-02-09. Breakdown: 92.1% novel, 100.0% IMO 2024, 83.3% USAMO 2025. The leaderboard states no grading method or run count. The Aletheia paper reports 95.1% for the same system. | SourceLab-reported2026-02-09 |
| ModelGPT-5.5 Pro (xhigh)OpenAI | Score88.1% | ConfigurationIMO-Bench leaderboard, query date 2026-05-15. Breakdown: 80.2% novel, 100.0% IMO 2024, 100.0% USAMO 2025 — the two recycled sources are perfect while the novel problems are twenty points lower. Run by DeepMind; OpenAI does not report this benchmark. | SourceThird-party2026-05-15 |
| ModelGemini 3 Deep ThinkGoogle DeepMind | Score76.7% | ConfigurationIMO-Bench leaderboard, query date 2026-02-04. Breakdown: 75.4% novel, 73.8% IMO 2024, 83.3% USAMO 2025. | SourceLab-reported2026-02-04 |
| ModelGPT-5.5 (xhigh)OpenAI | Score71.9% | ConfigurationIMO-Bench leaderboard, query date 2026-05-15. Breakdown: 71.4% novel, 61.9% IMO 2024, 83.3% USAMO 2025. | SourceThird-party2026-05-15 |
| ModelGemini Deep Think (IMO Gold)Google DeepMind | Score65.7% | ConfigurationTable 6 of the IMO-Bench paper, human expert graded, query date 2025-08-02. Breakdown: 61.1% novel, 76.2% IMO 2024, 69.0% USAMO 2025 — the flattest source profile in that table. | SourceLab-reported2025-08-02 |
| ModelDeepSeekMath-V2 (Heavy)DeepSeek | Score61.9% | ConfigurationFigure 3 of the DeepSeekMath-V2 paper, graded by DeepSeek's own experts following the published grading guidelines rather than by DeepMind. Figure 3 is a chart image, so the decimals are not recoverable from the paper text, which says only that the model is "competitive on the advanced" set. The configuration — 64 proof samples with 64 verification analyses each, refined for up to 16 iterations — is described in the text but never named "Heavy" there. | SourceLab-reported2025-11-27 |
| ModelGemini 3.1 ProGoogle DeepMind | Score49% | ConfigurationIMO-Bench leaderboard, query date 2026-05-25. Breakdown: 60.3% novel, 23.8% IMO 2024, 40.5% USAMO 2025 — one of only two rows on the current board that score higher on the novel problems than on either recycled source. | SourceLab-reported2026-05-25 |
| ModelGPT-5.2 Thinking (high)OpenAI | Score35.7% | ConfigurationIMO-Bench leaderboard, query date 2026-01-14. Breakdown: 26.2% novel, 66.7% IMO 2024, 50.0% USAMO 2025. | SourceThird-party2026-01-14 |
| ModelClaude Opus 4.5Anthropic | Score23.8% | ConfigurationIMO-Bench leaderboard, query date 2026-01-14. Breakdown: 21.4% novel, 14.3% IMO 2024, 42.9% USAMO 2025. Anthropic has never published an IMO-ProofBench figure itself. | SourceThird-party2026-01-14 |
| ModelGrok 4 (heavy)xAI | Score23.3% | ConfigurationTable 6, query date 2025-07-12. 76.2% on USAMO 2025 against 11.1% on novel problems, the widest source gap published. Three problems were scored 0 after repeated query failures. | SourceThird-party2025-07-12 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelDeepSeekMath-V2 (Heavy)DeepSeek | Score99% | ConfigurationFigure 3 of the DeepSeekMath-V2 paper, graded by DeepSeek's own experts. The highest basic-subset figure published by anyone, ten points above DeepMind's IMO-gold model — but read off a chart image; the text claims only that it "outperforms DeepThink (IMO Gold) on the basic set". | SourceLab-reported2025-11-27 |
| ModelGemini Deep Think (IMO Gold)Google DeepMind | Score89% | ConfigurationTable 6, human expert graded, query date 2025-08-02. The benchmark's basic set has had no published score for any model released after September 2025. | SourceLab-reported2025-08-02 |
| ModelGemini 2.5 Deep Think (IMO lite)Google DeepMind | Score83.8% | ConfigurationTable 6, human expert graded, query date 2025-08-20. | SourceLab-reported2025-08-20 |
| ModelGemini 2.5 Pro + Huang & Yang harnessGoogle DeepMind | Score69.5% | ConfigurationTable 6, query date 2025-07-14. An external open-source agentic scaffold around Gemini 2.5 Pro, not a single model call — it lifts the same base model from 55.2% to 69.5% on basic and from 17.6% to 24.8% on advanced. | SourceThird-party2025-07-14 |
| ModelGPT-5OpenAI | Score59% | ConfigurationTable 6, human expert graded, query date 2025-09-18. Advanced score for the same run was 20.0%. | SourceThird-party2025-09-18 |
| ModelGemini 2.5 ProGoogle DeepMind | Score55.2% | ConfigurationTable 6, query date 2025-08-04. This is also the model ProofAutoGrader is built on, so the benchmark's automatic grader is a system that scores 17.6% on the advanced problems it grades. | SourceLab-reported2025-08-04 |
| Modelo3OpenAI | Score54.8% | ConfigurationTable 6, query date 2025-08-04. Advanced 20.5%. | SourceThird-party2025-08-04 |
| ModelGrok 4xAI | Score46.7% | ConfigurationTable 6, query date 2025-08-20. Advanced 18.6%. | SourceThird-party2025-08-20 |
| ModelClaude Sonnet 4Anthropic | Score27.1% | ConfigurationTable 6, query date 2025-09-17. One problem was scored 0 after the query failed three times. | SourceThird-party2025-09-17 |
| ModelClaude Opus 4Anthropic | Score11.9% | ConfigurationTable 6, query date 2025-08-04. The lowest basic score in the paper; advanced 2.9%. | SourceThird-party2025-08-04 |
Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Near ceiling — The top models are close enough to the ceiling to crowd.