Math · Reasoning
FrontierMath
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Unpublished research-level math problems with automatically checkable answers.
- Released
- 2024
- Built by
- Epoch AI, with 60+ research mathematicians
- Size
- 338 tasks
- Status
- Near ceiling
- Reported by
- Epoch AI, OpenAI, Anthropic, Google DeepMind
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 40
- The top models crowd the ceiling; gaps are shrinking.
- Resolution
- 68
- 338 instances, so one item moves the score by 0.296 points.
- Fine print
- 50
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 71
- Reported by 4 labs; 7 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Python sandbox, effort-tiered
Models get a code-execution sandbox and are expected to use it; this is not optional, because many problems require real computation. Reasoning-effort setting dominates the result, and Epoch reports max, xhigh and high as separate rows for the same model because they differ by tens of points. A no-tools FrontierMath number would be a different benchmark.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Agentic, model may run code
- Tools exposed
- Python code execution
Where the tasks came from
338 problems after v2: 295 in Tiers 1-3 and 43 in Tier 4, where one problem is 2.3pp and the published error bars run to plus or minus 7 points.
Commissioned from 60+ research mathematicians and peer-reviewed. Twelve problems public, the rest private; an additional holdout set is withheld from the funder.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The v2 release corrected 123 problems in Tiers 1-3 and 12 in Tier 4, and removed 5 from Tiers 1-3 and 7 from Tier 4. That is errors touched in 42% of the dataset. Pre-v2 and post-v2 scores are measurements of different benchmarks, and a great deal of 2025 reporting cites v1 numbers without saying so.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5.6 Sol (max)OpenAI | Score89.1% | ConfigurationTiers 1-3, v2, 295 private problems, +/-1.8 | SourceThird-party2026-08-16 |
| ModelClaude Fable 5 (max)Anthropic | Score87.8% | ConfigurationTier 4, v2, 43 private problems, +/-5.2 | SourceThird-party2026-08-16 |
| ModelGPT-5.5 Pro (xhigh)OpenAI | Score87.7% | ConfigurationTiers 1-3, v2, 295 private problems, +/-1.9 | SourceThird-party2026-08-16 |
| ModelClaude Fable 5 (max)Anthropic | Score87% | ConfigurationTiers 1-3, v2, 295 private problems, +/-2.0 | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (max)OpenAI | Score82.9% | ConfigurationTier 4, v2, 43 private problems, +/-5.9 | SourceThird-party2026-08-16 |
| ModelAI co-mathematicianGoogle DeepMind | Score75.6% | ConfigurationTier 4, v2, 43 private problems, +/-6.7 | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (max)Anthropic | Score73.2% | ConfigurationTier 4, v2, 43 private problems, +/-7.0 | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Near ceiling — The top models are close enough to the ceiling to crowd.