All benchmarks

Math · Reasoning

FrontierMath

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Unpublished research-level math problems with automatically checkable answers.

Released
2024
Built by
Epoch AI, with 60+ research mathematicians
Size
338 tasks
Status
Near ceiling
Reported by
Epoch AI, OpenAI, Anthropic, Google DeepMind
Signal62Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
68
338 instances, so one item moves the score by 0.296 points.
Fine print
50
4 documented caveats, the heaviest being construct validity.
Adoption
71
Reported by 4 labs; 7 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Python sandbox, effort-tiered

Models get a code-execution sandbox and are expected to use it; this is not optional, because many problems require real computation. Reasoning-effort setting dominates the result, and Epoch reports max, xhigh and high as separate rows for the same model because they differ by tens of points. A no-tools FrontierMath number would be a different benchmark.

Single turn, no tools. The most reproducible setup there is.

Budget
Agentic, model may run code
Tools exposed
Python code execution

Where the tasks came from

338 problems after v2: 295 in Tiers 1-3 and 43 in Tier 4, where one problem is 2.3pp and the published error bars run to plus or minus 7 points.

Commissioned from 60+ research mathematicians and peer-reviewed. Twelve problems public, the rest private; an additional holdout set is withheld from the funder.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The v2 release corrected 123 problems in Tiers 1-3 and 12 in Tier 4, and removed 5 from Tiers 1-3 and 7 from Tier 4. That is errors touched in 42% of the dataset. Pre-v2 and post-v2 scores are measurements of different benchmarks, and a great deal of 2025 reporting cites v1 numbers without saying so.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGPT-5.6 Sol (max)OpenAIScore89.1%ConfigurationTiers 1-3, v2, 295 private problems, +/-1.8SourceThird-party2026-08-16
ModelClaude Fable 5 (max)AnthropicScore87.8%ConfigurationTier 4, v2, 43 private problems, +/-5.2SourceThird-party2026-08-16
ModelGPT-5.5 Pro (xhigh)OpenAIScore87.7%ConfigurationTiers 1-3, v2, 295 private problems, +/-1.9SourceThird-party2026-08-16
ModelClaude Fable 5 (max)AnthropicScore87%ConfigurationTiers 1-3, v2, 295 private problems, +/-2.0SourceThird-party2026-08-16
ModelGPT-5.6 Sol (max)OpenAIScore82.9%ConfigurationTier 4, v2, 43 private problems, +/-5.9SourceThird-party2026-08-16
ModelAI co-mathematicianGoogle DeepMindScore75.6%ConfigurationTier 4, v2, 43 private problems, +/-6.7SourceThird-party2026-08-16
ModelClaude Opus 5 (max)AnthropicScore73.2%ConfigurationTier 4, v2, 43 private problems, +/-7.0SourceThird-party2026-08-16

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Near ceiling The top models are close enough to the ceiling to crowd.