All benchmarks

Math · Reasoning

MathArena Apex

MathArena Apex

Twelve final-answer problems chosen because frontier models all failed them.

Released
2025
Built by
SRI Lab, ETH Zurich
Size
12 tasks
Status
Active
Reported by
MathArena (ETH Zurich)
Signal56Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
82
Still separates the frontier from everything below it.
Resolution
10
12 instances, so one item moves the score by 8.33 points.
Fine print
47
4 documented caveats, the heaviest being construct validity.
Adoption
35
Reported by 1 lab; 10 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Zero-shot, four runs, cost published

Standard MathArena protocol: zero-shot, no tools, four runs per problem averaged, with cost per problem and token counts published for every model. The cost column is where Apex is most instructive: Gemini 3.1 Pro reaches 60.9% at $0.41 per problem while GPT-5.4-Pro reaches 69.8% at $8.65, a twentyfold price difference for nine points.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn per problem

Where the tasks came from

12 problems, so one problem is 8.33pp. This is the noisiest benchmark in this collection; treat any sub-10-point gap as unresolved.

Curated by the MathArena team from recent competitions, selected specifically because contemporary frontier models scored zero on them.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The gap between GPT-5.5 (xhigh) at 80.21% and Claude-Opus-4.8 (max) at 81.25% is a fraction of one problem across four runs. Even the gap between fourth and sixth place is roughly one problem. Rank ordering here should be read as tiers, not as a list.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

AccuracyHigher is better
Accuracyhigher is better
ModelClaude-Opus-4.8 (max)AnthropicScore81.25%Configuration12 problems, 4 runs each, zero-shot, no toolsSourceThird-party2026-08-16
ModelGPT-5.5 (xhigh)OpenAIScore80.21%Configuration12 problems, 4 runs each, zero-shot, no toolsSourceThird-party2026-08-16
ModelGPT-5.4-ProOpenAIScore69.79%Configuration12 problems, 4 runs each, zero-shot, no toolsSourceThird-party2026-08-16
ModelKimi K3Moonshot AIScore65.62%Configuration12 problems, 4 runs each, zero-shot, no toolsSourceThird-party2026-08-16
ModelGemini 3.1 ProGoogle DeepMindScore60.94%Configuration12 problems, 4 runs each, zero-shot, no toolsSourceThird-party2026-08-16
ModelDeepSeek-v4-ProDeepSeekScore28.12%Configuration12 problems, 4 runs each, zero-shot, no toolsSourceThird-party2026-08-16
Cost per taskLower is better
Cost per tasklower is better
ModelGemini 3.1 ProGoogle DeepMindScore$0.41ConfigurationCost per problem at 60.94% accuracySourceThird-party2026-08-16
ModelKimi K3Moonshot AIScore$1.11ConfigurationCost per problem at 65.62% accuracySourceThird-party2026-08-16
ModelClaude-Opus-4.8 (max)AnthropicScore$4.59ConfigurationCost per problem at 81.25% accuracySourceThird-party2026-08-16
ModelGPT-5.4-ProOpenAIScore$8.65ConfigurationCost per problem at 69.79% accuracySourceThird-party2026-08-16

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.