Math · Reasoning
MathArena Apex
MathArena Apex
Twelve final-answer problems chosen because frontier models all failed them.
- Released
- 2025
- Built by
- SRI Lab, ETH Zurich
- Size
- 12 tasks
- Status
- Active
- Reported by
- MathArena (ETH Zurich)
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 10
- 12 instances, so one item moves the score by 8.33 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 35
- Reported by 1 lab; 10 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Zero-shot, four runs, cost published
Standard MathArena protocol: zero-shot, no tools, four runs per problem averaged, with cost per problem and token counts published for every model. The cost column is where Apex is most instructive: Gemini 3.1 Pro reaches 60.9% at $0.41 per problem while GPT-5.4-Pro reaches 69.8% at $8.65, a twentyfold price difference for nine points.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn per problem
Where the tasks came from
12 problems, so one problem is 8.33pp. This is the noisiest benchmark in this collection; treat any sub-10-point gap as unresolved.
Curated by the MathArena team from recent competitions, selected specifically because contemporary frontier models scored zero on them.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The gap between GPT-5.5 (xhigh) at 80.21% and Claude-Opus-4.8 (max) at 81.25% is a fraction of one problem across four runs. Even the gap between fourth and sixth place is roughly one problem. Rank ordering here should be read as tiers, not as a list.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude-Opus-4.8 (max)Anthropic | Score81.25% | Configuration12 problems, 4 runs each, zero-shot, no tools | SourceThird-party2026-08-16 |
| ModelGPT-5.5 (xhigh)OpenAI | Score80.21% | Configuration12 problems, 4 runs each, zero-shot, no tools | SourceThird-party2026-08-16 |
| ModelGPT-5.4-ProOpenAI | Score69.79% | Configuration12 problems, 4 runs each, zero-shot, no tools | SourceThird-party2026-08-16 |
| ModelKimi K3Moonshot AI | Score65.62% | Configuration12 problems, 4 runs each, zero-shot, no tools | SourceThird-party2026-08-16 |
| ModelGemini 3.1 ProGoogle DeepMind | Score60.94% | Configuration12 problems, 4 runs each, zero-shot, no tools | SourceThird-party2026-08-16 |
| ModelDeepSeek-v4-ProDeepSeek | Score28.12% | Configuration12 problems, 4 runs each, zero-shot, no tools | SourceThird-party2026-08-16 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGemini 3.1 ProGoogle DeepMind | Score$0.41 | ConfigurationCost per problem at 60.94% accuracy | SourceThird-party2026-08-16 |
| ModelKimi K3Moonshot AI | Score$1.11 | ConfigurationCost per problem at 65.62% accuracy | SourceThird-party2026-08-16 |
| ModelClaude-Opus-4.8 (max)Anthropic | Score$4.59 | ConfigurationCost per problem at 81.25% accuracy | SourceThird-party2026-08-16 |
| ModelGPT-5.4-ProOpenAI | Score$8.65 | ConfigurationCost per problem at 69.79% accuracy | SourceThird-party2026-08-16 |
Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.
At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.