All benchmarks

Math · Reasoning

AIME

American Invitational Mathematics Examination (2024/2025/2026 editions used as LLM benchmarks)

Thirty short-answer olympiad qualifier problems with integer answers 0-999.

Released
2024
Built by
Mathematical Association of America; LLM runs by MathArena (ETH Zurich)
Size
30 tasks
Status
Saturated
Reported by
MathArena (ETH Zurich), OpenAI, Anthropic, Google DeepMind, xAI, DeepSeek, Alibaba (Qwen), Moonshot AI, Mistral, Meta
Signal49Needs the fine print
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
22
30 instances, so one item moves the score by 3.33 points.
Fine print
54
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
100
Reported by 10 labs; 11 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Closed book, sampling varies wildly

Zero-shot with no tools in the canonical setup, but the sampling protocol is inconsistent across every reporter. MathArena runs 4 attempts per problem and averages. Artificial Analysis historically ran multiple repeats. Labs variously publish pass@1, maj@8, maj@32 or cons@64, and some publish a with-Python number, which is a different benchmark. MathArena also publishes cost per problem and average output tokens, which is where the 100x spread shows up.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn per problem

Where the tasks came from

30 problems, so one question is 3.33pp and one paper's 15 problems make it 6.67pp. Six 2026 models sit at 96.67%, exactly one question behind the leaders.

Written by the MAA problem-writing committee and released publicly right after each administration. No held-out set.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • One question is worth 3.33 percentage points on a full year and 6.67 on a single paper. Two models reported at 93.3% and 96.7% differ by exactly one question. At 50% accuracy on 30 items the binomial standard error is about 9.1pp. Essentially every AIME comparison between frontier models in 2025 and 2026 sits inside that band.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

AccuracyHigher is better
Accuracyhigher is better
ModelGPT-5.5 (xhigh)OpenAIScore100%ConfigurationAIME 2026, 30 problems, 4 runs per problem, no toolsSourceThird-party2026-08-16
ModelClaude-Opus-4.8 (max)AnthropicScore100%ConfigurationAIME 2026, 30 problems, 4 runs per problem, no toolsSourceThird-party2026-08-16
ModelGPT-5.2 (high)OpenAIScore100%ConfigurationAIME 2025, 30 problems, 4 runs per problem, no toolsSourceThird-party2026-08-16
ModelGPT-5.4 (xhigh)OpenAIScore99.17%ConfigurationAIME 2026, 30 problems, 4 runs per problem, no toolsSourceThird-party2026-08-16
ModelGemini 3.1 Pro PreviewGoogle DeepMindScore98.33%ConfigurationAIME 2026, 30 problems, 4 runs per problem, no toolsSourceThird-party2026-08-16
ModelStep 3.5 FlashStepFunScore96.67%ConfigurationAIME 2026, 30 problems, 4 runs per problem; 42,072 avg output tokensSourceThird-party2026-08-16
ModelKimi K3 (Think)Moonshot AIScore96.67%ConfigurationAIME 2026, 30 problems, 4 runs per problem, no toolsSourceThird-party2026-08-16
Cost per taskLower is better
Cost per tasklower is better
ModelStep 3.5 FlashStepFunScore$0.013ConfigurationAIME 2026 cost per problem at 96.67% accuracySourceThird-party2026-08-16
ModelGPT-5.2 (high)OpenAIScore$0.12ConfigurationAIME 2026 cost per problem at 98.33% accuracySourceThird-party2026-08-16
ModelGPT-5.5 (xhigh)OpenAIScore$0.16ConfigurationAIME 2026 cost per problem at 100% accuracy; 5,219 avg output tokensSourceThird-party2026-08-16
ModelClaude-Opus-4.8 (max)AnthropicScore$0.54ConfigurationAIME 2026 cost per problem at 100% accuracySourceThird-party2026-08-16

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.