Math
MATH-500
MATH-500, the 500-problem subset of MATH introduced in "Let's Verify Step by Step"
A 500-problem slice of MATH kept as a cheap standard subset.
- Released
- 2023
- Built by
- OpenAI
- Size
- 500 tasks
- Status
- Saturated
- Reported by
- OpenAI, DeepSeek, xAI, Anthropic, Artificial Analysis
- Judge
- 88
- Symbolic equivalence, so formatting differences do not change the score.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 82
- 500 instances, so one item moves the score by 0.200 points.
- Fine print
- 63
- 3 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 79
- Reported by 5 labs; 5 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, 0-shot CoT
0-shot chain of thought is the modern standard; 4-shot was standard historically. No tools. Because everything is above 99%, the remaining differences between models come from answer formatting and from how the equivalence check handles it, not from mathematics.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
500 items, giving about plus or minus 1pp at 95% accuracy. Precise, and pointless once the whole field is above 99%.
Sampled from the MATH test set by the OpenAI PRM800K team, chosen to be disjoint from their process-supervision training data.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
107 models have a MATH-500 score in the Artificial Analysis dataset and AA no longer runs it on new ones, so no 2026 frontier model (GPT-5.5/5.6, Claude Opus 5, Gemini 3.x, Grok 4.6) has a MATH-500 number at all. Its only remaining use is as a smoke test that a model has not regressed at basic mathematics.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5 (high)OpenAI | Score99.4% | ConfigurationArtificial Analysis run; AA no longer runs MATH-500 | SourceThird-party2026-08-16 |
| Modelo3OpenAI | Score99.2% | ConfigurationArtificial Analysis run; AA no longer runs MATH-500 | SourceThird-party2026-08-16 |
| ModelGrok 3 mini Reasoning (high)xAI | Score99.2% | ConfigurationArtificial Analysis run; AA no longer runs MATH-500 | SourceThird-party2026-08-16 |
| ModelClaude 4 Sonnet (Reasoning)Anthropic | Score99.07% | ConfigurationArtificial Analysis run; AA no longer runs MATH-500 | SourceThird-party2026-08-16 |
| ModelGrok 4xAI | Score99% | ConfigurationArtificial Analysis run; AA no longer runs MATH-500 | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.