All benchmarks

Math

MATH-500

MATH-500, the 500-problem subset of MATH introduced in "Let's Verify Step by Step"

A 500-problem slice of MATH kept as a cheap standard subset.

Released
2023
Built by
OpenAI
Size
500 tasks
Status
Saturated
Reported by
OpenAI, DeepSeek, xAI, Anthropic, Artificial Analysis
Signal63Read with context
Judge
88
Symbolic equivalence, so formatting differences do not change the score.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
82
500 instances, so one item moves the score by 0.200 points.
Fine print
63
3 documented caveats, the heaviest being possible training-set contamination.
Adoption
79
Reported by 5 labs; 5 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Closed book, 0-shot CoT

0-shot chain of thought is the modern standard; 4-shot was standard historically. No tools. Because everything is above 99%, the remaining differences between models come from answer formatting and from how the equivalence check handles it, not from mathematics.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn

Where the tasks came from

500 items, giving about plus or minus 1pp at 95% accuracy. Precise, and pointless once the whole field is above 99%.

Sampled from the MATH test set by the OpenAI PRM800K team, chosen to be disjoint from their process-supervision training data.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • 107 models have a MATH-500 score in the Artificial Analysis dataset and AA no longer runs it on new ones, so no 2026 frontier model (GPT-5.5/5.6, Claude Opus 5, Gemini 3.x, Grok 4.6) has a MATH-500 number at all. Its only remaining use is as a smoke test that a model has not regressed at basic mathematics.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGPT-5 (high)OpenAIScore99.4%ConfigurationArtificial Analysis run; AA no longer runs MATH-500SourceThird-party2026-08-16
Modelo3OpenAIScore99.2%ConfigurationArtificial Analysis run; AA no longer runs MATH-500SourceThird-party2026-08-16
ModelGrok 3 mini Reasoning (high)xAIScore99.2%ConfigurationArtificial Analysis run; AA no longer runs MATH-500SourceThird-party2026-08-16
ModelClaude 4 Sonnet (Reasoning)AnthropicScore99.07%ConfigurationArtificial Analysis run; AA no longer runs MATH-500SourceThird-party2026-08-16
ModelGrok 4xAIScore99%ConfigurationArtificial Analysis run; AA no longer runs MATH-500SourceThird-party2026-08-16

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.