Math
GSM8K
Grade School Math 8K
Grade-school word problems needing two to eight arithmetic steps.
- Released
- 2021
- Built by
- OpenAI
- Size
- 1,319 tasks
- Status
- Saturated
- Reported by
- OpenAI, Google DeepMind, Meta, Mistral, Alibaba (Qwen)
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 92
- 1,319 instances, so one item moves the score by 0.076 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 65
- Reported by 5 labs; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, sampling matters
Historically 8-shot with worked CoT exemplars; modern reasoning models run 0-shot. Enabling a calculator or Python changes the number substantially, and maj@k self-consistency was the standard trick of the 2022-2023 era, with maj@100 common in papers. A GSM8K number without its shot count, sampling and tool setting is not a measurement.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
1,319 test items. Large enough for tight error bars, which is not the constraint here: contamination and ceiling effects are.
Written from scratch by human problem writers hired by OpenAI under a 2-8 step constraint. 7,473 train / 1,319 test, fully public.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Scale AI built GSM1k, a fresh 1,250-problem set matched to GSM8K's distribution, and found accuracy drops of up to 8 points for some model families, with a Spearman correlation of roughly 0.36 (r-squared) between how likely a model was to emit GSM8K test items verbatim and the size of its performance gap. Frontier models showed minimal overfitting; smaller and heavily tuned ones did not.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Saturated — Scores are bunched at the ceiling; it no longer separates models.