Knowledge · Reasoning
MMLU-Pro
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Harder MMLU rebuild with ten options and reasoning-heavy questions.
- Released
- 2024
- Built by
- TIGER-Lab (University of Waterloo), with Toronto and CMU
- Size
- 12,032 tasks
- Status
- Saturated
- Reported by
- Alibaba (Qwen), DeepSeek, Mistral, Meta, Google DeepMind, Z.ai, Moonshot AI, Artificial Analysis
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 100
- 12,032 instances, so one item moves the score by 0.008 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 89
- Reported by 8 labs; 4 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, CoT expected
No tools, no retrieval. The paper's protocol is 5-shot CoT; nearly all 2025-2026 numbers are 0-shot CoT from reasoning models, and the two are not the same measurement. Some vendors report it with self-consistency or maj@k without labelling it. It is far less prompt-sensitive than MMLU (about 2% score variance versus 4-5%), but ten long generations across 12,032 items make it expensive enough that few groups rerun it.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
12,032 test items across 14 disciplines, with a 10% guessing floor rather than MMLU's 25%.
Filtered MMLU items plus STEM websites, TheoremQA and SciBench; GPT-4-generated distractors, human-verified. Fully public.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Artificial Analysis dropped MMLU-Pro from the Intelligence Index in v4.0 (January 2026) and now runs it only as an Additional Evaluation at one repeat. The practical consequence is that no 2026 frontier model (GPT-5.5/5.6, Claude Opus 5, Grok 4.6) has an MMLU-Pro score in that dataset at all, so any 2026 cross-model MMLU-Pro comparison mixes numbers measured a year apart.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGemini 3 Pro Preview (high)Google DeepMind | Score89.8% | ConfigurationArtificial Analysis run; Additional Evaluation, 1 repeat | SourceThird-party2026-08-16 |
| ModelGemini 3 Pro Preview (low)Google DeepMind | Score89.5% | ConfigurationArtificial Analysis run; Additional Evaluation, 1 repeat | SourceThird-party2026-08-16 |
| ModelClaude Opus 4.5 (Reasoning)Anthropic | Score89.5% | ConfigurationArtificial Analysis run; Additional Evaluation, 1 repeat | SourceThird-party2026-08-16 |
| ModelGemini 3 Flash Preview (Reasoning)Google DeepMind | Score89% | ConfigurationArtificial Analysis run; Additional Evaluation, 1 repeat | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.