All benchmarks

Multimodal · Reasoning

MMMU-Pro

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

MMMU rebuilt to block text-only shortcuts: ten options, questions hidden in screenshots.

Released
2024
Built by
MMMU Benchmark team with CMU
Size
1,730 tasks
Status
Near ceiling
Reported by
Google DeepMind
Signal62Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
92
1,730 instances, so one item moves the score by 0.058 points.
Fine print
56
4 documented caveats, the heaviest being construct validity.
Adoption
24
Reported by 1 lab; 4 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Four fixed prompts

Zero-shot, single-turn, no tools. The official prompts.yaml defines exactly four prompts: mode in {cot, direct} crossed with setting in {standard, vision}. Direct prompting and CoT prompting produce materially different scores, so the mode has to be stated for a number to mean anything. The paper also tested adding explicit OCR instructions in the vision setting and found they did not significantly alter performance, which is its main defence against the charge that Vision measures OCR rather than reasoning.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn
Actual promptSource
CoT / vision:
Write out the multiple-choice question in the image and then solve it. The last line of your response should be of the following format: "Answer: $LETTER" (without quotes) where LETTER is one of options. Think step by step before answering.

Direct / standard:
Answer with the option letter from the given choices directly.

Where the tasks came from

1,730 questions per config, three configs (standard 4-option, standard 10-option, vision), all in a single test split. The paper counts standard plus screenshot instances as 3,460 total.

Derived from MMMU by text-only-LLM filtering, then LLM-generated distractors filtered by Claude 3.5 and reviewed twice by human experts. Vision screenshots were captured manually. Apache-2.0.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Google's 81.0% for Gemini 3 Pro is the official Standard-10 plus Vision average. vals.ai's ~89-90% figures come from a four-option CoT variant over roughly 1,700 questions. Artificial Analysis runs its own AA-MMMU-Pro. Claude Opus 5 scores 89.88% on vals.ai and 84.7% on AA-MMMU-Pro. These are different benchmarks in practice and should never be placed on the same axis.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelClaude Opus 5AnthropicScore89.88%Configurationvals.ai variant: ~1,700 questions, 4 options, CoT prompting, vals.ai's own answer extraction. Not comparable to the official Standard+Vision average above.SourceThird-party2026-08-15
ModelGemini 3 ProGoogle DeepMindScore81%ConfigurationMMMU-Pro scores are averaged across the Standard (10 options) and Vision settings.SourceLab-reported2025-11
ModelGPT-5.1OpenAIScore76%ConfigurationStandard (10 options) and Vision averaged; computed by Google DeepMind for its comparison table rather than self-reported by OpenAI.SourceThird-party2025-11
ModelClaude Sonnet 4.5AnthropicScore68%ConfigurationStandard (10 options) and Vision averaged; computed by Google DeepMind, not self-reported by Anthropic.SourceThird-party2025-11

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Near ceiling The top models are close enough to the ceiling to crowd.