All benchmarks

Multimodal · Reasoning · Knowledge

MMMU

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

College-exam questions with images, across thirty subjects, needing expert knowledge.

Released
2023
Built by
MMMU Benchmark team (academic consortium, first author Xiang Yue)
Size
11,550 tasks
Status
Saturated
Reported by
Anthropic
Signal55Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
100
11,550 instances, so one item moves the score by 0.009 points.
Fine print
46
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
24
Reported by 1 lab; 4 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Zero-shot, no tools

The project states evaluation is zero-shot, with no fine-tuning and no few-shot demonstrations, using each model's own default multiple-choice prompt where one exists and a validation-tuned prompt otherwise. Images are passed inline at their placeholder positions. Chain of thought is not mandated by the harness, so labs run with whatever reasoning mode their model defaults to — which is a large part of why two published MMMU numbers for the same model can disagree. Test-set answers were released publicly in February 2026, so the split is now gradable locally rather than through a submission server.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn

Where the tasks came from

11,550 questions split dev 150 / validation 900 / test 10,500. Nearly all reported numbers are the 900-item validation split.

Hand-collected from college exams, quizzes, textbooks and lecture materials by 50+ subject-major students, including coauthors. Apache-2.0.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • MMStar reports that Gemini Pro reached 42.9% on MMMU with no visual input, and that a text-only Qwen1.5-72B reached 42.4% two-shot, against an average random baseline of about 18.7%. A large share of the benchmark is answerable from the question text and options alone, so an MMMU score is only partly a measure of perception.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGPT-5.1OpenAIScore85.4%ConfigurationMMMU validation split. Cited from the MMMU leaderboard in Anthropic's Claude Opus 4.5 System Card (Table 2.3.A), not measured by Anthropic and not self-reported by OpenAI.SourceThird-party2025-11
ModelGPT-5-20250807OpenAIScore81.8%ConfigurationMMMU_VAL, top of the OpenCompass OpenVLM leaderboard data file, snapshot 20250917132916SourceThird-party2025-09-17
ModelClaude Opus 4.5AnthropicScore80.72%Configurationan average of 5 trials with a 64k thinking budgetSourceLab-reported2025-11
ModelClaude Sonnet 4.5AnthropicScore77.8%ConfigurationMMMU validation split, Claude Opus 4.5 System Card Table 2.3.ASourceLab-reported2025-11

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.