Multimodal
MMAU
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Audio understanding across speech, environmental sound and music.
- Released
- 2024
- Built by
- University of Maryland and Adobe
- Size
- 10,000 tasks
- Status
- Saturated
- Reported by
- Alibaba (Qwen), Google DeepMind, Amazon, StepFun
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 100
- 10,000 instances, so one item moves the score by 0.010 points.
- Fine print
- 72
- 3 documented caveats, the heaviest being cross-lab comparability.
- Adoption
- 63
- Reported by 4 labs; 4 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Audio in, text out
Audio in, text out, single turn. The condition that decides comparability is which split was used: results exist for both test-mini and the full test split, and the same model differs by roughly two points between them. Most vendor reports quote a single number without saying which, so a MMAU figure taken from a technical report and one taken from the official leaderboard generally are not measuring the same thing.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single turn
Where the tasks came from
10,000 clips across 3 domains and 27 skills, reported on a test-mini split and a full test split whose scores differ by roughly 2 points for the same model.
Curated audio clips with human-annotated reasoning questions across speech, environmental sound and music.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The official leaderboard reports test-mini and full test separately — Nova 2 Omni at 77.8 test-mini and 75.28 test, Audio-Thinker at 77.7 and 75.98. Vendor technical reports quote a single figure with the split unstated, and the highest of them (82.2) is almost certainly test-mini. Placing a vendor figure and a leaderboard figure on the same axis silently compares two different measurements.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelQwen3.5-Omni-PlusAlibaba Qwen | Score82.2% | ConfigurationQwen3.5-Omni Technical Report, Table 5. The split is not stated in the source but is almost certainly test-mini, whose human baseline is 82.23. | SourceLab-reported2026-04 |
| ModelGemini 3.1 ProGoogle DeepMind | Score81.1% | Configurationsame Qwen3.5-Omni comparison table, split unstated; not self-reported by Google | SourceThird-party2026-04 |
| ModelNova 2 OmniAmazon | Score77.8% | Configurationtest-mini split on the official leaderboard; the same model scores 75.28 on the full test split | SourceCommunity2026-08-16 |
| ModelAudio-Thinker (8.4B)Academic | Score75.98% | Configurationfull test split — the highest verified score on the full test split; the same model scores 77.7 on test-mini | SourceCommunity2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.