Multimodal
Video-MME
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Video question answering across clip lengths from seconds to an hour.
- Released
- 2024
- Built by
- University of Science and Technology of China and collaborators
- Size
- 2,700 tasks
- Status
- Near ceiling
- Reported by
- Alibaba (Qwen), Shanghai AI Lab, Kuaishou
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 40
- The top models crowd the ceiling; gaps are shrinking.
- Resolution
- 100
- 2,700 instances, so one item moves the score by 0.037 points.
- Fine print
- 52
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 50
- Reported by 3 labs; 4 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Sampling left to you
The benchmark does not fix a frame-sampling policy. Evaluators sample anywhere from roughly 1 to 10 frames per second with different caps on total frames, and that choice materially moves the score, so two harnesses can report different numbers for the same model without either being wrong. The second uncontrolled dimension is whether subtitles and audio are supplied. A Video-MME number without both the sampling policy and the subtitle condition stated is not comparable to anything.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single turn
Where the tasks came from
2,700 manually annotated QA pairs over 900 videos (254 hours), split short / medium / long.
Videos collected across 6 domains and 30 subfields; all QA pairs manually annotated.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Leading models gain several points when subtitles are supplied, so a portion of this benchmark is a reading task wearing a video costume. TVBench generalises the critique to the whole video-benchmark family: static information from a single frame often suffices, the question plus candidate answers are frequently informative enough to answer without any visual input, and world knowledge alone answers many items — making them knowledge-replication tests rather than video reasoning tests.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGemini 3 ProGoogle DeepMind | Score88.6% | ConfigurationThird-party evaluation in the InternVideo3 paper, Table 2. The source does not state its subtitle setting; without-subtitles is inferred from convention and from alignment with vendor figures, so treat the condition as unconfirmed. | SourceThird-party2026-06-10 |
| ModelGemini 2.5 Pro ThinkingGoogle DeepMind | Score85.1% | Configurationwithout subtitles, explicitly stated; from the Qwen3-VL technical report comparison table | SourceThird-party2025-11 |
| ModelGPT-5-highOpenAI | Score84.7% | Configurationwithout subtitles, explicitly stated; from the Qwen3-VL technical report comparison table | SourceThird-party2025-11 |
| Modelvideo-SALMONN 2+Tsinghua and ByteDance | Score81.6% | Configurationwith subtitles; the maximum across all 51 rows of the official leaderboard, and the newest entry on it | SourceCommunity2025-09-28 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Near ceiling — The top models are close enough to the ceiling to crowd.