All benchmarks

Multimodal

Video-MME

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Video question answering across clip lengths from seconds to an hour.

Released
2024
Built by
University of Science and Technology of China and collaborators
Size
2,700 tasks
Status
Near ceiling
Reported by
Alibaba (Qwen), Shanghai AI Lab, Kuaishou
Signal66Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
100
2,700 instances, so one item moves the score by 0.037 points.
Fine print
52
4 documented caveats, the heaviest being construct validity.
Adoption
50
Reported by 3 labs; 4 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Sampling left to you

The benchmark does not fix a frame-sampling policy. Evaluators sample anywhere from roughly 1 to 10 frames per second with different caps on total frames, and that choice materially moves the score, so two harnesses can report different numbers for the same model without either being wrong. The second uncontrolled dimension is whether subtitles and audio are supplied. A Video-MME number without both the sampling policy and the subtitle condition stated is not comparable to anything.

Single turn, no tools. The most reproducible setup there is.

Budget
single turn

Where the tasks came from

2,700 manually annotated QA pairs over 900 videos (254 hours), split short / medium / long.

Videos collected across 6 domains and 30 subfields; all QA pairs manually annotated.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Leading models gain several points when subtitles are supplied, so a portion of this benchmark is a reading task wearing a video costume. TVBench generalises the critique to the whole video-benchmark family: static information from a single frame often suffices, the question plus candidate answers are frequently informative enough to answer without any visual input, and world knowledge alone answers many items — making them knowledge-replication tests rather than video reasoning tests.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGemini 3 ProGoogle DeepMindScore88.6%ConfigurationThird-party evaluation in the InternVideo3 paper, Table 2. The source does not state its subtitle setting; without-subtitles is inferred from convention and from alignment with vendor figures, so treat the condition as unconfirmed.SourceThird-party2026-06-10
ModelGemini 2.5 Pro ThinkingGoogle DeepMindScore85.1%Configurationwithout subtitles, explicitly stated; from the Qwen3-VL technical report comparison tableSourceThird-party2025-11
ModelGPT-5-highOpenAIScore84.7%Configurationwithout subtitles, explicitly stated; from the Qwen3-VL technical report comparison tableSourceThird-party2025-11
Modelvideo-SALMONN 2+Tsinghua and ByteDanceScore81.6%Configurationwith subtitles; the maximum across all 51 rows of the official leaderboard, and the newest entry on itSourceCommunity2025-09-28

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Near ceiling The top models are close enough to the ceiling to crowd.