Long context · Reasoning
LongBench v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
Hard multiple-choice questions over contexts up to two million words.
- Released
- 2024
- Built by
- Tsinghua University (THUDM)
- Size
- 503 tasks
- Status
- Active
- Reported by
- Z.ai, OpenAI
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 82
- 503 instances, so one item moves the score by 0.199 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 32
- Reported by 2 labs; 2 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Single prompt, sometimes too long
No tools and no interaction — but contexts at the two-million-word end exceed every deployed model's window, so systems have to truncate or bring retrieval. The benchmark permits that and reports it, which means a LongBench v2 score can reflect the quality of a retrieval pipeline as much as the quality of a context window.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn; contexts range from 8,000 words to 2 million words.
Where the tasks came from
503 questions across six categories, so per-category claims each rest on fewer than a hundred items.
Authored by nearly 100 highly educated annotators, then filtered automatically and manually to remove questions that were too easy.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The 53.7% human figure was produced under a 15-minute time limit on contexts up to two million words, which is not enough time to read the input. Beating it is a real result but it is not evidence of superhuman long-context comprehension — it is evidence of beating a human who was not given time to look.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| Modelo1-previewOpenAI | Score57.7% | ConfigurationWith extended reasoning; roughly four points above the 15-minute human expert baseline. | SourceThird-party2024-12 |
| ModelHuman expertsLongBench v2 baseline | Score53.7% | ConfigurationHuman experts working under a 15-minute time limit — a speed-reading baseline, not a competence ceiling. | SourceThird-party2024-12 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.