All benchmarks

Long context

MRCR

OpenAI MRCR (Multi-Round Co-reference Resolution), building on Google DeepMind's MRCR from the Michelangelo paper

Find the second poem about tapirs among fifty near-identical ones.

Released
2025
Built by
OpenAI (dataset); original task from Google DeepMind
Size
2,400 tasks
Status
Active
Reported by
OpenAI, Google DeepMind, Anthropic
Signal72Reads cleanly
Judge
58
Mixed grading, so part of the score inherits the weakest judge.
Headroom
82
Still separates the frontier from everything below it.
Resolution
100
2,400 instances, so one item moves the score by 0.042 points.
Fine print
53
4 documented caveats, the heaviest being construct validity.
Adoption
61
Reported by 3 labs; 8 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

One long prompt

Nothing but a single prompt — no tools, no interaction, no retrieval step the model can outsource to a scaffold. That is the whole point: it isolates context handling with almost no confounds, which is what makes it a comparatively honest number in a field where most agentic results are dominated by harness choices.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn; prompts range from about 4k to over 1M tokens.

Where the tasks came from

2,400 rows: 100 samples per bin across 8 context-length bins from 4,096 to 1,048,576 tokens.

Fully synthetic; assistant turns generated by GPT-4o, with needles drawn from the same distribution as the distractors.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Google's Gemini 3 Pro model card (results as of November 2025) reports 77.0% at 8-needle 128k against 61.6% for GPT-5.1, 58.0% for Gemini 2.5 Pro and 47.1% for Claude Sonnet 4.5 — but at 1M pointwise, Gemini 3 Pro drops to 26.3% and Gemini 2.5 Pro to 16.4%. Quoting the 128k number as the model's long-context ability is the single most common misreading of this benchmark.

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGemini 3 ProGoogleScore77%ConfigurationMRCR v2, 8-needle at 128k, average; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelClaude Opus 4.6AnthropicScore76%Configuration8-needle at 1M context — a far harder configuration than the 128k rows, and not comparable to them.SourceLab-reported2026-02 · Anthropic reporting for Claude Opus 4.6
ModelGPT-5.1OpenAIScore61.6%ConfigurationMRCR v2, 8-needle at 128k, average; measured by Google for the Gemini 3 Pro model card comparison table.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelGemini 2.5 ProGoogleScore58%ConfigurationMRCR v2, 8-needle at 128k, average; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelClaude Sonnet 4.5AnthropicScore47.1%ConfigurationMRCR v2, 8-needle at 128k, average; measured by Google for the Gemini 3 Pro model card comparison table.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelGemini 3 ProGoogleScore26.3%Configuration1M pointwise — the same model that scores 77.0 at 8-needle 128k. Competitor models are marked not supported at this length because they do not offer 1M context.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelClaude Sonnet 4.5AnthropicScore18.5%Configuration8-needle at 1M context, reported by Anthropic alongside Claude Opus 4.6.SourceLab-reported2026-02 · Anthropic reporting for Claude Opus 4.6
ModelGemini 2.5 ProGoogleScore16.4%Configuration1M pointwise; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.