All benchmarks

Agentic · Coding

PaperBench

PaperBench: Evaluating AI's Ability to Replicate AI Research

Reproduce a whole ICML paper from scratch, without the authors' code.

Released
2025
Built by
OpenAI (Preparedness)
Size
20 tasks
Status
Frontier
Reported by
OpenAI
Signal48Needs the fine print
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
10
20 instances, so one item moves the score by 5.00 points.
Fine print
47
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
16
Reported by 1 lab; 1 published score collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Long-running GPU container

A containerised compute environment with GPUs and a shell, running long enough to actually train the models a paper describes. The original repository is withheld throughout. Runs are among the most expensive of any published eval, which is itself a barrier — it limits how often anyone can re-run PaperBench, including to check someone else's result.

The model gets a shell. Scaffold quality and step budget move the score independently of the model.

Budget
Long-horizon: hours of GPU time per paper.
Tools exposed
ShellGPU computeThe paper PDF

Where the tasks came from

20 papers, graded through 8,316 rubric leaves — the resolution comes from the rubric, not the paper count.

ICML 2024 Spotlight and Oral papers, with hierarchical rubrics co-developed with the original authors.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • One paper is worth five percentage points of the average, and papers differ enormously in how much compute a faithful replication needs. Aggregate PaperBench scores move on which papers a run happened to make progress on.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelClaude 3.5 Sonnet (New)AnthropicScore21%ConfigurationAverage replication score; best model at release, measured by the PaperBench authors.SourceThird-party2025-04

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Frontier Nothing is close to solving it.