Agentic · Coding
PaperBench
PaperBench: Evaluating AI's Ability to Replicate AI Research
Reproduce a whole ICML paper from scratch, without the authors' code.
- Released
- 2025
- Built by
- OpenAI (Preparedness)
- Size
- 20 tasks
- Status
- Frontier
- Reported by
- OpenAI
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 10
- 20 instances, so one item moves the score by 5.00 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 16
- Reported by 1 lab; 1 published score collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Long-running GPU container
A containerised compute environment with GPUs and a shell, running long enough to actually train the models a paper describes. The original repository is withheld throughout. Runs are among the most expensive of any published eval, which is itself a barrier — it limits how often anyone can re-run PaperBench, including to check someone else's result.
The model gets a shell. Scaffold quality and step budget move the score independently of the model.
- Budget
- Long-horizon: hours of GPU time per paper.
- Tools exposed
- ShellGPU computeThe paper PDF
Where the tasks came from
20 papers, graded through 8,316 rubric leaves — the resolution comes from the rubric, not the paper count.
ICML 2024 Spotlight and Oral papers, with hierarchical rubrics co-developed with the original authors.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
One paper is worth five percentage points of the average, and papers differ enormously in how much compute a faithful replication needs. Aggregate PaperBench scores move on which papers a run happened to make progress on.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude 3.5 Sonnet (New)Anthropic | Score21% | ConfigurationAverage replication score; best model at release, measured by the PaperBench authors. | SourceThird-party2025-04 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.