Agentic
METR time horizon
Measuring AI Ability to Complete Long Tasks (the 50% task-completion time horizon)
How long a job can an AI finish? Measured in human hours.
- Released
- 2025
- Built by
- METR
- Size
- Not published
- Status
- Frontier
- Reported by
- METR
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 47
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 16
- Reported by 1 lab; 1 published score collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Containerised task environments
METR's own agent harness over containerised environments with a shell and, where the task needs it, GPUs. Notably METR has run the same measurement under several scaffolds — ReAct, Triframe, Claude Code and Codex — precisely because scaffold choice is a known confound. Their February 2026 comparison found Opus 4.5 with Claude Code beat ReAct in 50.7% of bootstrap samples and GPT-5 with Codex beat Triframe in 14.5%, neither statistically significant.
The model gets a shell. Scaffold quality and step budget move the score independently of the model.
- Budget
- Bounded by each task's time budget rather than by a step count.
- Tools exposed
- ShellGPU compute where the task requires it
Where the tasks came from
RE-Bench's 7 environments plus HCAST plus 66 short SWAA tasks; expanded again in Time Horizon 1.1 (January 2026).
METR-authored and pooled task suites, with human completion times measured from contracted professionals.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
METR renders its per-model chart as a graphic, so specific figures quoted widely in secondary coverage — an hours-and-minutes horizon for a named frontier model, or a claimed current doubling time — cannot be read out of METR's own pages as text. This page therefore quotes only the figures METR states in prose. Treat any specific per-model horizon as unverified unless you have read it off METR's chart yourself.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude 3.7 SonnetAnthropic | Score50 min | ConfigurationApproximately 50 minutes at the 50% time horizon, as reported in the original paper. Superseded by later analyses, including the March 2026 regularisation correction. | SourceThird-party2025-03 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.