Coding · Agentic · Tool use
Terminal-Bench 3.0
Terminal-Bench 3.0, formerly branded Frontier-Bench
Expert-level computer work across seven fields; the best model manages under half.
- Released
- 2026
- Built by
- Laude Institute and Stanford University
- Size
- 74 tasks
- Status
- Frontier
- Reported by
- Anthropic
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 36
- 74 instances, so one item moves the score by 1.35 points.
- Fine print
- 50
- 4 documented caveats, the heaviest being harness sensitivity.
- Adoption
- 13
- Reported by 1 lab; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Harbor-native, agent still pluggable
The same sandbox model as Terminal-Bench 2.x: Docker containers on Modal, Daytona, e2b or locally, with pluggable agent adapters. The leaderboard's columns are Model, Agent, Resolution Rate, Cost and Tokens — which confirms the score is again a model-plus-adapter pair and promotes budget from a footnote to a ranked axis. Anthropic's own reported run used the mini-SWE-agent harness on a GKE backend, so both harness and backend remain free parameters here exactly as they were in 2.x.
The model gets a shell. Scaffold quality and step budget move the score independently of the model.
- Budget
- not published
Where the tasks came from
Reported as 74 tasks across seven domains, but only via a secondary aggregator — the official dataset endpoint and leaderboard route both return 404, and tbench.ai still renders the page as '0 tasks / under construction'.
Contributed expert-level tasks collected March-May 2026, designed so the best models at release solve at most 30%.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Frontier-Bench and Terminal-Bench 3.0 are the same benchmark — frontierbench.ai says Terminal-Bench was formerly called Frontier-Bench, and tbench.ai links Terminal-Bench 3 straight to it. But Anthropic's Claude Opus 5 page cites 'Frontier-Bench v0.1', and nothing primary reconciles that tag with the 3.0 label. The most plausible reading is that 3.0 names the generation while v0.1 is a dataset release tag inside it; that reading is weakly supported by one secondary source and should be treated as an open question rather than a fact.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Frontier — Nothing is close to solving it.