All benchmarks

Coding · Agentic · Tool use

Terminal-Bench 1.0

Terminal-Bench (terminal-bench-core v0.1.1)

Can an agent finish real command-line jobs inside a Linux container?

Released
2025
Built by
Stanford University and the Laude Institute
Size
80 tasks
Status
Saturated
Reported by
Anthropic
Signal48Needs the fine print
Judge
94
Graded on the final state of the environment, not on prose.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
36
80 instances, so one item moves the score by 1.25 points.
Fine print
67
3 documented caveats, the heaviest being harness sensitivity.
Adoption
19
Reported by 1 lab; 2 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Keystrokes into a tmux pane

Each trial builds a fresh Docker container and attaches a tmux session to it. Commands are delivered as literal keystrokes typed into the pane, and the whole session is recorded as an asciinema cast. Terminus, the reference agent, runs outside the container and its only affordance is that pane — so it has to choose between an echo heredoc and launching vim to write a file, and it can never ask a question. Internet access is allowed, which agents use to install packages freely and which the team later flagged as a reproducibility hazard. The default agent timeout is 360 seconds.

The model gets a shell. Scaffold quality and step budget move the score independently of the model.

Budget
360-second default agent timeout rather than a step cap
Tools exposed
tmux terminal

Where the tasks came from

80 hand-crafted tasks, pinned as terminal-bench-core==0.1.1 for the leaderboard.

Written by the core team and open-source contributors via pull requests, each requiring an oracle solution that provably passes the tests.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The top entries are bespoke commercial harnesses — Apex2, Chaterm, Abacus AI Desktop, Ante, Droid — wrapped around a frontier model. Rows are therefore not model-comparable: what is being ranked is a product. Wall-clock timeout sensitivity compounds it, since a correct agent that works slowly scores as a failure.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelApex2 + claude-4-5-sonnetApexScore64.5%Configuration±1.1. A commercial agent harness wrapped around a frontier model, not a bare model score.SourceCommunity2025-10-15
ModelChaterm + claude-4-5-sonnetChatermScore63.7%ConfigurationCommercial agent harness.SourceCommunity2025-10-31

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.