All benchmarks

Agentic · Coding · Tool use

TheAgentCompany

TheAgentCompany — benchmarking LLM agents on consequential real-world tasks

175 workplace tasks inside a fake software company, scored checkpoint by checkpoint.

Released
2024
Built by
Carnegie Mellon University
Size
175 tasks
Status
Active
Signal56Read with context
Judge
58
Mixed grading, so part of the score inherits the weakest judge.
Headroom
82
Still separates the frontier from everything below it.
Resolution
52
175 instances, so one item moves the score by 0.571 points.
Fine print
45
5 documented caveats, the heaviest being possible training-set contamination.
Adoption
22
No lab reports it in a model card, so it rarely appears in a comparison.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

A whole intranet in Docker, plus coworkers who are LLMs

The environment is a self-hosted company: GitLab for code and wikis, ownCloud for files, Plane for issues and sprints, RocketChat for chat, all seeded with mock data and resettable, alongside a Linux workspace the agent has a shell in. The distinguishing feature is the colleagues — every simulated human has a name, a role, responsibilities and project affiliations, and is played by an LLM the agent has to talk to over RocketChat to get unblocked. In the paper's own runs those NPCs were all Claude 3.5 Sonnet; on the leaderboard the NPC model is a per-submission column.

The model gets a shell. Scaffold quality and step budget move the score independently of the model.

Budget
No fixed step cap. The top leaderboard entry averages 29.9 steps and $0.40 per task; the paper's OpenHands + Claude 3.5 Sonnet baseline averaged 29.2 steps and $6.34.
Tools exposed
bashIPythonfile editorweb browserRocketChatGitLabownCloudPlane

Where the tasks came from

175 tasks, so one task is worth 0.57 points of the headline. The task directories contain 527 checkpoint sections between them, about three per task, ranging from one to eight.

Hand-authored by the paper's team, with mock company data, personas and repositories built for the benchmark. Every task brief, checkpoint description and grader is public in the repository.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • S_partial = 0.5 × (points earned / points available) + 0.5 × S_full. Invert it on the top entry — 42.86% full completion and a 52.40 partial score — and the implied share of checkpoint points earned is 61.9%. So the agent that "scores 52" finished under 43% of the work and collected credit on roughly three fifths of the intermediate steps of everything else. Checkpoints are frequently cheap: in admin-ask-for-meeting-feedback the first point is awarded for the existence of any chat history with Huang Jie.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

full completionHigher is better
full completionhigher is better
ModelTTE-MatrixAgent + DeepSeek-V3.2TTEScore42.86%ConfigurationClosed-source scaffold, NPC colleagues run on Qwen Plus rather than the paper's Claude 3.5 Sonnet. 29.91 steps and $0.40 per task on average.SourceCommunity2025-11-10
ModelMUSE + Gemini 2.5 FlashMUSEScore41.14%ConfigurationOpen-source scaffold, NPC colleagues run on GPT-4o. Steps and cost not reported.SourceCommunity2025-10-13
ModelOpenHands-Versa + Claude Sonnet 4All Hands AI / OpenHands-VersaScore33.14%ConfigurationNPCs on Claude 3.5 Sonnet. 46.45 steps and $1.63 per task. The highest entry using a frontier model directly.SourceCommunity2025-06-14
ModelOpenHands + Gemini 2.5 ProAll Hands AI / OpenHands 0.28.1Score30.29%ConfigurationNPCs on Claude 3.5 Sonnet. 27.23 steps and $4.23 per task — ten times the cost per task of the top entry.SourceCommunity2025-05-10
ModelOpenHands + Claude 3.5 SonnetAll Hands AI / OpenHands 0.14.2Score24%ConfigurationThe paper's own strongest baseline at publication, December 2024. 29.17 steps and $6.34 per task.SourceThird-party2024-12-17
ModelOpenHands + GPT-4oAll Hands AI / OpenHands 0.14.2Score8.6%ConfigurationPaper baseline. Partial score 16.7, implying only 24.8% of checkpoint points earned — the partial-score gap is widest for weak agents.SourceThird-party2024-12-17
partial scoreHigher is better
partial scorehigher is better
ModelTTE-MatrixAgent + DeepSeek-V3.2TTEScore52.4%ConfigurationSame run. Implies 61.9% of available checkpoint points earned across the 175 tasks.SourceCommunity2025-11-10
ModelMUSE + Gemini 2.5 FlashMUSEScore51.78%ConfigurationSame run.SourceCommunity2025-10-13
ModelOpenHands-Versa + Claude Sonnet 4All Hands AI / OpenHands-VersaScore43.19%ConfigurationSame run.SourceCommunity2025-06-14
ModelOpenHands + Gemini 2.5 ProAll Hands AI / OpenHands 0.28.1Score39.28%ConfigurationSame run.SourceCommunity2025-05-10
ModelOpenHands + Claude 3.5 SonnetAll Hands AI / OpenHands 0.14.2Score34.4%ConfigurationSame run. The abstract's "30% of tasks" refers to later Gemini 2.5 Pro and Claude 3.7 Sonnet runs, not this one.SourceThird-party2024-12-17

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.