All benchmarks

Tool use · Agentic

MCP-Atlas

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

A thousand tool-use tasks against 36 real MCP servers, not mocks.

Released
2026
Built by
Scale AI
Size
1,000 tasks
Status
Active
Reported by
Anthropic, Scale AI
Signal63Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
82
Still separates the frontier from everything below it.
Resolution
92
1,000 instances, so one item moves the score by 0.100 points.
Fine print
45
5 documented caveats, the heaviest being possible training-set contamination.
Adoption
48
Reported by 2 labs; 12 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Containerised real servers, deliberately noisy toolbox

Servers run in containers with sandboxed filesystems and allow-listed network egress, version-pinned and restarted per run for a clean initial state. The harness mounts only the servers a task declares and exposes 10-25 tools to the model — 3-7 that the reference solution needs plus 5-10 distractors drawn from the same servers, so success depends on discovery and parameterisation rather than recognising a server name. All invocations must satisfy the server-declared JSON schema. Efficiency is logged but not enforced, which the paper flags as leaving room to brute-force discovery.

The model gets a shell. Scaffold quality and step budget move the score independently of the model.

Budget
20 turns and 10-25 exposed tools on the original leaderboard; Scale moved to a 100-tool-call budget in April 2026. Scale separately ran an extended 256-turn / 100-tool configuration for Opus 4.7.
Tools exposed
brave_searchddg_searchexagoogle-mapsairtablemongodbnotionslackarxivpubmedgithubtwelvedata

Where the tasks came from

1,000 tasks, of which 500 are released publicly and 500 are held out. Bucket mix is roughly Basic 30-35%, Productivity 20-25%, Coding 22%, Analytics 10-15%, Financial 10-15%.

Written by Scale AI authors against 36 real production MCP servers, then put through expert review, a prompt-naturalness review, automated claims verification and a random QC sample.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The Opus 4.6 system card reports 59.5% at max effort and 62.7% at high. The Opus 4.7 card reports Opus 4.6 at 75.8%. The Opus 4.8 card reports it at 76.8%. Nothing about the model changed: in April 2026 Scale refreshed the harness with an upgraded judge and retry handling for transient tool errors, moved from a 20-turn limit to a 100-tool-call budget, and re-scored the leaderboard. The Opus 4.7 card states it plainly — 'Prior-harness Opus results are not comparable' — and the Opus 4.7 launch post carries a one-line footnote saying the Opus 4.6 score was updated for revised Scale AI grading. Any MCP-Atlas figure quoted without saying which harness produced it is unusable.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

Refreshed harness (Scale, 2026)Higher is better
Refreshed harness (Scale, 2026)higher is better
ModelMuse Spark 1.1MetaScore88.1%ConfigurationTop of the Scale leaderboard, +/-1.95, refreshed harness. Read from the live leaderboard on 2026-08-17; the page footer still reads 'Updated April 8, 2026' while listing models released in June and July.SourceThird-party2026-08-17
ModelClaude Opus 5 (xhigh)AnthropicScore85.8%ConfigurationMean claim coverage 89.1%. The card warns that 'effort settings may have changed slightly for our production deployment and therefore some scores may not be precisely reproducible'. Scale's leaderboard lists the same figure against claude-opus-5 (xhigh), +/-2.10.SourceLab-reported2026-07-24 · Claude Opus 5 system card, section 8.13.2
Modelgemini-3.5-flash (high)Google DeepMindScore83.6%ConfigurationScale leaderboard, +/-2.30, refreshed harness. A Flash-tier model above every Claude except Opus 5, which is the sort of ordering that survives only because the intervals overlap — all of these rows share rank 2.SourceThird-party2026-08-17
ModelClaude Opus 4.8AnthropicScore82.2%ConfigurationRefreshed harness, mean claim coverage 86.2%, meaning most remaining failures were partial rather than complete. The card restates Opus 4.7 at 79.1% and Opus 4.6 at 76.8%, neither of which matches those models' own cards.SourceLab-reported2026-05-28 · Claude Opus 4.8 system card, section 8.13.4
Modelgpt-5.6 (sol)OpenAIScore81.8%ConfigurationScale leaderboard, +/-2.40, refreshed harness. OpenAI does not report MCP-Atlas in its own model documentation; this is Scale's measurement.SourceThird-party2026-08-17
ModelClaude Opus 4.7 (max effort)AnthropicScore77.3%ConfigurationRun by Scale AI, adaptive thinking at max effort, refreshed harness. Anthropic adds that in Scale's extended 256-turn / 100-tool configuration the same model reached 79.5% at max and 79.7% at high.SourceLab-reported2026-04-16 · Claude Opus 4.7 system card, section 8.10.3
ModelClaude Opus 4.6 (max effort)AnthropicScore76.8%ConfigurationRefreshed harness — upgraded judge, retry handling, 100-tool-call budget. Same model whose own system card reported 59.5%. 79.0% on the public 500.SourceThird-party2026-08-17
Original harness (paper and Opus 4.6 card)Higher is better
Original harness (paper and Opus 4.6 card)higher is better
ModelClaude Opus 4.6 (high effort)AnthropicScore62.7%ConfigurationOriginal harness. Not in Table 2.3.A — it appears only in footnote 3, which says Anthropic reports the max effort score 'to avoid cherry-picking'. Higher than the same model at max effort.SourceLab-reported2026-02
ModelClaude Opus 4.5AnthropicScore62.3%ConfigurationPaper Table 3, original harness, all 1,000 tasks. Mean coverage 78.5%. Top of the leaderboard at publication; the paper's headline finding is that the best model passes under two thirds.SourceThird-party2026-01-31
ModelClaude Opus 4.6 (max effort)AnthropicScore59.5%ConfigurationOriginal harness, Table 2.3.A. Adaptive thinking, max effort, default sampling, averaged over 5 trials. Anthropic notes this is 'slightly worse than Claude Opus 4.5's 62.3%'.SourceLab-reported2026-02
ModelGemini 3 ProGoogle DeepMindScore54.1%ConfigurationPaper Table 3, original harness. Mean coverage 73.2%. Same lightweight MCP client wrapper as every other model in the table.SourceThird-party2026-01-31
ModelGPT-5OpenAIScore44.5%ConfigurationPaper Table 3, original harness. Mean coverage 61.7%. Default configuration, no behaviour-shaping system prompt.SourceThird-party2026-01-31

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.