All benchmarks

Agentic · Tool use

tau2-bench

τ²-bench: Evaluating Conversational Agents in a Dual-Control Environment

Like tau-bench, but the simulated customer can also push buttons.

Released
2025
Built by
Sierra AI
Size
Not published
Status
Near ceiling
Reported by
Sierra AI, Google DeepMind, Anthropic, OpenAI
Signal64Read with context
Judge
94
Graded on the final state of the environment, not on prose.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
50
Instance count is not published, so per-item weight is unknown.
Fine print
65
3 documented caveats, the heaviest being construct validity.
Adoption
63
Reported by 4 labs; 4 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Dual-control shared world

Formally a dual-control Dec-POMDP: the agent and the user simulator both hold tools that mutate one shared simulated world — a telecom account plus a simulated phone. No real device, no internet. The user simulator carries a persona with a stated technical competence level and will misunderstand or misreport instructions accordingly, so a correct instruction that is badly explained still fails. A domain policy document constrains the agent throughout.

Another model plays the user, so its behaviour is part of the measurement.

Budget
Fixed turn cap per episode.
Tools exposed
Agent: read-only network diagnosticsUser: toggle airplane modeUser: check APN settingsUser: reboot the device

Where the tasks came from

About 114 telecom tasks generated from a device/network state machine, plus re-implemented τ-retail and τ-airline domains.

Compositional generator over a formal telecom domain model, filtered and validated for solvability.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The simulator is one vendor's LLM prompted with a persona, and its choice materially moves scores. Cross-lab τ²-bench numbers are therefore only loosely comparable, and model cards rarely state which simulator was used or at what temperature.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGemini 3 ProGoogleScore85.4%ConfigurationRow labelled τ2-bench (agentic tool use); results as of November 2025. Averaged over domains at pass^1.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelClaude Sonnet 4.5AnthropicScore84.7%ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelGPT-5.1OpenAIScore80.2%ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card
ModelGemini 2.5 ProGoogleScore54.9%ConfigurationRow labelled τ2-bench (agentic tool use); results as of November 2025.SourceLab-reported2025-11 · Gemini 3 Pro model card

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Near ceiling The top models are close enough to the ceiling to crowd.