Agentic · Tool use
tau2-bench
τ²-bench: Evaluating Conversational Agents in a Dual-Control Environment
Like tau-bench, but the simulated customer can also push buttons.
- Released
- 2025
- Built by
- Sierra AI
- Size
- Not published
- Status
- Near ceiling
- Reported by
- Sierra AI, Google DeepMind, Anthropic, OpenAI
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 40
- The top models crowd the ceiling; gaps are shrinking.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 65
- 3 documented caveats, the heaviest being construct validity.
- Adoption
- 63
- Reported by 4 labs; 4 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Dual-control shared world
Formally a dual-control Dec-POMDP: the agent and the user simulator both hold tools that mutate one shared simulated world — a telecom account plus a simulated phone. No real device, no internet. The user simulator carries a persona with a stated technical competence level and will misunderstand or misreport instructions accordingly, so a correct instruction that is badly explained still fails. A domain policy document constrains the agent throughout.
Another model plays the user, so its behaviour is part of the measurement.
- Budget
- Fixed turn cap per episode.
- Tools exposed
- Agent: read-only network diagnosticsUser: toggle airplane modeUser: check APN settingsUser: reboot the device
Where the tasks came from
About 114 telecom tasks generated from a device/network state machine, plus re-implemented τ-retail and τ-airline domains.
Compositional generator over a formal telecom domain model, filtered and validated for solvability.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The simulator is one vendor's LLM prompted with a persona, and its choice materially moves scores. Cross-lab τ²-bench numbers are therefore only loosely comparable, and model cards rarely state which simulator was used or at what temperature.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGemini 3 ProGoogle | Score85.4% | ConfigurationRow labelled τ2-bench (agentic tool use); results as of November 2025. Averaged over domains at pass^1. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
| ModelClaude Sonnet 4.5Anthropic | Score84.7% | ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
| ModelGPT-5.1OpenAI | Score80.2% | ConfigurationMeasured by Google for the Gemini 3 Pro model card comparison table; results as of November 2025. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
| ModelGemini 2.5 ProGoogle | Score54.9% | ConfigurationRow labelled τ2-bench (agentic tool use); results as of November 2025. | SourceLab-reported2025-11 · Gemini 3 Pro model card |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Near ceiling — The top models are close enough to the ceiling to crowd.