Agentic
Vending-Bench 2
Vending-Bench 2
A full simulated business year, with adversarial suppliers.
- Released
- 2026
- Built by
- Andon Labs
- Size
- Not published
- Status
- Frontier
- Reported by
- Andon Labs, Epoch AI, Google DeepMind, Anthropic
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 62
- 3 documented caveats, the heaviest being construct validity.
- Adoption
- 74
- Reported by 4 labs; 10 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Adversarial business world
A text-based business world, materially harsher than the first version. Suppliers differ in price and reliability, and some are explicitly adversarial — designed to try to exploit an AI operator. Deliveries fail stochastically, suppliers go bankrupt, and customers ask for refunds. Negotiation is required rather than optional: the top performers get better prices by persistently negotiating or by finding better suppliers, not by playing the demand curve.
Another model plays the user, so its behaviour is part of the measurement.
- Budget
- 365 simulated days per run, far longer than any model's context window, so the agent must manage its own notes.
- Tools exposed
- Supplier messaging and negotiationOrdering and inventory managementPrice settingCash and daily fee handling
Where the tasks came from
One parameterised year-long simulation, run five times per model and averaged.
Closed proprietary simulator built by Andon Labs; no public code or task list.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
A single unlucky supplier bankruptcy or a single late-year behavioural collapse can dominate a model's average. The characteristic failure is degradation: agents that start the year well stop acting consistently as the run lengthens, which produces heavy-tailed outcomes that five runs cannot pin down.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Opus 4.6Anthropic | Score$8017.59 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelClaude Sonnet 4.6Anthropic | Score$7204.14 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelGemini 3 ProGoogle | Score$5478.16 | ConfigurationMean end-of-year balance across five year-long runs; independently corroborated by Google's Gemini 3 Pro model card. | SourceThird-party2026-02 |
| ModelClaude Opus 4.5Anthropic | Score$4967.06 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelGLM-5Zhipu AI | Score$4432.12 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelClaude Sonnet 4.5Anthropic | Score$3838.74 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelGemini 3.1 ProGoogle | Score$3774.25 | ConfigurationRun with custom tools; mean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelGemini 3 FlashGoogle | Score$3634.72 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelGPT-5.2OpenAI | Score$3591.33 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
| ModelGLM-4.7Zhipu AI | Score$2376.82 | ConfigurationMean end-of-year balance across five year-long runs. | SourceThird-party2026-02 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.