All benchmarks

Agentic

Vending-Bench 2

Vending-Bench 2

A full simulated business year, with adversarial suppliers.

Released
2026
Built by
Andon Labs
Size
Not published
Status
Frontier
Reported by
Andon Labs, Epoch AI, Google DeepMind, Anthropic
Signal79Reads cleanly
Judge
94
Graded on the final state of the environment, not on prose.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
50
Instance count is not published, so per-item weight is unknown.
Fine print
62
3 documented caveats, the heaviest being construct validity.
Adoption
74
Reported by 4 labs; 10 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Adversarial business world

A text-based business world, materially harsher than the first version. Suppliers differ in price and reliability, and some are explicitly adversarial — designed to try to exploit an AI operator. Deliveries fail stochastically, suppliers go bankrupt, and customers ask for refunds. Negotiation is required rather than optional: the top performers get better prices by persistently negotiating or by finding better suppliers, not by playing the demand curve.

Another model plays the user, so its behaviour is part of the measurement.

Budget
365 simulated days per run, far longer than any model's context window, so the agent must manage its own notes.
Tools exposed
Supplier messaging and negotiationOrdering and inventory managementPrice settingCash and daily fee handling

Where the tasks came from

One parameterised year-long simulation, run five times per model and averaged.

Closed proprietary simulator built by Andon Labs; no public code or task list.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • A single unlucky supplier bankruptcy or a single late-year behavioural collapse can dominate a model's average. The characteristic failure is degradation: agents that start the year well stop acting consistently as the run lengthens, which produces heavy-tailed outcomes that five runs cannot pin down.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelClaude Opus 4.6AnthropicScore$8017.59ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelClaude Sonnet 4.6AnthropicScore$7204.14ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelGemini 3 ProGoogleScore$5478.16ConfigurationMean end-of-year balance across five year-long runs; independently corroborated by Google's Gemini 3 Pro model card.SourceThird-party2026-02
ModelClaude Opus 4.5AnthropicScore$4967.06ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelGLM-5Zhipu AIScore$4432.12ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelClaude Sonnet 4.5AnthropicScore$3838.74ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelGemini 3.1 ProGoogleScore$3774.25ConfigurationRun with custom tools; mean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelGemini 3 FlashGoogleScore$3634.72ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelGPT-5.2OpenAIScore$3591.33ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02
ModelGLM-4.7Zhipu AIScore$2376.82ConfigurationMean end-of-year balance across five year-long runs.SourceThird-party2026-02

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Frontier Nothing is close to solving it.