All benchmarks

Agentic · Computer use

Mind2Web 2

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

Long agentic-search tasks graded by a tree of judge agents.

Released
2025
Built by
OSU NLP Group
Size
130 tasks
Status
Frontier
Reported by
OSU NLP Group
Signal58Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
52
130 instances, so one item moves the score by 0.769 points.
Fine print
59
3 documented caveats, the heaviest being construct validity.
Adoption
13
Reported by 1 lab; 0 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Live web, both sides

The live internet, reached through whatever browsing and search stack the evaluated system brings. No step cap is imposed by the benchmark; systems are compared on the outcome and on the evidence they cited. Unusually, the grading side lives on the live web too — the judge agents open the cited pages themselves, so both the answer and the check are subject to the same drift.

The web changes underneath the eval, so a score is only valid for the day it was run.

Budget
No cap imposed by the benchmark.
Tools exposed
Search and browsing supplied by the evaluated systemBrowsing tools available to the judge agents

Where the tasks came from

130 tasks, so a single task is worth roughly 0.8% of the strict success rate.

Hand-authored by the OSU NLP Group; the authors report over 1,000 hours of human labour to build and validate the set.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Judge agents browse, which means grading costs real time and money and inherits the same flakiness as the task. When a judge marks a citation unsupported, checking whether the judge was right requires repeating its browsing session — so grader errors are cheap to make and expensive to find.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

No published scores collected for this benchmark yet.

Frontier Nothing is close to solving it.