Agentic · Computer use
Mind2Web 2
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
Long agentic-search tasks graded by a tree of judge agents.
- Released
- 2025
- Built by
- OSU NLP Group
- Size
- 130 tasks
- Status
- Frontier
- Reported by
- OSU NLP Group
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 52
- 130 instances, so one item moves the score by 0.769 points.
- Fine print
- 59
- 3 documented caveats, the heaviest being construct validity.
- Adoption
- 13
- Reported by 1 lab; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Live web, both sides
The live internet, reached through whatever browsing and search stack the evaluated system brings. No step cap is imposed by the benchmark; systems are compared on the outcome and on the evidence they cited. Unusually, the grading side lives on the live web too — the judge agents open the cited pages themselves, so both the answer and the check are subject to the same drift.
The web changes underneath the eval, so a score is only valid for the day it was run.
- Budget
- No cap imposed by the benchmark.
- Tools exposed
- Search and browsing supplied by the evaluated systemBrowsing tools available to the judge agents
Where the tasks came from
130 tasks, so a single task is worth roughly 0.8% of the strict success rate.
Hand-authored by the OSU NLP Group; the authors report over 1,000 hours of human labour to build and validate the set.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Judge agents browse, which means grading costs real time and money and inherits the same flakiness as the task. When a judge marks a citation unsupported, checking whether the judge was right requires repeating its browsing session — so grader errors are cheap to make and expensive to find.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Frontier — Nothing is close to solving it.