Agentic · Knowledge
BrowseComp
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Short questions whose answers are buried deep on the open web.
- Released
- 2025
- Built by
- OpenAI
- Size
- 1,266 tasks
- Status
- Active
- Reported by
- OpenAI, Anthropic, Google DeepMind, xAI, Perplexity
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 92
- 1,266 instances, so one item moves the score by 0.079 points.
- Fine print
- 54
- 3 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 65
- Reported by 5 labs; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
The open internet
Whatever search and browse tooling the evaluated system happens to have. There is no fixed corpus, no frozen index, and no step cap defined by the benchmark itself — the harness decides how many searches and page loads are allowed. That is the realism and also the reproducibility problem: two labs running BrowseComp are not running the same environment, and the search backend they sit on is part of what is being measured.
The web changes underneath the eval, so a score is only valid for the day it was run.
- Budget
- Undefined by the benchmark; the evaluating harness sets its own search and page-load budget.
- Tools exposed
- Web search supplied by the evaluated systemPage browsing supplied by the evaluated system
Where the tasks came from
1,266 questions, so a single item is worth about 0.08%.
Authored by OpenAI trainers through an inversion protocol, filtered against GPT-4o, o1, an early deep-research model, and a ten-minute human search.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Two BrowseComp runs differ in search backend, page-load budget, and the state of the web on the day. Google made the same point bluntly about the sibling browsing benchmark WebVoyager in its Gemini 2.5 Computer Use documentation: because different models are run at different points in time and exclude different tasks, self-reported results are difficult to compare. Treat any BrowseComp figure without a stated browsing stack and run date as an anecdote.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Active — Still spreads the field.