All benchmarks

Agentic · Knowledge

BrowseComp

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Short questions whose answers are buried deep on the open web.

Released
2025
Built by
OpenAI
Size
1,266 tasks
Status
Active
Reported by
OpenAI, Anthropic, Google DeepMind, xAI, Perplexity
Signal77Reads cleanly
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
82
Still separates the frontier from everything below it.
Resolution
92
1,266 instances, so one item moves the score by 0.079 points.
Fine print
54
3 documented caveats, the heaviest being possible training-set contamination.
Adoption
65
Reported by 5 labs; 0 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

The open internet

Whatever search and browse tooling the evaluated system happens to have. There is no fixed corpus, no frozen index, and no step cap defined by the benchmark itself — the harness decides how many searches and page loads are allowed. That is the realism and also the reproducibility problem: two labs running BrowseComp are not running the same environment, and the search backend they sit on is part of what is being measured.

The web changes underneath the eval, so a score is only valid for the day it was run.

Budget
Undefined by the benchmark; the evaluating harness sets its own search and page-load budget.
Tools exposed
Web search supplied by the evaluated systemPage browsing supplied by the evaluated system

Where the tasks came from

1,266 questions, so a single item is worth about 0.08%.

Authored by OpenAI trainers through an inversion protocol, filtered against GPT-4o, o1, an early deep-research model, and a ten-minute human search.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Two BrowseComp runs differ in search backend, page-load budget, and the state of the web on the day. Google made the same point bluntly about the sibling browsing benchmark WebVoyager in its Gemini 2.5 Computer Use documentation: because different models are run at different points in time and exclude different tasks, self-reported results are difficult to compare. Treat any BrowseComp figure without a stated browsing stack and run date as an anecdote.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

No published scores collected for this benchmark yet.

Active Still spreads the field.