All benchmarks

Knowledge

SimpleQA Verified

SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge

Cleaned, de-duplicated, topic-balanced thousand-question rebuild of SimpleQA.

Released
2025
Built by
Google DeepMind / Google Research
Size
1,000 tasks
Status
Active
Reported by
Google DeepMind, Artificial Analysis
Signal61Read with context
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
82
Still separates the frontier from everything below it.
Resolution
92
1,000 instances, so one item moves the score by 0.100 points.
Fine print
45
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
29
Reported by 2 labs; 1 published score collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Closed book, no search

No retrieval, no browsing, no tools. This is not incidental: a SimpleQA-with-browsing number measures retrieval quality, not parametric knowledge, and the two are not comparable. Abstention is an allowed and separately scored response, so a model's refusal policy is part of what is being measured.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn

Where the tasks came from

1,000 prompts, giving about plus or minus 1.6pp standard error at 50%. Better than most benchmarks here, still not enough to separate close models.

Filtered from OpenAI's 4,326-question SimpleQA by de-duplication, topic balancing and source reconciliation.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Original SimpleQA is 4,326 OpenAI-authored items; SimpleQA Verified is 1,000 Google-filtered items. The two F1 numbers are not on the same scale and are frequently conflated because of the names. Original SimpleQA also publishes three different headline numbers (accuracy, accuracy-given-attempted, F-score), which compounds the confusion.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGemini 2.5 ProGoogle DeepMindScore55.6%ConfigurationF1, 1,000 prompts, closed-book, gpt-4.1-2025-04-14 autorater; paper's own snapshotSourceLab-reported2025-09

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.