Knowledge
SimpleQA Verified
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
Cleaned, de-duplicated, topic-balanced thousand-question rebuild of SimpleQA.
- Released
- 2025
- Built by
- Google DeepMind / Google Research
- Size
- 1,000 tasks
- Status
- Active
- Reported by
- Google DeepMind, Artificial Analysis
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 92
- 1,000 instances, so one item moves the score by 0.100 points.
- Fine print
- 45
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 29
- Reported by 2 labs; 1 published score collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, no search
No retrieval, no browsing, no tools. This is not incidental: a SimpleQA-with-browsing number measures retrieval quality, not parametric knowledge, and the two are not comparable. Abstention is an allowed and separately scored response, so a model's refusal policy is part of what is being measured.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
1,000 prompts, giving about plus or minus 1.6pp standard error at 50%. Better than most benchmarks here, still not enough to separate close models.
Filtered from OpenAI's 4,326-question SimpleQA by de-duplication, topic balancing and source reconciliation.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Original SimpleQA is 4,326 OpenAI-authored items; SimpleQA Verified is 1,000 Google-filtered items. The two F1 numbers are not on the same scale and are frequently conflated because of the names. Original SimpleQA also publishes three different headline numbers (accuracy, accuracy-given-attempted, F-score), which compounds the confusion.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGemini 2.5 ProGoogle DeepMind | Score55.6% | ConfigurationF1, 1,000 prompts, closed-book, gpt-4.1-2025-04-14 autorater; paper's own snapshot | SourceLab-reported2025-09 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.