Knowledge · Reasoning
GPQA
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
PhD-level science questions that skilled web searchers still fail.
- Released
- 2023
- Built by
- NYU, Cohere, Anthropic
- Size
- 198 tasks
- Status
- Saturated
- Reported by
- OpenAI, Anthropic, Google DeepMind, xAI, Meta, DeepSeek, Alibaba (Qwen), Moonshot AI, Mistral, Z.ai, Artificial Analysis, Epoch AI
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 52
- 198 instances, so one item moves the score by 0.505 points.
- Fine print
- 45
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 92
- Reported by 12 labs; 5 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, no search
0-shot chain of thought, no tools and no search, since search access defeats the entire construction. Artificial Analysis runs the 198 Diamond items at 5 repeats with regex extraction and pass@1 scoring. Model cards vary: some report a single pass, some report maj@8 or maj@32 without saying so, and some agentic evaluations run it with search enabled and still call it GPQA.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
GPQA Diamond is 198 questions, so one item is 0.505pp and the binomial standard error is about 3.5pp at mid-range. Grok 4.6 at 94.95% and Claude Opus 5 at 93.74% differ by roughly one to two questions.
Written by 61 domain experts, validated by same-field experts and by searching non-experts. Main set 448, extended 546, Diamond 198. Gated on HuggingFace and password-protected in the repo to keep it out of crawls.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
At 50% accuracy the binomial standard error on 198 items is about 3.55pp; at 94% it is about 1.7pp. A 95% confidence interval on the difference between two models evaluated on the same 198 items is roughly 5 to 7 percentage points. The entire 2026 top tier spans under 1.5 points, so the ordering within it is not measurable.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGrok 4.6 (high)xAI | Score94.95% | ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extraction | SourceThird-party2026-08-16 |
| ModelGemini 3.7 Flash (high)Google DeepMind | Score94.55% | ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extraction | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (max)OpenAI | Score94.14% | ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extraction | SourceThird-party2026-08-16 |
| ModelGemini 3.1 Pro PreviewGoogle DeepMind | Score94.14% | ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extraction | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (Adaptive Reasoning, High/Xhigh Effort)Anthropic | Score93.74% | ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extraction | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.