All benchmarks

Knowledge · Reasoning

GPQA

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

PhD-level science questions that skilled web searchers still fail.

Released
2023
Built by
NYU, Cohere, Anthropic
Size
198 tasks
Status
Saturated
Reported by
OpenAI, Anthropic, Google DeepMind, xAI, Meta, DeepSeek, Alibaba (Qwen), Moonshot AI, Mistral, Z.ai, Artificial Analysis, Epoch AI
Signal53Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
52
198 instances, so one item moves the score by 0.505 points.
Fine print
45
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
92
Reported by 12 labs; 5 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Closed book, no search

0-shot chain of thought, no tools and no search, since search access defeats the entire construction. Artificial Analysis runs the 198 Diamond items at 5 repeats with regex extraction and pass@1 scoring. Model cards vary: some report a single pass, some report maj@8 or maj@32 without saying so, and some agentic evaluations run it with search enabled and still call it GPQA.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn

Where the tasks came from

GPQA Diamond is 198 questions, so one item is 0.505pp and the binomial standard error is about 3.5pp at mid-range. Grok 4.6 at 94.95% and Claude Opus 5 at 93.74% differ by roughly one to two questions.

Written by 61 domain experts, validated by same-field experts and by searching non-experts. Main set 448, extended 546, Diamond 198. Gated on HuggingFace and password-protected in the repo to keep it out of crawls.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • At 50% accuracy the binomial standard error on 198 items is about 3.55pp; at 94% it is about 1.7pp. A 95% confidence interval on the difference between two models evaluated on the same 198 items is roughly 5 to 7 percentage points. The entire 2026 top tier spans under 1.5 points, so the ordering within it is not measurable.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGrok 4.6 (high)xAIScore94.95%ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extractionSourceThird-party2026-08-16
ModelGemini 3.7 Flash (high)Google DeepMindScore94.55%ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extractionSourceThird-party2026-08-16
ModelGPT-5.6 Sol (max)OpenAIScore94.14%ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extractionSourceThird-party2026-08-16
ModelGemini 3.1 Pro PreviewGoogle DeepMindScore94.14%ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extractionSourceThird-party2026-08-16
ModelClaude Opus 5 (Adaptive Reasoning, High/Xhigh Effort)AnthropicScore93.74%ConfigurationGPQA Diamond, 198 questions x 5 repeats, regex extractionSourceThird-party2026-08-16

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.