All benchmarks

Reasoning

CritPt

CritPt: Complex Research using Integrated Thinking - Physics Test

Unpublished research-level physics problems written by working physicists.

Released
2025
Built by
CritPt team with 50+ physics researchers; graded with Artificial Analysis
Size
70 tasks
Status
Frontier
Reported by
Artificial Analysis, Epoch AI, CritPt team
Signal62Read with context
Judge
58
Mixed grading, so part of the score inherits the weakest judge.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
36
70 instances, so one item moves the score by 1.43 points.
Fine print
50
4 documented caveats, the heaviest being construct validity.
Adoption
53
Reported by 3 labs; 5 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Solve, then reformat

Artificial Analysis uses a two-step protocol: the first call asks the model to solve the challenge with full reasoning, the second formats the response into the expected code format for grading. Both calls' tokens and cost count toward the reported figures. 5 repeats per question, pass@1, on the 70 test-set challenges (the single worked example is excluded).

Single turn, no tools. The most reproducible setup there is.

Budget
Two calls per challenge

Where the tasks came from

70 scored challenges, so one challenge is 1.43pp. The 190 checkpoint tasks are a separate, higher-scoring denominator.

71 challenges (70 scored) newly written by 50+ active physics researchers from their own unpublished research; answer key held on a private grading server.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The answer key never leaves the official grading server, and API access is granted case by case. That is a genuine defence against contamination, and it also means published CritPt numbers cannot be independently verified by anyone the maintainers have not approved. The two-step solve-then-format protocol adds a second harness dependency, since a formatting failure scores as a wrong answer.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelGPT-5.6 Sol (max)OpenAIScore32.29%Configuration70 challenges x 5 repeats, official grading serverSourceThird-party2026-08-16
ModelGPT-5.5 Pro (xhigh)OpenAIScore30.57%Configuration70 challenges x 5 repeats, official grading serverSourceThird-party2026-08-16
ModelGPT-5.6 Terra (max)OpenAIScore30%Configuration70 challenges x 5 repeats, official grading serverSourceThird-party2026-08-16
ModelClaude Opus 5 (Adaptive Reasoning, Max Effort)AnthropicScore29.14%Configuration70 challenges x 5 repeats, official grading serverSourceThird-party2026-08-16
ModelGPT-5 (high)OpenAIScore5.7%ConfigurationBest base model at publication, September 2025SourceThird-party2025-09

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Frontier Nothing is close to solving it.