Reasoning
CritPt
CritPt: Complex Research using Integrated Thinking - Physics Test
Unpublished research-level physics problems written by working physicists.
- Released
- 2025
- Built by
- CritPt team with 50+ physics researchers; graded with Artificial Analysis
- Size
- 70 tasks
- Status
- Frontier
- Reported by
- Artificial Analysis, Epoch AI, CritPt team
- Judge
- 58
- Mixed grading, so part of the score inherits the weakest judge.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 36
- 70 instances, so one item moves the score by 1.43 points.
- Fine print
- 50
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 53
- Reported by 3 labs; 5 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Solve, then reformat
Artificial Analysis uses a two-step protocol: the first call asks the model to solve the challenge with full reasoning, the second formats the response into the expected code format for grading. Both calls' tokens and cost count toward the reported figures. 5 repeats per question, pass@1, on the 70 test-set challenges (the single worked example is excluded).
Single turn, no tools. The most reproducible setup there is.
- Budget
- Two calls per challenge
Where the tasks came from
70 scored challenges, so one challenge is 1.43pp. The 190 checkpoint tasks are a separate, higher-scoring denominator.
71 challenges (70 scored) newly written by 50+ active physics researchers from their own unpublished research; answer key held on a private grading server.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The answer key never leaves the official grading server, and API access is granted case by case. That is a genuine defence against contamination, and it also means published CritPt numbers cannot be independently verified by anyone the maintainers have not approved. The two-step solve-then-format protocol adds a second harness dependency, since a formatting failure scores as a wrong answer.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5.6 Sol (max)OpenAI | Score32.29% | Configuration70 challenges x 5 repeats, official grading server | SourceThird-party2026-08-16 |
| ModelGPT-5.5 Pro (xhigh)OpenAI | Score30.57% | Configuration70 challenges x 5 repeats, official grading server | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Terra (max)OpenAI | Score30% | Configuration70 challenges x 5 repeats, official grading server | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | Score29.14% | Configuration70 challenges x 5 repeats, official grading server | SourceThird-party2026-08-16 |
| ModelGPT-5 (high)OpenAI | Score5.7% | ConfigurationBest base model at publication, September 2025 | SourceThird-party2025-09 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.