Reasoning · Knowledge
HLE
Humanity's Last Exam
Expert-written questions at the frontier of human academic knowledge.
- Released
- 2025
- Built by
- Center for AI Safety and Scale AI
- Size
- 2,500 tasks
- Status
- Active
- Reported by
- OpenAI, Anthropic, Google DeepMind, xAI, Moonshot AI, DeepSeek, Z.ai, Artificial Analysis, Epoch AI
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 100
- 2,500 instances, so one item moves the score by 0.040 points.
- Fine print
- 45
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 92
- Reported by 9 labs; 5 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, no tools
The canonical protocol is closed-book with no tools and no search, and reasoning models run at high or max effort. This is the single biggest source of incomparable HLE numbers: several labs publish HLE with search, or HLE with a code interpreter, and those scores are dramatically higher. Text-only (2,158 items) versus full with images (2,500) is the second axis. Press coverage conflates all of it.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
The question count is itself a trap: 3,000 announced in January 2025, 2,700 in the paper, 2,500 in the May 2025 revision, and 2,158 in the text-only subset Artificial Analysis actually runs.
Roughly 1,000 contributing experts; accepted only if six frontier 2024 models all failed the item. CAIS retains a private held-out set to detect overfitting.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The canonical run is closed-book with no search. Several labs report search-enabled or code-interpreter-enabled HLE scores that are substantially higher, and headline comparisons between labs often differ on this axis alone. Add the text-only 2,158 subset versus the full 2,500 with images, and the judge model differing between harnesses, and three separate configuration choices sit behind one number.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | Score54.87% | Configuration2,158 text-only questions, no tools, pass@1, LLM equality checker | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (Xhigh)Anthropic | Score54.4% | Configuration2,158 text-only questions, no tools, pass@1, LLM equality checker | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (High)Anthropic | Score52.83% | Configuration2,158 text-only questions, no tools, pass@1, LLM equality checker | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (Medium)Anthropic | Score51.3% | Configuration2,158 text-only questions, no tools, pass@1, LLM equality checker | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (max)OpenAI | Score49.49% | Configuration2,158 text-only questions, no tools, pass@1, LLM equality checker | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Active — Still spreads the field.