All benchmarks

Reasoning · Knowledge

HLE

Humanity's Last Exam

Expert-written questions at the frontier of human academic knowledge.

Released
2025
Built by
Center for AI Safety and Scale AI
Size
2,500 tasks
Status
Active
Reported by
OpenAI, Anthropic, Google DeepMind, xAI, Moonshot AI, DeepSeek, Z.ai, Artificial Analysis, Epoch AI
Signal70Reads cleanly
Judge
44
A model grades the model. Update the grader and old scores move.
Headroom
82
Still separates the frontier from everything below it.
Resolution
100
2,500 instances, so one item moves the score by 0.040 points.
Fine print
45
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
92
Reported by 9 labs; 5 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Closed book, no tools

The canonical protocol is closed-book with no tools and no search, and reasoning models run at high or max effort. This is the single biggest source of incomparable HLE numbers: several labs publish HLE with search, or HLE with a code interpreter, and those scores are dramatically higher. Text-only (2,158 items) versus full with images (2,500) is the second axis. Press coverage conflates all of it.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn

Where the tasks came from

The question count is itself a trap: 3,000 announced in January 2025, 2,700 in the paper, 2,500 in the May 2025 revision, and 2,158 in the text-only subset Artificial Analysis actually runs.

Roughly 1,000 contributing experts; accepted only if six frontier 2024 models all failed the item. CAIS retains a private held-out set to detect overfitting.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The canonical run is closed-book with no search. Several labs report search-enabled or code-interpreter-enabled HLE scores that are substantially higher, and headline comparisons between labs often differ on this axis alone. Add the text-only 2,158 subset versus the full 2,500 with images, and the judge model differing between harnesses, and three separate configuration choices sit behind one number.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelClaude Opus 5 (Adaptive Reasoning, Max Effort)AnthropicScore54.87%Configuration2,158 text-only questions, no tools, pass@1, LLM equality checkerSourceThird-party2026-08-16
ModelClaude Opus 5 (Xhigh)AnthropicScore54.4%Configuration2,158 text-only questions, no tools, pass@1, LLM equality checkerSourceThird-party2026-08-16
ModelClaude Opus 5 (High)AnthropicScore52.83%Configuration2,158 text-only questions, no tools, pass@1, LLM equality checkerSourceThird-party2026-08-16
ModelClaude Opus 5 (Medium)AnthropicScore51.3%Configuration2,158 text-only questions, no tools, pass@1, LLM equality checkerSourceThird-party2026-08-16
ModelGPT-5.6 Sol (max)OpenAIScore49.49%Configuration2,158 text-only questions, no tools, pass@1, LLM equality checkerSourceThird-party2026-08-16

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Active Still spreads the field.