Aggregate index
HELM
Holistic Evaluation of Language Models (HELM), including HELM Capabilities and HELM Safety
Evaluate every model on every scenario, on seven axes, transparently.
- Released
- 2022
- Built by
- Stanford CRFM
- Size
- Not published
- Status
- Saturated
- Reported by
- Stanford CRFM
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 52
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 35
- Reported by 1 lab; 13 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Standardized prompting
Prompting, adaptation method, and metric computation are standardized per scenario so that otherwise incompatible datasets can be placed side by side — that standardization is the framework's actual product. HELM Capabilities uses zero-shot chain-of-thought for multiple choice. There is no agentic environment and no tool use in the Capabilities or Safety tracks.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single turn
Where the tasks came from
Varies by leaderboard. HELM Safety v1.17.0 totals 2,950 scored instances across five benchmarks at 87 models; HELM Capabilities v1.15.0 is MMLU-Pro 1,000, GPQA 446, IFEval 541, WildBench 1,000, Omni-MATH 1,000 at 68 models. A sixth Safety group, HarmBenchGCGTransfer, is declared but has 0 models — which is why the blurb says five benchmarks while six groups exist.
Existing public benchmarks re-run under a standardized harness; HELM contributes the harness, the taxonomy, and the aggregation, not the items.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Across 87 models in v1.17.0, the top ten run from 0.9864 (GPT-5 nano) to 0.9744 (Claude 4 Sonnet) — a total range of 0.012 on a 0-1 scale. At that spread the ordering is decided by a handful of individual judgements from the model annotators, and the difference between 'first' and 'tenth' carries no usable information about deployment safety. A saturated safety leaderboard is worse than a missing one, because it licenses the claim that everything on it is safe.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5 nanoOpenAI | Score0.9864% | ConfigurationHELM Safety v1.17.0 mean score, 87 models. Rank 1 — and 0.012 above rank 10. | SourceThird-party2025-11-24 |
| Modelo3 (2025-04-16)OpenAI | Score0.9816% | ConfigurationHELM Safety v1.17.0 mean score. Rank 2. | SourceThird-party2025-11-24 |
| Modelgpt-oss-120bOpenAI | Score0.9813% | ConfigurationHELM Safety v1.17.0 mean score. Rank 3. | SourceThird-party2025-11-24 |
| ModelClaude 4 Sonnet (extended thinking)Anthropic | Score0.9807% | ConfigurationHELM Safety v1.17.0 mean score. Rank 4, tied to four decimal places with GPT-5 at rank 5. | SourceThird-party2025-11-24 |
| ModelGPT-5OpenAI | Score0.9807% | ConfigurationHELM Safety v1.17.0 mean score. Rank 5. | SourceThird-party2025-11-24 |
| ModelGPT-5 miniOpenAI | Score0.9803% | ConfigurationHELM Safety v1.17.0 mean score. Rank 6. | SourceThird-party2025-11-24 |
| ModelKimi K2 InstructMoonshot AI | Score0.9795% | ConfigurationHELM Safety v1.17.0 mean score. Rank 7. | SourceThird-party2025-11-24 |
| ModelClaude 3.5 SonnetAnthropic | Score0.9767% | ConfigurationHELM Safety v1.17.0 mean score. Rank 8 — a model over a year older than most of the board, inside 0.01 of first place. | SourceThird-party2025-11-24 |
| Modelo1OpenAI | Score0.9758% | ConfigurationHELM Safety v1.17.0 mean score. Rank 9. | SourceThird-party2025-11-24 |
| ModelClaude 4 SonnetAnthropic | Score0.9744% | ConfigurationHELM Safety v1.17.0 mean score. Rank 10 of 87 — 0.012 below rank 1. | SourceThird-party2025-11-24 |
| ModelGPT-5 miniOpenAI | Score0.819% | ConfigurationHELM CAPABILITIES v1.15.0 mean score, 68 models. Rank 1. Note that this is a different leaderboard from HELM Safety and the two are not additive. | SourceThird-party2025-11-24 |
| ModelGemini 3 Pro (Preview)Google DeepMind | Score0.7993% | ConfigurationHELM Capabilities v1.15.0 mean score. Rank 5 of 68. | SourceThird-party2025-11-24 |
| ModelClaude 4 Opus (extended thinking)Anthropic | Score0.78% | ConfigurationHELM Capabilities v1.15.0 mean score. Rank 8 of 68 — a 0.039 gap to rank 1, roughly three times the entire spread of the Safety top ten. | SourceThird-party2025-11-24 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.