Safety
AILuminate
AILuminate: v1.0 of the AI Risk and Reliability Benchmark from MLCommons
An industry-consortium safety grade, with the real test set kept private.
- Released
- 2024
- Built by
- MLCommons
- Size
- 24,000 tasks
- Status
- Active
- Reported by
- MLCommons
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 100
- 24,000 instances, so one item moves the score by 0.004 points.
- Fine print
- 47
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 13
- Reported by 1 lab; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Whole deployed product
Single-turn prompt and response against the system under test as deployed, including its filters, system prompt, and refusal layer. This is a product evaluation, not a model evaluation, and the distinction changes what the grade means: the same weights behind a different safety stack are a different entrant.
Single turn, no tools. The most reproducible setup there is.
- Budget
- single turn
Where the tasks came from
24,000 prompts per language: 12,000 practice and 12,000 official private, 1,000 per hazard category, with only 10% of the practice set public. The 2026 site reports 59,624 test prompts and 477 test images across all suites and languages, with 109 models benchmarked.
Written by a 100+ organization consortium against a jointly agreed twelve-hazard taxonomy; the official half is never released.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The exact composition, weighting, and thresholds of the guard-model ensemble are not published. For a benchmark explicitly designed for procurement and public reporting, that is an unusual property: a vendor who receives a Poor grade cannot inspect which responses were marked violating or why, and no external party can reproduce the judgement. Human validation with 3 raters and majority vote is real evidence about the ensemble's average quality; it is not an audit trail for any individual grade.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Active — Still spreads the field.