Aggregate index
AA Intelligence Index
Artificial Analysis Intelligence Index
One number combining several public evals, reweighted whenever they saturate.
- Released
- 2024
- Built by
- Artificial Analysis
- Size
- Not published
- Status
- Active
- Reported by
- Google DeepMind, Anthropic
- Judge
- 58
- Mixed grading, so part of the score inherits the weakest judge.
- Headroom
- 82
- Still separates the frontier from everything below it.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 47
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 26
- Reported by 2 labs; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Fixed decoding config
Artificial Analysis pins the run configuration rather than the scaffold: non-reasoning models at temperature 0, reasoning models at temperature 0.6, and a maximum output of 16,384 tokens for non-reasoning models. Evaluation is text-only and English-only. The company states a 95% confidence interval within plus-or-minus 1%.
Single turn, no tools. The most reproducible setup there is.
- Budget
- per constituent eval
Where the tasks came from
Not a dataset. Nine constituent evals in four weighted categories as of Aug 2026; ten flatly averaged evals in v3.0.
Third-party evals selected and weighted editorially by Artificial Analysis, a commercial benchmarking company.
Every task is public, so there is no held-out split to check contamination against.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Between v3.0 and the current version the index went from a flat average over ten evals to a weighted average over nine in four categories. A model's Index score can rise or fall across that boundary without the model or its measured component scores changing at all, purely because the arithmetic changed. Quoting an Index number without its version is quoting a number whose definition is unstated.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Active — Still spreads the field.