All benchmarks

Aggregate index

AA Intelligence Index

Artificial Analysis Intelligence Index

One number combining several public evals, reweighted whenever they saturate.

Released
2024
Built by
Artificial Analysis
Size
Not published
Status
Active
Reported by
Google DeepMind, Anthropic
Signal57Read with context
Judge
58
Mixed grading, so part of the score inherits the weakest judge.
Headroom
82
Still separates the frontier from everything below it.
Resolution
50
Instance count is not published, so per-item weight is unknown.
Fine print
47
4 documented caveats, the heaviest being construct validity.
Adoption
26
Reported by 2 labs; 0 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Fixed decoding config

Artificial Analysis pins the run configuration rather than the scaffold: non-reasoning models at temperature 0, reasoning models at temperature 0.6, and a maximum output of 16,384 tokens for non-reasoning models. Evaluation is text-only and English-only. The company states a 95% confidence interval within plus-or-minus 1%.

Single turn, no tools. The most reproducible setup there is.

Budget
per constituent eval

Where the tasks came from

Not a dataset. Nine constituent evals in four weighted categories as of Aug 2026; ten flatly averaged evals in v3.0.

Third-party evals selected and weighted editorially by Artificial Analysis, a commercial benchmarking company.

Every task is public, so there is no held-out split to check contamination against.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Between v3.0 and the current version the index went from a flat average over ten evals to a weighted average over nine in four categories. A model's Index score can rise or fall across that boundary without the model or its measured component scores changing at all, purely because the arithmetic changed. Quoting an Index number without its version is quoting a number whose definition is unstated.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

No published scores collected for this benchmark yet.

Active Still spreads the field.