Knowledge
AA-Omniscience
AA-Omniscience (Artificial Analysis knowledge and hallucination benchmark)
Six thousand fact questions that penalize confident wrong answers.
- Released
- 2025
- Built by
- Artificial Analysis
- Size
- 6,000 tasks
- Status
- Frontier
- Reported by
- Artificial Analysis
- Judge
- 44
- A model grades the model. Update the grader and old scores move.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 100
- 6,000 instances, so one item moves the score by 0.017 points.
- Fine print
- 50
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 27
- Reported by 1 lab; 5 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Closed book, no tools
No search, no retrieval, no tools, one repeat per question. Because abstention is scored neutrally and hallucination is scored negatively, a model's refusal policy directly moves its number without any change in what it knows. That is a product decision showing up as a capability measurement.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn
Where the tasks came from
6,000 questions across 42 topics, run at a single repeat, so sampling noise is meaningful despite the large item count.
Authored by Artificial Analysis. Only a public subset is released; the scored set is proprietary.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The plus-one / minus-one / zero rule makes the Index highly sensitive to how readily a model declines to answer. A model tuned to refuse more will score better while knowing exactly the same amount. That is arguably the point, since reliability is what is being measured, but it means the Index is as much a measure of product tuning as of knowledge.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Opus 5 (Adaptive Reasoning, Max Effort)Anthropic | Score37.07 | ConfigurationAA-Omniscience Index 37.07 on a -100..100 scale, NOT a percentage; closed-book, 1 repeat | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (Xhigh)Anthropic | Score35.38 | ConfigurationAA-Omniscience Index 35.38 on a -100..100 scale, NOT a percentage; closed-book, 1 repeat | SourceThird-party2026-08-16 |
| ModelGemini 3.1 Pro PreviewGoogle DeepMind | Score31.88 | ConfigurationAA-Omniscience Index 31.88 on a -100..100 scale, NOT a percentage; closed-book, 1 repeat | SourceThird-party2026-08-16 |
| ModelGrok 4.6 (high)xAI | Score30.48 | ConfigurationAA-Omniscience Index 30.48 on a -100..100 scale, NOT a percentage; closed-book, 1 repeat | SourceThird-party2026-08-16 |
| ModelClaude Opus 4.8 (Max)Anthropic | Score28.75 | ConfigurationAA-Omniscience Index 28.75 on a -100..100 scale, NOT a percentage; closed-book, 1 repeat | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.