Reasoning
ARC-AGI-1
Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1
Colored-grid puzzles a person can solve without any training.
- Released
- 2019
- Built by
- François Chollet, now stewarded by the ARC Prize Foundation
- Size
- 400 tasks
- Status
- Saturated
- Reported by
- ARC Prize Foundation, OpenAI, Anthropic, Google DeepMind, xAI, Moonshot AI, DeepSeek, Z.ai, Thinking Machines
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 8
- Bunched at the ceiling. A win here no longer means much.
- Resolution
- 68
- 400 instances, so one item moves the score by 0.250 points.
- Fine print
- 52
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 100
- Reported by 9 labs; 8 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Single turn, cost tracked
For raw LLMs this is one turn per task with no tools, though the leaderboard also lists Refinement and Kaggle systems that wrap a base model in program synthesis, test-time training, or best-of-n search. The Kaggle track caps total compute at $50 for 120 tasks with no internet. ARC Prize publishes cost per task at retail API pricing next to every accuracy figure, and only lists systems whose full run cost under $10,000.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn per task, two attempts
Where the tasks came from
400 public evaluation tasks, but only 100-120 private tasks, so headline numbers carry the same small-n noise problem as GPQA.
Hand-designed by François Chollet. 400 public training, 400 public evaluation, 100 private evaluation, plus a semi-private set for API testing.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
The ARC Prize leaderboard lists Claude Fable 5 at 98.5%, against a Human Panel row of 98.0% at $17.00 per task and a STEM Grad row of 98.0% at $10.00. Average MTurkers score 77.0%. The benchmark still teaches well, but it no longer separates frontier systems, which is why ARC-AGI-2 and ARC-AGI-3 exist.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Fable 5 (Max)Anthropic | Score98.5% | Configurationpass@2, ARC Prize leaderboard, $5.45/task | SourceThird-party2026-08-16 |
| ModelHuman Panel (reference row)ARC Prize Foundation | Score98% | ConfigurationHuman baseline row on the ARC Prize leaderboard | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (High)Anthropic | Score97.5% | Configurationpass@2, ARC Prize leaderboard, $1.45/task | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (XHigh)OpenAI | Score97.5% | Configurationpass@2, ARC Prize leaderboard, $1.04/task | SourceThird-party2026-08-16 |
| ModelAverage MTurker (reference row)ARC Prize Foundation | Score77% | ConfigurationHuman baseline row on the ARC Prize leaderboard | SourceThird-party2026-08-16 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5.6 Sol (XHigh)OpenAI | Score$1.04 | ConfigurationRetail API pricing at 97.5% pass@2 | SourceThird-party2026-08-16 |
| ModelClaude Fable 5 (Max)Anthropic | Score$5.45 | ConfigurationRetail API pricing at 98.5% pass@2 | SourceThird-party2026-08-16 |
| ModelHuman Panel (reference row)ARC Prize Foundation | Score$17 | ConfigurationHuman baseline cost at 98.0% pass@2 | SourceThird-party2026-08-16 |
Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.
At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Saturated — Scores are bunched at the ceiling; it no longer separates models.