Reasoning
ARC-AGI-2
Abstraction and Reasoning Corpus for Artificial General Intelligence, version 2
Harder grid puzzles targeting symbolic, compositional, and contextual reasoning.
- Released
- 2025
- Built by
- ARC Prize Foundation
- Size
- 120 tasks
- Status
- Near ceiling
- Reported by
- ARC Prize Foundation, OpenAI, Anthropic, Google DeepMind, xAI, Moonshot AI, DeepSeek, Z.ai, Thinking Machines
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 40
- The top models crowd the ceiling; gaps are shrinking.
- Resolution
- 52
- 120 instances, so one item moves the score by 0.833 points.
- Fine print
- 54
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 100
- Reported by 9 labs; 8 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Single turn, cost mandatory
Same run conditions as version 1, except that from ARC-AGI-2 onward the efficiency metric is mandatory: no score is published without its cost. The stated rationale is that unlimited brute-force search could eventually solve ARC, and that would not be intelligence. Kaggle contest entries run under a $50-per-120-tasks compute cap and are not comparable to API-priced frontier runs on the same table.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single turn per task, two attempts
Where the tasks came from
120 tasks per evaluation set, so one task is 0.83pp and the noise band is wide. Three separate 120-task sets exist, calibrated to within about 1pp of each other.
1,000 public training tasks; 120 public, 120 semi-private and 120 private evaluation tasks. Human-calibrated in a 2025 live study.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
GPT-5.6 Sol (Max) sits at 92.5% and Claude Opus 5 (Max) at 90.4%, both above the 85% target, against a Human Panel row of 100%. For contrast, Gemini 3 Pro scored 31.1% at launch in November 2025 and the best 2024 Kaggle custom system managed 2.5%. The gap closed in under a year.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelHuman Panel (reference row)ARC Prize Foundation | Score100% | ConfigurationHuman baseline row on the ARC Prize leaderboard | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (Max)OpenAI | Score92.5% | Configurationpass@2, ARC Prize leaderboard, $1.44/task | SourceThird-party2026-08-16 |
| ModelClaude Opus 5 (Max)Anthropic | Score90.4% | Configurationpass@2, ARC Prize leaderboard, $2.06/task | SourceThird-party2026-08-16 |
| ModelClaude Fable 5 (Max)Anthropic | Score89.2% | Configurationpass@2, ARC Prize leaderboard, $5.45/task | SourceThird-party2026-08-16 |
| ModelGemini 3 Deep Think (2/26)Google DeepMind | Score84.6% | Configurationpass@2, ARC Prize leaderboard, $13.62/task | SourceThird-party2026-08-16 |
| ModelGemini 3 Pro (at launch)Google DeepMind | Score31.1% | Configurationpass@2, ARC Prize leaderboard, launch-day figure | SourceThird-party2025-11-18 |
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelGPT-5.6 Sol (Max)OpenAI | Score$1.44 | ConfigurationRetail API pricing at 92.5% pass@2 | SourceThird-party2026-08-16 |
| ModelGemini 3 Deep Think (2/26)Google DeepMind | Score$13.62 | ConfigurationRetail API pricing at 84.6% pass@2 | SourceThird-party2026-08-16 |
Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.
At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Near ceiling — The top models are close enough to the ceiling to crowd.