Reasoning · Agentic
ARC-AGI-3
Abstraction and Reasoning Corpus for Artificial General Intelligence, version 3 (Interactive Reasoning Benchmark)
Novel video games an agent must learn with no instructions.
- Released
- 2025
- Built by
- ARC Prize Foundation
- Size
- Not published
- Status
- Frontier
- Reported by
- ARC Prize Foundation
- Judge
- 94
- Graded on the final state of the environment, not on prose.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 50
- Instance count is not published, so per-item weight is unknown.
- Fine print
- 50
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 30
- Reported by 1 lab; 6 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Multi-turn game loop
This is the structural break from ARC-AGI-1 and 2: the model acts in a loop over many steps with memory across the episode, rather than answering once. The leaderboard tracks a separate Cost (V3) column measured in thousands of dollars per full evaluation, not dollars per task, because running ARC-AGI-3 costs roughly ten thousand times what ARC-AGI-2 costs. Total run costs land between $2K and $25K per system.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Many steps per game, memory across the episode
- Tools exposed
- Game action spaceFrame observations
Where the tasks came from
The number of private game environments is not published, so it is recorded as null rather than guessed. Very few models have any ARC-AGI-3 score at all.
Hand-built games, all verified human-solvable, with public preview games and held-out environments for the leaderboard and the Kaggle track.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Claude Opus 5 (High) scores 30.2% while the next best is GPT-5.6 Sol (Max) at 7.8% and most entries are under 3%. That is not the top of a smooth curve, it is an outlier. With a field that sparse, a single system's scaffolding advantage is indistinguishable from a capability jump.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Opus 5 (High)Anthropic | Score30.2% | ConfigurationTotal evaluation cost $20.7K (not per task) | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (Max)OpenAI | Score7.8% | ConfigurationTotal evaluation cost $25.1K (not per task) | SourceThird-party2026-08-16 |
| ModelGPT-5.6 Sol (XHigh)OpenAI | Score7% | ConfigurationTotal evaluation cost $19.2K (not per task) | SourceThird-party2026-08-16 |
| ModelGrok 4.6 (XHigh)xAI | Score2.1% | ConfigurationTotal evaluation cost $5.6K (not per task) | SourceThird-party2026-08-16 |
| ModelClaude Opus 4.8 (High)Anthropic | Score1.5% | ConfigurationTotal evaluation cost $10.0K (not per task) | SourceThird-party2026-08-16 |
| ModelGPT-5.5 (High)OpenAI | Score0.4% | ConfigurationTotal evaluation cost $10.0K (not per task) | SourceThird-party2026-08-16 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.