All benchmarks

Reasoning · Agentic

ARC-AGI-3

Abstraction and Reasoning Corpus for Artificial General Intelligence, version 3 (Interactive Reasoning Benchmark)

Novel video games an agent must learn with no instructions.

Released
2025
Built by
ARC Prize Foundation
Size
Not published
Status
Frontier
Reported by
ARC Prize Foundation
Signal72Reads cleanly
Judge
94
Graded on the final state of the environment, not on prose.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
50
Instance count is not published, so per-item weight is unknown.
Fine print
50
4 documented caveats, the heaviest being construct validity.
Adoption
30
Reported by 1 lab; 6 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Multi-turn game loop

This is the structural break from ARC-AGI-1 and 2: the model acts in a loop over many steps with memory across the episode, rather than answering once. The leaderboard tracks a separate Cost (V3) column measured in thousands of dollars per full evaluation, not dollars per task, because running ARC-AGI-3 costs roughly ten thousand times what ARC-AGI-2 costs. Total run costs land between $2K and $25K per system.

Single turn, no tools. The most reproducible setup there is.

Budget
Many steps per game, memory across the episode
Tools exposed
Game action spaceFrame observations

Where the tasks came from

The number of private game environments is not published, so it is recorded as null rather than guessed. Very few models have any ARC-AGI-3 score at all.

Hand-built games, all verified human-solvable, with public preview games and held-out environments for the leaderboard and the Kaggle track.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Claude Opus 5 (High) scores 30.2% while the next best is GPT-5.6 Sol (Max) at 7.8% and most entries are under 3%. That is not the top of a smooth curve, it is an outlier. With a field that sparse, a single system's scaffolding advantage is indistinguishable from a capability jump.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelClaude Opus 5 (High)AnthropicScore30.2%ConfigurationTotal evaluation cost $20.7K (not per task)SourceThird-party2026-08-16
ModelGPT-5.6 Sol (Max)OpenAIScore7.8%ConfigurationTotal evaluation cost $25.1K (not per task)SourceThird-party2026-08-16
ModelGPT-5.6 Sol (XHigh)OpenAIScore7%ConfigurationTotal evaluation cost $19.2K (not per task)SourceThird-party2026-08-16
ModelGrok 4.6 (XHigh)xAIScore2.1%ConfigurationTotal evaluation cost $5.6K (not per task)SourceThird-party2026-08-16
ModelClaude Opus 4.8 (High)AnthropicScore1.5%ConfigurationTotal evaluation cost $10.0K (not per task)SourceThird-party2026-08-16
ModelGPT-5.5 (High)OpenAIScore0.4%ConfigurationTotal evaluation cost $10.0K (not per task)SourceThird-party2026-08-16

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Frontier Nothing is close to solving it.