All benchmarks

Reasoning

ARC-AGI-2

Abstraction and Reasoning Corpus for Artificial General Intelligence, version 2

Harder grid puzzles targeting symbolic, compositional, and contextual reasoning.

Released
2025
Built by
ARC Prize Foundation
Size
120 tasks
Status
Near ceiling
Reported by
ARC Prize Foundation, OpenAI, Anthropic, Google DeepMind, xAI, Moonshot AI, DeepSeek, Z.ai, Thinking Machines
Signal63Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
52
120 instances, so one item moves the score by 0.833 points.
Fine print
54
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
100
Reported by 9 labs; 8 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Single turn, cost mandatory

Same run conditions as version 1, except that from ARC-AGI-2 onward the efficiency metric is mandatory: no score is published without its cost. The stated rationale is that unlimited brute-force search could eventually solve ARC, and that would not be intelligence. Kaggle contest entries run under a $50-per-120-tasks compute cap and are not comparable to API-priced frontier runs on the same table.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn per task, two attempts

Where the tasks came from

120 tasks per evaluation set, so one task is 0.83pp and the noise band is wide. Three separate 120-task sets exist, calibrated to within about 1pp of each other.

1,000 public training tasks; 120 public, 120 semi-private and 120 private evaluation tasks. Human-calibrated in a 2025 live study.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • GPT-5.6 Sol (Max) sits at 92.5% and Claude Opus 5 (Max) at 90.4%, both above the 85% target, against a Human Panel row of 100%. For contrast, Gemini 3 Pro scored 31.1% at launch in November 2025 and the best 2024 Kaggle custom system managed 2.5%. The gap closed in under a year.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

pass@2Higher is better
pass@2higher is better
ModelHuman Panel (reference row)ARC Prize FoundationScore100%ConfigurationHuman baseline row on the ARC Prize leaderboardSourceThird-party2026-08-16
ModelGPT-5.6 Sol (Max)OpenAIScore92.5%Configurationpass@2, ARC Prize leaderboard, $1.44/taskSourceThird-party2026-08-16
ModelClaude Opus 5 (Max)AnthropicScore90.4%Configurationpass@2, ARC Prize leaderboard, $2.06/taskSourceThird-party2026-08-16
ModelClaude Fable 5 (Max)AnthropicScore89.2%Configurationpass@2, ARC Prize leaderboard, $5.45/taskSourceThird-party2026-08-16
ModelGemini 3 Deep Think (2/26)Google DeepMindScore84.6%Configurationpass@2, ARC Prize leaderboard, $13.62/taskSourceThird-party2026-08-16
ModelGemini 3 Pro (at launch)Google DeepMindScore31.1%Configurationpass@2, ARC Prize leaderboard, launch-day figureSourceThird-party2025-11-18
Cost per taskLower is better
Cost per tasklower is better
ModelGPT-5.6 Sol (Max)OpenAIScore$1.44ConfigurationRetail API pricing at 92.5% pass@2SourceThird-party2026-08-16
ModelGemini 3 Deep Think (2/26)Google DeepMindScore$13.62ConfigurationRetail API pricing at 84.6% pass@2SourceThird-party2026-08-16

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Near ceiling The top models are close enough to the ceiling to crowd.