All benchmarks

Reasoning

ARC-AGI-1

Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1

Colored-grid puzzles a person can solve without any training.

Released
2019
Built by
François Chollet, now stewarded by the ARC Prize Foundation
Size
400 tasks
Status
Saturated
Reported by
ARC Prize Foundation, OpenAI, Anthropic, Google DeepMind, xAI, Moonshot AI, DeepSeek, Z.ai, Thinking Machines
Signal58Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
8
Bunched at the ceiling. A win here no longer means much.
Resolution
68
400 instances, so one item moves the score by 0.250 points.
Fine print
52
4 documented caveats, the heaviest being construct validity.
Adoption
100
Reported by 9 labs; 8 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Single turn, cost tracked

For raw LLMs this is one turn per task with no tools, though the leaderboard also lists Refinement and Kaggle systems that wrap a base model in program synthesis, test-time training, or best-of-n search. The Kaggle track caps total compute at $50 for 120 tasks with no internet. ARC Prize publishes cost per task at retail API pricing next to every accuracy figure, and only lists systems whose full run cost under $10,000.

Single turn, no tools. The most reproducible setup there is.

Budget
Single turn per task, two attempts

Where the tasks came from

400 public evaluation tasks, but only 100-120 private tasks, so headline numbers carry the same small-n noise problem as GPQA.

Hand-designed by François Chollet. 400 public training, 400 public evaluation, 100 private evaluation, plus a semi-private set for API testing.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • The ARC Prize leaderboard lists Claude Fable 5 at 98.5%, against a Human Panel row of 98.0% at $17.00 per task and a STEM Grad row of 98.0% at $10.00. Average MTurkers score 77.0%. The benchmark still teaches well, but it no longer separates frontier systems, which is why ARC-AGI-2 and ARC-AGI-3 exist.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

pass@2Higher is better
pass@2higher is better
ModelClaude Fable 5 (Max)AnthropicScore98.5%Configurationpass@2, ARC Prize leaderboard, $5.45/taskSourceThird-party2026-08-16
ModelHuman Panel (reference row)ARC Prize FoundationScore98%ConfigurationHuman baseline row on the ARC Prize leaderboardSourceThird-party2026-08-16
ModelClaude Opus 5 (High)AnthropicScore97.5%Configurationpass@2, ARC Prize leaderboard, $1.45/taskSourceThird-party2026-08-16
ModelGPT-5.6 Sol (XHigh)OpenAIScore97.5%Configurationpass@2, ARC Prize leaderboard, $1.04/taskSourceThird-party2026-08-16
ModelAverage MTurker (reference row)ARC Prize FoundationScore77%ConfigurationHuman baseline row on the ARC Prize leaderboardSourceThird-party2026-08-16
Cost per taskLower is better
Cost per tasklower is better
ModelGPT-5.6 Sol (XHigh)OpenAIScore$1.04ConfigurationRetail API pricing at 97.5% pass@2SourceThird-party2026-08-16
ModelClaude Fable 5 (Max)AnthropicScore$5.45ConfigurationRetail API pricing at 98.5% pass@2SourceThird-party2026-08-16
ModelHuman Panel (reference row)ARC Prize FoundationScore$17ConfigurationHuman baseline cost at 98.0% pass@2SourceThird-party2026-08-16

Each table above measures a different quantity, so rows are only ever ordered within their own table. A model can appear more than once, and its numbers are not meant to be added or averaged.

At least one table here is lower is better — a harm score, lying rate, or attack success rate. The best result is at the top of every table, which means the smallest number leads where the direction is inverted.

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Saturated Scores are bunched at the ceiling; it no longer separates models.