All benchmarks

Agentic

GDPval

GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Real deliverables from 44 jobs, judged blind against human experts.

Released
2025
Built by
OpenAI
Size
1,320 tasks
Status
Frontier
Reported by
OpenAI, Epoch AI, Anthropic
Signal73Reads cleanly
Judge
62
Human graders. Not reproducible run to run, and expensive to repeat.
Headroom
100
Nothing is close to solving it, so the spread is real.
Resolution
92
1,320 instances, so one item moves the score by 0.076 points.
Fine print
46
4 documented caveats, the heaviest being construct validity.
Adoption
50
Reported by 3 labs; 4 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Files in, files out

There is no interactive environment. The model receives the prompt and the reference files and returns files, including spreadsheet, CAD and media formats it must generate directly. All of the expensive, interesting machinery in GDPval sits in the grading loop rather than in the environment, which is unusual for a benchmark about doing work.

Single turn, no tools. The most reproducible setup there is.

Budget
Single request per task; OpenAI sampled each model three times per prompt.

Where the tasks came from

1,320 tasks across 44 occupations; a 220-task gold subset is open-sourced and the rest is held back.

Written by industry professionals with an average of 14 years in the occupation, based on work they actually perform.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • Human inter-rater agreement on the gold subset is 71% and the automated grader manages 66%. That bounds how precise any GDPval figure can be: two models separated by a few points are indistinguishable, and a leaderboard reordering within that band is measurement noise rather than progress.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

ModelClaude Opus 4.1AnthropicScore47.6%ConfigurationWins plus ties against the human expert deliverable; best model reported in the paper.SourceLab-reported2025-10
ModelGPT-5OpenAIScore39%ConfigurationWins plus ties against the human expert deliverable.SourceLab-reported2025-10
Modelo3OpenAIScore35.2%ConfigurationWins plus ties against the human expert deliverable.SourceLab-reported2025-10
Modelo4-miniOpenAIScore29.1%ConfigurationWins plus ties against the human expert deliverable.SourceLab-reported2025-10

Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.

Frontier Nothing is close to solving it.