Agentic
GDPval
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Real deliverables from 44 jobs, judged blind against human experts.
- Released
- 2025
- Built by
- OpenAI
- Size
- 1,320 tasks
- Status
- Frontier
- Reported by
- OpenAI, Epoch AI, Anthropic
- Judge
- 62
- Human graders. Not reproducible run to run, and expensive to repeat.
- Headroom
- 100
- Nothing is close to solving it, so the spread is real.
- Resolution
- 92
- 1,320 instances, so one item moves the score by 0.076 points.
- Fine print
- 46
- 4 documented caveats, the heaviest being construct validity.
- Adoption
- 50
- Reported by 3 labs; 4 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Files in, files out
There is no interactive environment. The model receives the prompt and the reference files and returns files, including spreadsheet, CAD and media formats it must generate directly. All of the expensive, interesting machinery in GDPval sits in the grading loop rather than in the environment, which is unusual for a benchmark about doing work.
Single turn, no tools. The most reproducible setup there is.
- Budget
- Single request per task; OpenAI sampled each model three times per prompt.
Where the tasks came from
1,320 tasks across 44 occupations; a 220-task gold subset is open-sourced and the rest is held back.
Written by industry professionals with an average of 14 years in the occupation, based on work they actually perform.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
Human inter-rater agreement on the gold subset is 71% and the automated grader manages 66%. That bounds how precise any GDPval figure can be: two models separated by a few points are indistinguishable, and a leaderboard reordering within that band is measurement noise rather than progress.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
| Model | Score | Configuration | Source |
|---|---|---|---|
| ModelClaude Opus 4.1Anthropic | Score47.6% | ConfigurationWins plus ties against the human expert deliverable; best model reported in the paper. | SourceLab-reported2025-10 |
| ModelGPT-5OpenAI | Score39% | ConfigurationWins plus ties against the human expert deliverable. | SourceLab-reported2025-10 |
| Modelo3OpenAI | Score35.2% | ConfigurationWins plus ties against the human expert deliverable. | SourceLab-reported2025-10 |
| Modelo4-miniOpenAI | Score29.1% | ConfigurationWins plus ties against the human expert deliverable. | SourceLab-reported2025-10 |
Ordering is not a ranking. Rows with different configurations are not measuring the same thing — a number produced with tools, extended thinking, or majority voting is not comparable to one produced without.
Frontier — Nothing is close to solving it.