Agentic · Tool use · Multimodal
GAIA
GAIA: A Benchmark for General AI Assistants
Questions easy for humans, hard for AI, needing tools and multi-step work.
- Released
- 2023
- Built by
- Meta AI, HuggingFace, AutoGPT
- Size
- 466 tasks
- Status
- Near ceiling
- Reported by
- Meta, HuggingFace, OpenAI
- Judge
- 80
- Deterministic, but it can only grade a final answer and not the work.
- Headroom
- 40
- The top models crowd the ceiling; gaps are shrinking.
- Resolution
- 68
- 466 instances, so one item moves the score by 0.215 points.
- Fine print
- 45
- 4 documented caveats, the heaviest being possible training-set contamination.
- Adoption
- 39
- Reported by 3 labs; 0 published scores collected here.
Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.
How the evaluation works
Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.
Stage 3 · What the model can do while measured
Bring your own scaffold
GAIA defines the questions and refuses to define the agent. Systems bring their own browser, code interpreter, file reader and search tools, and answering usually needs the live internet plus local file processing. There is no step cap and no prescribed tool set, which is precisely why two scaffolds wrapped around the same base model post very different GAIA numbers.
The web changes underneath the eval, so a score is only valid for the day it was run.
- Budget
- None. The benchmark specifies no step cap.
- Tools exposed
- Browser and web search supplied by the evaluated systemCode interpreter supplied by the evaluated systemFile readers for spreadsheets, PDFs, images and audio
Where the tasks came from
466 questions: 300 in a private test set with unreleased answers, 166 in the public validation set.
Hand-authored by the paper's authors and collaborators; each question has a verified unique answer and a recorded human solution trace.
A held-out or private split exists, which makes contamination easier to detect than on a fully public set.
What the number does not tell you
The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.
After two years of blog posts, notebooks and repos publishing worked solutions, a validation-set number says very little. Only the private test set is meaningful, and only through the leaderboard.
Source
Published scores
Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.
No published scores collected for this benchmark yet.
Near ceiling — The top models are close enough to the ceiling to crowd.