All benchmarks

Agentic · Tool use · Multimodal

GAIA

GAIA: A Benchmark for General AI Assistants

Questions easy for humans, hard for AI, needing tools and multi-step work.

Released
2023
Built by
Meta AI, HuggingFace, AutoGPT
Size
466 tasks
Status
Near ceiling
Reported by
Meta, HuggingFace, OpenAI
Signal57Read with context
Judge
80
Deterministic, but it can only grade a final answer and not the work.
Headroom
40
The top models crowd the ceiling; gaps are shrinking.
Resolution
68
466 instances, so one item moves the score by 0.215 points.
Fine print
45
4 documented caveats, the heaviest being possible training-set contamination.
Adoption
39
Reported by 3 labs; 0 published scores collected here.

Weighted from facts already on the record — not an editorial rating. A low score means a number from this benchmark needs more context to read, not that the benchmark is bad.

How the evaluation works

Every benchmark on this site is broken onto the same five stages. Select a stage to see what it actually involves.

Stage 3 · What the model can do while measured

Bring your own scaffold

GAIA defines the questions and refuses to define the agent. Systems bring their own browser, code interpreter, file reader and search tools, and answering usually needs the live internet plus local file processing. There is no step cap and no prescribed tool set, which is precisely why two scaffolds wrapped around the same base model post very different GAIA numbers.

The web changes underneath the eval, so a score is only valid for the day it was run.

Budget
None. The benchmark specifies no step cap.
Tools exposed
Browser and web search supplied by the evaluated systemCode interpreter supplied by the evaluated systemFile readers for spreadsheets, PDFs, images and audio

Where the tasks came from

466 questions: 300 in a private test set with unreleased answers, 166 in the public validation set.

Hand-authored by the paper's authors and collaborators; each question has a verified unique answer and a recorded human solution trace.

A held-out or private split exists, which makes contamination easier to detect than on a fully public set.

What the number does not tell you

The most useful part of any benchmark is the fine print. These are the reasons a published score can mislead.

  • After two years of blog posts, notebooks and repos publishing worked solutions, a validation-set number says very little. Only the private test set is meaningful, and only through the leaderboard.

    Source

Published scores

Collected from model cards and independent evaluators. The configuration column is the part that decides whether two numbers can be compared at all.

No published scores collected for this benchmark yet.

Near ceiling The top models are close enough to the ceiling to crowd.