While LLMs were beating humans on professional exams, GAIA went the other way: everyday questions demanding reasoning + browsing + tools + multimodal understanding — humans 92%, GPT-4 with plugins 15%.
2023's paradox: models aced bar exams and failed errands.
GAIA's questions are engineered to be effortless for humans, hard for AI — the reverse of exam-style benchmarks. The recipe: real-world questions whose difficulty is not deep knowledge but orchestration — a bit of reasoning, a file to look at, a website to check, a tool to run, in sequence. Humans do this transparently (92%). For 2023 models, each hop is a failure surface, and GPT-4 even WITH plugins managed only 15%. The design insight: interaction, not intelligence, was the bottleneck — exactly what the agent era needed to measure.
Exam-style evals measured knowledge; assistants needed orchestration.
Exam benchmarks are trivia night: depth of recall, one domain, no tools allowed. GAIA is an errand list on a rainy Tuesday: 'check the museum site for Tuesday hours, look at the photo of the ticket, tell me if we can make the 6pm showing' — no single hard step, just a chain of small competences that a person does without noticing and a 2023 model did without surviving.
What a GAIA question actually demands — and why the braid is the point.
The post-mortem that steered agent design for two years.
Plugin-equipped models failed at the seams, not the parts: wrong tool selected, browsing queries that never find the page, image content ignored under text instructions, intermediate results dropped between steps. The lesson the agent frameworks absorbed: the orchestration layer is the product — ReAct-style interleaving (entry #53), planner-executor splits, and tool-routing curricula all cite this gap. GAIA's second-order effect: the field stopped treating "model + a tool bag" as an assistant and started engineering the loop.
The number pair that re-labeled 'AI progress' for a public audience.
| Benchmark axis | Exam-style (bar, MMLU-class) | GAIA |
|---|---|---|
| Difficulty source | specialist knowledge | orchestration of everyday abilities |
| Tools/browsing | forbidden | required |
| Human performance | below model | 92% — far above 2023 models |
| Grading | multiple choice / rubric | short answers, programmatic verification |
| Measures | knowledge ceiling | assistant readiness |
The inversion table — GAIA measures the complementary axis to knowledge benchmarks.
GAIA gave the agent era its north-star metric and its humility.
Check your understanding of the key concepts from GAIA.
Everything you need to remember about this paper.