History Problem Core Idea Anatomy Results Impact Quiz Takeaways
Interactive Paper Explainer

Questions Humans Find Easy
GAIA

While LLMs were beating humans on professional exams, GAIA went the other way: everyday questions demanding reasoning + browsing + tools + multimodal understanding — humans 92%, GPT-4 with plugins 15%.

Start Learning Read the Paper ↗
466
Questions
92% vs 15%
Humans vs GPT-4+plugins
4
Ability families
2023
Mialon et al.
History

The Inverted Benchmark

2023's paradox: models aced bar exams and failed errands.

2023
Humans surpassed — on paper
GPT-4-class results on professional exams flip the usual benchmark story: machines above people on specialist tests.
2023
The everyday-errand gap
Meanwhile assistants still fumble: counting letters, browsing for a fact, reading an attached image — tasks a teenager handles.
Nov 2023
🚀 GAIA
Mialon et al. (Meta + HuggingFace + others): 466 questions that are conceptually SIMPLE for humans but require orchestrated abilities — the inversion made rigorous.
2023-24
The agent-development compass
GAIA levels 1-3 become the standard progress meter for general assistants; plugin/tool systems target its gap explicitly.
2024-25
Closing, then saturating
Tool-equipped agents climb level 1 → level 3; the benchmark's role shifts to component diagnosis as successors ( browse-comp, deep-research evals) push further.
Design for the Inversion

GAIA's questions are engineered to be effortless for humans, hard for AI — the reverse of exam-style benchmarks. The recipe: real-world questions whose difficulty is not deep knowledge but orchestration — a bit of reasoning, a file to look at, a website to check, a tool to run, in sequence. Humans do this transparently (92%). For 2023 models, each hop is a failure surface, and GPT-4 even WITH plugins managed only 15%. The design insight: interaction, not intelligence, was the bottleneck — exactly what the agent era needed to measure.

Chapter 01

Benchmarks Pointing Up the Wrong Mountain

Exam-style evals measured knowledge; assistants needed orchestration.

📈
The Specialist-Exam Trap
  • Professional-exam benchmarks reward memorized expertise — models win, everyone applauds, nothing about 'assistant capability' is measured
  • Questions with a single retrieval step flatter lookups, not workflows
  • No metric for the everyday composite: browse a page, view an image, reason, answer precisely
  • The public narrative ('surpasses humans') diverges from product reality ('can't book a table')
🔄
The GAIA Answer
  • 466 questions, non-gameable, requiring fundamental abilities TOGETHER: reasoning, multimodality, web browsing, tool proficiency
  • Deliberately human-easy: 92% average (non-expert respondents, a few minutes per question)
  • Machine-hard: 15% for GPT-4 + plugins (2023) — the inversion quantified
  • Three difficulty levels; automatic verifiers (short answers, programmatic checks) keep grading objective
Analogy — The Errand List vs the Trivia Night

Exam benchmarks are trivia night: depth of recall, one domain, no tools allowed. GAIA is an errand list on a rainy Tuesday: 'check the museum site for Tuesday hours, look at the photo of the ticket, tell me if we can make the 6pm showing' — no single hard step, just a chain of small competences that a person does without noticing and a 2023 model did without surviving.

Chapter 02

The Ability Braid

What a GAIA question actually demands — and why the braid is the point.

🧠 Reasoning
A few steps of logic/arithmetic — nothing graduate-level, but sequencing that survives the next hop's distraction.
🖼 Multimodality
Attached images, PDFs, or video frames carrying part of the answer — reading them is mandatory, not decorative.
🌐 Browsing
Live web navigation to find a fact that appears nowhere in training data (recent, obscure, or computed).
🔧 Tool use
Calculators, file readers, code execution — the assistant must CHOOSE the tool and sequence it correctly.
The measurement (2023)
  • Humans (non-expert): 92% average across levels
  • GPT-4 with plugins: 15% — plugins help but orchestration still fails
  • Level 1 → 3: increasing steps/hops, decreasing human-ease — the difficulty ladder is workload, not obscurity
  • Auto-verification: short factual answers checked programmatically
The design rules
  • Answers must be simple, verifiable, and non-gameable — no "write an essay" grading
  • Questions resist retrieval-only solving: the fact is entangled or simply not on the static web
  • Full reproducibility: everything needed ships with the question (files) or is live-web stable
Interactive Demo — Anatomy of Three Questions

Tab through GAIA-style questions — each easy for a person, each a different failure braid for a model.

Chapter 03

What the 15% Taught

The post-mortem that steered agent design for two years.

Failure Anatomy

Plugin-equipped models failed at the seams, not the parts: wrong tool selected, browsing queries that never find the page, image content ignored under text instructions, intermediate results dropped between steps. The lesson the agent frameworks absorbed: the orchestration layer is the product — ReAct-style interleaving (entry #53), planner-executor splits, and tool-routing curricula all cite this gap. GAIA's second-order effect: the field stopped treating "model + a tool bag" as an assistant and started engineering the loop.

Interactive Demo — One Question, Two Solvers

Follow the same multi-hop question through a human and a 2023 plugin-agent — watch where the 15% comes from.

Chapter 05

92 vs 15

The number pair that re-labeled 'AI progress' for a public audience.

HUMANS
92%
non-experts, minutes per question
GPT-4 + PLUGINS (2023)
15%
tools present, orchestration missing
QUESTIONS
466
3 difficulty levels, verified answers
ABILITIES BRAID
4
reasoning · multimodal · browsing · tools
Interactive Demo — Why Not Just Ask Harder Questions?

The benchmark could have escalated difficulty instead of inverting. Press reveal for the design reasoning.

Benchmark axisExam-style (bar, MMLU-class)GAIA
Difficulty sourcespecialist knowledgeorchestration of everyday abilities
Tools/browsingforbiddenrequired
Human performancebelow model92% — far above 2023 models
Gradingmultiple choice / rubricshort answers, programmatic verification
Measuresknowledge ceilingassistant readiness

The inversion table — GAIA measures the complementary axis to knowledge benchmarks.

Legacy

Legacy — The Assistant Axis

GAIA gave the agent era its north-star metric and its humility.

🧭 The agent progress meter
Levels 1-3 became the standard 'general assistant' scoreboard — framework releases (AutoGen-era, deep-research systems) reported GAIA deltas as the capability claim.
🔧 Orchestration as the product
The 15%-with-plugins autopsy legitimized planner-executor designs and tool-routing research: the loop, not the model, was the unit of engineering.
📉 The public narrative correction
92-15 punctured 'AI surpasses humans' hype with a concrete, quotable inversion — the benchmark as communication device.
🧪 Verifiable short-answer design
Programmatic checking of simple answers (no judge, no rubric) joined the objective-grading doctrine — GAIA, GPQA, LiveBench as one family.
⚠️ What it did NOT solve
Live-web dependence makes some questions drift; 466 questions saturate as agents improve; and 'everyday' skews Western, English, connected-world errands — culture and access bias the everyday.
🛤 Read next
The assistants chain: WebArena · BrowseComp · ReAct
Test Yourself

Quick Quiz

Check your understanding of the key concepts from GAIA.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ GAIA = 466 real-world questions braiding reasoning, multimodality, browsing, and tool use.
✅ The inversion: humans 92% (non-experts, minutes) vs GPT-4-with-plugins 15%.
✅ Difficulty = orchestration, not obscurity — assistant readiness as the measured axis.
✅ Short verifiable answers + programmatic checks keep grading judge-free and objective.
✅ The autopsy: failures live at tool seams, not in knowledge — orchestration became the product.
✅ Read it as the benchmark that gave the agent era both a compass and a humility check.