History Problem Core Idea Practice Results Impact Quiz Takeaways
Interactive Paper Explainer

The Autopsy of Multi-Agent Systems
Why Do MAS Fail?

Multi-agent benchmarks show minimal gains over single agents. The answer why: 1600+ real failure traces, 14 failure modes, 3 categories — and the finding that most failures are systemic, not modelic.

Start Learning Read the Paper ↗
1600+
Annotated traces
14
Failure modes
3
Categories
0.88
Annotator kappa
History

Hype Meets Autopsy

The 2024-25 arc: enthusiasm, minimal gains, and finally a diagnosis.

2023-24
The multi-agent boom
CAMEL, MetaGPT, AutoGen (entries #105-107): society-of-agents as the frontier — applications everywhere.
2024-25
The gains question
Benchmark reality: MAS often barely beat a good single agent — enthusiasm without mechanism.
Mar 2025
🚀 Why Do MAS Fail?
Cemri et al.: MAST-Data — 1600+ annotated failure traces across 7 popular frameworks — and the MAST taxonomy: 14 modes, 3 categories, expert-validated (kappa 0.88).
2025
The diagnostic tool
An LLM-as-a-Judge pipeline (human-validated) classifies new failures automatically; the field gains its instrument for building reliable MAS.
Failures Live in the Wiring

The taxonomy's three categories locate blame precisely: (i) specification and system design — the human wrote bad rules (underspecified goals, wrong orchestration logic); (ii) inter-agent misalignment — the agents disagree productively never (contradictions, role confusion, information asymmetry, circular reasoning); (iii) task verification — nobody checks the work (unverified outputs, premature termination). The meta-finding: most failures trace to system design and inter-agent misalignment — not to model capability. The models are fine; the conversations around them are broken. Reliability is an engineering discipline, and now it has a bug tracker.

Chapter 01

Vibes Over Verification

The 2025 status: MAS deployed on enthusiasm, failing silently, debugged by folklore.

🌀
The Diagnostic Vacuum
  • Multi-agent gains on benchmarks are often minimal — the cause was unstudied
  • Failure debugging is folklore: read transcripts, guess, re-prompt
  • No shared vocabulary: 'it went off the rails' is not a failure class
  • Frameworks multiply; reliability knowledge does not
📋
The MAST Answer
  • MAST-Data: 1600+ annotated failure traces across 7 popular MAS frameworks
  • MAST taxonomy: 14 failure modes in 3 categories — specification/design, inter-agent misalignment, verification
  • Expert-validated: inter-annotator agreement kappa 0.88
  • LLM-as-a-Judge classifier with high human agreement — failures now diagnose at scale
Analogy — The Air-Crash Investigation

Single-model failures are engine failures — swap the engine (bigger model), problem often solved. MAS failures are air-traffic-control failures: every plane airworthy, the choreography lethal — two agents cleared onto the same task, one landing while another takes off, nobody watching the runway (verification). MAST is the crash-investigation board: a taxonomy of accidents that mostly blame the tower, not the jets.

Chapter 02

The Three Categories

Where blame lands — the taxonomy's top level.

1️⃣ Specification & System Design
The human's fault: underspecified goals, incorrect orchestration logic, missing context in prompts — the system as designed cannot work.
2️⃣ Inter-Agent Misalignment
The conversation's fault: agents contradict each other, roles blur, information fails to propagate, loops and groupthink emerge.
3️⃣ Task Verification
The audit's absence: outputs unverified, termination premature — work ships un-checked.
📊 The meta-finding
Across 1600+ traces: design and misalignment dominate — model capability is rarely the binding constraint.
Representative modes (of 14)
  • Goal underspecification — the objective lacks success criteria
  • Role confusion — agents act outside their defined responsibilities
  • Circular reasoning / loops — agents defer to each other forever
  • Information asymmetry — one agent lacks what another knows
  • Unverified output — no agent checks the deliverable
The methodology
  • Taxonomy built from rigorous analysis of 150 traces, guided by expert annotators
  • Validated at kappa 0.88 agreement — a reproducible classification standard
  • LLM-as-a-Judge pipeline extends annotation to scale, matching humans closely
  • MAST-Data released: the field's first failure corpus for MAS
Interactive Demo — One Failed Run, Three Diagnoses

Tab through a failed multi-agent coding session — the same transcript, three candidate root causes.

Chapter 03

The Practice Guide

What the taxonomy prescribes for builders.

Prescriptions by Category

Design failures demand explicit specifications: success criteria, role charters, and orchestration logic as reviewable artifacts. Misalignment failures demand interface contracts: what each agent must send, receive, and own — plus loop detection and divergence monitoring. Verification failures demand dedicated verifier agents and termination conditions tied to checked work, not message counts. The quiet revolution: these are software-engineering disciplines applied to conversations — MAS reliability as a testable property, with MAST as the bug taxonomy.

Interactive Demo — Watch a Failure Unfold

Follow one multi-agent run as it walks through three failure modes in sequence — the anatomy of a bad day.

Chapter 05

Engineering, Not Models

The diagnosis that reframed multi-agent reliability.

TRACES
1600+
across 7 popular frameworks
TAXONOMY
14 modes
3 categories, kappa 0.88
DOMINANT CAUSE
systemic
design + misalignment over model limits
TOOLING
LLM-judge
human-matched automated diagnosis
Interactive Demo — The Constructive Surprise

A paper about failure that ends in a tool. Press reveal.

CategoryBlameExample modeFix discipline
Specification & designthe buildergoal underspecificationexplicit success criteria
Inter-agent misalignmentthe wiringrole confusion, loopsinterface contracts + monitoring
Task verificationthe auditunverified outputverifier agents + real termination tests
Model capabilityrarely the cause——

The decomposition that redirected multi-agent research from 'better models' to 'better engineering'.

Legacy

Legacy — The Discipline Turn

Multi-agent research got its failure science.

📋 The bug taxonomy
MAST became the shared vocabulary for MAS debugging — papers and postmortems now cite failure modes by name, not anecdote.
🔧 Reliability engineering
The prescriptions (specs, contracts, verifiers) reframed multi-agent building as software engineering — the field's maturity event.
🧰 The diagnostic tooling
The human-validated LLM-judge classifier made failure analysis scalable — continuous reliability monitoring became conceivable.
⚠️ What it did NOT solve
Taxonomies describe, they don't prevent; annotation lives within 7 frameworks' idioms; benchmarks for GOOD multi-agent behavior remain nascent; and the LLM-judge inherits its own blind spots — the instrument needs its own calibration culture.
🛤 Read next
Test Yourself

Quick Quiz

Check your understanding of the key concepts from Why MAS Fail.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ MAST: 14 failure modes in 3 categories — design, misalignment, verification.
✅ MAST-Data: 1600+ annotated traces across 7 frameworks; expert agreement kappa 0.88.
✅ Meta-finding: failures are mostly systemic — models are rarely the binding constraint.
✅ Prescriptions: explicit specs, interface contracts, dedicated verifiers, real termination tests.
✅ LLM-as-a-Judge diagnosis at scale — the taxonomy as living tooling.
✅ Read it as multi-agent systems getting their reliability-engineering discipline.