Multi-agent benchmarks show minimal gains over single agents. The answer why: 1600+ real failure traces, 14 failure modes, 3 categories — and the finding that most failures are systemic, not modelic.
The 2024-25 arc: enthusiasm, minimal gains, and finally a diagnosis.
The taxonomy's three categories locate blame precisely: (i) specification and system design — the human wrote bad rules (underspecified goals, wrong orchestration logic); (ii) inter-agent misalignment — the agents disagree productively never (contradictions, role confusion, information asymmetry, circular reasoning); (iii) task verification — nobody checks the work (unverified outputs, premature termination). The meta-finding: most failures trace to system design and inter-agent misalignment — not to model capability. The models are fine; the conversations around them are broken. Reliability is an engineering discipline, and now it has a bug tracker.
The 2025 status: MAS deployed on enthusiasm, failing silently, debugged by folklore.
Single-model failures are engine failures — swap the engine (bigger model), problem often solved. MAS failures are air-traffic-control failures: every plane airworthy, the choreography lethal — two agents cleared onto the same task, one landing while another takes off, nobody watching the runway (verification). MAST is the crash-investigation board: a taxonomy of accidents that mostly blame the tower, not the jets.
Where blame lands — the taxonomy's top level.
What the taxonomy prescribes for builders.
Design failures demand explicit specifications: success criteria, role charters, and orchestration logic as reviewable artifacts. Misalignment failures demand interface contracts: what each agent must send, receive, and own — plus loop detection and divergence monitoring. Verification failures demand dedicated verifier agents and termination conditions tied to checked work, not message counts. The quiet revolution: these are software-engineering disciplines applied to conversations — MAS reliability as a testable property, with MAST as the bug taxonomy.
The diagnosis that reframed multi-agent reliability.
| Category | Blame | Example mode | Fix discipline |
|---|---|---|---|
| Specification & design | the builder | goal underspecification | explicit success criteria |
| Inter-agent misalignment | the wiring | role confusion, loops | interface contracts + monitoring |
| Task verification | the audit | unverified output | verifier agents + real termination tests |
| Model capability | rarely the cause | — | — |
The decomposition that redirected multi-agent research from 'better models' to 'better engineering'.
Multi-agent research got its failure science.
Check your understanding of the key concepts from Why MAS Fail.
Everything you need to remember about this paper.