MultiAgentBench: Evaluating
Collaboration & Competition
Moving beyond static single-agent exams to evaluate LLM agents navigating complex social incentives, strategic deception, bargaining, and competitive game theory.
Collaboration vs Competition: The Two Testing Arenas
Single-agent benchmarks like MMLU only measure passive knowledge. Real-world autonomous agents must navigate divergent economic incentives and strategic game play.
Shared Payoff Optimization
Agents must pool disjoint partial observations, formulate joint plans, and synchronize actions toward a collective payoff without stepping on each other.
Zero-Sum & Incomplete Information
Agents must protect private reservation values, detect adversarial bluffing, negotiate surplus splits, and disguise intentions under social deduction games.
Bilateral Bargaining Simulation
Observe Buyer and Seller agents negotiate a deal under private reservation values.
Key Findings across Frontier LLMs
How state-of-the-art models perform when thrown into complex game-theoretic multi-agent environments.
Smaller and open-source models (Llama-3-8B, Mistral-7B) frequently leak their private reservation price directly in round 1 ("My budget is $100, but I offer $60"), collapsing their economic surplus to zero.
- Deception Detection: GPT-4 detects werewolf imposters with 68% accuracy vs 24% for random guessing.
- Role Adherence: Open models often break character under pressure, apologizing or explicitly revealing their hidden card when directly challenged.
- Pareto Efficiency: Collaborative teams achieve >82% Pareto-optimal resource trades when communication protocols are structured.
MultiAgentBench Quiz
Assess your understanding of competitive game theory and collaborative benchmarking for LLM agents.