← All Papers Dual Axes Negotiation Lab Benchmark Results Quiz
Benchmark · Game Theory & Cooperation

MultiAgentBench: Evaluating
Collaboration & Competition

Moving beyond static single-agent exams to evaluate LLM agents navigating complex social incentives, strategic deception, bargaining, and competitive game theory.

Focus: Game-Theoretic Multi-Agent Benchmark
Environments: Werewolf, Bargaining, Auction, Collaborative Science
Key Metrics: Surplus Extraction, Deception Detection, Pareto Efficiency

Collaboration vs Competition: The Two Testing Arenas

Single-agent benchmarks like MMLU only measure passive knowledge. Real-world autonomous agents must navigate divergent economic incentives and strategic game play.

Axis 1: Pure & Mixed Collaboration

Shared Payoff Optimization

Agents must pool disjoint partial observations, formulate joint plans, and synchronize actions toward a collective payoff without stepping on each other.

Environments: Minecraft Collaborative Building, Joint Software Debugging, Cross-Discipline Scientific Paper Synthesis.
Axis 2: Strategic Competition & Deception

Zero-Sum & Incomplete Information

Agents must protect private reservation values, detect adversarial bluffing, negotiate surplus splits, and disguise intentions under social deduction games.

Environments: Multi-Round Double Auction, Bilateral Price Bargaining, Werewolf / Mafia (Social Deception).

Bilateral Bargaining Simulation

Observe Buyer and Seller agents negotiate a deal under private reservation values.

BUYER CEILING: $100 | SELLER FLOOR: $60 | BARGAINING SURPLUS: $40

Key Findings across Frontier LLMs

How state-of-the-art models perform when thrown into complex game-theoretic multi-agent environments.

Strategic Bluffing & Information Leakage

Smaller and open-source models (Llama-3-8B, Mistral-7B) frequently leak their private reservation price directly in round 1 ("My budget is $100, but I offer $60"), collapsing their economic surplus to zero.

Frontier Superiority: Models like GPT-4 and Claude 3.5 consistently exploit revealed information, anchoring initial offers favorably and executing strategic concessions.
Werewolf Social Deception Rates
  • Deception Detection: GPT-4 detects werewolf imposters with 68% accuracy vs 24% for random guessing.
  • Role Adherence: Open models often break character under pressure, apologizing or explicitly revealing their hidden card when directly challenged.
  • Pareto Efficiency: Collaborative teams achieve >82% Pareto-optimal resource trades when communication protocols are structured.

MultiAgentBench Quiz

Assess your understanding of competitive game theory and collaborative benchmarking for LLM agents.