History Problem Core Idea Patterns Results Impact Quiz Takeaways
Interactive Paper Explainer

Agents That Talk It Out
AutoGen

One abstraction — the conversable agent — with modes spanning LLMs, humans, and tools. Program their interactions in natural language and code, and complex applications become conversation patterns.

Start Learning Read the Paper ↗
1
Core abstraction
3+
Agent modes
NL + code
Programming
2023
Wu et al. (Microsoft)
History

From Pipelines to Conversations

The framework bet: applications are conversations, not graphs.

2023
Chain-of-prompt frameworks
LangChain-style pipelines chain prompt calls — powerful, rigid: control flow hand-wired, humans bolted on awkwardly.
2023 · spring
Multi-agent experiments
CAMEL role-play (entry #105) and MetaGPT SOPs (entry #106) show conversation-as-computation — as bespoke research systems.
Aug 2023
🚀 AutoGen
Microsoft (Wu et al.): the general framework — conversable agents with customizable modes, interaction programming in NL + code, human-in-the-loop as a first-class mode.
2023-24
The ecosystem standard
AutoGen becomes the most-used multi-agent framework; debate culture (ChatDev lineage), math, and coding apps ship on it.
2025
The reckoning
Why Do Multi-Agent Systems Fail (entry #111) audits exactly this ecosystem — AutoGen's flexibility is the failure surface.
Everything Is a Conversable Agent

The design bet: collapse LLMs, humans, tools, and grouped teams into ONE interface — an agent that can receive, reply, and generate messages. Autonomy is then a dial: an AssistantAgent (LLM-powered, autonomous-ish), a UserProxyAgent (executes code, calls functions, or IS a human typing), or custom hybrids. Applications become conversation programs: who talks to whom, in what order, with what termination conditions — specified in natural language (conversation-level behavior) plus ordinary code (control flow). The framework handles message passing, execution plumbing, and checkpointing; the developer handles the interaction design.

Chapter 01

Pipelines Can't Negotiate

The 2023 framework gap: complex LLM apps needed improvisation, not just chains.

📊
The Rigidity Problem
  • Chained pipelines hard-wire control flow — every branch hand-coded
  • Humans-in-the-loop are awkward bolt-ons, not participants
  • Multi-agent ideas existed as research code — no shared substrate
  • Tool use, code execution, and conversation lived in different abstractions
💬
The AutoGen Answer
  • ConversableAgent: one interface for LLM agents, humans, tool-executors, and teams
  • Conversation programming: NL instructions define agent behavior; code defines orchestration
  • Built-in patterns: two-agent dialogue, group chat (manager-moderated), nested teams
  • Code execution, tool/function calling, and human input as agent MODES, not plugins
Analogy — The Conference Call Abstraction

A prompt pipeline is a fax chain: each office sends one document forward, no replies. AutoGen is a conference call: any seat can be an AI assistant, a human expert, or a robot that runs code; anyone can reply to anyone; the moderator (your code) decides who speaks next and when the call ends. Applications stop being forms to fill and become conversations to convene.

Chapter 02

The Abstraction

ConversableAgent — the one-class design and its mode dial.

The interface
  • Every agent: receive a message → decide to reply (or not) → generate reply
  • Reply generation is pluggable: LLM call, code execution, function call, human input, or custom logic
  • Conversability is symmetric — agents talk TO each other, not just THROUGH a pipeline
The stock agents
  • AssistantAgent: LLM-powered planner/writer — proposes solutions and code
  • UserProxyAgent: the doer — executes code, calls tools, or relays a human's keystrokes
  • GroupChat: a manager agent moderates multi-party conversations
  • Custom agents: any reply logic you can write
Conversation programming

Two layers of control: natural language — each agent's system prompt defines its conversational behavior (when to ask, when to verify, what to refuse); code — the orchestrating script defines who converses with whom, iteration limits, and termination conditions. The canonical pattern: assistant proposes code → proxy executes it → error messages flow BACK to the assistant → repair loop until done or budget spent. That repair conversation, emergent rather than scripted, is what chain-pipeline architectures could not express.

Interactive Demo — The Repair Conversation

Follow a canonical AutoGen coding run — proposal, execution, failure, repair — the loop pipelines couldn't express.

Chapter 03

The Patterns

Three interaction structures that ship with the framework.

1️⃣ Two-agent chat
Assistant ↔ executor: the write-run-repair loop — the workhorse pattern for coding and math applications.
2️⃣ Group chat
A manager agent selects speakers per turn — team conversations with roles, without hand-wired turn-taking.
3️⃣ Nested composition
Teams as agents: a group chat can participate in an outer conversation — hierarchical orchestration.
👤 Human-in-the-loop
A conversable human: approval gates, live steering, or full pairing — the human is a peer, not a bolt-on.
Demonstrated applications (from the paper)

Math problem solving (conversational chess), question answering with retrieval-augmented debate, decision-making, and the flagship: code generation with execution-grounded repair — agents whose conversations include real program runs and their failures. The framework's open-source release made it the substrate for thousands of applications and the reference point the failure-analysis study (entry #111) later audited.

Interactive Demo — One Abstraction, Four Seats

Tab through conversable-agent modes — the same interface wearing different hats.

Chapter 05

The Substrate Standard

AutoGen's success metric was adoption, not a benchmark score.

ABSTRACTION
1 class
ConversableAgent covers LLM/human/tool/team
PROGRAMMING
NL + code
behavior and orchestration split cleanly
PATTERNS
3+
two-agent, group chat, nested + human-in-the-loop
ADOPTION
ecosystem
the most-used multi-agent framework of its era
Interactive Demo — The Flexibility Bill

Conversation-first is powerful — and unstructured. Press reveal for the cost the next paper measured.

Framework axisChained pipelinesAutoGen
Control flowhand-wired graphconversation program: NL behavior + code orchestration
Human rolebolt-on approvalconversable peer with modes
Code executionexternal pluginan agent mode with reply semantics
Improvisationnone — scriptedemergent repair loops

The design contrasts that made conversation-first the mainstream multi-agent architecture.

Legacy

Legacy — The Mainstream Substrate

AutoGen made multi-agent an application pattern, not a research niche.

🏗 The framework standard
Conversation-first multi-agent became the default application architecture — AutoGen's abstractions appear across successor frameworks (AG2, CrewAI lineage concepts).
👤 Human-in-the-loop as peer
Humans as conversable agents legitimized interactive AI workflows — approval gates and live steering became product features, not afterthoughts.
🔁 The execution-repair pattern
Code-running agents whose failures re-enter the conversation became the standard coding-agent loop (OpenHands-era scaffolds descend from it).
⚠️ What it did NOT solve
Conversation freedom without verification discipline produces the failure modes entry #111 cataloged; token costs of chatty agents; and debugging story — reading long transcripts — remains weak.
🛤 Read next
The family: CAMEL · Why Do MAS Fail? · MetaGPT
Test Yourself

Quick Quiz

Check your understanding of the key concepts from AutoGen.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ AutoGen: one abstraction (ConversableAgent) for LLMs, humans, executors, and teams.
✅ Conversation programming: NL defines agent behavior; code defines orchestration.
✅ The execution-repair loop: failed runs re-enter the conversation as messages.
✅ Human-in-the-loop as a conversable peer — approval gates and live steering.
✅ Group chat + nested teams: hierarchy by composition.
✅ Read it as the substrate that mainstreamed multi-agent applications.