One abstraction — the conversable agent — with modes spanning LLMs, humans, and tools. Program their interactions in natural language and code, and complex applications become conversation patterns.
The framework bet: applications are conversations, not graphs.
The design bet: collapse LLMs, humans, tools, and grouped teams into ONE interface — an agent that can receive, reply, and generate messages. Autonomy is then a dial: an AssistantAgent (LLM-powered, autonomous-ish), a UserProxyAgent (executes code, calls functions, or IS a human typing), or custom hybrids. Applications become conversation programs: who talks to whom, in what order, with what termination conditions — specified in natural language (conversation-level behavior) plus ordinary code (control flow). The framework handles message passing, execution plumbing, and checkpointing; the developer handles the interaction design.
The 2023 framework gap: complex LLM apps needed improvisation, not just chains.
A prompt pipeline is a fax chain: each office sends one document forward, no replies. AutoGen is a conference call: any seat can be an AI assistant, a human expert, or a robot that runs code; anyone can reply to anyone; the moderator (your code) decides who speaks next and when the call ends. Applications stop being forms to fill and become conversations to convene.
ConversableAgent — the one-class design and its mode dial.
Two layers of control: natural language — each agent's system prompt defines its conversational behavior (when to ask, when to verify, what to refuse); code — the orchestrating script defines who converses with whom, iteration limits, and termination conditions. The canonical pattern: assistant proposes code → proxy executes it → error messages flow BACK to the assistant → repair loop until done or budget spent. That repair conversation, emergent rather than scripted, is what chain-pipeline architectures could not express.
Three interaction structures that ship with the framework.
Math problem solving (conversational chess), question answering with retrieval-augmented debate, decision-making, and the flagship: code generation with execution-grounded repair — agents whose conversations include real program runs and their failures. The framework's open-source release made it the substrate for thousands of applications and the reference point the failure-analysis study (entry #111) later audited.
AutoGen's success metric was adoption, not a benchmark score.
| Framework axis | Chained pipelines | AutoGen |
|---|---|---|
| Control flow | hand-wired graph | conversation program: NL behavior + code orchestration |
| Human role | bolt-on approval | conversable peer with modes |
| Code execution | external plugin | an agent mode with reply semantics |
| Improvisation | none — scripted | emergent repair loops |
The design contrasts that made conversation-first the mainstream multi-agent architecture.
AutoGen made multi-agent an application pattern, not a research niche.
Check your understanding of the key concepts from AutoGen.
Everything you need to remember about this paper.