History Problem Environments Protocol Results Failure Impact Deep Dive Quiz
Interactive Paper Explainer

From Answers to Actions
AgentBench

A visual, step-by-step guide to the paper that built the first systematic multi-environment benchmark for LLMs as agents — 8 live environments where a model must run commands, query databases, browse the web, and win games, with every action checked against reality.

Start Learning Read the Paper ↗
8
Interactive Environments
29
LLMs Evaluated
78%
GPT-4's Best Success Rate
2023
Year Released
History

From Answering to Acting

LLMs became "agents" in the popular imagination almost overnight. AgentBench arrived in August 2023 to find out — measurably — what that actually meant.

2020 · May
GPT-3 (Brown et al.)
175B parameters, few-shot question answering. Astonishing at answering — but never asked to act, observe feedback, or recover from a mistake.
2022 · Oct
ReAct (Yao et al.)
Interleaves reasoning traces with actions in one loop — the "Think then Act" template AgentBench later adopts inside every environment.
2023 · Feb
Toolformer (Schick et al.)
LLMs teach themselves to call APIs mid-sentence. Tools enter the picture — and tool-using agents need tool-based exams.
2023 · Mar
The AutoGPT wave
AutoGPT, BabyAGI and AgentGPT go viral: autonomous agents promise everything — and loop, stall, and hallucinate in public. Hype outruns measurement.
2023 · Aug
🚀 AgentBench (Liu et al.)
8 live environments, 29 LLMs, verifiable success criteria — the first systematic benchmark for LLMs as agents (arXiv 2308.03688).
2024 →
Agent evaluation becomes a field
Published at ICLR 2024. GAIA, SWE-bench, OSWorld, AgentDojo and successors turn "measure the agent" into a research area of its own.
Key Insight

Answering a question and acting in an environment are different skills. A chat benchmark grades a single response; an environment grades the whole loop — plan, act, read feedback, adapt, and know when to commit. In mid-2023, no benchmark graded that loop.

TWO DIFFERENT EXAMS
static QA  : question ──► answer ✔/✘
AgentBench: goal ──► act ──► feedback ──► act ──► … ──► success?
The agent only passes if the environment's checking script agrees.
Chapter 01

The Problem — Grading Answers, Not Actions

By 2023, models aced static QA benchmarks while agent apps built on the very same models looped, misformatted, and lost the plot. The exams were measuring the wrong skill.

📝
Static QA Benchmarks
  • One question in, one answer out — zero interaction rounds
  • No environment feedback to read, parse, or recover from
  • No consequence: nothing checks whether the plan actually executed
  • Long-horizon planning and memory are never exercised
  • Fluent text can hide a completely broken agent
🕹️
AgentBench's Solution
  • 8 live environments that respond to every action
  • Multi-round interaction: 5–50 rounds to solve one problem
  • Verifiable success — checking scripts, exact answers, hashed DB states
  • Spans the terminal, SQL, knowledge graphs, games, and the web
  • Per-environment success metrics + one weighted overall score
The Analogy — Written Test vs Road Test

A static benchmark is the written driving test: you can ace it and still stall at the first intersection. AgentBench is the road test — the Ubuntu container actually runs your commands, MySQL actually rejects your malformed SQL, the simulated shop actually ships the wrong product. In 2023 the industry was handing out licenses based on the written exam alone.

written test: "What does this sign mean?" ──► pick A/B/C/D
road test  : "Park here." ──► signal ──► mirror ──► steer ──► curb check ──► pass/fail
Chapter 02

Eight Worlds to Act In

AgentBench groups its environments into three families — code-grounded, game-grounded, and web-grounded. Each has its own action space, its own feedback, and its own way of deciding "success".

🖥️
Operating System
Ubuntu Docker · real bash
Code · SR
🗄️
Database
MySQL · SELECT / INSERT / UPDATE
Code · SR
🧠
Knowledge Graph
Freebase · 5 query tools
Code · F1
🐟
Digital Card Game
Aquawar · turn-based fish battle
Game · Reward
🐢
Lateral Thinking Puzzles
riddles via yes/no questions
Game · Progress
🏠
House Holding
ALFWorld · text household tasks
Game · SR
🛒
Web Shopping
WebShop · ~1M real products
Web · Reward
🌐
Web Browsing
Mind2Web · click / type / select
Web · Step SR
Interactive Demo — Environment Atlas

Pick an environment to see what it tests, a sample of what an interaction actually looks like, and how the strongest API model compares with a famous open one. All scores are the paper's real test-set results (Table 3, standard setting).

Chapter 03

The Protocol — One Loop, Strict Rules

Every environment runs the same contract: the model is an agent inside a partially observable world, acting one round at a time until it succeeds, breaks the format, or runs out of rounds.

AgentBench task = ( 𝒮, 𝒜, 𝒯, ℛ, 𝒰, 𝒪 )

Formally, interactive LLM-as-Agent evaluation is a partially observable Markov decision process (POMDP):

𝒮
State space
Everything the environment can be — files, tables, pages, game boards.
𝒜
Action space
What the agent may emit: bash, SQL, tool calls, clicks, game moves.
𝒯
Transition
How actions change state — the environment executes and moves on.
ℛ
Reward
The checking script: success rate, F1, win rate, or shopping reward.
𝒰
Instruction space
The natural-language goal: "count the files in /etc", "buy blue shoes".
𝒪
Observation space
What the agent gets back — often partial, truncated, or noisy.
The Interaction Loop
<USER>  instruction + last environment feedback
<AGENT> Think: plan  ·  Act: one action
   ↓ environment executes the action
   ↓ observation returns → next round
loop until: Complete ✔ · round limit · invalid output

Prompts follow the ReAct-style "Thought + Action" pattern in a single round, with a 1-shot CoT example so models learn the format. Decoding is temperature=0 (greedy) for reproducibility.

Formatting Is Part of the Exam
Act: answer(220)  ✔ valid
"the answer is 220"    ✘ invalid format
"…you will be judged as FAIL immediately."

Each environment demands an exact output pattern — Think/Act blocks, one SQL per code block, actions only from the provided list. The Database prompt literally warns: "If your response cannot match any pattern I mentioned earlier, you will be judged as FAIL immediately." Long histories are trimmed to ~3,500 tokens ("[NOTICE] 2r messages are omitted.").

AVG ROUNDS / PROBLEM
5–35
Web Shopping 5 · OS 8 · House Holding 35 (solving rounds 5–50)
PROBLEMS
1,283
269 dev + 1,014 test ≈ 11k inference calls — MMLU-scale
DECODING
temp 0
greedy decoding on all tasks, for reproducible runs
CONTEXT BUDGET
~3,500
tokens of trimmed interaction history per inference
COMPLETED RUNS
6.0
median rounds in successful trajectories (median 1,850 tokens)
OVERALL SCORE
OA
each environment's mean resized to 1, then weighted average
Interactive Demo — Be the Agent (Centerpiece)

You are the LLM. This is a real Operating System task from the paper's appendix — with its real round limit of 8. Pick an action at each step; wrong picks burn a round and trigger classic AgentBench failure modes. Can you finish before the limit?

Round 1 / 8 · Step 1 of 3
Chapter 04

The Results — A Clear Pecking Order

29 models ran the gauntlet. One topped the table — and a chasm opened between the API frontier and the open-source field of 2023.

OPERATING SYSTEM
42.4
GPT-4 success rate (%)
GPT-3.5: 32.6 · best OSS: 2.8
DATABASE
32.0
GPT-4 success rate (%)
GPT-3.5: 36.7 · best OSS: 14.0
KNOWLEDGE GRAPH
58.8
GPT-4 answer F1
GPT-3.5: 25.9 · best OSS: 23.5
DIGITAL CARD GAME
74.5
GPT-4 reward score
GPT-3.5: 33.7 · best OSS: 8.4
LATERAL THINKING
16.6
GPT-4 game progress
GPT-3.5: 10.5 · best OSS: 0.7
HOUSE HOLDING
78.0
GPT-4 success rate (%)
GPT-3.5: 16.0 · best OSS: 4.0
WEB SHOPPING
61.1
GPT-4 shopping reward
GPT-3.5: 64.1 · best OSS: 52.1
WEB BROWSING
29.0
GPT-4 step success (%)
GPT-3.5: 20.0 · best OSS: 20.0
The API-vs-Open Gap (Table 3, test set — "OA" = weighted overall AgentBench score)
ModelTypeOAStandoutWeakest
gpt-4 (0613)API4.01best on 6 of 8 envs · HH 78.0LTP 16.6
claude-3-opusAPI3.11DB 51.7 — best of all modelsLTP 14.3
glm-4API2.89KG 46.3, WS 61.6LTP 14.2
claude-2API2.49DCG 55.5, WS 61.4WB 0.0 (!)
gpt-3.5-turbo (0613)API2.32WS 64.1 — beats GPT-4 hereHH 16.0
text-davinci-003API1.71KG 34.9DCG 3.0
codellama-34b (instruct)OSS0.96WS 52.1 — best OSS overallLTP 0.7
llama-2-70b-chatOSS0.78DCG 21.3LTP 0.0

Every API-based LLM scored above 1.00 overall; OSS models averaged 0.51 vs the API average of 2.32. The paper restricts OSS entries to models ≤ 70B. claude-2's 0.0 on Web Browsing is real — it never produced a valid choice on Mind2Web's adapted setup.

Finding — Strong at the Top, Not Enough

GPT-4 was best on 6 of 8 environments and hit a 78% success rate on House Holding — "indicating its practical usability in this scenario," as the paper puts it. But it still failed more than half of Operating System and Database problems, and scored just 16.6 on Lateral Thinking Puzzles. The authors' blunt summary: even the strongest GPT-4 "is not qualified as a practically usable agent."

Finding — Open Models Were Nowhere Close

Open-source models (≤70B) averaged OA 0.51 vs 2.32 for API models — and often scored near zero in the hardest environments (llama-2-70b-chat: LTP 0.0, HH 2.0, WS 5.6). The most capable OSS model, codellama-34b (0.96), still fell far short of gpt-3.5-turbo. Alignment mattered: vicuna-13b, tuned on high-quality ShareGPT conversations, beat llama-2-13b and rivaled a 3× larger codellama-34b. Code tuning was ambivalent — great for procedural Web Shopping, worse for strategic game play.

Interactive Demo — The Gap

Pick an environment (or the overall score), press Run, and watch the 2023 gap between API frontier models and open models animate into view. Real Table 3 numbers, standard test setting.

Chapter 05

How Agents Fail

AgentBench didn't just score models — it dissected every trajectory into five finish reasons: Complete, Invalid Format, Invalid Action, Task Limit Exceeded, and Context Limit Exceeded. Here is what actually breaks.

🔁 Repetition loops (TLE)
The #1 failure cause. Over 90% of Task-Limit-Exceeded trajectories contain near-identical repeated responses (Rouge-L ≥ 0.8 in the last 10 rounds), averaging 25.5 rounds before being cut off.
📋 Invalid Format
Breaking the strict output pattern: 53.3% of Database outcomes, 38.5% of Digital Card Game, 17.2% of Web Shopping. Instant FAIL in strict environments.
❌ Invalid Action
Well-formed but outside the allowed action space: 64.1% of House Holding outcomes, 10.2% of DCG, 8.4% of Web Browsing — the "Nothing happened" trap.
🧭 Goal drift
The paper's case study: gpt-3.5-turbo decomposed House Holding tasks fine, but "gradually lost sight of the original plan" as failed attempts piled up.
🩹 Feedback blindness
Models that can't read their own errors. In Database, models that self-correct their SQL after error messages "significantly outscore others" — claude-2 is the paper's showcase self-corrector.
📜 Context Limit Exceeded
Interaction history overflows the context window (only the 2,048-token text-davinci-002/003 models). Rare — 0–3.5% of outcomes — but always fatal; AgentBench's history-trimming strategy exists to contain it.
A Real Loop, From the Paper (gpt-3.5-turbo, House Holding)
task: put a clean soapbar in countertop
agent: THOUGHT: …examine cabinet 2 for a clean soapbar.
       ACTION: examine cabinet 2
user:  The cabinet 1 is closed.
agent: THOUGHT: …try examining cabinet 1 again…
       ACTION: examine cabinet 1
user:  The cabinet 1 is closed.
agent: THOUGHT: …try opening cabinet 1 again…
↻ the same 4-state cycle repeats…

A real trajectory pattern from the paper's planning case study: the model cycles open → close → examine → open while the environment stops changing. AgentBench judges a task failed after the round limit — or after three identical outputs in a row.

What AgentBench Did NOT Solve
  • Primitive strategies only: plain CoT, no reflection, search, or retries — so scores are a floor, not a ceiling, for clever agent frameworks.
  • Text surrogates: the OS is a Docker sandbox, the web is HTML-as-text; real desktops and live sites stayed out of reach (OSWorld's 2024 job).
  • Single agent, single environment: no multi-agent cooperation, no cross-environment tool transfer.
  • A 2023 snapshot: gpt-4-0613-era results; API models keep changing, and their training data is unknowable.
  • Narrow open-model range: only OSS models ≤ 70B were tested — the biggest open weights never ran.
Legacy

Impact — Agent Evaluation Becomes a Field

"Put the model in an environment and measure verifiable success" went from one paper's design to an entire evaluation ecosystem.

🧪 An evaluation field emerges
GAIA (2023), SWE-bench (2023), OSWorld (2024), AgentDojo (2024) and more followed the same recipe: live environments, real tasks, verifiable outcomes — agent testing became its own research area.
🏛️ The multi-environment pattern
Eight diverse worlds — code, game, web — under one protocol and one weighted score became the design template for judging general-purpose agents instead of single-task specialists.
📦 Reproducible infrastructure
Docker-encapsulated, environment-isolated tasks plus an API-centric toolkit, released open-source (THUDM/AgentBench) — anyone can stand up the exact same eight exams.
📈 Training signal for open models
The 0.51-vs-2.32 gap and the diagnosis (instruction following, multi-round alignment data) gave open-model teams a concrete target; agentic ability became a fine-tuning objective, not an afterthought.
🔍 Failure taxonomy that stuck
Invalid Format / Invalid Action / Task Limit / Context Limit gave the community a shared vocabulary for how agents break — repetition loops and format errors are still the first things practitioners check.
💡 "Agents ≠ chatbots"
The core lesson of the numbers: fluent conversation does not imply competent action. Acting well needs instruction following, long-horizon planning, memory, and feedback parsing — a different skill profile than static QA.
Deep Dive

The Chat-to-Agent Cliff

AgentBench's most important chart isn't the leaderboard — it's the gap between talking and doing. The same GPT-4 that explains database theory fails over half of the actual database tasks. The cliff has a measured shape: strict output formats, multi-step feedback, and environments that punish drifting attention.

🏙️
Eight Worlds, One Protocol
  • OS, Database, Knowledge Graph, Card Game, Lateral Puzzles, ALFWorld households, WebShop, Mind2Web browsing
  • Every task a POMDP loop: Think → Act → environment feedback, with round limits and verifiable success
  • GPT-4 topped the table — overall 4.01, best on 6 of 8 worlds, 78% on House Holding
  • API models averaged 2.32 vs 0.51 for open models: the capability gap was the story of 2023
🪨
Failure Modes With Names
  • Invalid Format — 53.3% of DB failures: the SQL was right-shaped, the wrapper was wrong
  • Invalid Action — 64.1% of House Holding failures: legal-looking actions that the environment rejects
  • Repetition loops — over 90% of time-limit failures: the agent re-issues the same failing move
  • Goal drift and feedback blindness: instructions lose to momentum as turns accumulate
Interactive Demo — Failure-Mode Autopsy

A real-shaped Database task, replayed as the paper's failure taxonomy. Step through the loop and watch a strong LLM die by a thousand format errors — then replay with the same model under format-repair prompting and watch the failure class change, not vanish.

TASK · DATABASE WORLD · "list the titles of all orders placed by customer 42, sorted by date"
failure class counters live at the end of each run
VERDICT
Strongest ≠ usable agent
A model can be the best in the world at describing SQL and still fail the task because it wrapped the query in the wrong envelope — three times in a row. The paper's prescription became the field's to-do list: format-constrained decoding, loop detection, feedback-grounded retries. The systems that later passed these worlds — see Magentic-One — are exactly this list, engineered. Safety in the same worlds is a different axis again: Agent-SafetyBench measures what agents do wrong on purpose.
📮 Format is a skill
Over half of DB failures were envelope errors, not SQL errors. Instruction-following under strict schemas is a distinct capability from knowing the content — later solved by constrained decoding, not by smarter SQL.
🔁 The repetition trap
When an action fails, the most likely next action is… the same action. Over 90% of time-limit failures were loops — evidence that agents need explicit loop breakers, a component every modern orchestrator now ships.
🧩 POMDP discipline
Framing every environment as a partially-observable loop made results comparable across worlds and made failure modes countable — the benchmark's methodological gift.
📉 2.32 vs 0.51
The API-vs-open gap on overall score was the raw material for a year of open-model agent work — closing it became the explicit target of the open-source agent stacks of 2024.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the AgentBench paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ AgentBench (ICLR 2024) is the first systematic multi-environment benchmark for LLMs as agents: 8 interactive environments, 29 LLMs.
✅ The 8 worlds: Operating System, Database, Knowledge Graph, Digital Card Game, Lateral Thinking Puzzles, House Holding (ALFWorld), Web Shopping (WebShop), Web Browsing (Mind2Web).
✅ Every task is a POMDP loop — Think + Act, environment feedback, strict formats, verifiable success (SR / F1 / reward), round limits.
✅ GPT-4 topped the table (OA 4.01, best on 6 of 8, 78% on House Holding) — yet still failed over half of OS/DB problems. Strongest ≠ usable agent.
✅ The gap: API models averaged OA 2.32 vs 0.51 for open models; best OSS (codellama-34b, 0.96) trailed gpt-3.5-turbo (2.32) — while some scored near zero on KG/DCG/HH.
✅ Failure modes have names: repetition loops (>90% of TLE), Invalid Format (53.3% of DB), Invalid Action (64.1% of House Holding), goal drift, feedback blindness.