A visual, step-by-step guide to the paper that built the first systematic multi-environment benchmark for LLMs as agents — 8 live environments where a model must run commands, query databases, browse the web, and win games, with every action checked against reality.
LLMs became "agents" in the popular imagination almost overnight. AgentBench arrived in August 2023 to find out — measurably — what that actually meant.
Answering a question and acting in an environment are different skills. A chat benchmark grades a single response; an environment grades the whole loop — plan, act, read feedback, adapt, and know when to commit. In mid-2023, no benchmark graded that loop.
By 2023, models aced static QA benchmarks while agent apps built on the very same models looped, misformatted, and lost the plot. The exams were measuring the wrong skill.
A static benchmark is the written driving test: you can ace it and still stall at the first intersection. AgentBench is the road test — the Ubuntu container actually runs your commands, MySQL actually rejects your malformed SQL, the simulated shop actually ships the wrong product. In 2023 the industry was handing out licenses based on the written exam alone.
AgentBench groups its environments into three families — code-grounded, game-grounded, and web-grounded. Each has its own action space, its own feedback, and its own way of deciding "success".
Every environment runs the same contract: the model is an agent inside a partially observable world, acting one round at a time until it succeeds, breaks the format, or runs out of rounds.
Formally, interactive LLM-as-Agent evaluation is a partially observable Markov decision process (POMDP):
Prompts follow the ReAct-style "Thought + Action" pattern in a single round, with a 1-shot CoT example so models learn the format. Decoding is temperature=0 (greedy) for reproducibility.
Each environment demands an exact output pattern — Think/Act blocks, one SQL per code block, actions only from the provided list. The Database prompt literally warns: "If your response cannot match any pattern I mentioned earlier, you will be judged as FAIL immediately." Long histories are trimmed to ~3,500 tokens ("[NOTICE] 2r messages are omitted.").
29 models ran the gauntlet. One topped the table — and a chasm opened between the API frontier and the open-source field of 2023.
| Model | Type | OA | Standout | Weakest |
|---|---|---|---|---|
| gpt-4 (0613) | API | 4.01 | best on 6 of 8 envs · HH 78.0 | LTP 16.6 |
| claude-3-opus | API | 3.11 | DB 51.7 — best of all models | LTP 14.3 |
| glm-4 | API | 2.89 | KG 46.3, WS 61.6 | LTP 14.2 |
| claude-2 | API | 2.49 | DCG 55.5, WS 61.4 | WB 0.0 (!) |
| gpt-3.5-turbo (0613) | API | 2.32 | WS 64.1 — beats GPT-4 here | HH 16.0 |
| text-davinci-003 | API | 1.71 | KG 34.9 | DCG 3.0 |
| codellama-34b (instruct) | OSS | 0.96 | WS 52.1 — best OSS overall | LTP 0.7 |
| llama-2-70b-chat | OSS | 0.78 | DCG 21.3 | LTP 0.0 |
Every API-based LLM scored above 1.00 overall; OSS models averaged 0.51 vs the API average of 2.32. The paper restricts OSS entries to models ≤ 70B. claude-2's 0.0 on Web Browsing is real — it never produced a valid choice on Mind2Web's adapted setup.
GPT-4 was best on 6 of 8 environments and hit a 78% success rate on House Holding — "indicating its practical usability in this scenario," as the paper puts it. But it still failed more than half of Operating System and Database problems, and scored just 16.6 on Lateral Thinking Puzzles. The authors' blunt summary: even the strongest GPT-4 "is not qualified as a practically usable agent."
Open-source models (≤70B) averaged OA 0.51 vs 2.32 for API models — and often scored near zero in the hardest environments (llama-2-70b-chat: LTP 0.0, HH 2.0, WS 5.6). The most capable OSS model, codellama-34b (0.96), still fell far short of gpt-3.5-turbo. Alignment mattered: vicuna-13b, tuned on high-quality ShareGPT conversations, beat llama-2-13b and rivaled a 3× larger codellama-34b. Code tuning was ambivalent — great for procedural Web Shopping, worse for strategic game play.
AgentBench didn't just score models — it dissected every trajectory into five finish reasons: Complete, Invalid Format, Invalid Action, Task Limit Exceeded, and Context Limit Exceeded. Here is what actually breaks.
A real trajectory pattern from the paper's planning case study: the model cycles open → close → examine → open while the environment stops changing. AgentBench judges a task failed after the round limit — or after three identical outputs in a row.
"Put the model in an environment and measure verifiable success" went from one paper's design to an entire evaluation ecosystem.
AgentBench's most important chart isn't the leaderboard — it's the gap between talking and doing. The same GPT-4 that explains database theory fails over half of the actual database tasks. The cliff has a measured shape: strict output formats, multi-step feedback, and environments that punish drifting attention.
Check your understanding of the key concepts from the AgentBench paper.
Everything you need to remember about this paper.