A visual, step-by-step guide to the benchmark paper that drops agents into real terminal containers — compile, debug, serve, recover — with hidden tests and a time bar set by human experts. Covers the benchmark through its current 2.0 release: 89 curated tasks.
The industry's most important interface — the command line — spent years outside agent benchmarks. Terminal-Bench put it back at the center.
The terminal is the actual operating layer of computing: compiling code, training models, setting up servers, editing configs, recovering broken systems. It is unstructured (no friendly buttons), stateful (previous commands matter), and verifiable (did the server actually start?). Any agent that truly "does computer work" must pass through here.
Agents were acing stylized environments while remaining unable to do an afternoon of real sysadmin work.
Most agent benchmarks are an IKEA app: the parts are pre-sorted, the steps are numbered, and the app congratulates you. The terminal is a workshop with the lights half-off: a goal ("make the CNC machine cut this part"), scattered tools, and a finished-part gauge nobody lets you see until you claim you're done.
Every one of the 89 tasks ships as a self-contained container with three artifacts: an instruction, a human-written solution, and a test suite.
"Hard, realistic" is enforced by curation: tasks are kept only when they resist trivial solution and reflect genuine workflow patterns. The result: frontier models and agents score below 65% — a benchmark that still separates the field years after release, and whose scores improved as much through harness engineering as through model releases.
What the agent actually experiences: a prompt, a shell, and no hints.
Verification is outcome-based, hidden, and binary — the same anti-gaming lineage as SWE-bench's test grading, transplanted to general system tasks.
Each task carries a human-expert solve time from its authoring process. This grounds difficulty in effort units rather than model-percentile units: "tasks that take a professional about an hour" is a stable yardstick that survives model releases — the same move as SWE-Bench Pro's hours-to-days calibration.
The paper's headline: frontier models and agents score less than 65% on the benchmark — followed by a failure taxonomy.
Terminal-Bench became the field's proxy for "can agents do real computer work?" — and its leaderboard taught the community about scaffolds.
Terminal-Bench 2.0's leaderboard made something undeniable: measured capability lives in the model-harness-environment triad, and "the model's score" is a category error.
Terminal-Bench 2.0's 89 containers quietly reset the field's unit of analysis. When the same weights swing tens of points with scaffolding and strategy, "model capability" stops being a scalar. The benchmark's biggest export isn't a leaderboard — it's the discipline of reading agent claims as claims about systems: which harness, how many runs, what recovery logic, and what the hidden tests actually measured.
Check your understanding of the key concepts from the Terminal-Bench 2.0 paper.
Everything you need to remember about this paper.