History Problem Core Idea Tasks Verification Results Impact Deep Dive Quiz
Interactive Paper Explainer

The Command Line Is the Final Exam
Terminal-Bench

A visual, step-by-step guide to the benchmark paper that drops agents into real terminal containers — compile, debug, serve, recover — with hidden tests and a time bar set by human experts. Covers the benchmark through its current 2.0 release: 89 curated tasks.

Start Learning Read the Paper ↗
89
Tasks (2.0 release)
<65%
Best Frontier Score
1 env + tests
Per Task
2025
Benchmark Released
History

The Terminal Comes Back

The industry's most important interface — the command line — spent years outside agent benchmarks. Terminal-Bench put it back at the center.

2023
Agent benchmarks go web-and-API
Shopping, browsing, tool-calling — polished interfaces, scripted environments, narrow action spaces.
2025 · Sep
Terminal-Bench
Laude Institute + Stanford: agents in real containers, graded by hidden tests — v1.0 of a general-purpose terminal eval.
2025 · Q4
🚀 Terminal-Bench 2.0
89 carefully curated hard tasks, each with a unique environment, human-written solution, and comprehensive tests; frontier agents land under 65%.
2026 · Jan
The formal paper
Merrill, Shaw, Carlini, et al. publish the benchmark paper with error analysis — plus a leaderboard ecosystem (Harbor harness) that revealed how much scaffolds, not models, drive scores.
Why the Terminal?

The terminal is the actual operating layer of computing: compiling code, training models, setting up servers, editing configs, recovering broken systems. It is unstructured (no friendly buttons), stateful (previous commands matter), and verifiable (did the server actually start?). Any agent that truly "does computer work" must pass through here.

🧭 Benchmark family
Repo bugs: SWE-bench · enterprise scale: SWE-Bench Pro · general computer work: this page.
Chapter 01

App-Sandbox Success Theater

Agents were acing stylized environments while remaining unable to do an afternoon of real sysadmin work.

🖱
The Stylized-Environment Trap
  • Curated action spaces (click buttons, call listed APIs) reward following instructions, not solving problems
  • Web-shopping tasks are shallow: few steps, gentle feedback, no system state to reason about
  • High benchmark scores coexisted with agents that can't compile a project or debug a daemon
  • Some benchmarks became easy enough that "hard" meant only "prompt is long"
⌨
The Terminal Answer
  • A real shell: thousands of possible commands, no action-space guardrails
  • Containerized tasks with genuine system state — packages, services, files, processes
  • Problems drawn from real workflows: build systems, model training, server setup, recovery
  • Binary, hidden verification: tests run after the agent says "done"
Analogy — The Workshop vs The IKEA App

Most agent benchmarks are an IKEA app: the parts are pre-sorted, the steps are numbered, and the app congratulates you. The terminal is a workshop with the lights half-off: a goal ("make the CNC machine cut this part"), scattered tools, and a finished-part gauge nobody lets you see until you claim you're done.

Chapter 02

Anatomy of a Task

Every one of the 89 tasks ships as a self-contained container with three artifacts: an instruction, a human-written solution, and a test suite.

📄 Instruction
Natural-language goal in English — real-workflow flavored, deliberately not a checklist.
🧑 Human solution
A worked reference path written by the task author — the calibration anchor and study material.
🧪 Comprehensive tests
Hidden verification scripts checking outcomes, not process — the agent never sees them while working.
🐳 Unique environment
Each task runs in its own container: deps pinned, files staged, services pre-installed in a broken state.
Task Categories (flavor, not silos)
  • Build & compile: get a broken build pipeline producing correct artifacts
  • Debug & recover: diagnose failing services, corrupted configs, dead processes
  • Deploy & serve: stand up servers and daemons that survive verification probes
  • Data & training: run or repair small model-training and data-processing jobs
  • Systems & tooling: cron, permissions, packaging, cross-tool plumbing
The Hardness Bar

"Hard, realistic" is enforced by curation: tasks are kept only when they resist trivial solution and reflect genuine workflow patterns. The result: frontier models and agents score below 65% — a benchmark that still separates the field years after release, and whose scores improved as much through harness engineering as through model releases.

Chapter 03

A Task From the Inside

What the agent actually experiences: a prompt, a shell, and no hints.

Interactive Demo — Agent Terminal Session

Follow an agent working a server-recovery task: state, commands, and the final hidden verification.

Chapter 04

Grading You Can't Sweet-Talk

Verification is outcome-based, hidden, and binary — the same anti-gaming lineage as SWE-bench's test grading, transplanted to general system tasks.

The Verification Contract
  • Tests execute after the agent's session ends — no feedback loop to exploit
  • Checks probe the outcome: is the service up, is the artifact correct, does the training run complete
  • Many tasks include negative checks — the agent must not have broken unrelated systems
  • Binary scoring per task; leaderboard statistics aggregate over runs
The Human Time Bar

Each task carries a human-expert solve time from its authoring process. This grounds difficulty in effort units rather than model-percentile units: "tasks that take a professional about an hour" is a stable yardstick that survives model releases — the same move as SWE-Bench Pro's hours-to-days calibration.

Interactive Demo — Outcome Probe

Two agents claim "the server is fixed." Run the hidden probe on each — text is identical, worlds differ.

Chapter 05

The Under-65% Ceiling

The paper's headline: frontier models and agents score less than 65% on the benchmark — followed by a failure taxonomy.

FRONTIER SCORE
<65%
best models and agents, on 89 hard tasks
TASK COUNT
89
curated for hardness and realism — quality over volume
ERROR ANALYSIS
taxonomy
the paper categorizes where models and agents fail — a roadmap for improvement
HARNESS EFFECT
large
the same model scores very differently under different scaffolds — evaluation is model×harness
Interactive Demo — Model × Harness Grid

Same three models under three harness tiers. Click a cell to see why harness engineering became its own discipline.

Legacy

Impact — The General-Work Benchmark

Terminal-Bench became the field's proxy for "can agents do real computer work?" — and its leaderboard taught the community about scaffolds.

🏆 The leaderboard era
Frontier labs now report TB scores alongside SWE-bench — with run-count and harness disclosure increasingly expected.
🧰 Harness science
Big score swings from scaffolds alone (planning loops, recovery strategies, context management) legitimized harness engineering as a research area.
📦 Containerized task standard
Instruction + solution + tests + Dockerfile became a reusable task format far beyond this benchmark.
🧪 Error-analysis culture
The paper's failure taxonomy — not just the score — guides agent development.
⚠️ What it did NOT solve
89 tasks is statistically thin; binary outcomes hide partial progress; containerization can't capture GUI work or long-running multi-day ops.
🧭 Read next
Reliability over time: τ-bench — different axis, same discipline.
Deep Dive

The Agent Is a System, Not a Model

Terminal-Bench 2.0's leaderboard made something undeniable: measured capability lives in the model-harness-environment triad, and "the model's score" is a category error.

🏷
The Model-Credit Habit
  • Leaderboards print MODEL scores — so harness work is invisible in the number
  • A model + a better scaffold can leapfrog a better model — and often did
  • Scores vary with run count, context strategy, and recovery policy — rarely disclosed
  • Capability claims become unaudable marketing
🧩
The System-Level Read
  • Report model × harness × runs — the full configuration is the subject
  • Hard benchmarks with hidden tests make scaffold gains legitimate (not gaming) and visible
  • Error taxonomies localize: model failures vs harness failures vs task ambiguity
  • Evaluation matured the same way software did: blame the system, study the system
Interactive Demo — Score Anatomy

One headline number, decomposed into model capability, harness contribution, and run-count luck. Press decompose.

Verdict

Terminal-Bench 2.0's 89 containers quietly reset the field's unit of analysis. When the same weights swing tens of points with scaffolding and strategy, "model capability" stops being a scalar. The benchmark's biggest export isn't a leaderboard — it's the discipline of reading agent claims as claims about systems: which harness, how many runs, what recovery logic, and what the hidden tests actually measured.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the Terminal-Bench 2.0 paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 89 hard, realistic tasks in command-line environments, inspired by real workflows.
✅ Each task: unique containerized environment + human-written solution + comprehensive hidden tests.
✅ Frontier models and agents score less than 65% — a durable ceiling with room to separate systems.
✅ Verification is outcome-based and hidden until the agent finishes — no feedback loop to exploit.
✅ Harnesses matter enormously: the same model varies widely across scaffolds and strategies.
✅ Read agent claims as system claims: model × harness × run policy — the benchmark's deepest lesson.