History Problem Core Idea Doctrine Results Impact Quiz Takeaways
Interactive Paper Explainer

The Interface Is the Capability
SWE-agent

Same model, same problems, same benchmark — 3.8% solved. Design the tools around the model's needs (an Agent-Computer Interface), and the identical setup solves 12.5%. Interface engineering as capability engineering.

Start Learning Read the Paper ↗
12.5%
SWE-bench resolved (vs 3.8%)
4
Core ACI tools
~92%
Human level reference
2024
Yang et al. (Princeton)
History

Tools Were Scaffolding

The quiet assumption: interfaces are plumbing. This paper measured them as the product.

2023
SWE-bench arrives
Jimenez et al. (entry #76) publish real GitHub-issue resolution: models alone solve a few percent — the gap is framed as a model problem.
2023
ReAct-era tool use
Agents get tool access — but tools are designed for humans: full-file reads, terminal firehoses, cryptic errors (entries #53, #54).
May 2024
🚀 SWE-agent
Yang et al.: treat LMs as a new user class with distinct needs — design the file viewer, editor, search, and actions around LM failure modes. The ACI thesis, measured.
2024+
Interface engineering era
SWE-agent-style interfaces (linting editors, bounded views, clean errors) become standard scaffolding for OpenHands, Devin-class products, and the coding-agent industry.
LMs Are a New User Class

The paper's framing: humans needed IDEs — bounded file views, syntax-aware editing, searchable context, readable errors — because raw terminals exceed human cognition. LMs have the same problem, differently parameterized: context windows that drown in 2,000-line files, no proprioception for cursor state, catastrophic edits that humans would never make (breaking 20 functions at once). The Agent-Computer Interface (ACI) is the IDE re-designed for this user: browsable 100-line windows, guarded edits with lint checks, bounded search, minimal-action grammar — plus guardrails (lint-test before submission, confirm-file-exists). The measurement: an agent with a designed ACI solves 12.5% of SWE-bench; the same agent without it, 3.8%.

Chapter 01

Tools Built for Somebody Else

The default: give models human interfaces, then blame the model.

🧰
The Mismatched Toolkit
  • cat dumps 2,000 lines into a context built for 8k tokens — attention drowns
  • sed/regex editing: one bad pattern silently breaks a whole file — no lint, no preview
  • Terminal errors arrive as walls of stack trace — the actionable line buried
  • Humans compensate with IDEs and habits; agents inherit the raw tools and the failure modes
🎛
The ACI Answer
  • Design tools around documented LM failure modes — the agent as a new end-user class
  • Bounded views: 100-line file windows with on-demand scrolling/search
  • Guarded edits: edit-with-line-search, automatic linter pass, edit-window context
  • Minimal grammar + guardrails: few actions, every action confirmable and reversible-looking
  • Measured: 3.8% → 12.5% on SWE-bench — interface design as capability
Analogy — The Left-Handed Keyboard

An interface built for a different user isn't neutral — it's a tax. Type on a keyboard designed for someone else's hand shape and every keystroke costs accuracy; hand the same person a keyboard molded to THEIR hands and speed appears 'out of nowhere'. SWE-agent's claim: 9 points of SWE-bench accuracy were hiding in the keyboard, not the typist.

Chapter 02

The ACI Design Rules

Four principles, each mapping to a measured failure mode.

1️⃣ Bounded observation
Files open in ~100-line windows; search returns bounded snippets. The agent sees a desk, not the warehouse.
2️⃣ Guarded actions
Edits require line-level targeting; a linter runs on every change — broken code is caught at the moment of creation, not at test time.
3️⃣ Minimal grammar
A handful of composite actions (open, search, edit, submit) instead of the full shell space — fewer degrees of freedom, fewer catastrophic moves.
4️⃣ Feedback loops
Every action returns an informative, compact result — errors the agent can act on immediately. Interface as conversation, not firehose.
The agent loop
  • Issue prompt → agent explores: search, open windows, read context
  • Localize the fault → guarded edit → linter feedback → iterate
  • Run tests when confident → submit a diff (git-style patch)
  • Everything through the 4-tool ACI — no raw shell
The measured result
  • GPT-4 + SWE-agent ACI: 12.5% of SWE-bench issues resolved
  • Same model, no custom interface: 3.8% — the interface delta isolated
  • Ablations: removing lint-on-edit or bounded views degrades measurably — every rule earns its place
  • Approaches the human-engineer resolution reference on the same benchmark slice
Interactive Demo — Same Edit, Two Interfaces

Tab through one buggy-file repair under the raw terminal and the ACI — watch the failure modes flip.

Chapter 03

The Thesis Generalized

Why ACI outgrew SWE-bench — the idea became a discipline.

From Scaffold to Doctrine

The ACI concept reframed agent engineering: the tool layer is a design surface with measurable capability consequences — not plumbing, not a constant. It explains WebArena's brittleness (entry #74), OSWorld's grounding wars (entry #78), and the harness effects of Terminal-Bench (entry #87): when the same model scores wildly differently across interfaces, "model capability" was never the right unit. Post-SWE-agent, every serious agent framework ships a designed ACI — and interface ablations became standard practice.

Interactive Demo — The Agent Loop, Through the ACI

Follow one issue from prompt to submitted patch — every step through the four-tool interface.

Chapter 05

3.8 → 12.5: Interface Points

The ablation that made interface engineering a discipline.

NO CUSTOM INTERFACE
3.8%
same model, same benchmark
WITH SWE-AGENT ACI
12.5%
+8.7 points from interface design
ACI PRINCIPLES
4
bounded views · guarded edits · minimal grammar · feedback
TOOLSET SIZE
4 tools
search · view · edit · submit
Interactive Demo — The Interface Ablation

Press run for the isolated interface effect — the same model with and without a designed tool layer.

Interface elementRaw terminal (default)SWE-agent ACI
File viewingcat (whole file)100-line windows + scroll/search
Editingsed/regex, blindtargeted edit + lint-on-change
Searchgrep firehosebounded, ranked snippets
Action spacefull shellminimal composite grammar
SWE-bench resolved3.8%12.5%

Each row is a design decision; the last row is their measured, isolated sum.

Legacy

Legacy — Interface Engineering

ACI became a permanent word in the agent vocabulary — and its measured doctrine.

🧰 The coding-agent skeleton
Lint-guarded editors, bounded viewers, and minimal action grammars shipped into OpenHands, Devin-class products, and every serious SWE scaffold — SWE-agent's tools, industrialized.
📐 The measurement methodology
Interface ablations (same model, different tools) became the standard way to attribute agent performance — the paper's experimental template.
🧠 The 'new user class' frame
Designing for LM cognition as a distinct user profile spread beyond code — computer-use agents (entry #78) and harness engineering (entry #87) are ACI thinking on other surfaces.
⚠️ What it did NOT solve
12.5% is still far from human-level on the full benchmark; the ACI is hand-designed (the paper open-questions automated ACI design); and interface gains can mask model limits — a designed crutch is still a crutch.
🛤 Read next
Test Yourself

Quick Quiz

Check your understanding of the key concepts from SWE-agent.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ ACI = tools designed for LM cognition as a new user class: bounded, guarded, minimal, informative.
✅ The measured delta: 3.8% → 12.5% SWE-bench resolution from interface design alone.
✅ Core tools: search, 100-line viewer, lint-guarded edit, submit — no raw shell.
✅ Every rule maps to a documented LM failure mode (context drowning, silent breakage).
✅ Interface ablation became the standard attribution method for agent performance.
✅ Read it as the moment scaffolding became product engineering.