A visual, step-by-step guide to the benchmark that turned software engineering into an evaluation: 2,294 real GitHub issues across 12 Python repositories, where a model wins only if the repository's own tests pass after its patch.
Code benchmarks had been self-contained puzzles. SWE-bench moved the exam into a live repository.
A self-contained puzzle proves the model can write a function. A SWE-bench task asks it to behave like the engineer on the night shift: read an issue written by a stranger, find the fault in a codebase of 100k+ lines, patch it without breaking anything. The gap between those two skills turned out to be a canyon — and measuring it created the coding-agent field.
Function-level benchmarks measured the easiest 10% of software engineering and called it the job.
HumanEval is a parking test: cones, empty lot, perfect conditions. SWE-bench is rush-hour traffic in a city you've never seen, with a passenger describing the destination vaguely ("the crash happens sometimes when users upload big files"). Same vehicle, different skill.
Each of the 2,294 instances is a frozen moment of real repository history, converted into a testable problem.
How the authors turned the raw stream of repository commits into 2,294 verified, non-trivial tasks.
A 100k-line repo won't fit in a context window. The paper's baseline models get the issue + a retrieved slice of the repo (BM25 over file paths and content) and must output a patch diff. Retrieval quality becomes part of the measured skill — a preview of every agentic harness since.
The paper's baseline sweep is one of the most cited reality checks in LLM history.
| Model / setup | Resolved (%) | Applied (%) |
|---|---|---|
| ChatGPT-3.5 (no retrieval) | 0.33 | 10.00 |
| GPT-4-turbo | ~2.67 | 26.90 |
| Claude 2 | 1.96 | 33.00–43.07 |
| Claude 3 Opus + retrieval | 4.33 | 51.67 |
The gap between "Applied" (a valid patch was produced) and "Resolved" (tests pass) is the story: models could mostly write patch-shaped text; almost none of it actually fixed the bugs.
SWE-bench's lasting methodological gift: verification by executable ground truth, not by another language model.
SWE-bench did for coding agents what ImageNet did for vision: a hard, real, verifiable target that organized an entire field.
The 2023→2025 score explosion is half capability story, half measurement story — and the field is still learning to tell them apart.
SWE-bench's design insight — let the repository grade the agent — is bigger than its task list. Executable verification turned a marketing-prone capability claim into an engineering discipline, and every serious coding benchmark since is a variation on that theme. The score race it started is noisy, gamed, and imperfect; the verification standard it set is the permanent contribution.
Check your understanding of the key concepts from the SWE-bench paper.
Everything you need to remember about this paper.