History Problem Core Idea Pipeline Results Verification Impact Deep Dive Quiz
Interactive Paper Explainer

Real Repos, Real Bugs
SWE-bench

A visual, step-by-step guide to the benchmark that turned software engineering into an evaluation: 2,294 real GitHub issues across 12 Python repositories, where a model wins only if the repository's own tests pass after its patch.

Start Learning Read the Paper ↗
2,294
Task Instances
12
Python Repositories
1.96%
Best Model (Claude 2)
2024
Year Published
History

From Code Quizzes to Real Repos

Code benchmarks had been self-contained puzzles. SWE-bench moved the exam into a live repository.

2021
HumanEval (Chen et al.)
Function-level Python puzzles: clean signatures, visible unit tests — the standard for two years.
2022–23
Repo-level tasks emerge
RepoBench and friends study cross-file completion — but completion is still not fixing.
2023 · Oct
🚀 SWE-bench (Jimenez, Princeton)
2,294 task instances mined from real issue→PR pairs in 12 popular Python repos; resolution verified by the repo's own tests.
2024 →
The coding-agent industry
SWE-agent, Devin, Agentless, OpenHands — a benchmark became an industry's scoreboard. Verified/Lite splits and SWE-Bench Pro followed.
The Night-Shift Test

A self-contained puzzle proves the model can write a function. A SWE-bench task asks it to behave like the engineer on the night shift: read an issue written by a stranger, find the fault in a codebase of 100k+ lines, patch it without breaking anything. The gap between those two skills turned out to be a canyon — and measuring it created the coding-agent field.

🧭 Study path
Agent skills in general: AgentBench → SWE-bench → SWE-Bench Pro.
Chapter 01

Why Puzzles Lied

Function-level benchmarks measured the easiest 10% of software engineering and called it the job.

🧩
The Puzzle Illusion
  • Self-contained: all needed context sits in the prompt — real bugs are buried in cross-file interactions
  • Tests are visible and unit-shaped — real regressions hide in integration behavior
  • One file, one function — real fixes span classes, configs, and imports
  • High puzzle scores coexisted with models that couldn't navigate a repo at all
🐛
Ground Truth From the Wild
  • Mine real issue→merged-PR pairs from popular repositories — the community already solved each task once
  • The merged patch is the reference answer; the repo's test suite is the grader
  • Tasks require multi-function, multi-file reasoning — by construction, not design
  • Messy, ambiguous issue text included: exactly what engineers actually face
Analogy — Parking vs Driving

HumanEval is a parking test: cones, empty lot, perfect conditions. SWE-bench is rush-hour traffic in a city you've never seen, with a passenger describing the destination vaguely ("the crash happens sometimes when users upload big files"). Same vehicle, different skill.

Chapter 02

Anatomy of a Task Instance

Each of the 2,294 instances is a frozen moment of real repository history, converted into a testable problem.

📦 Codebase snapshot
The repository state before the fix — the model edits this.
📝 Issue text
The real issue (and linked PR discussion, where available) — often incomplete, sometimes misleading.
🧪 FAIL→PASS tests
Tests that failed before the merged patch and passed after — these must pass for the model to win.
🛡 PASS→PASS tests
Tests that passed before and must still pass — the "don't break the world" check.
resolved ⇔ FAIL→PASS tests all pass  ·  PASS→PASS tests all still pass
FAIL→PASS
The fix proof
Behavior the issue complained about, pinned as executable checks.
PASS→PASS
The regression guard
Everything else must survive the patch — gaming one test isn't enough.
2,294
Instances
From 12 popular Python repos: django, sympy, scikit-learn, astropy, flask, requests…
gold patch
Reference solution
The actually-merged human patch — retrieval target and study material.
Chapter 03

Mining Signal From History

How the authors turned the raw stream of repository commits into 2,294 verified, non-trivial tasks.

Interactive Demo — Instance Mining Pipeline

Walk one pull request through the pipeline: candidate → test-linked → environment-built → non-trivial → task instance.

Filters That Bited
  • Test linkage: the PR must change test outcomes — most commits don't, and are discarded
  • Environment creation: each instance runs in a Docker environment that must build and reproduce the before/after test behavior
  • Non-trivial resolution: applying the gold patch's tests alone (without the fix) must not already pass — a guard against self-grading patches
  • Repo popularity floor: the 12 repos are libraries with thousands of GitHub stars — real, maintained software
The Retrieval Wrinkle

A 100k-line repo won't fit in a context window. The paper's baseline models get the issue + a retrieved slice of the repo (BM25 over file paths and content) and must output a patch diff. Retrieval quality becomes part of the measured skill — a preview of every agentic harness since.

Chapter 04

The Humbling Table

The paper's baseline sweep is one of the most cited reality checks in LLM history.

Baseline Results (paper-era models)
Model / setupResolved (%)Applied (%)
ChatGPT-3.5 (no retrieval)0.3310.00
GPT-4-turbo~2.6726.90
Claude 21.9633.00–43.07
Claude 3 Opus + retrieval4.3351.67

The gap between "Applied" (a valid patch was produced) and "Resolved" (tests pass) is the story: models could mostly write patch-shaped text; almost none of it actually fixed the bugs.

BEST RESOLUTION
4.33%
Claude 3 Opus with retrieval — the paper's strongest configuration
CLAUDE 2 ALONE
1.96%
the best-performing model in the original paper's headline framing
PATCH vs FIX GAP
~30 pts
of applied patches, only a few percent resolved — syntax ≠ engineering
TWO YEARS LATER
40–70%+
on SWE-bench Verified — agent scaffolds, retrieval, and training closed much of the gap
Interactive Demo — Resolution Machine

Send issues through three eras of system: raw model (2023), +retrieval, agent harness. Each task slot shows pass/fail after verification.

Chapter 05

Why the Tests Are the Paper

SWE-bench's lasting methodological gift: verification by executable ground truth, not by another language model.

The Anti-Gaming Stack
  • FAIL→PASS: the patch must actually change the broken behavior
  • PASS→PASS: the patch must not trade the fix for regressions elsewhere
  • Non-trivial guard: gold test files applied alone must not resolve the issue
  • Deterministic grading: the same patch scores the same every run — no judge drift
Failure Analysis Findings
  • Models often locate the right file but patch symptoms, not causes
  • Long, multi-hunk fixes span more of the repo than retrieved context covers
  • Specification ambiguity in issues transfers into plausible-but-wrong patches
  • Later community work (SWE-bench Verified) found label noise in the remaining fraction — and the benchmark's own verification style let it be cleaned publicly
Legacy

Impact — The Scoreboard of an Industry

SWE-bench did for coding agents what ImageNet did for vision: a hard, real, verifiable target that organized an entire field.

🤖 The agent-harness wave
SWE-agent, Devin, Agentless, OpenHands: scaffolds that navigate, edit, and run tests — the benchmark measured their birth.
✅ Verified & Lite splits
A human-validated 500-instance subset (Verified) and a 300-instance Lite set made evaluation cheaper and cleaner.
🏢 Enterprise follow-ups
SWE-Bench Pro extends to long-horizon, commercial-scale problems; MLE-bench and terminal tasks broaden the genre.
🧪 Contamination-resistant design
Task instances are (repo, commit) pairs — harder to memorize than static puzzles, though issue text can still leak.
📈 Training signal
RL on SWE-bench-style tasks became a standard post-training stage for coding models — the benchmark became curriculum.
⚠️ What it did NOT solve
Python-only at launch; test coverage = grading ceiling; "resolved" says nothing about code quality, security, or maintainability.
Deep Dive

When the Benchmark Becomes the Curriculum

The 2023→2025 score explosion is half capability story, half measurement story — and the field is still learning to tell them apart.

🪞
Ways to Rise Without Rising
  • Issue text from public repos is in pretraining corpora — memorization inflates scores
  • Test-gaming: patches that special-case the visible test files
  • Scaffold overfitting: agents tuned to this benchmark's repo layout quirks
  • Score-washing: reporting Lite or subset numbers without labels
🔬
Ways the Design Fights Back
  • Executed tests beat judge models — the strongest anti-gaming tool in common use
  • Verified subset: human-validated labels, published cleaning methodology
  • New unseen repos in Pro's held-out/commercial splits: memorization-resistant by construction
  • The benchmark's publicness made its own flaws fixable in public
Interactive Demo — The Score Decomposer

A headline "60% resolved" is several stacked capabilities. Toggle each component off to see how the number decomposes.

Verdict

SWE-bench's design insight — let the repository grade the agent — is bigger than its task list. Executable verification turned a marketing-prone capability claim into an engineering discipline, and every serious coding benchmark since is a variation on that theme. The score race it started is noisy, gamed, and imperfect; the verification standard it set is the permanent contribution.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the SWE-bench paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 2,294 task instances from real issue→merged-PR pairs in 12 popular Python repositories.
✅ Resolution = all FAIL→PASS tests pass AND all PASS→PASS tests still pass — the repo grades the patch.
✅ Paper-era best: Claude 2 resolved 1.96%; the strongest retrieval setup (Claude 3 Opus) reached 4.33%.
✅ The Applied-vs-Resolved gap exposed that models wrote patch-shaped text that rarely fixed anything.
✅ Non-triviality guard: applying gold tests without the fix must not already pass.
✅ Spawned Verified/Lite splits, SWE-agent harnesses, and SWE-Bench Pro — the coding-agent industry's scoreboard.