A visual, step-by-step guide to the benchmark that took coding evaluation into the enterprise: 1,865 long-horizon problems from 41 repositories, split into public, held-out, and commercial sets — with real-world scale as contamination control.
SWE-bench measured whether agents could patch a bug. Pro asks whether they can do the day-sized — and week-shaped — work companies actually pay for.
Three moves at once: scale up the task (long-horizon, multi-file feature and refactor work), scale up the world (business apps, B2B services, dev tools — not just Python libraries), and lock the test bank (held-out + commercial splits that agents can never train on). Each move attacks a specific way its predecessor's scores stopped meaning what they said.
SWE-bench's 2024-2025 score explosion was real progress wrapped in three measurement problems.
SWE-bench was a great interview loop: sharp, well-specified problems in a familiar codebase. Pro is the actual onboarding ticket: an unfamiliar 500k-line monorepo, a vague JIRA-shaped description, three days of work, and a reviewer who runs the tests. Passing interviews ≠ shipping software — the benchmark measures the gap.
The paper's signature design: a three-way partition that makes the most important test data structurally un-trainable.
Pro inherits the family's core: executable test verification (FAIL→PASS plus regression guards) and Docker-based environments. The grading philosophy — let the repository's own tests judge the agent — is unchanged; the tasks and worlds grew around it.
Problems are built from the development histories of maintained repositories — including changes with broader blast radius than bug fixes. Each ships with tests and environments reproduced up-front, so evaluation stays deterministic even as tasks get long.
Why the split architecture is the paper's real product.
The repository pool spans the actual shape of industrial software.
A model that aces Python libraries may be pattern-matching Python library conventions. Business apps, services, and tooling stress different muscles: reading unfamiliar architecture, respecting invariants nobody wrote down, and producing changes that survive contact with integration tests — the actual spectrum of professional software.
Pro's central empirical finding: performance drops sharply as tasks stretch — even for frontier agents.
Pro's influence is institutional: contamination control, economic framing, and commercial testbeds became normal vocabulary.
Pro's deepest lesson: at the frontier, the benchmark's governance design matters as much as its task design.
SWE-Bench Pro reads best as a governance paper wearing a benchmark costume. Its tasks are harder, its worlds are broader — but the durable contribution is the proof that evaluation can be structured so that trust doesn't decay with data availability. In a field whose numbers keep getting gamed, that is the rarest kind of result.
Check your understanding of the key concepts from the SWE-Bench Pro paper.
Everything you need to remember about this paper.