History Problem Core Idea Splits Domains Results Impact Deep Dive Quiz
Interactive Paper Explainer

Enterprise Scale, Locked Away
SWE-Bench Pro

A visual, step-by-step guide to the benchmark that took coding evaluation into the enterprise: 1,865 long-horizon problems from 41 repositories, split into public, held-out, and commercial sets — with real-world scale as contamination control.

Start Learning Read the Paper ↗
1,865
Problems
41
Repositories
3-way
Public · Held-out · Commercial
2025
Year Published
History

From Bug Fixes to Feature Work

SWE-bench measured whether agents could patch a bug. Pro asks whether they can do the day-sized — and week-shaped — work companies actually pay for.

2023
SWE-bench
2,294 real GitHub issues → verified patches. The coding-agent era's original scoreboard.
2024–25
The saturation scramble
Agent harnesses race up SWE-bench; Verified subset cleans labels; scores inflate; contamination worries grow.
2025 · Sep
🚀 SWE-Bench Pro (Scale AI · CAIS)
1,865 problems, 41 repositories, long-horizon tasks requiring hours-to-days of professional effort — partitioned to resist gaming.
2026 →
Enterprise-eval genre
Commercial splits, AI-for-code economic studies, and domain-specific (Verilog, AI-research) tracks become standard practice.
What "Pro" Actually Changes

Three moves at once: scale up the task (long-horizon, multi-file feature and refactor work), scale up the world (business apps, B2B services, dev tools — not just Python libraries), and lock the test bank (held-out + commercial splits that agents can never train on). Each move attacks a specific way its predecessor's scores stopped meaning what they said.

🧭 Read first
SWE-bench — this page builds on it directly.
Chapter 01

Three Ways the Old Score Lied

SWE-bench's 2024-2025 score explosion was real progress wrapped in three measurement problems.

⏱
Problem 1 — Task Scale
  • Issue-fixing is the "small" end of software work: localized root cause, bounded patch
  • Real enterprise engineering: multi-day features, migrations, refactors across services
  • High SWE-bench scores told you little about long-horizon autonomy
  • "Resolved" a bug ≠ "delivered" a feature with tests and docs
🏗
Problem 2 — World Domain
  • 12 popular open-source Python libraries: visible, well-documented, heavily discussed online
  • Enterprise code: proprietary, sprawling, boring, undocumented — the actual deployment surface
  • Popular-repo tasks live in pretraining corpora (issues, PRs, blogs) → memorization risk
  • Pro sources 41 actively maintained repos, including commercial ones never seen in training
Analogy — The Demo vs The Job

SWE-bench was a great interview loop: sharp, well-specified problems in a familiar codebase. Pro is the actual onboarding ticket: an unfamiliar 500k-line monorepo, a vague JIRA-shaped description, three days of work, and a reviewer who runs the tests. Passing interviews ≠ shipping software — the benchmark measures the gap.

Chapter 02

Contamination Control by Construction

The paper's signature design: a three-way partition that makes the most important test data structurally un-trainable.

1,865 problems = 11 public repos + 12 held-out repos + 18 commercial repos
Public
Open problems
Sourced from 11 repositories, freely available — the comparison and development split.
Held-out
Unseen repos
12 repositories whose problems are NOT publicly accessible — agents cannot have memorized them.
Commercial
Proprietary code
18 proprietary repos under formal partnership agreements; only results are released, never the tasks.
long-horizon
Task length
Problems may require hours to days for a professional software engineer — feature-scale, not bug-scale.
Same Verification DNA

Pro inherits the family's core: executable test verification (FAIL→PASS plus regression guards) and Docker-based environments. The grading philosophy — let the repository's own tests judge the agent — is unchanged; the tasks and worlds grew around it.

Task Construction

Problems are built from the development histories of maintained repositories — including changes with broader blast radius than bug fixes. Each ships with tests and environments reproduced up-front, so evaluation stays deterministic even as tasks get long.

Chapter 03

The Three-Bowl Design

Why the split architecture is the paper's real product.

Interactive Demo — Train / Report / Trust

Move an agent team through the three bowls: develop on public, report public+held-out, trust only commercial. Click each bowl.

🧪 Public split
Leaderboards and iteration: open problems, reproducible by anyone with the harness.
🔒 Held-out split
Unseen repos: contamination-resistant comparison — agents can't have their issues in pretraining.
🏢 Commercial split
Tasks never leave the partners' environments; only aggregate results publish. The trust anchor.
📈 Result deltas
Public-vs-held-out score gaps expose contamination and overfitting directly — a built-in honesty check.
Chapter 04

Not Just Python Anymore

The repository pool spans the actual shape of industrial software.

Repository Categories
  • Business applications: enterprise CRUD-scale systems with long histories and migration-shaped changes
  • B2B services: multi-service codebases where changes ripple across boundaries
  • Developer tools: build systems, CLIs, and platforms — tools that build tools
  • (Domain tracks) the Pro line extends into specialized tracks — e.g., hardware description (Verilog) and AI-research code
Why Domain Spread Matters

A model that aces Python libraries may be pattern-matching Python library conventions. Business apps, services, and tooling stress different muscles: reading unfamiliar architecture, respecting invariants nobody wrote down, and producing changes that survive contact with integration tests — the actual spectrum of professional software.

Interactive Demo — Task Length Ladder

Compare a SWE-bench-style fix to three Pro-shaped tasks. Advance to see the horizon stretch.

Chapter 05

The Long-Horizon Wall

Pro's central empirical finding: performance drops sharply as tasks stretch — even for frontier agents.

SHARP DROP
vs SWE-bench
frontier agents that score high on issue-fixing fall hard on Pro's long-horizon tasks
HOURS → DAYS
effort scale
task difficulty calibrated to professional engineer time — a new unit of measure
41 REPOS
3-way
public / held-out / commercial splits with score gaps revealing overfitting
ECONOMICS
$ studied
AI-for-code framed in cost-of-human-effort terms — benchmarks as economic instruments
Interactive Demo — Horizon vs Resolution

Slide along task-horizon; watch resolution decay for two agent generations (illustrative of the paper's qualitative finding).

Legacy

Impact — Evaluation Grows Up

Pro's influence is institutional: contamination control, economic framing, and commercial testbeds became normal vocabulary.

🔒 Held-out eval culture
Public-vs-held-out score deltas are now a standard honesty metric for agent claims.
🏢 Private testbeds
Benchmark tasks that never publish — results only — became an accepted trust pattern for enterprise claims.
💰 Economic benchmarks
Measuring AI labor in human-engineer-hours/cost reframed "how good" as "how much of the job."
🧭 Domain specialization
Verilog and AI-research tracks showed the template ports beyond Python — Terminal-Bench 2.0 carries the same spirit to CLI work.
🚧 What it did NOT solve
Task count is small; commercial results depend on partner trust; "resolves tests" still doesn't measure code review quality or maintainability.
🧪 Open harness
Public split + harness let anyone reproduce the open part — the reproducibility half of the trust bargain.
Deep Dive

Benchmarks as Institutions

Pro's deepest lesson: at the frontier, the benchmark's governance design matters as much as its task design.

🔓
The Open-Data Endgame
  • Any fully public benchmark eventually enters training corpora
  • Scores then measure memorization-plus-optimization, not capability
  • The community's decontamination rituals are leaky and unverifiable post-hoc
  • Trust in numbers decays on a schedule set by crawler reach
🏛
The Institutional Answer
  • Move the trust-critical data out of reach: held-out + commercial splits
  • Publish methodology and public results; keep tasks private
  • Let score gaps between splits diagnose contamination mathematically
  • Benchmarks become auditable institutions, not just datasets
Interactive Demo — Contamination Auditor

Compare an agent's public-split vs held-out-split performance. A wide gap is the contamination fingerprint.

Verdict

SWE-Bench Pro reads best as a governance paper wearing a benchmark costume. Its tasks are harder, its worlds are broader — but the durable contribution is the proof that evaluation can be structured so that trust doesn't decay with data availability. In a field whose numbers keep getting gamed, that is the rarest kind of result.

Test Yourself

Quick Quiz

Check your understanding of the key concepts from the SWE-Bench Pro paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ 1,865 problems from 41 actively maintained repositories — business apps, B2B services, dev tools.
✅ Three-way split: 11 public repos, 12 held-out (unseen), 18 commercial (proprietary, results-only).
✅ Long-horizon: problems that may take a professional engineer hours to days — feature-scale, not bug-scale.
✅ Same executable-verification DNA as SWE-bench: repo tests judge the agent, in Docker environments.
✅ Frontier agents drop sharply versus their SWE-bench-era scores — the long-horizon wall is real.
✅ Public-vs-held-out score gaps operationalize contamination auditing for coding agents.