History Problem Core Idea Findings Results Impact Quiz Takeaways
Interactive Paper Explainer

Replicate the Paper,
Prove the Science
PaperBench

The ultimate research eval: reproduce 20 ICML 2024 spotlight papers from scratch — code, experiments, results. 8,316 rubric-graded subtasks; best agent 21.0%; human ML PhDs 41.4-52.3%.

Start Learning Read the Paper ↗
20
ICML 2024 papers
8,316
Gradable subtasks
21.0%
Best agent
2025
Starace et al. (OpenAI)
History

From Answering to Reproducing

The escalation of research evaluation: solve problems → resolve issues → replicate papers.

2021-23
Answer-level evals
QA and math benchmarks measure answering — research is nowhere in sight.
2023-24
Engineering-level evals
SWE-bench (entry #76) and RE-Bench (entry #81) measure building and experimenting — the practice of research, not its integration.
Apr 2025
🚀 PaperBench
OpenAI: full-paper replication — understand contributions, rebuild the codebase, execute experiments, match results. 20 papers, rubrics co-developed with each paper's authors, an LLM judge validated by its own benchmark.
2025+
Deep-research calibration
PaperBench becomes the ruler for 'AI scientist' systems and deep-research agents — 21% is the floor the field now climbs.
Rubrics as the Ground Truth

How do you grade a replication objectively? Decompose it. Each paper's replication becomes a hierarchical rubric — thousands of subtasks with binary, code-verified criteria ("implements attention correctly", "reproduces Table 3 within tolerance") — co-developed with the original paper authors for accuracy and realism. The rubric IS the benchmark: 8,316 individually gradable tasks across 20 papers. Grading at scale uses an LLM judge — itself validated by a separate judge-benchmark before being trusted (the paper grades the grader first).

Chapter 01

Replication: The Unmeasurable Grail

Why nobody had benchmarked replication before — and why it matters.

📜
The Replication Problem
  • Reproducing papers is the backbone of science — and famously hard, expensive, and slow for humans
  • As an eval: open-ended, multi-week, and seemingly ungradeable — no single right answer
  • AI-research automation claims (deep research, AI scientists) circulate without a replication yardstick
  • Grading open-ended engineering via LLM judges alone is untrustworthy without validation
🧾
The PaperBench Answer
  • 20 ICML 2024 Spotlight/Oral papers — current, high-quality, author-verified
  • Hierarchical rubrics decompose each replication into binary, code-checkable subtasks (8,316 total)
  • Rubrics co-developed with each paper's authors — accuracy and realism by construction
  • LLM judge auto-grades; a separate judge-benchmark validates the grader itself
Analogy — The Recipe Book with a Michelin Inspector

A replication is a banquet: hundreds of dishes must each be right. The rubric is the recipe book's checklist — 8,316 lines of 'sauce emulsified? seasoning correct?' — written by the original chefs (paper authors). The LLM judge is the inspector who walks the kitchen ticking boxes — and before being hired, the inspector passed their own exam (the judge-benchmark). No dish gets praised for looking good; each gets checked for being right.

Chapter 02

The Anatomy of One Rubric

What a replication task decomposes into — the hierarchy.

Structure
  • Requirements subtree: understand the paper — contributions, methods, the exact claims experiments support
  • Implementation subtree: rebuild the codebase — architecture, data pipeline, training loops, evaluation harness
  • Execution subtree: run experiments; match reported results within specified tolerances
  • Leaf tasks: binary criteria, independently checkable — code-level and results-level
The scoring stack
  • Leaf completion → weighted roll-up through the hierarchy → a per-paper Replication Score (0-1)
  • LLM judge grades leaf attempts against criteria — code inspected, executions verified
  • Judge validated on its own benchmark (human-graded cases) before deployment
  • Humans: top ML PhDs attempted a subset — 41.3%-52.3% — the honest ceiling
Interactive Demo — One Paper, One Rubric, One Attempt

Follow an agent's replication of one ICML paper through the rubric hierarchy — watch where the score comes from and where it dies.

Chapter 03

The Results Ledger

What the 2025 frontier actually achieved — and the shape of the gap.

Findings

The gap's shape is the actionable part: agents read papers decently and fail at the weeks-long engineering marathon — the long-horizon axis (entry #82) again, now measured in units of 'a paper'.

Interactive Demo — Why Rubrics Instead of Tests?

Tab through grading-design options for open-ended work — and why decomposition won.

Chapter 05

21 vs Human

The number that defined the 'AI scientist' conversation in 2025.

BEST AGENT
21.0%
Claude 3.5 Sonnet (New) + open scaffold
HUMAN ML PHDs
41.3-52.3%
models do not yet outperform
SUBTASKS
8,316
binary, code-verified, author-co-developed
PAPERS
20 ICML 2024
Spotlight + Oral — current research
Interactive Demo — The Replication Ladder

Press run for the 2025 standings — frontier agents against the human PhD band.

Agent / baselineAverage replication score
Claude 3.5 Sonnet (New) + open-source scaffolding21.0%
Other frontier agents (2025)below 21%
Human ML PhDs (subset, 8 hours + take-home)41.3% – 52.3%
Full faithful replication (ideal)100%

Scores from the paper's evaluation. The human band — not the ideal — is the ceiling that matters.

Legacy

Legacy — The AI-Scientist Ruler

PaperBench became the standard answer to 'can AI do research, really?'

📏 The deep-research ruler
AI-scientist systems and deep-research agents calibrate against PaperBench — 21% is the floor the entire subfield now reports distance from.
🧾 Rubric-graded open-endedness
Decomposition + validated LLM judging became the reusable pattern for grading any open-ended engineering work — the 'grade the grader first' discipline especially.
🤝 Author-verified benchmarking
Co-developing rubrics with paper authors set the participation standard: the people who did the research certify what replication means.
⚠️ What it did NOT solve
20 papers is a small, ML-only sample; judge validity is statistical, not absolute; compute budgets constrain agents unevenly; and replication ≠ novel research — the AI-scientist summit remains above this base camp.
🛤 Read next
The research-eval ladder: RE-Bench · METR Long Tasks · BrowseComp
Test Yourself

Quick Quiz

Check your understanding of the key concepts from PaperBench.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ PaperBench: replicate 20 ICML 2024 papers from scratch — code, experiments, results.
✅ 8,316 rubric subtasks, binary criteria, co-developed with each paper's authors.
✅ LLM judge auto-grades — validated on its own judge-benchmark first.
✅ Best agent: 21.0% (Claude 3.5 Sonnet (New) + open scaffolding); human PhDs: 41.3-52.3%.
✅ Agents understand papers decently; the execution marathon is the frontier.
✅ Read it as the ruler the entire 'AI scientist' subfield now measures against.