The ultimate research eval: reproduce 20 ICML 2024 spotlight papers from scratch — code, experiments, results. 8,316 rubric-graded subtasks; best agent 21.0%; human ML PhDs 41.4-52.3%.
The escalation of research evaluation: solve problems → resolve issues → replicate papers.
How do you grade a replication objectively? Decompose it. Each paper's replication becomes a hierarchical rubric — thousands of subtasks with binary, code-verified criteria ("implements attention correctly", "reproduces Table 3 within tolerance") — co-developed with the original paper authors for accuracy and realism. The rubric IS the benchmark: 8,316 individually gradable tasks across 20 papers. Grading at scale uses an LLM judge — itself validated by a separate judge-benchmark before being trusted (the paper grades the grader first).
Why nobody had benchmarked replication before — and why it matters.
A replication is a banquet: hundreds of dishes must each be right. The rubric is the recipe book's checklist — 8,316 lines of 'sauce emulsified? seasoning correct?' — written by the original chefs (paper authors). The LLM judge is the inspector who walks the kitchen ticking boxes — and before being hired, the inspector passed their own exam (the judge-benchmark). No dish gets praised for looking good; each gets checked for being right.
What a replication task decomposes into — the hierarchy.
What the 2025 frontier actually achieved — and the shape of the gap.
The gap's shape is the actionable part: agents read papers decently and fail at the weeks-long engineering marathon — the long-horizon axis (entry #82) again, now measured in units of 'a paper'.
The number that defined the 'AI scientist' conversation in 2025.
| Agent / baseline | Average replication score |
|---|---|
| Claude 3.5 Sonnet (New) + open-source scaffolding | 21.0% |
| Other frontier agents (2025) | below 21% |
| Human ML PhDs (subset, 8 hours + take-home) | 41.3% – 52.3% |
| Full faithful replication (ideal) | 100% |
Scores from the paper's evaluation. The human band — not the ideal — is the ceiling that matters.
PaperBench became the standard answer to 'can AI do research, really?'
Check your understanding of the key concepts from PaperBench.
Everything you need to remember about this paper.