Seven real ML research-engineering tasks, 61 human experts, 71 timed attempts — the first frontier-R&D evaluation with direct human comparison. Agents win short budgets 4×; humans win the long ones; agents iterate 10×+ cheaper.
Frontier safety policies name AI-R&D automation as the capability to anticipate — RE-Bench made it measurable.
Each environment is a genuine 2-32 hour research-engineering problem — with a scoreable objective (kernel speedup, loss-landscape quality, prediction accuracy) so humans and agents compete on the same yardstick. The paper's key design: time-budget sweeps. Give both 2 hours: the best agents score 4× the human experts. Give 8: humans edge ahead. Give 32 (across attempts): humans double the agents. The crossover shape — agents fast, humans compounding — is the benchmark's signature finding, and the reason it reframes automation timelines.
The gap: everyone must anticipate AI-driven R&D; nobody has realistic numbers.
Agents are sprinters with a motorcycle: unbeatable over the first two hours, cheap per lap, and endlessly willing to re-try. Human experts are marathoners: slower starts, but their pace compounds with time — by hour 32, the human's coherent research program outruns the agent's scattered brilliance. RE-Bench is the first race with both runners, same track, clocked at every distance.
Open-ended but scoreable — the design tension, resolved.
Why the 4×-at-2h / 2×-at-32h shape matters for timelines.
A single-number leaderboard would have hidden the structure: agent advantage shrinks with time horizon. Two implications. For capability forecasting: the crossover point's movement IS the automation timeline — as models improve, the budget where humans still win migrates upward (the trend entry #82 quantified across benchmarks). For safety: cheap, fast, decent-quality R&D labor already exists — the bottleneck is long-horizon coherence, exactly the axis to monitor. The paper's own framing is careful: these are research-engineering environments, not full research taste — the frontier question is what happens when the crossover passes the workday.
Where agents and humans trade the lead as budgets grow.
| Total time budget | Human experts | Best AI agents | Winner |
|---|---|---|---|
| 2 hours | baseline | 4× human score | Agents |
| 8 hours | narrowly ahead | high | Humans |
| 32 hours (across attempts) | 2× top agent | high | Humans |
| Cost per attempt | salary-hours | ~10×+ cheaper/faster iteration | Agents |
The time-budget sweep from the paper's comparisons — the crossover shape is the finding.
RE-Bench gave safety and forecasting communities a shared empirical object.
Check your understanding of the key concepts from RE-Bench.
Everything you need to remember about this paper.