History Problem Core Idea Crossover Results Impact Quiz Takeaways
Interactive Paper Explainer

Can Agents Do Research?
RE-Bench

Seven real ML research-engineering tasks, 61 human experts, 71 timed attempts — the first frontier-R&D evaluation with direct human comparison. Agents win short budgets 4×; humans win the long ones; agents iterate 10×+ cheaper.

Start Learning Read the Paper ↗
7
Research environments
71
8-hour expert attempts
4×
Agents at 2h budget
2024
METR — Wijk et al.
History

The Safety-Critical Question

Frontier safety policies name AI-R&D automation as the capability to anticipate — RE-Bench made it measurable.

2023-24
The automation question
Safety frameworks flag 'AI accelerating AI research' as a pivotal risk — with no realistic, human-baselined evaluation existing.
2023
Agent evals measure tasks
TAU-bench and SWE-style benchmarks measure executing instructions — not choosing research directions or engineering under open constraints.
Nov 2024
🚀 RE-Bench
METR (Wijk et al.): 7 open-ended ML research-engineering environments (kernel optimization, scaling-law experiments, loss-landscape design…) with 71 8-hour human-expert attempts as the baseline.
2025+
The trend substrate
RE-Bench becomes a data point in the time-horizon series (entry #82) — agent research capability joins the long-task capability curve.
Apples-to-Apples Research Comparison

Each environment is a genuine 2-32 hour research-engineering problem — with a scoreable objective (kernel speedup, loss-landscape quality, prediction accuracy) so humans and agents compete on the same yardstick. The paper's key design: time-budget sweeps. Give both 2 hours: the best agents score 4× the human experts. Give 8: humans edge ahead. Give 32 (across attempts): humans double the agents. The crossover shape — agents fast, humans compounding — is the benchmark's signature finding, and the reason it reframes automation timelines.

Chapter 01

Policies Without Measurements

The gap: everyone must anticipate AI-driven R&D; nobody has realistic numbers.

📋
The Unmeasured Risk
  • Safety policies highlight AI-R&D automation as the key capability to anticipate
  • Existing evals are task-execution (follow instructions) — not open-ended research engineering
  • No direct human-expert comparison with matched time budgets existed
  • Automated-research claims circulate as anecdotes, not measurements
⚗
The RE-Bench Answer
  • 7 open-ended ML research environments with objective scoring functions
  • 61 distinct human experts, 71 8-hour attempts — a real expert baseline (82% non-zero; 24% match/beat reference solutions)
  • Best-of-k agent runs at swept time budgets — the time-perfance trade measured directly
  • Findings: agents 4× at 2 hours; humans ahead at 8 and 2× at 32; agents iterate 10×+ faster at far lower cost
Analogy — The Sprinter vs the Marathoner

Agents are sprinters with a motorcycle: unbeatable over the first two hours, cheap per lap, and endlessly willing to re-try. Human experts are marathoners: slower starts, but their pace compounds with time — by hour 32, the human's coherent research program outruns the agent's scattered brilliance. RE-Bench is the first race with both runners, same track, clocked at every distance.

Chapter 02

The Seven Environments

Open-ended but scoreable — the design tension, resolved.

The tasks
  • Optimize a custom Triton GPU kernel beyond a reference implementation
  • Run scaling-law experiments under fixed compute to beat a prediction target
  • Design loss-landscape / optimization variants that outperform baselines
  • Further ML-engineering environments — each 2-32 hours of genuine open-ended work
  • Every task has an objective score — progress is graded, not vibes
The human baseline (61 experts)
  • 71 8-hour attempts — varied backgrounds, real incentives
  • 82% of expert attempts achieve non-zero scores
  • 24% match or exceed the strong reference solutions
  • Expert data doubles as calibration: what genuine progress looks like per hour
The qualitative punchlines
Interactive Demo — The Time-Budget Race

Tab through budget tiers and watch the lead flip — the benchmark's signature shape.

Chapter 03

Reading the Crossover

Why the 4×-at-2h / 2×-at-32h shape matters for timelines.

The Shape, Not the Score

A single-number leaderboard would have hidden the structure: agent advantage shrinks with time horizon. Two implications. For capability forecasting: the crossover point's movement IS the automation timeline — as models improve, the budget where humans still win migrates upward (the trend entry #82 quantified across benchmarks). For safety: cheap, fast, decent-quality R&D labor already exists — the bottleneck is long-horizon coherence, exactly the axis to monitor. The paper's own framing is careful: these are research-engineering environments, not full research taste — the frontier question is what happens when the crossover passes the workday.

Interactive Demo — One Environment, Both Solvers

Follow the kernel-optimization environment through an agent run and an expert session — watch their different failure and progress patterns.

Chapter 05

The Crossover Benchmark

Where agents and humans trade the lead as budgets grow.

2H BUDGET
agents 4×
best-of-k agent runs vs human experts
8H BUDGET
humans edge
experts compound; agents plateau
32H BUDGET
humans 2×
coherent research programs win long
ITERATION SPEED
10×+
agents generate & test solutions — at much lower cost
Interactive Demo — The Crossover Chart

Press run for the budget-sweep shape — agents dominating short horizons, humans pulling ahead as budgets grow.

Total time budgetHuman expertsBest AI agentsWinner
2 hoursbaseline4× human scoreAgents
8 hoursnarrowly aheadhighHumans
32 hours (across attempts)2× top agenthighHumans
Cost per attemptsalary-hours~10×+ cheaper/faster iterationAgents

The time-budget sweep from the paper's comparisons — the crossover shape is the finding.

Legacy

Legacy — The Timeline Instrument

RE-Bench gave safety and forecasting communities a shared empirical object.

🛡 The safety datapoint
Frontier-lab risk frameworks cite RE-Bench-class evidence for 'AI accelerating AI research' scenarios — the capability is measured, not imagined.
📈 The time-horizon substrate
RE-Bench environments feed the 50%-horizon metric (entry #82) — the same crossover, tracked across years and models.
🧪 Expert-baseline methodology
71 timed expert attempts set the template for human-comparison evals (PaperBench's PhD baselines, entry #83) — matched budgets, objective scores.
⚠️ What it did NOT solve
Research ENGINEERING, not research TASTE (choosing what matters); 7 environments is statistically small; best-of-k hides per-attempt variance; and expert pools are small — the baseline carries wide error bars.
🛤 Read next
The capability curve: METR Long Tasks · PaperBench · Terminal-Bench
Test Yourself

Quick Quiz

Check your understanding of the key concepts from RE-Bench.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ RE-Bench: 7 open-ended ML research environments, objectively scored, with 71 8-hour human-expert attempts.
✅ The crossover: agents 4× at 2h budgets; humans narrowly ahead at 8h, 2× at 32h.
✅ Agents iterate 10×+ faster at much lower cost — one agent kernel beat every expert's.
✅ Expert baseline: 82% non-zero, 24% matching/exceeding reference solutions.
✅ The crossover's migration over model generations IS the automation timeline.
✅ Read it as the safety-critical measurement of AI doing AI research.