History Problem Core Idea Drivers Results Impact Quiz Takeaways
Interactive Paper Explainer

The Capability Clock
METR Time Horizons

One number per model: how long a task can be, with AI still succeeding half the time. That horizon has been doubling roughly every 7 months since 2019 — from seconds to about an hour today.

Start Learning Read the Paper ↗
~59 min
Claude 3.7 Sonnet horizon
~7 mo
Doubling time
2019 → now
Seconds → hour
2025
Kwa et al. (METR)
History

Benchmarks Don't Translate

The translation problem: what does '87% on SWE-bench' mean in human terms?

2023-24
Percentages pile up
SWE-bench, OSWorld, RE-Bench — every release quotes scores, but nobody can convert them into 'how long a task can this thing do?'
2024
RE-Bench's crossover clue
Entry #81 shows agent-vs-human comparisons flipping with time budget — the horizon axis appears implicitly.
Mar 2025
🚀 The horizon metric
METR (Kwa et al.): define 50%-task-completion time horizon — time humans take on tasks AI completes 50% of — and measure it across models, years, and task suites.
2025
The trend lands
~59 minutes for Claude 3.7 Sonnet; doubling ~7 months since 2019 (possibly accelerating in 2024); extrapolation: month-long tasks within ~5 years — if it generalizes.
2025+
Horizon as industry metric
Model cards and policy briefs quote time horizons; 'can it do a workday?' becomes the product question the clock answers.
One Unit for Autonomy

The metric inverts the usual question. Instead of 'what score on task X?', ask: how long are the tasks where this model succeeds 50% of the time? Procedure: time domain-expert humans on many tasks (RE-Bench, HCAST, and 66 novel shorter tasks); for each model, find the human-duration band where its success rate crosses 50%; that duration is its horizon. The genius is the unit — human-minutes — because it translates directly into labor automation language: a 59-minute horizon means hour-long human tasks are a coin-flip.

Chapter 01

Scores Without Units

The translation gap between benchmark points and real-world meaning.

📊
The Incommensurable Zoo
  • Every benchmark has its own scale — 72% here, 4.5 Elo there, 14/78 elsewhere
  • None convert to the question everyone asks: how long a task can it handle?
  • Trend claims ('improving fast') lack a principled axis to trend on
  • Policymakers and planners cannot price autonomy from leaderboard deltas
⏱
The Horizon Answer
  • Define: 50%-completion horizon = human time on tasks the AI completes 50% of
  • Measure: humans timed on RE-Bench + HCAST + 66 novel tasks; models run on the same
  • Fit: horizon per model over 2019-2025 — doubling every ~7 months (2024 may be faster)
  • Translate: Claude 3.7 Sonnet ≈ 59 minutes — a workday is 8 coin-flips long... for now
Analogy — The Depth Rating

A watch rated 50m and one rated 300m — you don't ask its 'swimming benchmark score'; you ask the depth it survives. The horizon metric gives AI a depth rating in human-minutes: dive to the model's rating and it works; past it, expect floods. And the rating has been doubling ~7 months — the tide is rising on a schedule.

Chapter 02

The Measurement Recipe

How a horizon is extracted — the three steps.

1️⃣ Time the humans
Domain experts are timed on a task battery (RE-Bench + HCAST + 66 novel tasks) — each task gets a representative human duration.
2️⃣ Run the models
Each model attempts the same tasks; success recorded per task (objective checks, environment graded).
3️⃣ Find the 50% crossing
Sort tasks by human duration; locate the band where the model's success rate crosses 50% — that human-duration is the model's time horizon.
The findings
  • Claude 3.7 Sonnet: ~59 minute 50% horizon (with extended thinking)
  • Across suites: horizon doubles ~every 7 months since 2019 — with possible 2024 acceleration
  • Drivers: reliability and adaptation-to-mistakes, plus reasoning and tool use — not raw model scale alone
  • Extrapolation (if general): month-long tasks within ~5 years
The honest caveats (the paper's own)
  • External validity: benchmark tasks ≠ real job tasks — the conversion is the assumption
  • 50% success is a low bar for deployment (would you accept a 50% colleague?)
  • Task suites skew software-ish; horizons may differ by domain
  • Trend extrapolation is a hypothesis, not a law — regimes can break
Interactive Demo — Watch the Clock Advance

Slide through the years and watch the horizon grow — from seconds to the hour, on the ~7-month doubling schedule.

Frames track the fitted doubling (~7 months per 2×). Per-model horizons are measured, not assumed — the trend is the fit across them.
Chapter 03

Why Reliability Is the Story

The driver analysis — what actually moves the clock.

The Decomposition

Why do newer models complete longer tasks? The paper's attribution: primarily greater reliability and the ability to adapt to mistakes — long tasks die by accumulated error, not by initial competence (a 99%-per-step success rate halves over 70 steps). Better reasoning and tool use contribute, but the compounding-error arithmetic is the villain — which reframes progress: the horizon doubles because per-step reliability inches upward, and the exponential forgives generously. That's also the sobering part: to reach month-long tasks, per-step reliability must reach levels where a 30-day chain survives — the extrapolation assumes the trend keeps beating the arithmetic.

Interactive Demo — What Doubles, Exactly

Tab through the candidate explanations for the doubling — and their evidence status.

Chapter 05

The 7-Month Clock

One trend line that reframed every capability conversation.

CURRENT HORIZON
~59 min
Claude 3.7 Sonnet, 50% success
DOUBLING TIME
~7 months
since 2019; 2024 possibly faster
DRIVERS
reliability
error adaptation + reasoning + tools
EXTRAPOLATION
month tasks
within ~5 years, IF it generalizes
Interactive Demo — The Extrapolation Debate

Month-long tasks within ~5 years — the claim everyone argues about. Press reveal for both sides.

EraRepresentative capability50% horizon (order)
2019-20GPT-2/3-era promptingseconds
2021-22Codex-era code completion~1-10 minutes
2023GPT-4 + plugins/tools~10-20 minutes
2024-25Claude 3.7 Sonnet (extended thinking)~59 minutes
Extrapolated ~2030—month-long tasks (if trend generalizes)

Orders of magnitude consistent with the paper's fitted doubling (~7 months). The clock, not any single row, is the contribution.

Legacy

Legacy — The Unit Everyone Quotes

Time horizons became the capability unit of the policy and product world.

⏱ The translation layer
Horizon-in-human-minutes gave executives, regulators, and researchers a common tongue — 'can it do a workday?' is answerable without knowing any benchmark.
📉 The trend as planning input
Safety timelines and automation forecasts now anchor on the ~7-month doubling — the number that made 'capability trend' a boardroom discussion.
🧩 Methodology export
The 50%-crossing procedure (time humans, run models, locate the band) became the standard design for human-comparable capability metrics.
⚠️ What it did NOT solve
External validity remains an assumption; 50% is not deployment-grade; software-heavy suites skew the estimate; and one metric cannot see quality-of-work, only task completion — the horizon measures duration, not value.
🛤 Read next
The measurement chain: RE-Bench · PaperBench · SWE-Bench Pro
Test Yourself

Quick Quiz

Check your understanding of the key concepts from METR Long Tasks.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Horizon metric: human-time on tasks the AI completes 50% of — capability in human-minutes.
✅ Doubling ~every 7 months since 2019; Claude 3.7 Sonnet lands around a 59-minute horizon.
✅ Driver: reliability + mistake adaptation (compounding-error arithmetic), plus reasoning and tools.
✅ Extrapolation: month-long software tasks within ~5 years, IF the trend generalizes.
✅ The caveats are load-bearing: benchmark≠job, 50%≠deployable, software-heavy suites.
✅ Read it as the translation layer between benchmarks and the autonomy conversation.