One number per model: how long a task can be, with AI still succeeding half the time. That horizon has been doubling roughly every 7 months since 2019 — from seconds to about an hour today.
The translation problem: what does '87% on SWE-bench' mean in human terms?
The metric inverts the usual question. Instead of 'what score on task X?', ask: how long are the tasks where this model succeeds 50% of the time? Procedure: time domain-expert humans on many tasks (RE-Bench, HCAST, and 66 novel shorter tasks); for each model, find the human-duration band where its success rate crosses 50%; that duration is its horizon. The genius is the unit — human-minutes — because it translates directly into labor automation language: a 59-minute horizon means hour-long human tasks are a coin-flip.
The translation gap between benchmark points and real-world meaning.
A watch rated 50m and one rated 300m — you don't ask its 'swimming benchmark score'; you ask the depth it survives. The horizon metric gives AI a depth rating in human-minutes: dive to the model's rating and it works; past it, expect floods. And the rating has been doubling ~7 months — the tide is rising on a schedule.
How a horizon is extracted — the three steps.
The driver analysis — what actually moves the clock.
Why do newer models complete longer tasks? The paper's attribution: primarily greater reliability and the ability to adapt to mistakes — long tasks die by accumulated error, not by initial competence (a 99%-per-step success rate halves over 70 steps). Better reasoning and tool use contribute, but the compounding-error arithmetic is the villain — which reframes progress: the horizon doubles because per-step reliability inches upward, and the exponential forgives generously. That's also the sobering part: to reach month-long tasks, per-step reliability must reach levels where a 30-day chain survives — the extrapolation assumes the trend keeps beating the arithmetic.
One trend line that reframed every capability conversation.
| Era | Representative capability | 50% horizon (order) |
|---|---|---|
| 2019-20 | GPT-2/3-era prompting | seconds |
| 2021-22 | Codex-era code completion | ~1-10 minutes |
| 2023 | GPT-4 + plugins/tools | ~10-20 minutes |
| 2024-25 | Claude 3.7 Sonnet (extended thinking) | ~59 minutes |
| Extrapolated ~2030 | — | month-long tasks (if trend generalizes) |
Orders of magnitude consistent with the paper's fitted doubling (~7 months). The clock, not any single row, is the contribution.
Time horizons became the capability unit of the policy and product world.
Check your understanding of the key concepts from METR Long Tasks.
Everything you need to remember about this paper.