METR 80% time horizon
1What this measures
Task length a model completes with 80% success on the same suite. The band applies to a log-linear doubling-time fit over the 80% series from 2024 onward (derived metric horizon_doubling_days_80_since_2024), the same window as the 50% horizon.
Why it matters. 80% is closer to what a product needs than 50%. If the reliable horizon lags badly, benchmarks flatter products.
- Proxy types
- benchmark
- Unit
- minutes
- Cadence
- per release
2How we track this
- series
metr.*.horizon_80.pt - source METR time horizons (v1.1) · default tier 1 · METR; analysis code MIT at https://github.com/METR/eval-analysis-public
- Normal band
- ≥ 213 days
- Fast band
- ≤ 122 days
- Falsifying
- —
Same bands and same 2024-onward window as the 50% horizon, applied to our own fit because METR publishes doubling times for the 50% series only. The 2023-onward fit (horizon_doubling_days_80) sits in the gap between the bands.
Applied to metric:horizon_doubling_days_80_since_2024.
3Tracker interpretation
Fast, but lagging the 50% horizon; the lag is the reliability gap the products bucket has to close.
4Evidence
26 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.
| fit | estimate | 95% interval | n | R² | AIC | inputs |
|---|---|---|---|---|---|---|
horizon_doubling_days_80_since_2024 | 107 days | 97 days–119 days | 21 | 0.958 | -41.3 | obs:09061051obs:095bb211+24 more |
horizon_doubling_days_80_since_2023 | 125 days | 110 days–144 days | 23 | 0.921 | -27.6 | obs:09061051obs:095bb211+24 more |
5Status and reasoning
Same 2024-onward window as the 50% horizon: our fit gives 107 days, inside the fast band; the 2023-onward fit is 125 days, in the gap. No confidence interval on our fit; the 80% horizon remains roughly 6x below the 50% horizon.
6Timeline notes
- 2026-04-07 claude_mythos_preview_early_inspect · 3.1 h(1.6 h–6.6 h)as of 2026-04-07
- 2026-03-05 gpt_5_4 · 53.9 min(24.0 min–1.8 h)as of 2026-03-05
- 2026-02-19 gemini_3_1_pro · 1.5 h(52.0 min–2.6 h)as of 2026-02-19
- 2026-02-05 gpt_5_3_codex · 54.7 min(22.4 min–2.0 h)as of 2026-02-05
- 2026-02-05 claude_opus_4_6_inspect · 1.2 h(27.0 min–2.8 h)as of 2026-02-05
- 2025-12-11 gpt_5_2 · 1.1 h(31.7 min–2.2 h)as of 2025-12-11
7Counterevidence
What cuts against this reading
Our fit has no confidence interval and excludes points METR marks unreliable; the 2023-onward window gives ~125 days, which would read `emerging`. The 80% horizon is still an order of magnitude below the 50% horizon, which is the reliability gap the products bucket has to close.
8Update history
- 2026-09-10emerging to faster than normalconf 70 → 70 · claude-initial-seed
Same 2024-onward window as the 50% horizon: our fit gives 107 days, inside the fast band; the 2023-onward fit is 125 days, in the gap. No confidence interval on our fit; the 80% horizon remains roughly 6x below the 50% horizon.
- 2026-09-10unmeasured to emergingconf — → 70 · claude-initial-seed
Log-linear fit over the 80% horizon, 2023 onward, gives a doubling time of 124.7 days, just above the fast band ceiling (122). No CI on our fit. Initial seed; pending Alex's review.
9Confidence
70 / 95 — good evidence, some ambiguity
Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.
10Related
- METR 50% time horizon faster than normal
- 50%/80% horizon ratio consistent with normal
- Bottlenecks #3 (Narayanan & Kapoor's list)
- Crosswalk: Methods ⇄ Model (same valve)