Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Methods · Model

METR 80% time horizon

faster than normal70/95 confidence, good evidence, some ambiguitygrade Aleading

1What this measures

Task length a model completes with 80% success on the same suite. The band applies to a log-linear doubling-time fit over the 80% series from 2024 onward (derived metric horizon_doubling_days_80_since_2024), the same window as the 50% horizon.

Why it matters. 80% is closer to what a product needs than 50%. If the reliable horizon lags badly, benchmarks flatter products.

Proxy types
benchmark
Unit
minutes
Cadence
per release

2How we track this

  • series metr.*.horizon_80.pt
  • source METR time horizons (v1.1) · default tier 1 · METR; analysis code MIT at https://github.com/METR/eval-analysis-public
Normal band
≥ 213 days
Fast band
≤ 122 days
Falsifying

Same bands and same 2024-onward window as the 50% horizon, applied to our own fit because METR publishes doubling times for the 50% series only. The 2023-onward fit (horizon_doubling_days_80) sits in the gap between the bands.

Applied to metric:horizon_doubling_days_80_since_2024.

3Tracker interpretation

Fast, but lagging the 50% horizon; the lag is the reliability gap the products bucket has to close.

4Evidence

Latest point
3.1 h(1.6 h6.6 h)as of 2026-04-07
claude_mythos_preview_early_inspect
Value the bands apply to
107 days(97 days119 days)as of 2026-04-07(26 obs)

26 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.

Model comparison (lower AIC fits better; both on ln(horizon) residuals)
fitestimate95% intervalnAICinputs
horizon_doubling_days_80_since_2024107 days97 days119 days210.958-41.3obs:09061051obs:095bb211+24 more
horizon_doubling_days_80_since_2023125 days110 days144 days230.921-27.6obs:09061051obs:095bb211+24 more

5Status and reasoning

faster than normalsince 2026-09-10 · claude-initial-seed

Same 2024-onward window as the 50% horizon: our fit gives 107 days, inside the fast band; the 2023-onward fit is 125 days, in the gap. No confidence interval on our fit; the 80% horizon remains roughly 6x below the 50% horizon.

6Timeline notes

  • 2026-04-07 claude_mythos_preview_early_inspect · 3.1 h(1.6 h6.6 h)as of 2026-04-07
  • 2026-03-05 gpt_5_4 · 53.9 min(24.0 min1.8 h)as of 2026-03-05
  • 2026-02-19 gemini_3_1_pro · 1.5 h(52.0 min2.6 h)as of 2026-02-19
  • 2026-02-05 gpt_5_3_codex · 54.7 min(22.4 min2.0 h)as of 2026-02-05
  • 2026-02-05 claude_opus_4_6_inspect · 1.2 h(27.0 min2.8 h)as of 2026-02-05
  • 2025-12-11 gpt_5_2 · 1.1 h(31.7 min2.2 h)as of 2025-12-11

7Counterevidence

What cuts against this reading

Our fit has no confidence interval and excludes points METR marks unreliable; the 2023-onward window gives ~125 days, which would read `emerging`. The 80% horizon is still an order of magnitude below the 50% horizon, which is the reliability gap the products bucket has to close.

8Update history

  1. 2026-09-10emerging to faster than normalconf 7070 · claude-initial-seed

    Same 2024-onward window as the 50% horizon: our fit gives 107 days, inside the fast band; the 2023-onward fit is 125 days, in the gap. No confidence interval on our fit; the 80% horizon remains roughly 6x below the 50% horizon.

  2. 2026-09-10unmeasured to emergingconf 70 · claude-initial-seed

    Log-linear fit over the 80% horizon, 2023 onward, gives a doubling time of 124.7 days, just above the fast band ceiling (122). No CI on our fit. Initial seed; pending Alex's review.

9Confidence

70 / 95 — good evidence, some ambiguity

Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.

10Related