Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Methods · Model

50%/80% horizon ratio

consistent with normal70/95 confidence, good evidence, some ambiguitygrade Aleading

1What this measures

METR 50% horizon divided by 80% horizon for the same model (derived metric horizon_ratio_80_50). ~7-10x recently, up from ~5x.

Why it matters. The ratio is the reliability gap in one number. If it collapses toward 2x while the 80% horizon exceeds eight hours, the normal-technology thesis loses its productisation argument.

Proxy types
benchmark
Unit
ratio
Cadence
per release
Valve
invention to product

2How we track this

  • derived metric horizon_ratio_80_50 (formula in the semantic layer)
Normal band
≥ 5.00×
Fast band
≤ 2.00×
Falsifying

Normal = ratio ≥ 5 (gap holding or widening); fast = ratio ≤ 2. The falsifying condition (≤ 2 AND 80% horizon > 8 h) is compound and lives in the thesis monitor, not the band.

Applied to metric:horizon_ratio_80_50.

3Tracker interpretation

The gap is widening, which is pro-normal. It is the paper's reliability argument made measurable.

4Evidence

Latest point
6.34×as of 2026-03-05(2 obs)
Value the bands apply to
6.34×as of 2026-03-05(2 obs)

50 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.

Derived rows (25)
as ofdimsvalueinputs
2026-03-05gpt_5_46.34×obs:11a8bbd2obs:858c619b
2026-02-19gemini_3_1_pro4.28×obs:2eb7c897obs:5eb08488
2026-02-05gpt_5_3_codex6.39×obs:959db6ddobs:e466aae4
2026-02-05claude_opus_4_6_inspect10.3×obs:39c8213aobs:8f73b56b
2025-12-11gpt_5_25.34×obs:09061051obs:8a2d6997
2025-11-24claude_opus_4_5_inspect5.93×obs:72bcb9d8obs:75c93027
2025-11-19gpt_5_1_codex_max_inspect4.42×obs:8de1effaobs:a2aa2e4d
2025-11-18gemini_3_pro4.14×obs:3466aea4obs:833be5d2
2025-08-07gpt_5_2025_08_07_inspect5.30×obs:261c1f51obs:47db79bd
2025-08-05claude_4_1_opus_inspect4.28×obs:ae527806obs:fecd07e2
2025-05-22claude_4_opus_inspect4.91×obs:1f69c372obs:6af93ac0
2025-04-16o3_inspect3.99×obs:6b19d566obs:d1bb5ffe
2025-02-24claude_3_7_sonnet_inspect4.99×obs:095bb211obs:68864ed5
2024-12-05o1_inspect5.48×obs:37e97a6fobs:9450442a
2024-10-22claude_3_5_sonnet_20241022_inspect7.91×obs:0fb6ae10obs:19631991
2024-09-12o1_preview4.60×obs:734779b5obs:b5c1303b
2024-06-20claude_3_5_sonnet_20240620_inspect6.82×obs:145115d4obs:efa58e1c
2024-05-13gpt_4o_inspect5.52×obs:2dc1a032obs:7c977050
2024-04-09gpt_4_turbo_inspect4.02×obs:0d21fe65obs:deb6a1ab
2024-03-04claude_3_opus_inspect6.19×obs:f1409c69obs:fb7435cd
2023-11-06gpt_4_1106_inspect5.17×obs:29446d6cobs:cd50d564
2023-03-14gpt_44.48×obs:96d016faobs:a53da98c
2022-03-15gpt_3_5_turbo_instruct2.35×obs:4a682b8cobs:5e6df781
2020-05-28davinci_0022.56×obs:2325ad87obs:58c5e677
2019-02-14gpt24.20×obs:1c22fa34obs:64f0dec4

5Status and reasoning

consistent with normalsince 2026-09-10 · claude-initial-seed

Latest non-disputed model ratio is 6.3x (>= 5 normal band); the Mythos point is excluded as beyond the 16h suite ceiling. Matches the brief's expectation. Initial seed; pending Alex's review.

6Timeline notes

  • 2026-03-05 gpt_5_4 · 6.34×as of 2026-03-05(2 obs)
  • 2026-02-19 gemini_3_1_pro · 4.28×as of 2026-02-19(2 obs)
  • 2026-02-05 gpt_5_3_codex · 6.39×as of 2026-02-05(2 obs)
  • 2026-02-05 claude_opus_4_6_inspect · 10.3×as of 2026-02-05(2 obs)
  • 2025-12-11 gpt_5_2 · 5.34×as of 2025-12-11(2 obs)
  • 2025-11-24 claude_opus_4_5_inspect · 5.93×as of 2025-11-24(2 obs)

7Counterevidence

What cuts against this reading

Ratio is sensitive to the task mix and to the exclusion of models beyond the suite ceiling; a single new model can move it.

8Update history

  1. 2026-09-10unmeasured to consistent with normalconf 70 · claude-initial-seed

    Latest non-disputed model ratio is 6.3x (>= 5 normal band); the Mythos point is excluded as beyond the 16h suite ceiling. Matches the brief's expectation. Initial seed; pending Alex's review.

9Confidence

70 / 95 — good evidence, some ambiguity

Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.

10Related