Continual-learning ladder (highest production rung)
1What this measures
The highest rung of the continual-learning ladder (L0 static snapshot to L6 six-month-employee test) with a shipped production system on record, from dated rows citing the vendor's or lab's own description. Derived metric continual_learning_level.
Why it matters. The crux of the timelines argument. If systems learn on the job, deployment feeds invention automatically and the return arrow stops depending on human-data vendors; Narayanan and Kapoor's speed limits assume it stays unsolved.
- Proxy types
- product, model_release
- Unit
- rung
- Cadence
- quarterly
- Valve
- return arrow
2How we track this
- derived metric
continual_learning_level(formula in the semantic layer)
- Normal band
- ≤ 3.00 rung
- Fast band
- ≥ 5.00 rung
- Falsifying
- —
Appendix E's rule: L0-L3 (memory, population-level online learning, per-customer adapters) is normal; L5-L6 (persistent per-session weight updates, the six-month-employee test) is fast; L4 (organisation knowledge in weights) is emerging.
Applied to metric:continual_learning_level.
3Tracker interpretation
Rows are vendors describing their own systems; the rung is what they claim to ship, not what an independent evaluation found.
4Evidence
2 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.
Derived rows (2)
| as of | dims | value | inputs |
|---|---|---|---|
| 2026-05-27 | 3.00 rung | obs:1b7d08dbobs:dc983a81 | |
| 2025-09-12 | 2.00 rung | obs:dc983a81 |
5Status and reasoning
Highest production rung on record is L3: Trajectory's per-customer LoRA adapters refreshed hourly and A/B-routed behind provenanced endpoints, per Baseten's co-authored post of 27 May 2026; Cursor's Tab model sits at L2 (one online-trained model for all users, 1.5-2 hour checkpoint cycles). L4 (firm knowledge in weights, Engram + Harvey) and L5 (rank-1 LoRA merge, Nested Learning, TTT, self-distillation) exist only as research and do not move the level. L3 is inside the normal band, but every production row is a vendor describing its own system (tier 7), so the status is capped at emerging. Initial seed.
The tracker's prior expectation was consistent with normal; the evaluator reads emerging. The evaluator wins until a reviewed override.
6Timeline notes
7Counterevidence
What cuts against this reading
Every production row is tier 7 (self-described); research results are kept off the level; no public benchmark exists for L6, so the top rung cannot be reached by construction until one does.
8Update history
- 2026-09-10unmeasured to emergingconf — → 45 · evaluate
Highest production rung on record is L3: Trajectory's per-customer LoRA adapters refreshed hourly and A/B-routed behind provenanced endpoints, per Baseten's co-authored post of 27 May 2026; Cursor's Tab model sits at L2 (one online-trained model for all users, 1.5-2 hour checkpoint cycles). L4 (firm knowledge in weights, Engram + Harvey) and L5 (rank-1 LoRA merge, Nested Learning, TTT, self-distillation) exist only as research and do not move the level. L3 is inside the normal band, but every production row is a vendor describing its own system (tier 7), so the status is capped at emerging. Initial seed.
9Confidence
45 / 95 — limited or vague evidence
Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.
10Related
- Agent-workdays per human workday in frontier research (self-reported) emerging
- Expert-data market run-rate (Mercor) faster than normal
- Crosswalk: Return arrow ⇄ Training input (return arrow)