Agent-workdays per human workday in frontier research (self-reported)
1What this measures
Agent effort used by a frontier lab's research organisation, in standard eight-hour workdays, per workday of human labour. Today only OpenAI reports it, about itself.
Why it matters. The return arrow made numeric. If deployment inside the lab is feeding invention at more than one agent-day per human-day, the loop Narayanan and Kapoor draw but do not instrument is running; whether it substitutes for external bottlenecks is the thesis question.
- Proxy types
- deployment, behaviour
- Unit
- ratio
- Cadence
- quarterly
- Valve
- return arrow
2How we track this
- series
openai_blog.openai.rsi_agent_workdays_per_human.pt - source OpenAI research posts · default tier 7 · OpenAI; short quotation
- Normal band
- ≤ 1.00×
- Fast band
- ≥ 3.00×
- Falsifying
- —
Below one agent-day per human-day, agents are tools inside a human-run process (normal); at three or more, most research effort is agent effort, the regime the AI 2027 'R&D multiplier' assumes (fast). Between is emerging. Tier 7 evidence caps the status at emerging whatever the number says.
Applied to openai_blog.openai.rsi_agent_workdays_per_human.pt.
3Tracker interpretation
OpenAI says 3.1 as of mid-August 2026, having crossed 1.0 in June; Anthropic says its overall progress multiplier is below 2x. Lab statements about themselves, not outcome measures; the number to watch is an independent replication, which METR is funded to build.
4Evidence
5Status and reasoning
OpenAI's 'Research acceleration' post (6 Sep 2026, via the Internet Archive snapshot because openai.com blocks the fetcher): 3.1 agent-workdays per human workday as of mid-August 2026, having crossed 1.0 in June. Above the fast band (3 or more), but a lab's statement about itself is tier 7 and caps the status at emerging. METR's modelling note puts Anthropic's researcher uplift from coding agents at over 2x; Anthropic's own system card says overall R&D uplift is well short of a doubling. No independent measurement exists yet. Initial seed.
6Timeline notes
- 2026-08-15 openai · 3.10×as of 2026-08-15
7Counterevidence
What cuts against this reading
Self-measured and self-published by the lab whose timeline it supports; 'effort' counts agent time, not delivered research; the same post says over half of successful four-to-eight-hour tasks still needed human intervention. Anthropic's September 2026 system card reports staff self-estimating about 4x productivity uplift while the company puts its overall progress multiplier below 2x and observes no sustained AI-attributable 2x acceleration. METR's technical-worker survey finds a self-reported 1.4–2x value uplift, and METR's own RCT measured developers slower with AI.
8Update history
- 2026-09-10unmeasured to emergingconf — → 30 · evaluate
OpenAI's 'Research acceleration' post (6 Sep 2026, via the Internet Archive snapshot because openai.com blocks the fetcher): 3.1 agent-workdays per human workday as of mid-August 2026, having crossed 1.0 in June. Above the fast band (3 or more), but a lab's statement about itself is tier 7 and caps the status at emerging. METR's modelling note puts Anthropic's researcher uplift from coding agents at over 2x; Anthropic's own system card says overall R&D uplift is well short of a doubling. No independent measurement exists yet. Initial seed.
9Confidence
30 / 95 — limited or vague evidence
Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.
10Related
- Human interventions on 4–8 hour agent tasks (self-reported) emerging
- METR 80% time horizon faster than normal
- Developer productivity uplift (METR RCTs) consistent with normal
- Bottlenecks #86 (Narayanan & Kapoor's list)