Agent reliability across repeated runs (tau-bench pass^k)
1What this measures
Share of tau-bench tasks an agent passes on every one of k attempts. The 2024 paper's gpt-4o figures (pass^1 61%, pass^8 under 25%) set the baseline; the 2026 point is the leaderboard leader's pass^1 and pass^4 as tabulated by Automation Anywhere.
Why it matters. The 70% problem in one number - solving a task once is not solving it reliably. Deployment needs pass^k, not pass^1, and the gap between them is where the normal-technology speed limit lives.
- Proxy types
- benchmark
- Unit
- share
- Cadence
- quarterly
- Valve
- invention to product
2How we track this
- series
automation_anywhere.tau_bench_leaderboard_1.pass_hat_4.pt - series
automation_anywhere.tau_bench_leaderboard_1.pass_hat_1.pt - series
arxiv.tau_bench_retail_gpt4o.pass_hat_1.pt - series
arxiv.tau_bench_retail_gpt4o.pass_hat_8.pt - source arXiv abstract pages · default tier 6 · arXiv; abstract quotation
- source Automation Anywhere blog · default tier 7 · Automation Anywhere; short quotation
- Normal band
- ≤ 70.0%
- Fast band
- ≥ 90.0%
- Falsifying
- —
Normal = under 70% of tasks pass all four runs, so the best agent still fails one attempt in three; fast = 90% or more at k of four or higher, the reliability a workflow can be built on. Between is `emerging`.
Applied to automation_anywhere.tau_bench_leaderboard_1.pass_hat_4.pt.
3Tracker interpretation
Reliability still collapses with k - 70% once, 56% four times running - two years after the paper that named the problem.
4Evidence
5Status and reasoning
First scoring. The tau-bench leaderboard leader passes 70.2% of tasks once and 56.2% four times running (Automation Anywhere table, 18 May 2026), inside the normal band; the 2026 point is vendor-tabulated (tier 7) so the status is capped at emerging, with the 2024 paper (gpt-4o pass^1 61%, pass^8 under 25%) as the second source.
The tracker's prior expectation was consistent with normal; the evaluator reads emerging. The evaluator wins until a reviewed override.
6Timeline notes
- 2026-05-18 tau_bench_leaderboard_1 · 56.2%as of 2026-05-18
7Counterevidence
What cuts against this reading
The 2026 point is vendor-tabulated and the model is unnamed; tau-bench domains are narrow; pass^k penalises benign variation.
8Update history
- 2026-09-10unmeasured to emergingconf — → 55 · evaluate
First scoring. The tau-bench leaderboard leader passes 70.2% of tasks once and 56.2% four times running (Automation Anywhere table, 18 May 2026), inside the normal band; the 2026 point is vendor-tabulated (tier 7) so the status is capped at emerging, with the 2024 paper (gpt-4o pass^1 61%, pass^8 under 25%) as the second source.
9Confidence
55 / 95 — mixed or hard to operationalise
Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.
10Related
- Frontier models completing expert legal tasks end to end (Harvey LAB) emerging
- 50%/80% horizon ratio consistent with normal