Indicators / series
arxiv.tau_bench_retail_gpt4o.pass_hat_8.pt
Source: arXiv abstract pages (arXiv) · arXiv; abstract quotation · CSV
arXiv
| id | text | grade | as of date | published date | tier | audited vs reported | extraction method | review status | retrieved at | http status | content hash | supersedes id | url | flags |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| f6285e213b487ef4 | pass^8 <25% in retail | grade B, evidence tier 6 | 2024-06-17 | 2024-06-17 | 6 | reported | manual | approved | 2026-09-10T07:01:42.079101Z | 200 | d82d45164e1b | — | source page |
Raw snippets and dispute text
- f6285e213b487ef4 even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail)