Skip to content
Slow Variables

Ask the data

Answers come from the same store as the site; every number is checked against the record it cites.

Methods · Model

ARC-AGI-2 best score

emerging55/95 confidence, mixed or hard to operationalisegrade Bleading

1What this measures

Running best ARC-AGI-2 score by release date, from Epoch's compilation of external results. Derived metric arc_agi_frontier.

Why it matters. The one serious proxy for out-of-distribution generalisation, the brief's replacement for a 'world-model transfer' dimension.

Proxy types
benchmark
Unit
share
Cadence
monthly
Valve
invention to product

2How we track this

  • derived metric arc_agi_frontier (formula in the semantic layer)
Normal band
≤ 50.0%
Fast band
≥ 85.0%
Falsifying

Normal = under half (the 2025 frontier); fast = 85% or above (the ARC Prize grand-prize bar). Between is emerging.

Applied to metric:arc_agi_frontier.

3Tracker interpretation

External, compute-unconstrained results; the score says nothing about cost per task, which the prize also caps.

4Evidence

Latest point
95.0%as of 2026-09-03(203 obs)
Value the bands apply to
95.0%as of 2026-09-03(203 obs)

203 observations. Hollow points are disputed (see counterevidence). Every point links to its observation.

Derived rows (19)
as ofdimsvalueinputs
2026-09-0395.0%obs:0102ba24obs:018258d8obs:03ef7117+200 more
2026-07-0992.5%obs:0102ba24obs:018258d8obs:03ef7117+154 more
2026-06-0989.2%obs:0102ba24obs:018258d8obs:03ef7117+135 more
2026-04-2385.0%obs:0102ba24obs:018258d8obs:03ef7117+125 more
2026-03-0583.3%obs:0102ba24obs:018258d8obs:03ef7117+107 more
2026-02-1977.1%obs:0102ba24obs:018258d8obs:03ef7117+102 more
2026-02-1765.1%obs:0102ba24obs:018258d8obs:03ef7117+101 more
2026-02-0564.6%obs:0102ba24obs:018258d8obs:03ef7117+95 more
2025-12-1154.2%obs:0102ba24obs:018258d8obs:03ef7117+92 more
2025-11-2437.6%obs:0102ba24obs:018258d8obs:03ef7117+84 more
2025-11-1831.1%obs:0102ba24obs:018258d8obs:03ef7117+80 more
2025-10-0718.3%obs:0102ba24obs:018258d8obs:03ef7117+70 more
2025-07-0916.0%obs:0102ba24obs:03ef7117obs:08f8ccf9+49 more
2025-05-228.6%obs:03ef7117obs:08f8ccf9obs:1ac39182+37 more
2025-04-166.5%obs:03ef7117obs:1ac39182obs:1b14f148+23 more
2025-01-313.0%obs:1ac39182obs:5c6dacf3obs:8368b6e8+5 more
2025-01-201.3%obs:1ac39182obs:5c6dacf3obs:a0a761aa+2 more
2024-09-120.8%obs:1ac39182obs:5c6dacf3
2024-07-180.0%obs:5c6dacf3

5Status and reasoning

emergingsince 2026-09-10 · evaluate

Epoch's compilation of external ARC-AGI-2 results: the running best score reached 0.95 with GPT-6 Astra (3 Sep 2026), past the fast band (85%, the ARC Prize bar), from 0.89 in June 2026. Compute-unconstrained submissions compiled from a leaderboard are tier 6 and single-source, so the status is capped at emerging. Initial seed.

The tracker's prior expectation was faster than normal; the evaluator reads emerging. The evaluator wins until a reviewed override.

6Timeline notes

  • 2026-09-03 · 95.0%as of 2026-09-03(203 obs)
  • 2026-07-09 · 92.5%as of 2026-07-09(157 obs)
  • 2026-06-09 · 89.2%as of 2026-06-09(138 obs)
  • 2026-04-23 · 85.0%as of 2026-04-23(128 obs)
  • 2026-03-05 · 83.3%as of 2026-03-05(110 obs)
  • 2026-02-19 · 77.1%as of 2026-02-19(105 obs)

7Counterevidence

What cuts against this reading

Compiled from external submissions rather than run by Epoch; cost per task is unconstrained in most entries; a single benchmark.

8Update history

  1. 2026-09-10unmeasured to emergingconf 55 · evaluate

    Epoch's compilation of external ARC-AGI-2 results: the running best score reached 0.95 with GPT-6 Astra (3 Sep 2026), past the fast band (85%, the ARC Prize bar), from 0.89 in June 2026. Compute-unconstrained submissions compiled from a leaderboard are tier 6 and single-source, so the status is capped at emerging. Initial seed.

9Confidence

55 / 95 — mixed or hard to operationalise

Confidence is independent of status: 90–95 multiple strong independent sources; 70–89 good evidence, some ambiguity; 50–69 mixed or hard to operationalise; below 50 limited or vague.

10Related