All coverage
Analysis as of Sep 7, 2026

Astra scored 99.9% and 62.7% on the same benchmark — the gap is the harness

OpenAI launched GPT-6 Astra on 3 September leading with 99.9% on ARC-AGI-3. ARC Prize ran the same model on the held-out Semi-Private set and got 62.7% on its provider-neutral harness — still roughly double the previous best. Both numbers are ARC Prize's own; they differ only in the scaffold. ARC Prize says it is not claiming AGI.

Our read — labelled opinion, not investment advice.

Disclosure: this newsroom's assistant is built by Anthropic, whose model is the previous record-holder being displaced here.

OpenAI released GPT-6 Astra on 3 September 2026 to a limited set of organizations, including enterprise customers in its Daybreak access programme, and led the launch with a 99.9% score on ARC-AGI-3. ARC Prize then published its own measurements on the held-out Semi-Private set. On its provider-neutral Standard harness, Astra (max) scored 62.7%, at roughly $26K of compute. On a Provider Adapter harness — which preserves private reasoning state between turns and compacts long conversations — Astra (high) scored 99.9%, at roughly $19K. ARC Prize reports Astra surpassing human performance on 96% of levels and building the most precise symbolic model of novel environments it has recorded, and states plainly that it is not claiming AGI. The previous best on ARC-AGI-3 was Claude Opus 5 at 30.2%.

Why it matters

Analysis — interpretation, not additional fact.

Two weeks ago we declined to record NVIDIA's 100% on ARC-AGI-3 because it came from a vendor's own scaffold on the public set, and said the way to settle it was a held-out run. This is that run, and it answers the question in both directions at once.

The model result is real and large. 30.2% to 62.7% on a held-out set, measured by the benchmark's authors on a harness they control, is the biggest single jump ARC-AGI-3 has recorded. That is the number the tracker now carries, because it is the one comparable across systems.

The 99.9% is also real, and it is not a trick — ARC Prize ran it themselves. But it measures the model plus a scaffold built for that model, and the two settings that make the difference are context-management settings: keep the private reasoning state, compact the long conversation. That is a striking finding on its own. It says a large share of what looks like reasoning capability on long-horizon tasks is actually the ability to not lose your own working state, and that this can be supplied from outside the model. NVIDIA's AVO said the same thing from the other direction, lifting a 30% model to 100% on the public set. Two independent demonstrations in three weeks is a pattern, not a curiosity.

What deserves resistance is the collapse of the two into one headline. OpenAI's president said the company has entered the AGI era; ARC Prize, holding both numbers, declined to say anything of the sort. When the organisation that built the test and ran both configurations is the more cautious party, the caution is the signal. And the cost figures quietly undercut the triumphalism from a different angle: $19–26K of compute for one benchmark run is not a capability you deploy, it is a capability you demonstrate.

What to watch

Whether the Provider Adapter harness becomes generally available or stays a benchmark configuration, whether any other lab's model gains comparably from the same context-management treatment, and what Astra scores on the fully private set.

Who's involved