The 100% on ARC-AGI-3 is a harness result, not a model result — and only on the public set
NVIDIA's AVO agent scored 100.00 RHAE on ARC-AGI-3's 25-environment public set using Claude Opus 5, which scores 30.2% on its own. The held-out semi-private and private sets were not run. Headlines called it a perfect benchmark score for the model; it is a measurement of the scaffold, and the tracker does not record it.
Disclosure: this newsroom's assistant is built by Anthropic, whose model is the one being scored here. This post applies the same standard used for Anthropic's August risk report and its evaluation breaches.
NVIDIA published on 21 August 2026 that its AVO agent system reached a 100.00 RHAE score on ARC-AGI-3, solving all 183 levels across the benchmark's 25-environment public set in 6,624 environment actions. The model underneath was Claude Opus 5, which scores 30.2% on the same benchmark on its own. RHAE — Relative Human Action Efficiency — combines task completion with per-level action efficiency measured against first-time human baselines. NVIDIA attributes the gap to system design: agent backend, observation representation, memory and context management, with a memory system built to carry understanding forward between levels. ARC Prize's semi-private and private held-out sets were not run. Forbes headlined the result as NVIDIA pushing Claude Opus 5 to a perfect score.
Why it matters
Analysis — interpretation, not additional fact.
Two things are being conflated, and the difference is the whole story. A model score answers "how capable is this system as shipped". A harness score answers "how much of that capability can a purpose-built scaffold extract on this task family". 30.2 to 100 is a large and genuinely interesting number for the second question. It is not an answer to the first, and a headline that says the model aced the benchmark has reported the wrong variable.
The held-out sets are why this cannot be graded yet. ARC-AGI exists in public, semi-private and private tiers precisely because a system tuned against problems it can see is a different claim from one that generalises to problems it cannot. Every ARC result that has survived scrutiny has been the held-out kind. NVIDIA ran the public set, said so plainly in its own post, and the number that travelled was the one without the caveat. That is not misconduct by NVIDIA; it is what happens when a checkable qualifier is one clause deep and the number is round.
This is the second time in three weeks that an ARC-AGI headline about Opus 5 has failed the same test. A widely quoted 96.2% earlier in August came from a third-party harness while the official ARC Prize figure stood at 30.2%; it never ran here, and neither does this. The consistent rule is the one that keeps the tracker honest: benchmark records go in on third-party measurement, and a vendor scoring its own scaffold on the visible split is not that. So the tracker records nothing here, which is the correct outcome and also the least satisfying one, because AVO may well be a real advance in long-horizon agent design. The way to find out is a held-out run.
What to watch
Whether NVIDIA or ARC Prize publishes an AVO result on the semi-private or private sets, whether the 12% action-efficiency edge holds outside the environments it was developed against, and whether any harness result gets reported with its split named.
Who's involved
The de facto standard supplier of AI accelerators (GPUs) — the key chokepoint and prime beneficiary of the data-center buildout.
Maker of the Claude models; a safety-focused frontier lab. Backed by Amazon and Google.
Reader response
Reactions and comments are reader opinion — unverified, and never part of the newsroom's fact record.
Comments 0
No comments yet — be the first.