Milestone notes
Milestone note Reached Jul 2026

ARC-AGI-3 jumps from 7.8% to 30.2%

Anthropic (Claude Opus 5), score administered by ARC Prize — ARC Prize independently administered Claude Opus 5 on ARC-AGI-3 and recorded 30.2% — close to four times the previous best of 7.8% (GPT-5.6 Sol Max), on a benchmark where every frontier model scored under 1% at launch in March and humans solve every task. Opus 5 cleared several environments no model had beaten. Note the number that circulated more widely: a 96.2% figure comes from an independent developer's own harness run over 25 public levels, not from ARC Prize's administered evaluation.

Auto-drafted from a verified measurement, then human-checked.

More on this