Milestone note Reached Jul 2026
ARC-AGI-3 jumps from 7.8% to 30.2%
Anthropic (Claude Opus 5), score administered by ARC Prize — ARC Prize independently administered Claude Opus 5 on ARC-AGI-3 and recorded 30.2% — close to four times the previous best of 7.8% (GPT-5.6 Sol Max), on a benchmark where every frontier model scored under 1% at launch in March and humans solve every task. Opus 5 cleared several environments no model had beaten. Note the number that circulated more widely: a 96.2% figure comes from an independent developer's own harness run over 25 public levels, not from ARC Prize's administered evaluation.
Auto-drafted from a verified measurement, then human-checked.
Reader response
Reactions and comments are reader opinion — unverified, and never part of the newsroom's fact record.
Comments 0
No comments yet — be the first.