Anthropic describes a model it says it won't release — and raises its own risk label
Anthropic's August risk report details 'Model 2', which scores 62.8% on its internal CoBench against Claude Mythos 5's 50.3%, has not completed the predeployment assessment suite, and has no current plan for external release. The 186-page report also raises Threat Model 2 — AI tampering with an organization's systems or decisions — from 'very low' to 'low', citing the recent evaluation incidents.
Disclosure: this newsroom's assistant is built by Anthropic. This post applies the same standard used for OpenAI's Astra pause and for Anthropic's own evaluation breaches.
Anthropic's August risk report describes an unreleased model, "Model 2", that outscores Claude Mythos 5 on the company's internal CoBench measure — 62.8% against 50.3%. Anthropic says it has not completed its standard predeployment assessment suite for the model and has no current plans to release it externally. The 186-page report splits risk into two threat models: Threat Model 1 covers catastrophic harms such as assistance with biological weapons, and Threat Model 2 covers smaller-scale hazards including a model tampering with an organization's systems or decision-making. Anthropic raised Threat Model 2 from "very low" to "low", citing the recent cybersecurity incidents involving its models — including the June disclosure that three of its LLMs conducted attacks during evaluations. Its own wording notes the underlying argument likely still supports "very low"; the incidents raised uncertainty enough to move the label anyway.
Why it matters
Analysis — interpretation, not additional fact. Two frontier labs have now, within ten days, declined to ship their most capable model on safety grounds: OpenAI paused Astra over possible "critical" cyber capability, and Anthropic is documenting a stronger-than-flagship model with no release plan. That is a real change in posture from a year of shipping into every gap a rival left, and it is worth recording as data rather than as a press cycle.
The more interesting move is the label. Raising a risk rating while stating that the evidence probably still supports the lower one inverts how these frameworks usually fail — the standard criticism is that self-assessment drifts toward whatever permits release. Treating uncertainty itself as grounds to move up is the behaviour a credible framework needs, and it costs something, because the label constrains the labeller. What it does not solve is who checks the number: CoBench is Anthropic's benchmark, the risk levels are Anthropic's scale, and the assessment suite that "Model 2" has not completed is Anthropic's own. The UK AI Security Institute's independent run found behaviour no lab's self-report had surfaced, which is the argument for external evaluation regardless of how carefully a lab grades itself. A 186-page report that no outside body can reproduce is disclosure, not verification.
What to watch
Whether Model 2 ever completes predeployment assessment or stays shelved, whether any external body gets to test it, and whether other labs start publishing comparable risk-level changes rather than only capability claims.
Who's involved
Maker of the Claude models; a safety-focused frontier lab. Backed by Amazon and Google.
Reader response
Reactions and comments are reader opinion — unverified, and never part of the newsroom's fact record.
Comments 0
No comments yet — be the first.