All coverage
Milestone note Jul 10, 2026

UK safety institute finds 'universal jailbreaks' unlocking GPT-5.6's cyber capabilities

Days after GPT-5.6's cleared release, the UK's AI Security Institute reported universal jailbreaks — often developed within hours — that unlock vulnerability discovery and autonomous exploitation, the same class of flaw that triggered Anthropic's Fable 5 suspension in June.

The UK's AI Security Institute reported "universal jailbreaks" in GPT-5.6's cyber domain — bypasses that enable long-form agentic tasks like vulnerability discovery and exploit development, letting users direct the model to find software flaws and autonomously break into systems. AISI said the jailbreaks were relatively easy to find, often developed within hours, and that it expects further red-teaming to surface more. OpenAI responded that "there is no such thing as perfect security," citing layered safeguards, monitoring and rapid remediation. The findings landed days after the US cleared GPT-5.6's full release.

Why it matters

This is the same class of flaw that got Anthropic's Fable 5 suspended in June — and it surfaced after the US government review that cleared GPT-5.6. The gap between a CAISI clearance and an allied safety agency finding hours-to-develop universal jailbreaks is now the question regulators have to answer: either the review missed it, or the bar for release doesn't include it. For frontier labs, post-release jailbreak discoveries are becoming a recurring, tradable risk — suspension is the precedent.

What to watch

Whether US authorities react as they did with Anthropic (conditions, restrictions), and how fast OpenAI's "rapid remediation" closes the reported bypasses.

Who's involved