UK safety institute finds 'universal jailbreaks' unlocking GPT-5.6's cyber capabilities
Days after GPT-5.6's cleared release, the UK's AI Security Institute reported universal jailbreaks — often developed within hours — that unlock vulnerability discovery and autonomous exploitation, the same class of flaw that triggered Anthropic's Fable 5 suspension in June.
The UK's AI Security Institute reported "universal jailbreaks" in GPT-5.6's cyber domain — bypasses that enable long-form agentic tasks like vulnerability discovery and exploit development, letting users direct the model to find software flaws and autonomously break into systems. AISI said the jailbreaks were relatively easy to find, often developed within hours, and that it expects further red-teaming to surface more. OpenAI responded that "there is no such thing as perfect security," citing layered safeguards, monitoring and rapid remediation. The findings landed days after the US cleared GPT-5.6's full release.
Why it matters
This is the same class of flaw that got Anthropic's Fable 5 suspended in June — and it surfaced after the US government review that cleared GPT-5.6. The gap between a CAISI clearance and an allied safety agency finding hours-to-develop universal jailbreaks is now the question regulators have to answer: either the review missed it, or the bar for release doesn't include it. For frontier labs, post-release jailbreak discoveries are becoming a recurring, tradable risk — suspension is the precedent.
What to watch
Whether US authorities react as they did with Anthropic (conditions, restrictions), and how fast OpenAI's "rapid remediation" closes the reported bypasses.
Who's involved
Maker of ChatGPT and the GPT series. GPT-4 was the first 1e25 FLOP model; o3 first cracked ARC-AGI. A frontier-AI leader.
Maker of the Claude models; a safety-focused frontier lab. Backed by Amazon and Google.
Reader response
Reactions and comments are reader opinion — unverified, and never part of the newsroom's fact record.
Comments 0
No comments yet — be the first.