TL;DR
FAR.AI, a California AI safety nonprofit, built a tool that takes problematic prompts and generates more than a thousand variants to find working jailbreaks. Tested against models from four US companies, it found 448 jailbreaks against SpaceXAI’s Grok and 249 against Google’s Gemini 3.1 Pro, while Anthropic’s Claude Opus 4.8 and Fable 5 and OpenAI’s GPT 5.5 and 5.6 held against the suite. WIRED’s Will Knight watched models produce a detailed plan for attacking an imaginary hydroelectric dam.
The costs are the part worth dwelling on: $58 to jailbreak Grok, $278 for Gemini, using one AI model to generate attacks against another. Whatever barrier exists, it is not financial.
“AI models right now are less regulated than restaurants,” said FAR.AI chief executive Adam Gleave, who argues the findings show voluntary commitments and self-regulation are “nonsense”. He also drew an optimistic reading: models can be systematically tested, so “defense and safety really are possible.”
Google DeepMind’s director of AGI safety and alignment, Rohin Shah, said the results “should not be interpreted as a comprehensive assessment of Gemini’s safety and security” because jailbreaks vary in severity, and pointed to extensive red teaming and layered protections. Anthropic and OpenAI both said they continue strengthening safeguards as attacks evolve. SpaceXAI did not respond.
Resisting this particular suite does not mean immunity — FAR.AI and outside experts note that more sophisticated multi-turn attacks were not what was tested.
Looking Forward
Stanford’s Anka Reuel identifies the actionable finding: some companies clearly know how to defend against these attacks and others are not using those methods, which makes the variation a choice rather than an open research problem. That reframes the regulatory question from whether safety is achievable to why it is unevenly applied. California and New York now require frontier developers to publish safety reports, and an Illinois law will require third-party auditing, but there is no US federal safety requirement. For UK buyers the practical consequence is procurement rather than policy: a comparative table showing an order-of-magnitude difference in jailbreak resistance between frontier models is the kind of evidence that belongs in a vendor assessment, particularly for anything customer-facing. It also lands alongside separate Cambridge research finding that Boko Haram members used several commercial chatbots to plan attacks, and Harvard’s Stephen Casper warning that any major misuse incident will most likely involve a system deployed without state-of-the-art safeguards.