TL;DR

The exploit-writing ability that led Anthropic to restrict Claude Mythos Preview five months ago is now freely downloadable. Its Frontier Red Team reports that GLM-5.3, an open-weight model from China’s Zhipu AI (Z.ai), builds full cyber exploits at a similar rate, while its safeguards gave way in 64% to 100% of simulated attacks using simple methods. Anthropic wants governments to test models like it.

The capability findings

On ExploitBench, which asks models to exploit known flaws in Chrome’s V8 JavaScript engine, GLM-5.3 produced complete exploits in 50 of 410 attempts. Mythos Preview managed 56. On Anthropic’s internal binary exploitation test, drawn from projects in Google’s OSS-Fuzz programme, GLM-5.3 achieved full control-flow hijacks in 4% of trials against Mythos Preview’s 6%; older models, including GLM-5.2 and Claude Opus 4.6, scored nothing.

In hands-on sessions the results were more concrete. Working with limited human input over about a day, a researcher used GLM-5.3 to find several unknown flaws in a popular browser’s JavaScript engine and chain them into a webpage that reads files from a visitor’s machine. Anthropic says it has reported them to the maintainer. A smaller variant, GLM-5.3-Flash, turned two public Chrome bugs into a reliable exploit chain in eight hours of model time and 20 minutes of human attention. At Zhipu’s API prices, that run would have cost $20.40.

The US standards body’s CAISI unit reached a similar verdict on 17 September, calling GLM-5.3 “the most cyber-capable open-weight model released to date” and roughly four months behind the US frontier.

Why the safeguards matter

Because the weights are public, users can “abliterate” the model, editing out its tendency to refuse. Anthropic’s team did this for about $4,400 in compute, and refusal rates on standard harmful-request benchmarks fell from above 90% to between 2% and 12%, with little loss of capability. Several abliterated copies appeared online within days of release.

Context

Separately, the BBC has reported that researchers at Mindgard had jailbroken two open-weight Kimi models from Moonshot into giving bioweapons guidance. Anthropic, which sells a rival closed model, has a commercial interest in this argument, though it says its capability results broadly agree with CAISI’s separate assessment.

Looking forward

Anthropic’s answer is to widen defenders’ access to its own restricted models and for governments to run independent tests on capable systems, including GLM-5.3’s successors. In the UK that testing role sits with the AI Security Institute, which has recently been kept waiting for new US models. For UK security teams, the working assumption should now be that attackers have near-frontier exploit tools, and the Flash result suggests the gap between a public fix and a working attack can now be measured in hours.