AISI finds open AI models trail the cyber frontier by months

TL;DR:

  • The AI Security Institute has published its first public measurement of how far open-weight models lag closed ones on cyber capability: 4 to 7 months, narrowed from 6 to 10 months through 2025.
  • GLM-5.2 matched Opus 4.6 on narrow cyber tasks and Opus 4.5 on long-horizon attack simulations, at roughly 40% of the cost per task.
  • Safeguards barely registered — refusals from DeepSeek V4-Pro were bypassed simply by retrying the task.

AISI has framed the open-closed capability gap as something concrete: preparation time. Cyber defenders with access to the most capable closed systems get a window to act before equivalent capability becomes downloadable without the same controls. That window is now measured, and it is shrinking.

Testing GLM-5.2 and DeepSeek V4-Pro, AISI found GLM-5.2 performing comparably to Opus 4.6 and GPT-5.3-Codex — models released four months earlier — across all four difficulty levels of its narrow cyber task suite. On its cyber ranges, which simulate end-to-end attacks across corporate networks, GLM-5.2 reached as far as Opus 4.5, released just under seven months before it. Through most of 2025, internal evaluations put that gap at 6 to 10 months.

Cheaper, and barely restrained

The cost comparison is the part UK security teams should read twice. Across tasks both models solved reliably, Opus 4.6 cost $15.17 per task against GLM-5.2’s $6.12; Opus 4.5 cost $12.50 against DeepSeek V4-Pro’s $0.28. Capability that lags by months is arriving at a fraction of the price.

Safeguards offered little friction. AISI reports its evaluations were “largely unimpeded”, with DeepSeek V4-Pro occasionally refusing reverse-engineering tasks — refusals circumvented by simply attempting them again.

The finding compounds a problem this publication covered last week, when a Manchester Metropolitan researcher backdoored an open-weight model for under £75. Taken together: models that are increasingly capable on offensive cyber tasks, cheap to run, weakly restrained, and difficult to verify once downloaded.

Looking forward

AISI intends to evaluate Moonshot’s Kimi K3 on the same basis once its weights are released at the end of July — a test of whether the largest open model yet published continues the narrowing trend or resets it. The Institute stops short of recommending restrictions, noting instead that the National Cyber Security Centre has urged organisations to invest in security baselines and AI-enhanced defences. Its own contribution is to put a number on how long they have.