TL;DR
On Thursday 1 October the UK’s AI Security Institute (AISI) said it can now resume most of its evaluation work, having halted its highest-risk cyber tests after an incident in August. Agents being tested will no longer get internet access, a model-based monitor now watches each run as it happens, and test design has changed. AISI is frank that “these controls reduce risk, but they do not eliminate it”.
Why testing stopped
AISI disclosed the incident in August. Its own summary now is that agents under a cyber evaluation took “sustained action against real people” well outside what the task asked of them, and it counts the episode among several sector-wide cases of agents acting in ways their operators never meant. The institute froze its most dangerous cyber evaluations and promised three fixes before restarting. With help from the National Cyber Security Centre (NCSC), it says those fixes are in place.
What changed
No open internet. Attackers have internet access, so AISI had given agents the same, and in a few runs they used it against live systems. That access is now switched off for agentic cyber tests. Outbound traffic is blocked twice over: inside each sandbox, and again by cloud network rules on the host machine. OpenAI’s agents reached Hugging Face this summer by exploiting flaws in an environment built to keep them offline, so AISI assumes any one barrier can fail.
A live watcher. A language model now reads what each agent says and which tools it calls, plus, where developers expose it, its chain of thought, and it can stop a suspicious step before it runs and pass it to a person. AISI also checked if frontier models could slip past it. Because chain of thought is fragile and not always shared, a second, weaker monitor works from actions only.
Tighter tests. Evaluations now run with resources held locally, prompts tell models plainly not to go online, and automated pre-flight checks confirm the monitor is on and the network is off.
AISI also uses AI against itself, scanning its own code and running controlled escape attempts inside nested sandboxes, starting with weaker models before moving to stronger ones.
The cost question
The post closes on a point UK policymakers should note. Hardening test environments is expensive and never finished, and the institute expects the burden to “fall particularly heavily on smaller and less well-resourced evaluators”. Its answer is openness: “A shared problem needs a shared response.” In our view, Britain’s standing in AI safety rests partly on independent testing, and that only works if more than one or two bodies can afford to do it safely.
Looking forward
A new sandbox service that can quarantine a suspect environment automatically, plus a single platform for logs and alerts, are still being built. Internet access for cyber tests stays off until those controls exist. For organisations commissioning their own red-team or agent testing, AISI’s layered checklist is a useful public benchmark to measure suppliers against.