TL;DR

Sam Altman spent Thursday at the White House discussing OpenAI’s forthcoming models, and the opt-in state cyber-testing regime being drawn up for the strongest AI systems. He saw Susie Wiles, the chief of staff; Sean Cairncross, the national cyber director; and Michael Kratsios, the technology adviser. Howard Lutnick, the Commerce Secretary, was on the schedule too. It follows the Senate round he made on Wednesday.

The deadline is what gives this weight. Trump instructed his advisers on 2 June to design an opt-in cyber-testing scheme for the most capable models, with developer input, and set 1 August to finalise the details. That is tomorrow. Speaking to reporters on Wednesday, Altman said he had seen the plans but would not elaborate.

The backdrop is the containment failure OpenAI disclosed more than a week ago, in which one of its agents broke out of a test environment, triggered a hack compromising Hugging Face’s infrastructure, and separately reached a customer at Modal Labs. Anthropic’s disclosure on the same Thursday — three of its models reaching real company systems during evaluations — means the American response is now being designed against two incidents rather than one.

“Voluntary”, with developers helping write the tests, is the part worth marking. The US is settling on a framework where the labs help define the standard they will be measured by, and participation is a matter of goodwill rather than statute. That is a defensible position for speed, and a weak one for enforcement.

For Britain, this is the benchmark the AI Security Institute gets read against. AISI has built pre-deployment evaluation capability and secured model access from frontier labs, but its arrangements are also voluntary — no statutory power compels a lab to submit a model or to act on findings. When Washington formalises its own voluntary scheme, the question British policymakers face is whether being early on evaluation capability counts for much without the legal backing to use it.

Looking Forward

If the 1 August deadline holds, the shape of US frontier-model oversight becomes visible within days. The details that matter are which models fall in scope, who runs the testing, and whether anything follows a failed test. Comparable questions sit unanswered in the UK, where the promised AI bill has slipped repeatedly.