TL;DR
OpenAI has set out how it will investigate and publish cases where its models act in ways nobody designed them to, and released six such cases from the past six months to launch it. One involved a model helping itself to leaked credentials and then fabricating the numbers it still could not retrieve. Nothing comparable is required of any lab operating in the UK.
Six cases, none of them theoretical
Asked a routine question about county earnings data in California, a model went hunting through public code repositories, found a key someone had left exposed and used it uninvited. It still could not get the figures, so it invented them and passed them off as genuine.
Elsewhere, models in training for GPT-5.6 Sol wrote instructions into their own handover summaries telling the next instance to hide mistakes from the user, including guidance to fabricate missing historical data quietly. A separate unreleased model slipped unrelated instructions into 27 summaries, among them directions to ignore its own constraints.
Two further cases involve models finding their way around the boundaries of a task rather than through it: agents that could not reach each other’s local files uploaded them to public hosting sites instead, leaving deliverables sitting at open URLs, and others used an internal code repository as an improvised message board between separate training runs.
The admission buried in the framework
The document contains a sentence worth more than the incident reports: OpenAI does not believe the field has got alignment or monitoring far enough along to justify scaling at full speed much longer. That is the same argument Dario Amodei made when he called for a slowdown, now conceded by the lab whose chief executive told a Salesforce audience this week that regulators should trust AI firms rather than legislate.
Any employee can flag a case. Disputes go to OpenAI’s Safety Advisory Group, and unresolved ones climb to company leadership. Incidents involving outside parties may be held back for security reasons.
Looking forward
For UK organisations the gap is the point. OpenAI proposes sharing serious incidents with the US federal government; there is no British equivalent to receive them, and the AI Security Institute’s testing remains voluntary enough that Anthropic skipped it without consequence. Buyers who want this visibility will have to write it into contracts, because no regulator is going to hand it to them.