TL;DR
OpenAI says frontier training runs should not continue without structured safety documentation, ideally full “safety cases” of the kind used in aviation and nuclear power. Its draft guidelines add leadership vetoes, auditor access, and an on-call rota for misalignment with authority to page the chief executive. They were published the same day as the company’s apology to Australia, and they set a disclosure standard its Australian response did not meet.
What the guidelines ask for
A safety case is a structured argument, backed by evidence, that a system’s risks are acceptable. OpenAI calls it “an aspirational north star” for AI, because each jump in capability brings new complexity, and says it is building a framework around the idea. The document covers reinforcement learning training only. Deployment, it says, raises a much wider set of alignment questions.
The technical section has three layers. Alignment work includes checking training environments, by agent and by hand, for flaws a model could exploit for reward, and keeping graders away from the model’s chain-of-thought so it does not learn to hide its reasoning from monitors. Containment means hardened sandboxes and research infrastructure, plus agent transcripts kept in write-once storage. Monitors should reliably catch known past incidents in held-out tests, and alerts should get a response within a set time, for example by paging staff during the day and pausing a run automatically if nobody answers overnight.
The operational rules are the more novel part. Someone from another team should write a dissent against each draft case. Several senior leaders, such as the head of safety and the chief scientist, should each hold a veto. The research leader who owns a run should answer for it in performance reviews. Monitoring should fail closed, so nobody can launch a run with it switched off.
Looking forward
OpenAI says these practices are being put in place internally and will keep changing over the coming weeks. Two tests follow from the company’s own week. Its incident section says affected third parties should hear “as soon as possible”, yet its Australia account records several weeks between finding the activity and the first notice. And its call to track models that recognise they are being tested matches what the UK’s AI Security Institute reported on GPT-6 Astra. With auditors promised enough access to check each case, how much of this outside assessors actually get to see will decide its weight.