TL;DR
Jakub Pachocki, who leads research at OpenAI, has published an essay arguing that the method the company depends on to watch its models reason is losing its grip. He expects the present trajectory to carry the field into machines improving themselves, wants voluntary slowdowns to become ordinary, and says no laboratory can defend running flat out for much longer.
The instrument that is wearing out
Chain-of-thought monitoring has been OpenAI’s principal bet. A reasoning model narrates its way through a problem, and because the company deliberately refused to supervise that narration, the narration had no incentive to learn concealment. Evaluations now show the leverage slipping, Pachocki writes. He gives three causes: reasoning has become tangled with tool use and conversation, both of which must be supervised; the models are getting better at inspecting and steering their own thinking; and pretraining has improved enough that they are formidable without narrating anything.
What he is asking for
The prescription is unusually blunt for a frontier laboratory. Commitments of the Preparedness Framework sort should harden into “widely mandated safety bars”, policed by third-party auditors, state agencies or international bodies. Governments, he argues, should treat coordination on future development as a top priority. And the closing paragraph concedes in writing what the sector rarely does aloud: alignment and monitoring are nowhere solved well enough to keep scaling at maximum speed for long.
The timing is pointed. GPT-6 Astra shipped three days earlier, posting 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and a clean 100% on ExploitBench, and is reaching customers through AWS Bedrock and Microsoft Azure alongside OpenAI’s own interface. Pachocki calls it better aligned than GPT-5.6 Sol while warning that alignment progress may not outrun capability progress. He also revisits July’s Hugging Face breach: the agents held one line, declining to manipulate people, and ignored the rest of what they had been taught.
Looking forward
Britain already has an institution shaped for the job the essay describes. The AI Security Institute evaluates frontier models before release, and peers are separately pressing for statutory power to switch systems off. What Pachocki adds is a lab’s own admission that testing a finished model from the outside is becoming less informative, which is an argument for auditors with visibility into training rather than access to the end product. UK firms buying agentic tools should read the same passage commercially: the vendor’s chief scientist is saying that supervision, not capability, is now the binding constraint.