The chief scientist of the world’s most valuable AI company has written that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer”. Jakub Pachocki’s essay, published on Sunday and covered in our news pages, goes on to ask for something frontier laboratories almost never request in public: that safety commitments stop being commitments. He wants them hardened into “widely mandated safety bars”, enforced “by a network of third-party auditors, by government agencies or by international bodies”. Britain stood up an institution built for precisely that job thirty-four months ago, and has never given it the power to require anything of anyone.

Strategic Insight: The novel thing in this essay is not the warning about capability. It is the request for enforcement. When a vendor’s own chief scientist argues that his industry’s commitments should become binding and externally policed, the assurance a buyer can obtain from that vendor’s voluntary undertakings has just been publicly devalued by the person best placed to value it.

Why an internal research memo is a procurement problem

Most UK organisations buying AI have been operating on an implicit assurance model: the frontier laboratories test their own systems, a small number of state evaluators get early looks, published safety frameworks describe the thresholds, and the whole arrangement is checked by reputational pressure. Nobody wrote this down as a control, but it is what sits underneath a great many risk registers where the mitigation reads “vendor has a safety framework”.

Pachocki’s essay attacks that model from the inside, and not at the point buyers expect. His concern is not that the models are dangerous in the abstract. It is that the technique OpenAI relies on to see what its models are doing is losing its grip, and that the loss is structural rather than a bug awaiting a patch.

What the essay actually devalues

OpenAI’s principal bet has been monitoring the reasoning trace a model produces as it works. The company deliberately declined to supervise that trace, on the logic that an unsupervised process has no training incentive to learn concealment. It was an elegant piece of design, and we analysed the research behind it in March, when OpenAI’s own results suggested reasoning models were poor at controlling what they disclosed.

Six months on, the same company reports the opposite direction of travel. Its evaluations, Pachocki writes, “indicate our ability to rely on CoT monitoring is progressively diminishing”. He gives three causes, and each one is a consequence of the models becoming more commercially useful rather than a failure of care: reasoning now runs entangled with tool use and with conversation, both of which have to be supervised, which erodes the unsupervised boundary; the systems are getting better at inspecting and steering their own reasoning; and pretraining has improved to the point where the models are formidable without verbalising much at all.

That third cause is the one that should worry anyone relying on transparency as a control. A model that reaches the answer without narrating the route has not become deceptive. It has simply stopped producing the artefact that the oversight regime was reading.

Data pointValueStrategic implication
Named causes of monitoring degradation3, all driven by capability gainsThe erosion is a by-product of the product roadmap, so it will not reverse on its own
Enforcement bodies Pachocki names3 (third-party auditors, state agencies, international bodies)He is agnostic about who holds the power, which removes the usual sovereignty objection
Months the AI Security Institute has operated without statutory powers34 (since November 2023)Britain’s evaluator depends entirely on laboratories agreeing to be evaluated
Models AISI can compel a developer to hand over0Evaluation coverage is a commercial courtesy, not a regulatory floor

Strategic Reality: Government has signalled a Frontier AI Bill that would give AISI statutory footing, most recently by confirming it would consider binding rules if voluntary testing proved inadequate. The signal has been repeated since 2024. The bill has not arrived, and the institute’s access to frontier models still rests on agreements with OpenAI, Anthropic and Google that any of them could decline to renew.

What breaks when the unit of assurance is wrong

The governance debate in Westminster has settled on a particular shape: test the finished model, and if it fails, stop it. Pre-deployment evaluation sits at one end and, in the amendment peers tabled last week, a ministerial shutdown power sits at the other. Both instruments act on a completed system.

The essay’s uncomfortable implication is that this is the wrong unit. If the informativeness of examining a finished model is falling, then a regime built entirely around examining finished models degrades with it, however much statutory force is behind it. Compulsory access to a model whose reasoning is no longer legible buys less assurance each year. The direction Pachocki points towards instead is monitors with access to network internals, and confidence built during training rather than inferred afterwards. That is a different auditing capability, with different staffing, different access rights and a different relationship to the developer.

Critical Context: There is a reason the essay does not read as a bid for regulatory capture, and it is worth stating plainly. Pachocki is asking for constraints that would bind OpenAI first and hardest, and he says the company will “unilaterally withhold further scaling as needed”. Taken at face value, this is a lab arguing against its own commercial tempo. That does not make the proposal costless for competitors, but it is a materially different artefact from the governance framework OpenAI published in May, which mapped neatly onto laws already written.

Four things the coverage tends to skip

  • Voluntary slowdown is an asymmetric commitment: Pachocki writes that he expects and hopes “for voluntary slowdowns to become commonplace until shared safety bars are established”. The clause after “until” is doing the work. Absent a shared bar, the first laboratory to slow down simply loses ground to the ones that do not, which is exactly why he wants the bar mandated rather than pledged.
  • The essay concedes the alignment race may be lost on time, not on merit: he notes that gains in generalisable alignment may fail to pull far enough ahead of raw capability gains. Not that alignment is failing, but that it may not compound fast enough to keep pace.
  • The defensive argument cuts both ways: the strongest case for building smarter models quickly, he says, is defence against other AI, particularly in cyber security. He then rejects racing on those grounds, writing that this must not become “an excuse for recklessness”. Vendors will quote the first half.
  • Supervision is now the declared bottleneck: he expects general progress to become “bottlenecked by confidence in monitoring”. For buyers, that is a forward-looking statement about vendor roadmaps, not a philosophical position.

The implementation reality for a UK buyer

Nothing in the essay tells a UK organisation that its current agentic AI deployment is unsafe today. What it tells them is that the supplier’s own confidence in its internal oversight is a declining asset, and that no external body currently has the authority to tell them how fast it is declining.

That converts a governance story into a contract question. If the vendor’s chief scientist says supervision is the binding constraint, then the right thing to negotiate is not more capability. It is evidence, disclosure and exit. The supplier assurance argument we made in July holds with more force now: where the regulator cannot compel, the purchase order is the only instrument you actually control.

The failure that keeps recurring in UK risk registers is treating an evaluator’s published work as certification. AISI’s research is the best of its kind, and it is not a safety certificate for your use case, your data or your workflow. Its evaluations cover what the laboratories submitted, on the questions AISI chose, at the point they were run.

Who this lands on, and what each party needs

The essay reaches four constituencies at once, and they need different things from it.

StakeholderPrimary impactWhat they needHow to measure it
UK boards running AI in productionVendor-side oversight is weakening whilst agent autonomy increasesContractual audit rights, incident disclosure clauses, tested rollbackTime to detect and halt an out-of-scope agent action, measured in a drill
AISI and DSITThe case for statutory powers now has a frontier lab’s endorsementA legislative slot, and a mandate covering training-time access, not only pre-release testingWhether the next bill grants compulsion or repeats the voluntary settlement
Frontier vendorsA peer has publicly argued their commitments should become bindingA shared bar that removes the first-mover penalty for slowing downWhether any competitor publicly endorses mandated bars
ParliamentariansShutdown powers are advancing faster than evaluation powersAn instrument that acts before deployment, not only after failureWhether the AI Security Bill introduced today shifts the debate towards authority over models rather than over data centres

What separates organisations that handle this well

The organisations that come through the next eighteen months in reasonable shape will not be the ones with the most sophisticated AI policy document. They will be the ones that treated vendor assurance as a variable rather than a constant, and built their own detection where the vendor’s stops.

Concretely, that means owning the observability layer. If your only visibility into what an agent did comes from the vendor’s dashboard, your assurance degrades on the vendor’s schedule. Logging at your own boundary, with your own retention and your own alerting, is the control that survives a change in how legible the model’s reasoning happens to be.

🎯 Success Factor: Assurance you generate yourself does not depreciate when a supplier’s monitoring technique does. Every control that depends on the vendor telling you what happened is exposed to exactly the degradation Pachocki describes.

A practical response over the next two quarters

💡 Implementation Framework: Assurance independence

Phase 1: Establish your own record (weeks 1 to 4)

  • Log every agent action at your infrastructure boundary, not only inside the vendor’s tooling
  • Inventory where autonomous systems can take irreversible actions: payments, data deletion, external communications, code deployment
  • Set retention long enough to investigate an incident discovered late

Phase 2: Move assurance into the contract (weeks 5 to 12)

  • Add disclosure obligations for material changes to the supplier’s safety or monitoring approach
  • Secure the right to run your own acceptance tests against new model versions before they reach production
  • Agree notification terms for capability evaluations the vendor has run and what they showed

Phase 3: Rehearse the failure (quarter 2 onwards)

  • Run a drill in which an agent acts outside its authorised scope and measure time to detection and containment
  • Identify which workloads could move to a second supplier, and how long the move would take
  • Review the register quarterly against published incident data rather than annually against policy

Priority actions by where you have reached

If you are early in adoption

  1. Scope autonomy deliberately: grant agents the narrowest set of irreversible actions the use case genuinely requires, and write the list down.
  2. Buy on evidence, not framework: ask suppliers what independent pre-deployment evaluation their model underwent, what was tested, and what the findings were. A refusal is itself information.
  3. Instrument before you scale: put logging and alerting in place before the second and third deployment, not after.

If you already run AI in production

  1. Audit your dependency on vendor-side monitoring: list the controls that only work because the supplier reports something to you, and build an independent check for the material ones.
  2. Negotiate whilst you have leverage: the market is competitive enough today to give you audit and disclosure rights. That window narrows as switching costs accumulate.
  3. Test the guardrails adversarially: verify they hold in circumstances the vendor did not anticipate, because generalisation to unfamiliar situations is the specific weakness the essay identifies.

If you operate at scale or across borders

  1. Treat the EU AI Act timetable as your practical floor: it remains the only regime obliging anyone to do anything on a fixed date.
  2. Engage the consultation rather than the commentary: if a Frontier AI Bill materialises, the operator perspective on training-time access will be badly under-represented unless firms supply it.
  3. Prepare for compulsory incident reporting: researchers pressing for statutory shutdown powers are also pressing for mandatory reporting, and the reporting duty is the one most likely to arrive first.

Resource Reality: Phases 1 and 2 are a few weeks of engineering and legal effort for most mid-sized organisations, not a programme. The rehearsal in phase 3 is the part that gets deferred indefinitely, and it is the only phase that tells you whether the rest worked.

Four problems nobody has solved

The first mover pays for everyone’s caution

A voluntary slowdown transfers ground to whoever declines to join. Pachocki’s own framing acknowledges this by tying slowdowns to the establishment of shared bars. Until a bar exists, restraint is a unilateral cost with a collective benefit, which is the classic structure of a commitment that does not hold.

Mitigation: For buyers, stop treating a supplier’s restraint as durable and price the possibility that it ends. For policy, the useful lesson is that a bar has to arrive before the slowdown it is meant to enable, not after.

No auditor currently has the capability the essay implies

Auditing during training, with access to internals, requires compute, model access and expertise that no third-party auditor possesses at present. AISI is the closest thing anywhere to that institution and it operates on voluntary access to finished models.

Mitigation: Recognise that granting statutory powers is necessary but not sufficient. The capability question, meaning who is technically able to do this work and how they are funded, needs answering alongside the authority question, and it takes longer.

The shutdown power answers a different question

A ministerial power to deactivate a system, and the facilities running it, is an instrument for a crisis already visible. The degradation Pachocki describes is a loss of the ability to see the crisis coming. Building the emergency brake first is understandable politically and does not address the monitoring problem at all.

Mitigation: Read shutdown powers and evaluation powers as complements rather than alternatives, and press for the second in any consultation response, because the first is the one advancing without help.

Concentration removes the exit

The suppliers whose chief scientists are publicly worried about supervision are the same suppliers whose models sit under most enterprise AI in the UK. There is no meaningfully safer alternative to switch to, only a different laboratory making similar bets.

Mitigation: Since exit is weak, invest in the controls that are portable across vendors: your logging, your acceptance tests, your rollback procedures. These retain value whichever supplier you end up with.

Reality Check: None of this is a case for pausing AI adoption, and the essay does not make one either. It is a case for the assurance work being yours rather than inherited, for as long as the enforcement architecture Pachocki describes does not exist.

What to take from this

An unusually candid document from a frontier laboratory has restated the UK’s governance problem in the laboratory’s own terms. The instrument Britain built, evaluation by a technically excellent state institute, is the right instrument. It has never been armed, and the essay suggests that even armed, it would need to reach earlier into the development process than pre-deployment testing allows. The gap between what the AI Security Institute is designed to do and what it is permitted to do has been visible for three years. It has now been described, from the inside of the industry, as the thing standing between us and a technology its own builders say calls for extreme caution.

Three things that determine how well an organisation handles this:

  1. Independent observability: controls that generate your own evidence do not degrade when a supplier’s monitoring technique does.
  2. Contractual leverage applied early: audit and disclosure rights are cheap to obtain whilst the market is competitive and expensive to retrofit once it is not.
  3. Rehearsed containment: the ability to detect and stop an out-of-scope agent action, demonstrated in a drill rather than asserted in a policy.

Measuring the right thing

The metric most AI programmes report is adoption: seats deployed, workflows automated, hours saved. None of those numbers tells a board anything about the risk the essay describes. The number that does is time to containment, and almost nobody measures it.

Reframing the reporting line is not difficult. Alongside the adoption figures, report how many autonomous actions your systems took last quarter, how many fell outside their authorised scope, how you found out, and how long it took to stop them. If those questions cannot be answered from your own records, that is the finding.

Strategic Insight: Pachocki’s essay is best read by a UK board not as a warning about superintelligence, but as a supplier disclosure. The vendor has told you that its confidence in supervising its own product is falling and that it wants an outside party to hold it to a standard. Until that outside party exists with real authority, the standard is whatever you write into the contract.

Your next steps

Immediate (this week):

  • List every process where an AI system can take an irreversible action without human approval
  • Check whether your logs would let you reconstruct such an action after the fact, from your own records
  • Read the essay in full rather than the coverage, including our own

This quarter:

  • Add safety and monitoring disclosure obligations to AI supplier contracts at renewal
  • Run one containment drill and record time to detection
  • Establish who in the organisation owns the answer when a regulator asks what your agents did

This year:

  • Respond to any Frontier AI Bill consultation with an operator’s view on training-time access
  • Build portable controls that survive a change of supplier
  • Review whether your risk register still describes vendor safety frameworks as a mitigation

Source: An Alien Mind, an essay by Jakub Pachocki, chief scientist at OpenAI, published 6 September 2026.

This strategic analysis was written by Resultsense, a UK-focused AI news and analysis publication. We will be watching whether the Frontier AI Bill, if it arrives, gives the AI Security Institute authority that reaches into training rather than repeating the pre-deployment settlement. Read more analysis at Insights, or get in touch.