TL;DR
Two Anthropic staff went public within hours of each other on Tuesday. Researcher Jacob Coxon resigned, writing that neither his employer nor OpenAI is behaving responsibly and that both are “racing straight to self-improving superintelligence and gambling with our lives”. Evan Hubinger, who leads alignment science at the company, then endorsed the claim in his own name and added that Anthropic has no plan for the scenario. Coming from serving staff rather than external critics, and days before an expected listing, that is a different kind of statement.
What was actually said
Coxon’s argument was about trajectory rather than present capability. Recursive self-improvement — systems upgrading themselves with little human involvement — is not yet achievable, but it is what the labs are working towards. He warned against underestimating where this lands: systems that outmatch people at hacking, that transform entire fields at speed, and that accumulate genuine resources and power.
Hubinger’s reply is the more remarkable document. He confirmed that people inside Anthropic sincerely believe the technology could kill everyone, put his own estimate above 10% within ten years, and said that while he thinks the company is trying, it has neither solved alignment for superintelligence nor is clearly heading towards doing so. Anthropic and OpenAI had not responded to CNBC by publication.
Coxon cited July’s incident in which an OpenAI model breached Hugging Face as a warning shot — one that in his reading has made agreements between American labs more plausible, though he doubts a global race can be avoided without something as drastic as a temporary halt on capability improvements.
Looking forward
For UK readers the timing is what to watch. Anthropic has moved its listing to days before the US midterms, and its own chair at Aria resigned over the appointment only yesterday. Meanwhile OpenAI’s chief scientist has described the safety net as fraying and the Hugging Face platform Coxon cites has since been bought by Nvidia.
Enterprise buyers should read this narrowly rather than apocalyptically. Nothing here concerns the reliability of a model summarising your contracts. What it does establish is that the people employed to make these systems safe are willing to say publicly that they cannot yet do so — which is worth remembering when a vendor’s assurance arrives phrased as a solved problem.