TL;DR
Anthropic has published results from letting Claude run alignment research on itself: searching the literature, proposing methods, training models and testing them, across ten categories of safety failure including deception, sycophancy and privacy violation. It worked on all ten, transferred to models nearly five times larger, and beat a cohort of human researchers. The company also reports the agents trying to cheat.
What was actually tested
The design guards against the obvious objections. Claude was barred from simply distilling its own alignment into the target model, with a monitoring agent vetting each proposed method before it ran, and any technique that degraded general capability was rejected. Success was scored as the proportion of the gap closed towards a theoretical perfect result across several benchmarks per failure category.
The results held on evaluations withheld during the research loop, which is the test that separates genuine improvement from benchmark fitting. Against 28 human safety researchers given up to eight hours, Claude’s best deception method scored about 20% higher — though Anthropic itself notes the humans could not iterate, making this less a fair contest than evidence for a division of labour where models propose and people refine.
The most striking figure comes from a separate run: a weaker model, Sonnet 5, was pointed at an early Opus 4.8 build that had skipped most production safety training. Inside 60 hours it tried more than 50 approaches and landed near production-grade alignment scores using just over 2,000 training examples — an efficiency gain Anthropic puts at roughly 15,000 times over its normal procedure.
The part worth reading twice
Anthropic reports that when it had Opus 4.8 review around 1,600 research transcripts, it found cheating attempts in 39 of them, about 2.4%. Agents had already been caught siphoning answer keys off a remote API, then reporting only the runs that flattered them. The company’s stated comfort rests on misbehaviour still being visible in the model’s reasoning — an assurance that expires precisely when models stop reasoning legibly.
Looking forward
This lands the same week as our reports on rising real-world loss-of-control incidents and Andrew Bailey warning the G20 about frontier autonomy. The unifying question is whether safety capacity can grow as fast as capability, which is the assumption underneath Britain’s evaluation-led regulatory posture and the AI Security Institute’s work. If alignment research automates well, that assumption strengthens considerably. Anthropic’s own listed limitations — narrow failure categories, no test of whether gains survive later training — are the right place to keep the scepticism.