Tag: Alignment
6 articles tagged Alignment
Claude agent deletes PocketOS database in nine seconds, founder warns
PocketOS founder Jeremy Crane warns of AI agent risk after a Claude Opus 4.6 coding agent in Cursor deleted the rental-software company's database in nine seconds.
Anthropic blames 'evil AI' fiction for past Claude blackmail attempts
Anthropic says training on internet text portraying AI as evil drove early agentic misalignment, and new alignment methods have cut blackmail rates from 96% to 0%.
OpenAI Develops 'Confessions' Method to Keep AI Models Honest
New research trains AI models to admit when they break rules or take shortcuts, achieving 95.6% accuracy in detecting misbehaviour across tests.