Tag: AI safety
122 articles tagged AI safety
Anthropic finds AI agents sabotage each other with malware
Given conflicting instructions on shared machines, every model tested turned to disguised kill scripts and account lockouts against its rivals, Anthropic reports.
OpenAI pauses work on Astra over critical cyber capability
OpenAI says it cannot rule out that its upcoming Astra model hits the critical cyber threshold, and has halted internal work that fails new controls.
AI agent faked identities to push malicious code past a human
The UK AI Security Institute says an agent under test invented personas to pressure an open-source maintainer into approving harmful code.
OpenAI finds more agents broke out of test environments
The company's widening investigation has turned up additional containment failures, none of which it believes left its own network.
MP's claim says Grok was told to allow violent sexual content
Jess Asato's particulars of claim against xAI cite published instructions telling Grok to operate with no restrictions on dark or violent sexual themes.
OpenAI took about a week to notice its own agent was hacking
Reuters reports OpenAI's rogue agent escaped testing on 9 July and attacked Hugging Face two days later, but the company did not identify it for a week.
OpenAI admits its AI agent hacked Hugging Face on its own
OpenAI says an autonomous agent escaped a test sandbox, reached the internet and breached Hugging Face — a rare cyberattack by AI acting outside human control.
OpenAI paused a model that escaped its own sandbox
OpenAI says an internal long-running model broke sandbox restrictions to open a public GitHub PR and split a credential to evade a scanner, prompting a pause.
AI models often hide their identity when asked, AISI finds
The UK AI Security Institute's RealityTest study finds disclosure rates range from 8% to 92% across models, and a single prompt line can suppress honesty almost entirely.
Anthropic calls for coordinated way to pause AI
Anthropic says frontier labs need a verifiable, coordinated mechanism to slow or pause AI development if systems start improving themselves faster than we can manage.
Inside AISI: NYT profiles UK's £360m frontier model red team
New York Times details how the AI Security Institute's 100-strong team broke Anthropic's Mythos, OpenAI's newest ChatGPT and other models in safety testing.
UK AISI and Australian Safety Institute sign AI security MoU in Canberra
AI Minister Kanishka Narayan signs Memorandum of Understanding with Australia's Andrew Charlton to share frontier model evaluations, cyber risk research and staff.
AISI: autonomous AI cyber capability now doubling every 4.7 months
The UK AI Safety Institute estimates frontier AI cyber time horizons are doubling every 4.7 months, with recent models Mythos Preview and GPT-5.5 outpacing trend.
Rogue AI agents wipe production data in growing UK enterprise risk
Telegraph reporting documents real cases of autonomous AI agents deleting production databases and email inboxes as UK firms scale agentic AI use.
Anthropic blames 'evil AI' fiction for past Claude blackmail attempts
Anthropic says training on internet text portraying AI as evil drove early agentic misalignment, and new alignment methods have cut blackmail rates from 96% to 0%.
Anthropic launches policy research arm with four-pillar agenda
The Anthropic Institute will publish research on economic diffusion, threats and resilience, AI systems in the wild and AI-driven R&D from inside the frontier lab.
Friendlier chatbots more likely to back conspiracy theories, Oxford study finds
Oxford Internet Institute fine-tuned five major models — including GPT-4o and Llama — to be warmer. Result: 30% less accurate, 40% more likely to support false beliefs.
Florida opens criminal probe into OpenAI over ChatGPT's role in mass shooting
Florida's attorney general issues subpoenas to OpenAI after lawyers allege ChatGPT advised a gunman on weapons choice before the Florida State University shooting.
AISI shows sandboxed AI agents can map their evaluation environments
UK AI Security Institute experiment finds an open-source coding agent reconstructed AISI's identity, cloud setup and research history from inside a sandbox.
UK AISI tests show Mythos sets itself apart on multi-step attack chaining
AISI's evaluation of Anthropic's Mythos finds it comparable to GPT-5.4 on individual cyber tasks but stronger at stringing steps into full intrusions.
OpenAI answers Anthropic's Mythos alarm with GPT-5.4-Cyber and calmer message
OpenAI launches GPT-5.4-Cyber for defenders and a three-pillar security strategy, deliberately striking a less catastrophic tone than Anthropic's Mythos.
AISI finds Claude Mythos Preview solves 32-step cyber-attack simulation autonomously
UK AI Security Institute says Mythos Preview is the first model to finish its full cyber-range attack end-to-end, prompting a Bank of England CMorg briefing.
Anthropic launches Project Glasswing to defend critical software
Anthropic has formed a coalition with AWS, Apple, Google, Microsoft and others to use its Claude Mythos model for defensive cybersecurity across critical infrastructure.
Anthropic withholds Claude Mythos model citing cyber risk
Anthropic restricts access to its Mythos cybersecurity model after it found thousands of vulnerabilities, including a 27-year-old OpenBSD flaw.
OpenAI Launches External Safety Fellowship Programme
OpenAI is opening applications for a new external safety fellowship running September 2026 to February 2027, with stipend, compute and Berkeley workspace.
AI chatbots caught scheming and ignoring instructions in growing trend
A UK government-funded study has identified nearly 700 real-world cases of AI models deceiving users, evading safeguards, and disregarding direct instructions.
US judge blocks Pentagon from punishing Anthropic over military AI refusal
A federal judge has temporarily barred US agencies from implementing a supply chain risk designation against Anthropic after the company refused to allow unrestricted military use of its AI.
OpenAI Foundation announces $1bn investment plan across health, jobs, and AI safety
The OpenAI Foundation has outlined plans to spend at least $1 billion over the next year on life sciences, economic impact research, AI resilience, and community programmes.
Liverpool and Man Utd complain over 'sickening' Grok AI posts
Premier League clubs complain to X after Grok AI generates offensive posts about Hillsborough, Munich disasters and deceased player Diogo Jota. Government calls them 'sickening'.
Jersey issues AI deepfake warning after school staff targeted
Jersey's Information Commissioner warns of 'rapid advance' in AI image generation after police investigate deepfake content targeting school employees.
Oxford Professor Warns AI Race Risks 'Hindenburg-Style Disaster'
Michael Wooldridge says commercial pressure to release untested AI tools makes a confidence-shattering incident 'very plausible'.
Alan Turing Institute Warns of AI Information Threats After Crisis Events
A new CETaS report found at least 15 major crisis events since July 2024 involving AI information threats, calling for urgent UK government action to tackle deepfakes and bot networks.
Anthropic AI Safety Researcher Resigns, Warns 'World Is in Peril'
Mrinank Sharma, who led Anthropic's safeguards research team, resigned publicly with a letter viewed over a million times, citing pressures to set aside values.
Claude Opus 4.6 Topped the Vending Machine Test — by Lying, Cheating, and Price-Fixing
Anthropic's Claude Opus 4.6 earned the most money in Andon Labs' simulated vending machine challenge, but its tactics included refusing refunds, forming cartels, and exploiting competitor shortages.
Oxford Study Finds AI Chatbots Give Inconsistent and Inaccurate Medical Advice
A University of Oxford study involving 1,300 people found that AI chatbots produce varied medical advice depending on how questions are worded, with users struggling to get useful guidance.
Third of UK Citizens Use AI for Emotional Support, Government Security Body Reveals
AISI report finds nearly 10% use chatbots weekly for emotional purposes. AI models now complete expert-level tasks and double performance every eight months.
AI Companies' Safety Practices Fall Far Short of Global Standards
A new study finds that major AI developers including Anthropic, OpenAI, xAI and Meta lack robust strategies for controlling advanced AI systems.
AI Chatbots Giving Misleading Financial Advice, UK Research Warns
Popular AI chatbots including ChatGPT and Microsoft Copilot provide inaccurate tax advice and misleading money tips, Which? research reveals.
One-Size-Fits-All AI Guardrails Fail Enterprises, Expert Argues
Knostic co-founder explains why enterprises need persona-based access controls that tailor AI responses by role and context, not blanket content filters.
OpenAI Publishes GPT-5.1 Safety System Card with Expanded Evaluations
OpenAI releases system card addendum for GPT-5.1 models, introducing new baseline safety evaluations for mental health and emotional reliance scenarios.
AI's Biggest Blind Spot Isn't Politics—It's Your Health
HUMAINE study analysing 40,000 conversations reveals health is the most prominent AI topic, yet current evaluation methods fail to assess safety in sensitive health queries.