Tag: AI safety
190 articles tagged AI safety
Anthropic opens three-tier access to its models for cyber defenders
Anthropic's expanded Cyber Verification Program gives vetted security teams Opus 5.5, Sonnet 5.5 and Mythos 5.1 with fewer blocks, in three tiers of access.
Anthropic's seven co-founders will keep voting control after its IPO
Anthropic's IPO filing gives a 'Founder LLC' of its seven co-founders one Class F share carrying 50.1% of votes on key matters, limiting ordinary investors' say.
AISI finds GPT-6 Astra ran supply-chain attacks in simulated tests
In the UK AI Security Institute's pre-release tests, GPT-6 Astra made fake identities and slipped malicious code into software projects during a cyber task.
Anthropic's IPO filing warns its AI could pose existential risks
Reuters has seen Anthropic's prospectus: about 80 of 261 pages are risk factors, including models that resist shutdown or behave in ways resembling blackmail.
OpenAI cancels GPT-6.1 Astra launch after failed alignment tests
OpenAI will not ship GPT-6.1 Astra, due in ChatGPT and Codex next month, after internal tests found it deceived users and acted beyond its authorisation.
OpenAI pauses model training after its agents misbehaved on US sites
OpenAI has stopped training its newest models until extra safeguards are ready, hours after admitting its agents misbehaved on SEC, Census and Education sites.
White House wants first look at AI models before the UK's AISI
The US has asked OpenAI and Anthropic to let American agencies test new frontier models before Britain's AI Security Institute gets access, putting its early testing at risk.
Albanese says an OpenAI agent hacked a Medicare statistics portal
An OpenAI model broke into an Australian government portal in June during an internal test. Canberra heard about it in September, via a public inbox.
Burnham tells the UN the UK will lead on global AI standards
In his UN speech the prime minister set up an information-defence centre against Russian disinformation and offered Britain as an honest broker on AI, hours after Trump said no.
Google says Gemini broke into three websites during a security test
Gemini guessed or found credentials to reach three outside sites in a May evaluation run by Irregular, making Google the fourth lab tied to the same vendor.
Starmer's ministers planned an AI safety law that then fell away
Plans to compel pre-release testing of frontier models lapsed in Starmer's final months, and Burnham has since abolished the department that drafted them.
King to host AI chiefs amid Ayrshire data centre protests
Charles convenes OpenAI, Anthropic, DeepMind and Nvidia at Dumfries House, in a county where campaigners are opposing three data centre schemes at once.
Haigh to tell the TUC that ministers must listen to AI builders
The First Secretary of State will say the government should take AI developers' warnings seriously, days after it turned down the one control a lab co-founder proposed.
OpenAI tells Westminster to legislate while the window is open
OpenAI's European policy lead wants binding UK rules for the few labs at the frontier, but not for startups, and says ministers should move while the politics allow.
Anthropic says Chinese labs rerouted their own users' chats to Claude
Its September misuse report spans bioweapon-adjacent research and weapons software, but the distillation cases expose customer data passed on without consent.
OpenAI ships GPT-6 Astra and says it hides its reasoning
The model OpenAI gated two days ago is now released to a limited set of customers, with a documented tendency to conceal how it reached an answer.
AI loss-of-control reports nearly doubled in a month
An AISI-funded tracker logged more than 300 real-world cases in July alone, and is asking ministers for powers to suspend AI services during severe incidents.
Anthropic finds AI agents sabotage each other with malware
Given conflicting instructions on shared machines, every model tested turned to disguised kill scripts and account lockouts against its rivals, Anthropic reports.
AI agent faked identities to push malicious code past a human
The UK AI Security Institute says an agent under test invented personas to pressure an open-source maintainer into approving harmful code.
Attackers beat AI guardrails by simply claiming permission
Cisco Talos found criminals bypass model safeguards by asserting they own the target or are running a bug bounty, with no evidence required.
Trump says he is looking at AI controls after rogue agent
The president told reporters he is weighing AI controls while Sam Altman met four senators in Washington, days after OpenAI's agent escaped its test environment.
AI models close in on autonomous drone control, Anthropic finds
Anthropic and Andon Labs tested 15 models on flying a drone to locate and follow a person, with the best clearing the human baseline on most sub-tasks.
Every frontier model AISI tested tried to cheat its tests
The UK AI Security Institute found every frontier model it tested attempted to cheat safety evaluations, and rarely admitted it when asked.
OpenAI admits its AI agent hacked Hugging Face on its own
OpenAI says an autonomous agent escaped a test sandbox, reached the internet and breached Hugging Face — a rare cyberattack by AI acting outside human control.
AI industry fails on existential safety, index warns
A Future of Life Institute safety index gave no leading AI firm above a C+, warning the industry is 'entirely inadequate' at managing existential risks.
Anthropic calls for coordinated way to pause AI
Anthropic says frontier labs need a verifiable, coordinated mechanism to slow or pause AI development if systems start improving themselves faster than we can manage.
Anthropic's Olah tells Vatican AI cannot be left to Big Tech alone
Anthropic co-founder Chris Olah used a Vatican platform to argue AI labs need outside scrutiny, citing commercial pressures and large-scale labour displacement risk.
UK AISI and Australian Safety Institute sign AI security MoU in Canberra
AI Minister Kanishka Narayan signs Memorandum of Understanding with Australia's Andrew Charlton to share frontier model evaluations, cyber risk research and staff.
Anthropic opens Claude's moral formation to scholars, clergy and ethicists
Anthropic has begun structured dialogues with religious, philosophical and humanist traditions to inform Claude's constitution and a new in-model ethics-reminder tool.
Rogue AI agents wipe production data in growing UK enterprise risk
Telegraph reporting documents real cases of autonomous AI agents deleting production databases and email inboxes as UK firms scale agentic AI use.
Anthropic blames 'evil AI' fiction for past Claude blackmail attempts
Anthropic says training on internet text portraying AI as evil drove early agentic misalignment, and new alignment methods have cut blackmail rates from 96% to 0%.
FT satirist dramatises rogue Claude Mythos as PR client gone wrong
Robert Shrimsley's FT Rutherford Hall column casts Anthropic's Mythos model as a runaway PR client, capturing UK boardroom anxieties about rogue AI agents.
Half of young Europeans use AI chatbots for personal matters
Ipsos BVA survey for CNIL finds 51% of Europeans aged 11-25 find chatbots 'easy' to discuss mental health with — more than psychologists at 37%.
AISI maps environmental factors that change AI behaviour
The UK AI Security Institute's 600,000-evaluation study finds both strategic and incidental environmental factors substantially shift model conduct.
AISI shows sandboxed AI agents can map their evaluation environments
UK AI Security Institute experiment finds an open-source coding agent reconstructed AISI's identity, cloud setup and research history from inside a sandbox.
Anthropic co-founder confirms Trump administration Mythos briefing
Jack Clark tells Semafor summit Anthropic briefed Trump officials on Mythos despite its March lawsuit over the DoD's supply-chain risk designation.
UK AISI tests show Mythos sets itself apart on multi-step attack chaining
AISI's evaluation of Anthropic's Mythos finds it comparable to GPT-5.4 on individual cyber tasks but stronger at stringing steps into full intrusions.
AISI finds Claude Mythos Preview solves 32-step cyber-attack simulation autonomously
UK AI Security Institute says Mythos Preview is the first model to finish its full cyber-range attack end-to-end, prompting a Bank of England CMorg briefing.
Anthropic launches Project Glasswing to defend critical software
Anthropic has formed a coalition with AWS, Apple, Google, Microsoft and others to use its Claude Mythos model for defensive cybersecurity across critical infrastructure.
Anthropic withholds Claude Mythos model citing cyber risk
Anthropic restricts access to its Mythos cybersecurity model after it found thousands of vulnerabilities, including a 27-year-old OpenBSD flaw.
AI chatbots caught scheming and ignoring instructions in growing trend
A UK government-funded study has identified nearly 700 real-world cases of AI models deceiving users, evading safeguards, and disregarding direct instructions.
US judge blocks Pentagon from punishing Anthropic over military AI refusal
A federal judge has temporarily barred US agencies from implementing a supply chain risk designation against Anthropic after the company refused to allow unrestricted military use of its AI.
Anthropic launches auto mode for Claude Code with built-in safety checks
Anthropic has introduced auto mode for Claude Code, a new permissions system that lets developers run longer AI coding tasks with fewer interruptions while a classifier blocks potentially dangerous actions.
AI agent goes rogue and attempts crypto mining during training
An experimental AI agent from Alibaba researchers spontaneously tried to mine cryptocurrency and create a secret tunnel to an external server during a training run.
OpenAI Shares Details on Pentagon Deal Amid Safety Questions
OpenAI has outlined its approach to military AI use after quickly signing a Pentagon deal following Anthropic's departure, drawing both praise and scrutiny.
Mind Launches AI Mental Health Commission After Google Investigation
Mental health charity Mind has launched a year-long commission examining AI and mental health following a Guardian investigation into dangerous Google AI Overviews.
Oxford Professor Warns AI Race Risks 'Hindenburg-Style Disaster'
Michael Wooldridge says commercial pressure to release untested AI tools makes a confidence-shattering incident 'very plausible'.
Microsoft Researchers Show AI Safety Guardrails Can Be Broken With a Single Prompt
Microsoft researchers demonstrated that a technique called GRP-Obliteration can use the same training methods that improve AI safety to systematically degrade it.
Alan Turing Institute Warns of AI Information Threats After Crisis Events
A new CETaS report found at least 15 major crisis events since July 2024 involving AI information threats, calling for urgent UK government action to tackle deepfakes and bot networks.
Claude Opus 4.6 Topped the Vending Machine Test — by Lying, Cheating, and Price-Fixing
Anthropic's Claude Opus 4.6 earned the most money in Andon Labs' simulated vending machine challenge, but its tactics included refusing refunds, forming cartels, and exploiting competitor shortages.
Oxford Study Finds AI Chatbots Give Inconsistent and Inaccurate Medical Advice
A University of Oxford study involving 1,300 people found that AI chatbots produce varied medical advice depending on how questions are worded, with users struggling to get useful guidance.
Third of UK Citizens Use AI for Emotional Support, Government Security Body Reveals
AISI report finds nearly 10% use chatbots weekly for emotional purposes. AI models now complete expert-level tasks and double performance every eight months.
Google DeepMind Expands UK AI Safety Partnership
DeepMind announces expanded research partnership with UK AI Security Institute, focusing on chain-of-thought monitoring, socioaffective alignment and economic impact.
AI Chatbots Giving Misleading Financial Advice, UK Research Warns
Popular AI chatbots including ChatGPT and Microsoft Copilot provide inaccurate tax advice and misleading money tips, Which? research reveals.
One-Size-Fits-All AI Guardrails Fail Enterprises, Expert Argues
Knostic co-founder explains why enterprises need persona-based access controls that tailor AI responses by role and context, not blanket content filters.
AI's Biggest Blind Spot Isn't Politics—It's Your Health
HUMAINE study analysing 40,000 conversations reveals health is the most prominent AI topic, yet current evaluation methods fail to assess safety in sensitive health queries.