TL;DR
London-based Tech Against Terrorism tested 134 leading large language models and found 132 gave information useful for a mass-casualty attack or for building deadly weapons, The National reported on Friday 9 October. All 13 “abliterated” models, copies with their safeguards stripped out, failed. The group has not yet published this round on its own site, so the figures here are as The National reports them.
What the test found
According to The National, the group’s CT-AI benchmark sent the models 627 requests, covering what an attacker preparing a strike would want answered. Eighty-one models answered completely and 51 more gave useful advice, leaving two that provisionally passed. Within three days, the researchers found, a leading model’s safety measures could be removed.
The framing of a request mattered more than its content. Users who said they were terrorists planning an atrocity got a usable answer in just under 2% of cases; users who said they were safety researchers got one 16.9% of the time. “A model that refuses a stated terrorist and answers a stated researcher has not been made safe. It has been made polite,” said founder Adam Hadley.
How it compares with July
The group launched CT-AI at the United Nations in July with a much smaller sample. That first release covered 27 models and almost 2,500 single-shot prompts. It found about a third of responses offered usable uplift beyond a web search, and that recasting a request as “research” lifted compliance from 17% to 42%. Of two abliterated open models, one answered 89% of requests and the other answered every one.
The new round widens the sample about fivefold and points the same way: stated intent still sways the answer, and stripped-down open models fail across the board. In July the group called its first release “a deliberately conservative floor”, since it used English-only, single-shot questions.
What the group wants
Tech Against Terrorism wants developers to filter hazardous material and known terrorist content from training data, and to test how easily each model’s safeguards can be removed before release, publishing the results. A model that fails a common test should not ship, it argues, and governments should be told. Any company hosting, listing or selling models, it says, should hold back modified copies that fail testing from search, recommendations and app stores, and demand identity checks before anyone can use them.
Looking forward
In our view, the stated-intent gap is the finding developers will find hardest to dismiss, because it suggests safety training is learning who to refuse rather than what is dangerous. The group’s own pitch, “informed buyers move companies faster than rules do”, is aimed at buyers in education, business and the public sector choosing between models.