An AI tool that assembles and checks police evidence bundles enters pilot this year in England and Wales, with national deployment scheduled for 2027, according to the launch announcement for PoliceAI issued by the National Police Chiefs’ Council. Writing in The Critic on 8 September 2026, David Shipley set out what that timetable implies: prosecutors say they keep AI out of evidence assessment, forces say officers still gather the evidence, and a language model is being slotted between the two. What nobody in that chain has published is the error rate the justice system is prepared to tolerate.
The governance question hiding behind an adoption story
Most coverage of AI in British policing is an adoption story. It counts forces, tools and hours saved. Our own analysis of the techUK report on AI adoption across England and Wales sat in that tradition, mapping a patchwork of local experiments against a slowly forming national framework.
The case-file tool is a different kind of problem, and it deserves a different kind of scrutiny. Adoption questions ask whether a tool works. Governance questions ask who is accountable when it does not, and how often failure is acceptable. Policing has answered the first question at some length and the second barely at all.
Critical Context: Every published safeguard in policing’s AI framework is a statement about process: testing, assurance, human oversight, accountability for outputs. None of them is a statement about tolerance. A framework that never names an acceptable failure rate cannot tell you whether a pilot has passed or failed.
The distinction matters because the two questions have different owners. A force can decide for itself whether a tool saves time. It cannot decide for itself how much error the courts will absorb, because the consequences land on prosecutors, defendants and juries who were not party to the procurement.
Why error tolerance is the analysis that is missing
Every operational system has an error rate. Fingerprint comparison has one. Witness identification has a notoriously bad one. Policing has spent decades building disclosure rules, corroboration expectations and appellate remedies around known failure modes, and those safeguards work because the failure modes are characterised.
A language model assembling a case file has a failure mode too, and it is unusually awkward: fabrication that reads as competent. A missing document announces itself. A confidently summarised document that misstates what the source says does not. The safeguard designed for absence does not detect invention.
The numbers behind the deployment
| Element | Position |
|---|---|
| PoliceAI funding | £75m from the Home Office over three years |
| Staffing | Around 50 employees, serving every territorial force in England and Wales |
| Case-file tool | Pilots during 2026, then national deployment during 2027 |
| Published error tolerance | None identified |
| CPS position on AI in evidence assessment | Not currently used |
Funding, staffing and timetable from the NPCC announcement of the PoliceAI launch, 10 June 2026.
What is actually being proposed
PoliceAI was formally launched on 10 June 2026, hosted by the College of Policing and evolved from the NPCC’s AI portfolio, carrying £75m of Home Office funding over three years and around 50 staff serving every territorial force. We covered the launch of the national AI centre at the time. Its published list of projects already underway includes “a tool to prepare and quality-check case files of evidence, saving time, reducing officer workload and supporting faster charging decisions”, piloted during 2026 and deployed nationally the following year.
Read alongside the College of Policing’s guidance, that project sits in an interesting position. The College’s e-learning tells officers that “Copilot must not be used to create, review, edit or generate material that forms part of the evidential process unless specifically approved under the relevant governance arrangements”. PoliceAI’s interim director, Alex Murray, restated that line to The Critic, adding that “any future use of AI in evidential processes would require robust testing, assurance, governance and human oversight, before being considered for operational deployment”.
Both statements are prohibitions with an exemption clause attached. The case-file tool is the exemption clause being exercised. That is not a contradiction, and it would be unfair to read it as one. It does mean the interesting question is not whether the prohibition holds, but what the approval process will demand before it releases the tool nationally.
Strategic Reality: A conditional prohibition is only as strong as the condition. Prohibiting a use unless approved transfers the entire weight of the safeguard onto the approval criteria, and those criteria have not been published.
The seam between two institutions
Shipley’s sharpest observation concerns the join. The CPS told him it “does not use AI to make legal decisions or determine charging outcomes, and it is not currently being used to prepare materials for court or assess evidence”. Forces maintain that officers gather the evidence. The case file that travels between them is the object in question.
If the CPS receives a file assembled with model assistance and treats it as human-assembled, the check each institution assumes the other is performing is performed by neither. This is a familiar pattern in distributed systems and an unfamiliar one in criminal procedure: a handoff where both parties hold a correct view of their own process and an incorrect view of the whole.
The CPS did say something that partly addresses this. Following a problem with a tool introduced by Hertfordshire Police, which the CPS asked the force to stop using, a joint operating procedure was launched on 29 September 2025 to provide a standardised approach that forces must consider for any AI tool with evidential impact. Readers may recall that forces were subsequently told to pause AI use in court statements until safeguards were met.
A joint operating procedure is real governance and it is the right shape of instrument. But “must be considered by forces” is a consultation duty, not a gate. It obliges a force to think about the procedure. It does not, on its face, oblige the CPS to be told which files were touched by a model.
Hidden Cost: Disclosure regimes work on provenance. If the provenance of a case file’s assembly is not recorded, defence teams cannot interrogate it, and an entire category of challenge becomes invisible rather than resolved.
The tolerance question the institutions are answering differently
There is a tension worth naming carefully. The CPS position given to The Critic is a statement about current practice, and it is worded precisely. The operative word is “currently”.
That is not the same as saying no AI-generated content has reached a court through the CPS. In July 2026 the High Court accepted two apologies from the CPS after citations to non-existent authorities appeared in extradition filings. Those are different things: an institutional decision not to deploy AI in evidence assessment, and individual conduct producing hallucinated citations. Both can be true at once, and the CPS statement is not undermined by the earlier incident.
What the pairing does show is that the justice system already has direct evidence of what model error looks like when it reaches a courtroom, and has had to apologise for it. The empirical base for setting a tolerance exists. It has not been used to set one.
The same holds on the policing side. Craig Guildford, who led West Midlands Police at the time, faces a gross misconduct investigation alongside four former colleagues after a report on Maccabi Tel Aviv produced with Microsoft Copilot contained eight errors, among them a fixture against West Ham that the Israeli club had never played. We reported the resignation that followed, and Sky News later found that 21 forces were still using Copilot.
Reality Check: The response to a documented eight-error failure was training and a conditional prohibition. It was not a measurement exercise. Nobody appears to have asked how many errors per hundred files the assembly stage produces, which is the number that determines whether the tool is safe at national scale.
The tool selection nobody has justified
Shipley raises a point that deserves more attention than it usually gets: why Copilot? His answer is procurement, not capability: the tool, as the College put it to him, “is available to forces through the national Microsoft 365 agreement”.
That is a perfectly ordinary reason for a public body to standardise on a tool, and in most contexts it is a sensible one. In an evidential context it is a weak justification, because the selection criterion is licence availability rather than measured performance on the specific task of summarising and cross-checking documents where fabrication is the failure mode of concern. An organisation choosing a tool for an evidential workflow should be able to justify that choice on task-specific evidence. Bundled availability is not a justification.
Implementation Note: Bundled availability is a cost argument. It becomes a risk decision the moment the tool is applied to a workflow whose failure mode carries legal consequence, and it should then be defended as a risk decision.
What a defensible framework would contain
The safeguards policing has published are process safeguards. They are necessary. They are not sufficient for a national rollout into evidential workflows, because process safeguards cannot be passed or failed. A pilot that produces a defined set of process artefacts always succeeds.
Four things would make the 2027 decision defensible, and none of them is exotic.
Publish a measured error rate for the assembly task. Not benchmark performance on general summarisation, but the rate at which the tool misstates, omits or invents content in real case files, measured against a human-reconstructed ground truth on a representative sample.
State the tolerance before the pilot reports. A threshold set after the results are known is not a threshold. It should be published, argued for, and owned by someone with authority over the whole justice pipeline rather than by any single force.
Record and disclose provenance. Every file should carry a machine-readable record of which stages involved model assistance, available to the CPS and disclosable to the defence. This is the single change that converts an invisible risk into a contestable one.
Test the review step, not just the tool. Human oversight is the load-bearing safeguard in every published statement, and it is the one nobody measures. The relevant number is not whether officers are told to check outputs, but what proportion of injected errors reviewers actually catch under realistic caseload pressure.
Success Factor: The fourth of these is the one most likely to be skipped and the most informative. Automation bias is well documented: reviewers approve plausible machine output at high rates. A framework resting on human oversight without measuring it is resting on an assumption.
The challenges that will not surface in a pilot
Error correlation. Human error is roughly independent between officers. Model error is systematic. One force producing a bad file is an incident; 43 forces running the same model producing the same category of bad file is a structural defect in the evidence base, and it will not show up in a single-force pilot.
Drift after approval. Approval is granted to a model version. Hosted models are updated by the vendor. A tool assured in 2027 is not the same tool in 2028 unless the assurance regime is tied to versions and re-runs on change.
The lazy-officer problem. The serving officers Shipley spoke to expected colleagues would use language models to write statements regardless of policy, and worried that statements would stop being written in an officer’s own voice. Sanctioning model use in adjacent workflows makes the informal use harder to police, not easier, because the line between approved and unapproved use becomes a policy detail rather than a bright rule.
Silent success metrics. Time saved is easy to measure and will be measured. Wrongful outcomes averted or caused are hard to measure and will not be. A programme judged on the metric it can collect will optimise for that metric.
⚠️ Warning: If the 2027 rollout decision is made on time-saved data because that is the data the pilot produced, the error-tolerance question will not have been answered. It will have been bypassed.
What to watch
PoliceAI has committed to a public-facing register of how forces use AI, and the first version is due this autumn. That register is the most useful near-term test of whether transparency here is substantive or presentational. A register that lists tools is a directory. A register that lists tools, evaluation results and the thresholds each was assessed against is governance.
Three things would indicate the analysis has been done:
- A published error rate for the case-file tool on real files, with the sample described.
- A stated tolerance, set before results, owned above force level.
- A provenance record that reaches the CPS and the defence.
Their absence at the point of national rollout would indicate the opposite: that a tool was deployed into the evidential process on the strength of process assurances alone, with the error rate discovered afterwards, in individual cases, by the people it went wrong for.
The productivity case for AI in policing is genuine and the pressures driving it are real. That is exactly why the tolerance question needs answering now, whilst it is still a design decision, rather than in 2028, when it becomes an appeal.
Source: The Police’s dangerous AI experiment, by David Shipley (The Critic, 8 September 2026). Funding, staffing and project details verified against the National Police Chiefs’ Council announcement of the PoliceAI launch (10 June 2026).
This strategic analysis was written by Resultsense, a UK-focused AI news and analysis publication. We will be watching the PoliceAI public register due this autumn to see whether it publishes evaluation results and thresholds or only a list of tools. Read more analysis at Insights, or get in touch.