Seven widely used AI detectors, tested on 91 TOEFL essays written by people whose first language is not English, wrongly labelled 61.3% of it as machine-generated, according to research published in Patterns by Weixin Liang and colleagues at Stanford. The same tools handled essays by US eighth-graders almost perfectly. That is not a tool with an accuracy problem. It is a tool with an accuracy problem that has a postcode, and the buyers who deployed it never asked for the map.

Strategic Insight: A vendor sells you a headline accuracy figure. What you actually purchase is an error profile: a distribution of who the tool gets wrong, how often, and under what conditions. Those are different products, and only one of them appears in the procurement pack.

Writing for the Higher Education Policy Institute on 27 August, Anh Do, founder and chief executive of MAAS EdTech, walks UK universities through that evidence and reaches a conclusion that travels well beyond a marking queue. Detection tools, Do argues, can legitimately prompt a closer look. What they cannot do is stand in for the finding itself. Substitute any classifier a UK organisation is currently buying, and the argument holds without modification.

What does a one per cent error rate actually buy you?

The problem is not that classifiers make mistakes. Everyone knows classifiers make mistakes. The problem is that organisations treat the aggregate error rate as if it were a promise made individually to every person the system will process, and it is nothing of the sort.

Turnitin launched its AI detector claiming a false-positive rate of roughly one per cent. Vanderbilt University did the only sensible thing with that figure, which was to multiply it by its own volume. It had put 75,000 papers through Turnitin in 2022, so a one per cent rate implied that “around 750 student papers could have been incorrectly labeled as having some of it written by AI”. Vanderbilt switched the detector off. Several other US institutions stepped back over the same period, and UK universities have since followed.

The brochure figure is a population average, not a personal guarantee

A percentage is a statement about a distribution. It tells you nothing about which tail of that distribution your users occupy. One per cent across a homogeneous cohort of confident native-English writers is a very different instrument from one per cent across a cohort that is a quarter international, and the vendor’s benchmark almost never resembles the second.

This is where the Patterns numbers stop being an education story and become a procurement story. The detectors did not fail at random. They failed on a specific, identifiable, legally protected population, and they failed on it consistently. That is the signature of a proxy problem, and every classifier ever built has one.

MetricValueStrategic implication
Average false-positive rate, seven detectors on 91 human-written TOEFL essays61.3%The headline accuracy figure and the subgroup figure can differ by an order of magnitude
Human TOEFL essays flagged by all seven detectors19.8%Tools sharing a mechanism produce correlated errors, so unanimity is not corroboration
Human TOEFL essays flagged by at least one detector97.8%A “check it against a second tool” policy will manufacture a hit on almost anyone in this group
Turnitin’s stated launch false-positive rate applied to Vanderbilt’s annual volume~750 papersA rate that reads as trivial converts into hundreds of individual accusations
UK undergraduates using generative AI for assessed work, HEPI 202694%The base rate the classifier operates against has moved, which moves its precision with it

Strategic Reality: HEPI’s Student Generative AI Survey 2026, conducted by Savanta among 1,054 full-time UK undergraduates, found that the share putting AI-generated text directly into assessed work has climbed from 3% in 2024 to 12% in 2026. Genuine use is rising, which raises the pressure to detect. Respondents also told the survey they feared being blamed for misconduct they had not committed.

Why do the errors cluster instead of scattering?

Most detectors work by measuring text perplexity: how predictable the next word is to a language model. Varied, unexpected phrasing reads as human. Plain, repetitive, narrow-range phrasing reads as machine.

Now consider who writes in a narrow range. Somebody writing with care in a language they picked up later, drawing on a smaller stock of constructions precisely because they are being careful. That is the tell the system is trained to punish, and it is produced by diligence rather than deception. The Stanford team demonstrated the mechanism directly: using ChatGPT to enrich the vocabulary of the same TOEFL essays cut the false-positive rate by 49.7 percentage points, from 61.3% to 11.6%. Simplifying the vocabulary of the US essays pushed their misclassification up. Nothing about authorship changed in either direction. Only the linguistic texture did.

Critical Context: The paper’s own conclusion is a warning to buyers, not just to universities: “practitioners should exercise caution when using low perplexity as an indicator of AI-generated text, as such an approach could unintentionally exacerbate systemic biases against non-native authors within the academic community.” Read “practitioners” as anyone deploying the classifier, and the sentence becomes a procurement instruction.

The population at risk is wider than international students. A peer-reviewed review by Louie Giray, Hasan Nihat Caner and Abdolvahed Amoozegar, published in June 2026, finds that alongside non-native English writers, “high-achieving, original, and neurodivergent students are disproportionately accused”. The mechanism explains why. Formulaic structure, consistent register and repeated phrasing are common features of autistic and ADHD writing, and they are indistinguishable, to a perplexity metric, from the thing it was built to catch. A classifier that penalises linguistic uniformity will find it wherever uniformity comes from, including from a disability.

Four things classifier buyers routinely skip

  • Subgroup validation: A single accuracy figure is not a validation result. What matters is the conditional error rate for each population the tool will actually process, measured on data resembling yours rather than the vendor’s benchmark.
  • Error asymmetry: False positives and false negatives are not equivalent costs. Deciding which one you would rather absorb is a policy question, and it should be settled before the tool is switched on, not during the first appeal.
  • Mechanism transparency: Knowing that a detector measures perplexity is what makes its failure mode predictable. A classifier whose signal you cannot name is one whose errors you cannot anticipate.
  • Independence between tools: Two systems that measure the same proxy do not provide a second opinion. They provide the same opinion twice, with a spurious increase in confidence attached.

The steelman deserves a hearing, and it makes the point sharper

The fair conclusion is not that these tools never work. A 2024 study in Aloma compared 160 genuine student responses with 160 ChatGPT-generated responses on the same class task, and every human answer scored zero AI likelihood. Not one was wrongly flagged. Vendors have published rebuttals. Broader evaluations, including Elkhatat and colleagues in the International Journal for Educational Integrity, find performance swinging widely between tools rather than being uniformly poor.

That variance is the finding. No single accuracy figure is portable, because the rate at which a detector misfires shifts according to which product, which cohort and which threshold settings are in play. Which means an organisation cannot inherit an accuracy claim from a study, a vendor or a peer institution. It has to measure the rate in its own conditions, or accept that it does not know it.

⚠️ Warning: If your policy is to escalate when two detectors agree, the 97.8% figure should stop you. Nearly every non-native-English writer in the Stanford sample was flagged by at least one of the seven tools, and a fifth were flagged by all of them. Agreement between instruments that share a mechanism is a property of the mechanism, not evidence about the person.

Who actually carries the cost of a false positive?

The distribution of harm is the part that governance frameworks handle worst. A false positive is not an abstract accuracy decrement absorbed by the system. It is one identifiable person being asked to prove a negative, usually against an institution, usually with less information than the institution has, and usually with something material at stake.

Do’s sharpest observation is that an accusation of misconduct falls heaviest on exactly the group these tools flag most readily. That inversion, where the error lands hardest on whoever has the least capacity to contest it, is not unique to detection. It is what happens by default whenever a classifier’s failure mode correlates with a marker of disadvantage, and it is why “we reviewed the score before acting” is a weaker safeguard than it sounds.

Stakeholder groupPrimary impactSupport neededSuccess measure
Non-native English writers and international studentsDisproportionate flagging driven by linguistic texture rather than conductRight to see the evidence, contest it, and have the tool’s error profile disclosedFlag rate by cohort, tracked and published, not aggregate accuracy
Neurodivergent staff and studentsStructured or repetitive writing read as machine output; reasonable-adjustment questions raised lateAdjustment pathways established before deployment, not after an accusationZero cases where a disability-linked writing pattern is the sole basis for escalation
Frontline decision-makers (markers, HR screeners, case handlers)Asked to exercise judgement over a number whose derivation they cannot inspectTraining on what the score measures and what it cannot establishDocumented reasoning that stands on evidence other than the score
The deploying organisationEquality and data-protection exposure, plus reputational cost per overturned decisionSubgroup validation, published assumed error rate, auditable processOverturn rate on appeal, and defensibility of the evidence trail

The obligations already apply, and nobody has to wait for AI regulation

Two existing UK regimes bite here, and neither requires new legislation. Under the Equality Act 2010, a practice applied uniformly that puts people sharing a protected characteristic at a particular disadvantage is indirect discrimination unless it can be objectively justified. Race includes nationality and ethnic or national origins. Disability covers many neurodivergent conditions. A classifier with a documented differential error rate along either axis is not a neutral instrument applied equally; it is a practice with a measurable disparate effect, and the justification burden sits with the organisation applying it.

Second, Articles 22A to 22D of the UK GDPR, inserted by section 80 of the Data (Use and Access) Act 2025 and in force since 5 February 2026, turn on whether a significant decision was taken with meaningful human involvement. As we set out when examining agentic systems against that threshold, the phrase remains undefined, the Secretary of State’s power to define it is unused, and the regulator’s guidance has not arrived. A misconduct finding, a rejected application or a declined claim resting substantially on a classifier output is squarely inside the territory where that question will eventually be answered.

Neither regime asks whether the vendor’s benchmark was impressive. Both ask what the organisation knew about the tool’s behaviour on the person in front of it, and what it did with that knowledge.

How should a UK organisation buy a classifier?

💡 Implementation Framework: Error-Profile Procurement

Phase 1: Interrogate before signing (pre-contract)

  • Require the vendor to state the assumed false-positive rate and the population it was measured on
  • Name the underlying signal the model uses, and reason forward to which groups it will systematically misread
  • Multiply the stated rate by your own annual throughput, and read the answer as a count of people

Phase 2: Validate in your own conditions (first 90 days)

  • Run a shadow deployment against held-out data drawn from your actual population
  • Report error rates by cohort, not in aggregate, including the subgroups the mechanism predicts are at risk
  • Fix the escalation threshold against measured cost, not the vendor’s default

Phase 3: Build the process the score sits inside (ongoing)

  • Write down what evidence, other than the output, is required before any adverse decision
  • Log the human reasoning at the point of decision, because it cannot be reconstructed at appeal
  • Re-measure annually, and whenever the population or the base rate shifts

Priority actions by where you already are

Still evaluating a tool. First, ask the vendor for the subgroup breakdown and treat a refusal as the answer. Second, do the volume arithmetic in the business case, so the error count appears next to the licence cost. Third, decide in advance what the tool is permitted to trigger, and write it into the policy rather than leaving it to whoever reads the first amber score.

Already running one in production. First, publish your assumed false-positive rate internally and check whether anyone acting on the scores has ever seen it. Second, audit flag rates by cohort against your own population data; if you cannot produce that breakdown, that is the finding. Third, review every adverse decision from the last twelve months for cases where the score was effectively the sole evidence.

Running mature assurance. First, test whether your “second opinion” tools are genuinely independent or share a signal. Second, put the error profile into the equality impact assessment rather than the technical annexe, where it is a compliance artefact rather than a filed one. Third, treat base-rate drift as a monitoring obligation: as genuine AI use rises towards the 94% HEPI records, the meaning of a positive changes even when the model does not.

None of this is expensive. A shadow validation against held-out data is a few weeks of analyst time plus access to representative samples, which is small next to the cost of defending one indirect-discrimination claim, and it is the only route to an error rate anybody can actually stand behind.

Four problems that surface after the tool is live

Vendor accuracy figures do not transfer to your population

A benchmark result is a measurement taken somewhere else, on somebody else’s data, under settings you may not be running. It has no automatic validity for your cohort. The gap between 61.3% on TOEFL essays and near-perfect accuracy on US eighth-grade essays is the entire argument, produced by the same seven tools within one study.

Mitigation: Treat every external accuracy claim as a hypothesis about your deployment, not a property of it. Budget for local validation as a line item at procurement, and make the contract contingent on the vendor supplying the subgroup data needed to design that validation.

Correlated tools impersonate independent verification

Running a second detector feels like due diligence. When both measure perplexity, it is closer to asking the same person twice and recording two votes. The 19.8% unanimous-flag figure shows how far that correlation goes on the population where the mechanism misfires.

Mitigation: Before adopting a second instrument, establish that its signal differs. If it does not, the correct second source of evidence is not another model. It is process evidence: drafts, version history, and whether the person can explain their own reasoning aloud.

Base-rate drift quietly degrades precision

Classifier precision depends on how common the target behaviour actually is. As genuine generative AI use climbs towards near-universal, the population the detector is scanning changes shape, and the same model at the same threshold produces a different mix of true and false positives than it did at launch.

Mitigation: Schedule re-validation on a fixed cadence and tie threshold reviews to observed prevalence rather than to model updates. A tool that was calibrated in 2024 is not calibrated now, whether or not the vendor has shipped a new version.

The audit trail exists only if you build it beforehand

When a decision is challenged, the question is what evidence supported it at the time. A score is not a reason. Contemporaneous human reasoning is, but it cannot be manufactured retrospectively, and the absence of it is what turns a defensible judgement into an indefensible one.

Mitigation: Require the decision-maker to record, in the moment, what evidence other than the classifier output informed the outcome. Make that field mandatory in the workflow. It costs a minute per decision and it is the difference between a process you can explain and one you cannot.

Reality Check: None of this makes a classifier unusable. It makes an unexamined classifier unusable as evidence. The distinction is the whole of the argument, and it is the one that keeps getting collapsed at the point where the number appears on a screen.

The instrument is fine. The inference is the problem

Do’s formulation is the one worth carrying out of higher education and into every other sector buying classification: “As a prompt to ask a better question, they are defensible.” As the answer, they produce injustice, and they aim most of it at whoever a just process ought to be shielding first. The failure is not in the model. It is in the step where an organisation converts a probability into a verdict without ever having established what that probability is worth on its own population.

That step is happening right now in recruitment screening, where Stanford-led research has found AI hiring tools producing systemic racial disparities and women report being shut out of work by them. It is happening in policing, where the ICO has audited five forces on facial recognition. Detection in universities is simply the version with the best published evidence base, which makes it the cheapest place for everyone else to learn the lesson.

Three factors that separate the organisations that get this right

  1. They procure the error profile, not the accuracy figure. The question at contract stage is not how often the tool is right. It is who it is wrong about, and what that costs them.
  2. They keep the score subordinate to the process. The output opens an enquiry. Something else closes it, and that something else is written down.
  3. They measure locally and repeatedly. An accuracy number is a perishable good. Treating it as a permanent property of the tool is how a validated deployment becomes an unvalidated one without anyone noticing.

What a good deployment actually looks like

The tempting metric is detection accuracy, because vendors supply it and it goes in a slide. The better metrics are the uncomfortable ones: overturn rate on appeal, flag rate by cohort, and the proportion of adverse decisions where the classifier output was the only evidence on file. The third of those should be zero, and an organisation that cannot calculate it has already answered the question.

Do’s closing observation about UK higher education generalises without any strain at all. Britain hosts an enormous international student body, exactly the cohort these tools misjudge, which is why Do writes that “The fairness gap is not an import; it is already sitting in British marking queues.” Substitute the queue. British hiring pipelines, benefits assessments, fraud triage and content moderation are all queues, and every one of them contains people whose ordinary way of writing, speaking or behaving sits at the edge of what a model was trained to expect.

Strategic Insight: An unexamined classifier does not distribute its errors evenly and then apologise. It concentrates them on a predictable group and reports high confidence throughout. Knowing which group, before deployment, is the entire difference between a decision-support tool and a discrimination mechanism with a user interface.

Your next steps

Immediate actions (this week):

  • List every classifier making or informing decisions about people in your organisation
  • For each one, record the assumed false-positive rate and the population it was measured on
  • Multiply each rate by annual throughput and read the result as a headcount

Strategic priorities (this quarter):

  • Produce a flag-rate breakdown by cohort for at least your highest-volume classifier
  • Add the error profile to the relevant equality impact assessment
  • Make contemporaneous human reasoning a mandatory field before any adverse decision

Long-term considerations (this year):

  • Write subgroup-validation obligations into procurement templates and contract terms
  • Establish a re-validation cadence tied to observed prevalence, not vendor release notes
  • Test whether tools you treat as independent verification actually share a signal

Take Action: Three questions worth putting to any classifier in your organisation this month. What does it actually measure, as distinct from what it claims to detect? Whose ordinary behaviour does that measurement systematically misread? And if a decision it informed were challenged tomorrow, what would you show, other than the score?

Sources

This strategic analysis was written by Resultsense, a UK-focused AI news and analysis publication. We will be watching whether the ICO’s promised automated decision-making guidance addresses subgroup error rates, because that is the point at which “we had a human review the score” stops being an adequate answer. Read more analysis at Insights, or get in touch.