TL;DR
Researchers put 50 undergraduate bioscience essays through two versions of ChatGPT, scoring each against seven criteria under four prompting setups, and compared the results with the original human marks. The models were generous almost throughout, diverged by as much as 40 marks on a 100-point scale in one instance, and could not be relied on to indicate how a student had actually performed.
Consistent, but consistently wrong
The failure mode is more interesting than the headline. Given an identical prompt and an identical essay, the models produced stable marks — they are repeatable. What they are not is accurate: across essays of differing quality, the marks they awarded tracked the human marks poorly, and the study describes them as inadequate predictors of the grade a human assigned.
The error also has a shape. Weak essays were lifted; strong essays were pulled down; work in the middle of the range came closest to agreement. That is regression toward the mean rather than noise, and it damages exactly the judgements that carry consequences — identifying who is failing and who is excelling.
One further detail deserves attention from anyone tempted by an overall score. Aggregate marks lined up with human ones reasonably well on average, while the individual criteria underneath diverged considerably. Averaging concealed the disagreement rather than resolving it.
Beyond the seminar room
Cardiff’s William Kay, a co-author, was blunt that generative AI cannot at present assign marks to subjective written work comparably to a human, even after extensive tuning, and pointed separately at the ethics of putting student work into these tools without express consent.
The finding travels well past universities. Any organisation using a model to score human output — sifting applications, rating support tickets, reviewing written submissions — has bought the same leniency bias and the same compression at both ends. Yesterday’s Nuffield review of justice AI made a related point: systems get evaluated on throughput, not on whether the judgement they produce holds up.
Looking forward
The paper suggests sharper rubrics might help, since terms like “good” and “outstanding” give a model little to separate. Kay is doubtful alignment gets easy, partly because human markers disagree with each other too. For UK employers, the safer reading is that these tools rank the middle acceptably and misjudge the edges — which is where decisions actually get made.