TL;DR
Researchers randomised over 1,000 Bocconi University first-year undergraduates into four groups: tool access, reasoning training, both, or neither. On a five-point rubric, the ChatGPT group gained close to a whole point. Those given reasoning training scored no better — but produced ideas their peers had not thought of.
The design
Students tackled a genuine brief, developing marketing proposals for their university’s merchandise shop. Assignment to groups was by class period. The training arm had nothing to do with AI: it taught causal reasoning, linking cause to effect and articulating why a proposal might fail, through a game with worked examples and feedback.
Marking ran on two tracks. Human graders applied a conventional five-point rubric, while automated analysis counted how many ideas each submission contained, how varied they were, how much causal reasoning appeared, and how closely the work resembled submissions from three experts.
What separated the groups
The tool group produced work that read as more professional — richer in ideas, better structured, closer to expert output. Notably, students were not simply forwarding the brief to a chatbot; they still had to frame questions, judge the answers and decide what survived into the final piece.
The training group is the more instructive result. Their rubric scores did not move, because the rubric measured only whether recommendations addressed two set marketing objectives. The text analysis caught what the marking scheme could not: a wider spread of ideas, more distinct from what everyone else submitted. A well-organised conventional answer scores well. An answer nobody else thought of does not necessarily score at all.
Students who got both showed the variety of the training group and the polish of the tool group, with the broadest gains overall.
The caveats worth stating
This is vendor-adjacent research — OpenAI’s economic research team co-authored it — run at one institution, on one business case, using GPT-4o rather than a current model. The randomised design is genuine and the finding is not flattering to the sponsor’s product in isolation, but it is not independent evidence.
Looking forward
The implication for UK institutions is about assessment design, not AI policy. If a polished answer no longer distinguishes students, marking schemes that reward polish stop measuring anything useful. Resultsense reported this week that AI markers grade essays generously and fail hardest on the weakest work; here, a human rubric misses originality entirely. Both point the same way: the instrument needs revisiting before the debate about the tool is settled.