When the Bodleian Libraries announced their OpenAI partnership in March 2025, the pitch was access: out-of-copyright material, never before online, made searchable for students and researchers around the world. The word “train” does not appear in that announcement. It does not appear on the project’s own web page today either. Yet internal Oxford documents obtained by The Guardian describe the scanned material as used to “populate the OpenAI training set”. The gap between those two descriptions is the story for UK higher education. Universities and libraries are turning into suppliers of AI training data, and they are doing it through arrangements that carry other labels: digitisation, plagiarism checking, secondhand book sales.
Strategic Insight: The legal question here is mostly settled; the disclosure question is not. Oxford says everything shared is out of copyright. What the public record did not say, until a freedom of information request surfaced the minutes, is that model training was part of the purpose.
Our news report on the Bodleian documents covers what was scanned and what Oxford says. This analysis looks at the pattern the deal belongs to, and what it means for every UK institution holding text that AI developers want.
Why are AI companies coming to UK libraries now?
The Guardian’s reporting gives the reason plainly: the open web is filling up with text written by AI, which makes it less useful for training, so developers are turning to physical and often historical book collections instead. Libraries hold exactly that. So do universities, in the form of theses, dissertations and the essays students submit every term.
Three supply routes are now visible in the UK:
- Institutional partnerships. Oxford is the sole UK member of NextGenAI, the OpenAI programme whose US library and university members include MIT, Caltech, the University of Michigan and Boston Public Library. The Bodleian’s own project page says OpenAI funded the digitisation pilot as part of its five-year partnership with the university.
- Compulsory submission systems. Turnitin, used by 98% of UK universities according to the comparison site Uni Compare (as cited by the BBC), has proposed licence changes that students’ unions fear would let submitted work train AI.
- The secondhand market. Booksellers told The Guardian of a spate of unusual orders for little-known books, for example a guide to farm tools used in 18th-century Africa, or life stories of car drivers from the 1950s. Shop owners suspect the buyers want text that has never been digitised.
The critical numbers
| Figure | What it measures | Source |
|---|---|---|
| 3,500 | Global dissertations (1498 to 1884) named for digitisation in the March 2025 announcement | Bodleian Libraries |
| 125,000 | Images of Global Dissertations captured for digitisation, per the project page | Bodleian Libraries |
| 23 million | Items in the Bodleian collection; the OpenAI contract raises the prospect of digitising it en masse | The Guardian |
| $50 million | OpenAI’s commitment across the whole NextGenAI consortium, in grants, compute and API access | Bodleian Libraries |
| 98% | UK universities using Turnitin products, per Uni Compare | BBC News |
| 1 September 2027 | Earliest date Turnitin says its institutional agreement will be updated | BBC News |
Critical Context: No figure in the public record puts a price on the Bodleian material. The only money visible is OpenAI funding a research pilot, and a commitment spread across a consortium of 15 research institutions. Whatever the scans are worth as training data, that value is not being disclosed.
What is really happening in the Bodleian deal?
Read side by side, the public and internal accounts do not contradict each other so much as emphasise different halves of one arrangement. The March 2025 release describes a pilot to digitise public-domain material and make it available worldwide. Internal documents, as reported, describe the same scans flowing into OpenAI’s training set.
Oxford rejects the suggestion that anything was hidden. A university spokesperson told The Guardian that digitisation was the primary interest, that staff had been open about the project also contributing training data, and that the material is “modest in scale”, out of copyright and not exclusive to OpenAI. The library retains rights over the scans and plans to put them online openly within months.
Those are real protections, and they are better than much of what happens elsewhere. The collection stays intact. Anthropic, by contrast, has put tens of millions of dollars into buying books and cutting off their spines for scanning before pulping them, though it says it does not do this to rare or antiquarian books. 404 Media hid a tracker in a secondhand book order and followed it to a US Amazon site, where those books too were dismantled for scanning.
Where the openness claim is weakest
Oxford says the training element was not hidden from the public or from students. The public record is thinner. We checked the two public documents a student, researcher or member of the public would most likely find:
- the March 2025 announcement on the Bodleian website, and
- the Bodleian Digitisation Research project page, as it reads on 29 September 2026.
Neither contains the word “train” or the phrase “training data”. The project page lists seven research questions, covering scanning throughput, AI-enhanced metadata, collection audits and search. Model training is not among them.
Reality Check: Staff candour and public disclosure are different things. If the people whose work is being scanned, or the readers who use the library, would learn of the training element only through an FOI request, the institution has not really told them.
Scale is the open question
The spokesperson’s “modest in scale” describes what has happened so far. According to The Guardian, the contract raises the prospect of mass digitisation across the Bodleian’s 23 million items, and the minutes discuss an “Ask the Bod” chatbot. The project page is explicit that the pilot exists to test digitisation “at scale”. A modest pilot that proves a pipeline is a different proposition from the pipeline running across the whole collection, and the governance should be decided before that step, not after.
Who has a say, and who does not?
The three supply routes differ most in who can say no.
| Group | Material | Can they consent? | What they get |
|---|---|---|---|
| Authors of historical theses and ballads | Public-domain text | No role reported; Oxford says the material is out of copyright | Wider availability of their work once scans are published |
| Library readers and university members | Access to the collection | No say reported; staff concerns surfaced only in governance minutes released under FOI | Open online access; a possible chatbot |
| Students submitting coursework | In-copyright essays | Not meaningfully, their union argues, since submission is required for marking | Plagiarism checking; no royalty under the quoted licence |
| Secondhand booksellers | Physical copies | Yes, as sellers, though they can only speculate about the buyer’s purpose | The sale price |
| AI developers | All of the above | Not applicable | Fresh text that is scarce online |
The student case is the sharpest. Sam Dickinson of York Students’ Union put it to the BBC in terms that apply well beyond Turnitin: if a university uses the service and it uses student work to train AI, “you can’t meaningfully consent to that because you have to submit your work to get marked”.
Turnitin’s statement draws a careful line. It says it does “not use customer or student work to train the AI assistant within Turnitin Clarity”, but that it “may use anonymized student submissions to improve our detection and assessment tools”. Its licence, as quoted by Times Higher Education, grants the company a “non-exclusive, royalty-free, perpetual, worldwide, irrevocable” licence over papers submitted from outside the EU, limited to providing and improving the service.
Hidden Cost: “Royalty-free” and “perpetual” are the terms that matter for value. Whatever the essays are worth to an AI-enabled product, the licence as quoted carries no royalty for the people who wrote them.
Southampton has already decided not to renew its Turnitin contract after 2026-27, as we reported in August. Its reason, given to the BBC, was that even after the postponement Turnitin “has not fully clarified what future changes may be imposed”. Our coverage of the students’ unions’ campaign sets out where York, Lancaster, St Andrews, Reading and UEA stand.
What is the material actually worth?
Nobody in this story has published a number, and that is itself the finding. Thomas Lancaster, who holds a teaching fellowship at Imperial College London, told Times Higher Education that universities and students are “not recognising the value of that data”. He was talking about submitted essays, but the point transfers to library collections.
Consider what Oxford has and has not given away. The use is non-exclusive, and the Bodleian will publish the scans openly. In our view that is the right call for a public-interest library, but it has a consequence worth stating: once the scans are openly online, the Bodleian no longer controls who reads them, and that plausibly includes other developers’ training pipelines. Our reading is that OpenAI’s advantage lies less in owning the text than in getting it early: 125,000 dissertation images had been shared by June 2025, according to The Guardian, and open publication is still months away.
The bookseller evidence points the same way. If buyers are, as shop owners suspect, hunting titles precisely because they are not online, then rarity is what is being priced. A library’s undigitised holdings are rare by the same measure.
Strategic Reality: A UK library negotiating a digitisation deal is selling access to scarcity, whether or not the contract says so. Institutions that think of these as goodwill projects will underprice them.
What should UK institutions do now?
A three-phase approach
Phase 1: Disclose (this term)
- Audit every live agreement where text leaves the institution: digitisation partners, assessment platforms, repository licences
- Where training is a permitted or intended use, say so on the public project page, not only in committee minutes
- Tell students, in plain language, what their submission platform’s licence allows
Phase 2: Contract (before the next renewal)
- Write explicit training clauses: permitted, excluded, or permitted only for named products
- Keep the rights to scans and the right to publish them, as the Bodleian has
- Ask what “improve our detection and assessment tools” covers, and get the answer in the contract rather than a press statement
Phase 3: Price (before any scale-up)
- Treat training use as a distinct grant of value, separate from funding for digitisation equipment
- Decide in governance, before a pilot becomes a programme, what a collection-wide deal would require
Priority actions by starting point
For institutions with no AI data agreements yet
- Map which systems receive student or staff writing, and read their current licence terms
- Agree a default position on training use before a vendor proposes one
- Brief the students’ union early; York’s union raised concerns in August, ahead of Turnitin’s announcement of a delay
For institutions already in a digitisation partnership
- Check whether the public description of the project matches what the partner does with the scans
- Confirm rights retention and open-publication commitments in writing
- Record environmental and reputational assessments, as Oxford’s minutes show staff raising both
For institutions reviewing assessment platforms
- Use the gap before September 2027 to negotiate exclusions, as Turnitin says it wants time to discuss changes with customers
- Assess alternatives now, as Southampton is doing, so that walking away is a credible option
- Consider whether a consortium tool is realistic; Thomas Lancaster suggested universities could band together for one with more control of the data
Implementation Note: In our view, most of these steps need no new budget. They need someone to read the licences already signed and to publish what they say.
The hidden challenges
Challenge 1: The public record lags the private one
Internal records can run ahead of public pages. The Bodleian’s page still describes a digitisation research project with no mention of training.
Mitigation: Make the public description a condition of sign-off. If a use appears in the contract, it appears on the web page.
Challenge 2: Open access and training use are hard to separate
Publishing scans openly serves readers. It also puts the text within reach of anyone building a model, and it is hard to see how a library could open a collection to people while closing it to crawlers.
Mitigation: Decide deliberately whether open publication is intended to include machine use, and state that position in the collection’s terms of reuse.
Challenge 3: Compulsory systems cannot rely on consent
Where submission is a condition of assessment, the consent a student gives is not freely given in any useful sense, as York’s union argues.
Mitigation: Do not rely on consent at all. Exclude training use contractually, or offer a route to be marked that does not grant the licence.
Challenge 4: “Modest” pilots become large programmes
A pilot that proves the pipeline works invites the next phase. The contract already raises the prospect of a collection of 23 million items.
Mitigation: Require fresh governance approval, with public disclosure, before any pilot scales, and treat the scale-up as a new deal rather than an extension.
Warning ⚠️: Oxford staff flagged reputational risk in partnering with OpenAI at all. In our view the larger exposure is the public finding out second-hand. A disclosed arrangement can be defended; a discovered one has to be explained.
The strategic takeaway
UK universities and libraries hold what AI developers increasingly lack: large volumes of human-written text that is not already online. That makes them suppliers, whether or not they describe themselves that way. The Bodleian deal, with rights retained and collections intact, is one of the better-protected versions of this trade. Even so, its training element reached the public through an FOI request rather than the institution’s own pages.
Three things that decide whether institutions get this right
- Say what the deal is. If model training is a purpose, the public description names it.
- Keep control where it matters. Rights to scans, publication rights, and explicit training clauses on submitted work.
- Know what you are giving. Scarcity has a price, and a funded pilot is not the same thing as payment for training use.
Your next steps
Immediate actions (this month):
- List every agreement under which institutional or student text leaves your systems
- Read the licence terms of your assessment platform, including what “improve” covers
- Check public project pages against what partners actually do with the material
Short-term planning (before September 2027):
- Negotiate training exclusions or explicit terms with assessment vendors
- Add a disclosure requirement to digitisation partnership approvals
- Engage students’ unions on submission licences
Long-term considerations (this year):
- Set a governance route for any collection-wide digitisation offer
- Decide your position on machine use of openly published scans
- Watch whether other UK institutions join NextGenAI or sign similar deals
Source: Oxford lets OpenAI train its AI models on Bodleian Library (The Guardian, 26 September 2026), reported by Dan Milmo. Additional context from the Bodleian’s March 2025 announcement of the OpenAI collaboration and its digitisation research project page (read 29 September 2026), BBC News on students’ concerns over Turnitin, and Times Higher Education on Southampton’s decision.
This strategic analysis was written by Resultsense, a UK-focused AI news and analysis publication. We will be watching whether the Bodleian’s published scans and any scale-up come with a plain statement of training use. Read more analysis at Insights, or get in touch.