TL;DR

Historical texts that OpenAI digitised at Oxford’s Bodleian Library are being used to train its models, according to internal university documents seen by The Guardian. When Oxford announced the partnership in March 2025, it presented it as a way to open the collection up to researchers, with no reference to model training. The university says the material is out of copyright, the deal is not exclusive and staff were open about the training element. It is the second story in a week about UK academic writing becoming AI training material.

What went to OpenAI

Internal Oxford documents say the digitised material was used to “populate the OpenAI training set”. By June 2025 the Bodleian had passed OpenAI 125,000 scanned images of 19th- and 20th-century doctoral theses from European and American universities. Also scanned: some 10,000 Tudor “broadside ballads”, printed song sheets with lyrics and music. Candidates staff have discussed for scanning include the notebooks Dorothy Hodgkin kept on penicillin, 18th-century Irish state papers and private letters of the Irish novelist Maria Edgeworth.

Meeting minutes released under freedom of information record unease from staff, some of them on the library’s governance committee. They worried about the damage an OpenAI tie-up could do to Oxford’s name, and about how an energy-hungry technology squares with the university’s environmental pledges. The contract also raises the possibility of digitising far more of the Bodleian’s 23 million items, and an “Ask the Bod” chatbot was discussed.

Why old books matter to AI labs

Web text is increasingly polluted by AI-generated writing, so developers are turning to physical and historical collections. Oxford is the only UK institution in NextGenAI, OpenAI’s programme with US libraries including MIT and Boston Public Library. Secondhand booksellers report odd orders for obscure titles unlikely to exist online, which they suspect are being bought as fresh training data. Anthropic has bought books in bulk, at a cost running to tens of millions of dollars, and sliced off their spines for scanning, though it says it does not do this to rare or antiquarian volumes.

An Oxford spokesperson said the scanned material is “modest in scale”. The library retains its own rights over the digitised copies and expects to put them online for free within months. The collection itself stays intact.

Looking forward

Last week, students’ unions pushed back on Turnitin’s licence plans over fears that submitted essays could train AI. The Bodleian case is different in that the texts are out of copyright, but the question is the same one: who is told, and when, that academic material will end up inside a commercial model. In our view, other UK libraries weighing digitisation offers from AI companies should expect disclosure, not copyright, to draw the scrutiny.