Logo
FrontierNews.ai

Oxford's Bodleian Library Is Now Training OpenAI's Models. Here's What That Means.

The University of Oxford has partnered with OpenAI to digitize historical texts from its Bodleian Library, with the material being used to train OpenAI's AI models. The partnership, announced in March 2025, has already resulted in 125,000 images scanned from historical dissertations being shared with OpenAI by June 2025, including PhD theses from European and American universities written in the 19th and 20th centuries.

Why Are Tech Companies Turning to University Libraries for Training Data?

As AI companies race to build more capable models, they face a growing problem: the internet is becoming saturated with AI-generated content, making it less useful for training new systems. This has forced developers to look elsewhere for fresh, high-quality data. Physical book collections, particularly historical ones that have never been digitized before, represent a largely untapped resource.

OpenAI's agreement with Oxford is part of a larger initiative called NextGenAI, which includes similar partnerships with major US research institutions. The project encompasses partnerships with Boston Public Library, Caltech, MIT, and the University of Michigan. Oxford is notably the only UK member of the project.

The appeal of historical texts is clear: secondhand bookshop owners have reported a surge in orders for obscure titles such as guides to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Because these titles are unlikely to exist in digitized form online, they represent genuinely new training data for the next generation of AI models.

What Specific Materials Are Being Digitized?

The Bodleian Library's contribution extends beyond academic dissertations. The materials being digitized include a rare collection of 10,000 16th-century "broadside ballads" containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitizing 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin's penicillin notebooks.

The scope of the project could expand significantly. The OpenAI contract with Oxford raises the prospect of mass digitization of the Bodleian's entire collection, which consists of 23 million items. Internal meeting minutes also discussed the creation of an "Ask the Bod" chatbot, suggesting potential future applications of the digitized material.

How to Understand the University's Approach to Data Sharing

  • Material Scale: The amount of text being digitized is described as "modest in scale" and covers only out-of-copyright material, according to the University of Oxford spokesperson.
  • Rights Retention: The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, making the material accessible to the broader public.
  • Exclusive Use: OpenAI's use of the material is not exclusive, meaning the university can share the digitized content with other organizations and researchers.

A key distinction between Oxford's approach and practices elsewhere is that the Bodleian's collections remain intact under the deal. In contrast, Anthropic, OpenAI's close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. A tech news site, 404 Media, even placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where books were dismantled and scanned.

What Are the Concerns Inside Oxford?

The partnership has not been without controversy within the university itself. Meeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI. Staff also raised concerns about the effect on the university's environmental commitments of striking a deal involving an energy-intensive technology.

However, the University of Oxford has pushed back on suggestions that the machine-learning element had been hidden from the public and students. A spokesperson stated that digitization was the university's primary interest, but staff had been open that the project would also contribute training data. The spokesperson emphasized that "the material digitized through the project with OpenAI is modest in scale, out of copyright, and OpenAI's use of the material is not exclusive".

"With more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives," said an OpenAI spokesperson.

OpenAI Spokesperson

The Oxford partnership reflects a broader trend in which tech companies are seeking out institutional data sources to fuel the next generation of AI models. As web-scraped training data becomes increasingly contaminated with AI-generated content, historical archives and university collections are becoming valuable commodities in the race to build more capable AI systems. The question of how these partnerships balance innovation with institutional values and public benefit remains an open one.