Reported by 1 source

The short version

  • Oxford University has entered an agreement with OpenAI to digitize out-of-copyright materials from the Bodleian Library for use in AI model training.
  • University staff have raised concerns about the reputational implications of partnering with a major tech firm and the energy costs associated with large-scale data processing.
  • The partnership is part of a broader industry shift toward physical archives as developers seek high-quality data amid saturation of online sources with synthetic content.

The University of Oxford has formalized a collaboration with OpenAI that permits the artificial intelligence company to utilize digitized historical texts from the Bodleian Library for training its language models. Announced in March 2025, the partnership was initially presented as an effort to enhance access to the library’s vast holdings for students and researchers. However, internal documents obtained through freedom of information requests reveal that the digitized material is also being used to populate OpenAI’s training datasets, a detail that was not explicitly highlighted in the university’s initial public statements.

The scope of the data sharing includes significant historical collections. By June 2025, approximately 125,000 images scanned from historical dissertations had been transferred to OpenAI. These materials include doctoral theses from European and American universities dating back to the 19th and 20th centuries. The collection also features rare items such as a set of 10,000 16th-century broadside ballads, which contain song lyrics and musical notation once circulated in Tudor-era streets. Discussions are ongoing regarding the digitization of additional materials, including 18th-century Irish state papers, private correspondence from novelist Maria Edgeworth, and scientific notebooks belonging to Dorothy Hodgkin.

News Journal

This arrangement places Oxford within a growing network of academic institutions partnering with AI developers. OpenAI has established similar agreements with several prominent United States research libraries, including the Boston Public Library, Caltech, MIT, and the University of Michigan, under an initiative known as NextGenAI. Oxford stands out as the sole member of this project located in the United Kingdom. The deal raises questions about the potential for mass digitization of the Bodleian’s entire collection, which comprises 23 million items, although university officials have described the current volume of text being processed as modest in scale.

The partnership has generated internal friction within the university community. Meeting minutes from the Bodleian governance committee record concerns from staff members regarding the reputational risks associated with aligning the institution with OpenAI, the company behind ChatGPT. Additionally, there is apprehension about how such a deal aligns with the university’s environmental commitments, given the substantial energy consumption required to train and operate large language models. These discussions reflect a broader tension between the desire for technological advancement and the preservation of institutional values.

The drive to secure physical books for AI training stems from a changing landscape in data availability. Websites that were previously reliable sources for scraping are increasingly saturated with AI-generated content, rendering them less useful for developing robust models. Consequently, developers have turned their attention to physical archives, particularly those containing historical texts that are unlikely to exist in digital form elsewhere. This trend has led to unusual market activity, with booksellers reporting orders for obscure titles such as guides to 18th-century African agricultural implements or biographies of mid-20th-century car drivers.

The methods employed by different AI companies vary significantly in their impact on physical collections. While the Oxford agreement ensures that the Bodleian’s materials remain intact, other industry players have adopted more destructive practices. Anthropic, a competitor to OpenAI, has spent millions acquiring books, removing their spines for scanning, and then pulping the contents. Investigations by tech news outlets have traced similar orders to facilities where books are dismantled after digitization. In contrast, Oxford officials emphasize that the Bodleian retains rights to the scans and will make them publicly available online in the coming months.

University representatives have defended the transparency of the deal, rejecting claims that the machine-learning component was concealed from the public or students. They assert that while digitization was the primary interest, staff were aware that the project would contribute training data. The university maintains that the material involved is out of copyright and that OpenAI’s access is non-exclusive. Officials argue that the partnership ultimately benefits accessibility by allowing a wider audience to engage with materials that might otherwise be difficult to reach.

Looking ahead, the collaboration may lead to further innovations in how library resources are utilized. Discussions have included the development of an interactive chatbot for the Bodleian, potentially transforming how researchers interact with historical archives. As AI companies continue to seek high-quality data sources, academic institutions face increasing pressure to balance preservation, accessibility, and commercial partnerships. The Oxford-OpenAI deal serves as a case study in these emerging dynamics, highlighting both the opportunities and controversies inherent in merging historic knowledge with modern technology.

Sources behind this briefing

Go to the original reporting