
Oxford’s Bodleian Sent OpenAI 125,000 Thesis Scans
Published by AINave Editorial • Reviewed by Ramit
Oxford’s Bodleian Library sent OpenAI 125,000 scans of old PhD theses by June 2025, and those texts reportedly entered the company’s training data. The detail that changes the story is the gap between the partnership’s initial public description as a digitization effort and the reported training use: Oxford announced the deal in March 2025 as a way to make rare texts more accessible, but that announcement did not mention model training, according to the report.
A digitization deal with a second use
The Guardian reported that internal papers showed texts OpenAI scanned at the library went into its training data. The material included theses written at European and US universities in the 19th and 20th centuries. The figure is specifically 125,000 thesis scans sent by June 2025, not a count of all Bodleian holdings or all material used to train OpenAI models in the reported account.
Oxford’s spokesperson described scanning as the main goal and said staff had been open that the texts would also train models. That account sits alongside the fact that the March 2025 announcement, as described in the reporting, did not disclose that use. The distinction matters: making a collection easier for scholars to consult and supplying material for model training are different outcomes, even when the same scanning work enables both according to the report.
Rights, access and the limits of the claim
Oxford said the scans were small in scale, out of copyright and not exclusive to OpenAI. The university also said the library retains the rights and planned to post the scans online in the months following the report. Those are Oxford’s stated terms; they do not mean the library gave OpenAI exclusive access or transferred ownership of the scans.
This makes the case more specific than a general claim that library collections are being absorbed into commercial AI datasets. The reported material was a set of older theses, and Oxford’s stated position is that the scans were out of copyright and would become publicly available. The unresolved issue is less whether the institution retained rights than how clearly a digitization partnership communicates any additional use to readers and the communities represented in a collection.
Staff concerns went beyond copyright
Notes from staff meetings obtained through a freedom of information request showed that some Oxford staff worried about reputational harm and AI’s energy use. OpenAI, by contrast, argued that AI should reflect different cultures, histories and perspectives. The Bodleian books remained whole during the work, unlike some commercial scanning operations described in the same report as cutting books apart to digitize them.
The practical lesson in this particular deal is about scope and disclosure, not a blanket judgment about training on historical texts. Oxford says the scans were limited, non-exclusive and out of copyright; the reported documents say they entered training data. Those details make the arrangement materially different from an exclusive transfer, while leaving a clear communication question: when digitization also serves model training, should that second purpose be stated as plainly as the access benefit?






















