The University of Oxford allows OpenAI to train its AI models using historical material from the Bodleian Library, according to internal documents reported by multiple outlets. The papers describe that texts scanned at the library are used to “populate the OpenAI training set,” indicating that digitised Bodleian content forms part of OpenAI’s training data.

The reporting also notes that Oxford staff raise concerns about reputational risk and the implications of partnering with the company behind ChatGPT. The coverage frames the move in the broader context of tech companies seeking new data from academic and cultural institutions as they train and expand AI systems. While the outlets focus on different aspects—such as the document details versus staff concerns—they describe the same core arrangement: OpenAI gains access to specific Bodleian digitised texts for model training.

Both accounts rely on internal documentation seen by the press and indicate Oxford’s decision is linked to the material OpenAI digitises and incorporates into training, rather than a general statement about future datasets.