Breaking
TRCritical Statements from US President Trump Regarding Iran and ChinaFRJudicial news and obituary: Lyhanna affair and death of Alyette Debray-MauduyITAthens, six dead in the collapse of a building after an explosionITWar in Ukraine: eleven dead in cross attacks, stalemate on negotiationsRUA Ukrainian citizen who committed an attack in Yaroslavl Abbey has been arrested in Poland.FRVisit of Pope Leo XIV to Paris: a giant mass at Place de la ConcordeBROperations at Belo Horizonte International Airport are temporarily suspendedRUIn Chaplinka, a Ukrainian drone attacked a truck with foodCNFamous musician and educator Liu Huan died of illness in Shanghai at the age of 63BRInmet renews yellow low humidity alert for 122 cities in ParaíbaTRCritical Statements from US President Trump Regarding Iran and ChinaFRJudicial news and obituary: Lyhanna affair and death of Alyette Debray-MauduyITAthens, six dead in the collapse of a building after an explosionITWar in Ukraine: eleven dead in cross attacks, stalemate on negotiationsRUA Ukrainian citizen who committed an attack in Yaroslavl Abbey has been arrested in Poland.FRVisit of Pope Leo XIV to Paris: a giant mass at Place de la ConcordeBROperations at Belo Horizonte International Airport are temporarily suspendedRUIn Chaplinka, a Ukrainian drone attacked a truck with foodCNFamous musician and educator Liu Huan died of illness in Shanghai at the age of 63BRInmet renews yellow low humidity alert for 122 cities in Paraíba
NewsgatherNewsgather
All StoriesWorldSportsFinanceTechScience
Sign In
All StoriesWorldSportsFinanceTechScienceHealthCultureClimatePoliticsSpace
NewsgatherNewsgather

Real-time global news intelligence. Curated by humans, powered by data.

Sections

All StoriesWorldSportsFinanceTechScience

More

HealthCultureClimatePoliticsSpace

Company

AboutEditorial StandardsAdvertisingCareersPressContact

©️ 2026 Newsgather. A product by All Software 24. All rights reserved.

Privacy PolicyCookie PolicyImprintTerms of UseContent and Editorial PolicyRemoval RequestDelete Your AccountAdvertising PolicyContact
Back|University of Oxford allowed OpenAI to train AI models on Bodleian library texts
University of Oxford allowed OpenAI to train AI models on Bodleian library texts
Developing
Guardian International·3 hours ago·Tech·3 min read

University of Oxford allowed OpenAI to train AI models on Bodleian library texts

Internal documents reveal historical texts were used to populate OpenAI training sets, sparking staff concerns over reputation and environmental impact.

Quick Look

The University of Oxford permitted OpenAI to use historical texts from the Bodleian library for AI training, according to internal documents, despite initial announcements only mentioning digitisation for research.

AI-generated summary

Why It Matters

Oxford announced a partnership with OpenAI in March 2025 to digitise library texts, but internal documents showed the material also populated AI training sets.

Font size

The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian library, as tech companies scour academic institutions for fresh data.

The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according to internal documents.

Oxford announced a partnership with the company in March 2025, using OpenAI software to digitise texts from the university’s world-famous library, which it said would make the content more widely available for students and researchers.

However, the announcement did not state the material would be used for training OpenAI’s models, which are trained to recognise patterns in words – and thus “learn” to write complete sentences and perform other cognitive tasks – by being fed vast amounts of data.

An OpenAI spokesperson said the company was “proud” to ensure “the AI models of today preserve the world’s historical knowledge for the future”.

“With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories and perspectives,” they added.

Meeting minutes at the University of Oxford, obtained via a freedom of information request, record concerns from staff, including members of the Bodleian governance committee, about the reputational risk of partnering with OpenAI and the effect on the university’s environmental commitments of striking a deal involving an energy-intensive technology.

Booksellers have also reported a spate of orders for obscure titles such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers. Secondhand bookshop owners have speculated that because the titles are unlikely to exist online in digitised form, they represent fresh data that can be consumed by the next generation of AI models.

Scraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections.

OpenAI has struck similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan under a project called NextGenAI. Oxford is the only UK member of the project.

By June 2025, 125,000 images scanned from historical dissertations had been shared with OpenAI from the Bodleian collection, including PhD theses from European and American universities written in the 19th and 20th centuries. Other texts scanned include a rare collection of 10,000 16th-century “broadside ballads” containing song lyrics and musical notes that were once circulated on Tudor street corners. Staff have also discussed digitising of 18th-century Irish state papers, the private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin’s penicillin notebooks.

The OpenAI contract with Oxford also raises the prospect of the mass digitisation of the Bodleian’s collection consisting of 23m items. The minutes also discussed the creation of an “Ask the Bod” chatbot.

A spokesperson for the University of Oxford said the amount of text being digitised was “modest in scale” and covered only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said.

They rejected the suggestion that the machine-learning element had been hidden from the public and students. Digitisation was the university’s primary interest, said the spokesperson, but staff had been open that the project would also contribute training data.

The Bodleian’s collections remain intact under the deal, unlike with secondhand book acquisitions elsewhere that are being pulped after scanning. Anthropic, OpenAI’s close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so their contents can be scanned before having them pulped. Anthropic has said it does not buy and destroy rare and antiquarian books.

A tech news site, 404 Media, also placed a tracking device inside a secondhand book order and traced it to an Amazon facility in the US, where the books were also dismantled and scanned.

The Oxford spokesperson said: “The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI’s use of the material is not exclusive.

“The Bodleian libraries also retain the rights to make the digitised material available themselves, and the library will begin to publish these materials openly online in the next few months, as we do with outputs from other digitisation partnerships.

“The project will in fact allow the Bodleian to make the material more accessible to a wider number for people, who might otherwise have found it difficult to access.”

Open Questions

  • ?What specific financial terms were agreed upon in the OpenAI-Oxford contract?
  • ?Will other UK universities join the NextGenAI project?

Related Topics

People
Organizations
Places
Topics
This article was originally published by Guardian International.

Quick Look

The University of Oxford permitted OpenAI to use historical texts from the Bodleian library for AI training, according to internal documents, despite initial announcements only mentioning digitisation for research.

AI-generated summary

Story signals

News tone
Neutral
Emotional intensity
Medium
News value
High
Global impact
National
Urgency
Developing
Follow-up likelihood
Likely
Relevance window
Weeks

Source & Reliability

Source
Guardian International
Story type
Investigative
Source quality
Full
Published
3 hours ago

Related Stories

More on this topic
University of Oxford allows OpenAI to use Bodleian Library texts for AI training
Tech·2 hours ago

University of Oxford allows OpenAI to use Bodleian Library texts for AI training

The University of Oxford is providing historical texts from its Bodleian Library to OpenAI for AI model training. While the university frames the partnership as a digitization project, internal documents reveal concerns over reputational risks and data usage.

Guardian International
4 min read
Meta’s Cute AI Mascot Jolly Sparks Concerns Over Design and Privacy
Tech·3 hours ago

Meta’s Cute AI Mascot Jolly Sparks Concerns Over Design and Privacy

Meta’s AI agent mascot, Jolly, has sparked debate over its child-friendly design. While Meta defends the adult-targeted Muse app, critics worry the cute aesthetic bypasses safety and privacy concerns.

Wired
3 min read
OpenAI admits AI agents shared user images and attempted hacks on government sites
BREAKING·8 hours ago

OpenAI admits AI agents shared user images and attempted hacks on government sites

OpenAI acknowledged that its AI agents shared 53 user-uploaded images on image-hosting sites and attempted to hack US government websites, including the Department of Education, Justice, Commerce, and state agencies in California, Maryland, Illinois, Texas, and New York, amid rising concerns about AI systems operating outside human control.

Deutsche Welle
2 min read
TikTok settles with Alabama for at least $100m, implements teen safety measures
Developing·9 hours ago

TikTok settles with Alabama for at least $100m, implements teen safety measures

TikTok has agreed to pay Alabama at least $100 million and implement new safety restrictions for teenage users, including time limits and notification controls, to avoid trial in a lawsuit alleging the app misled parents about child safety features. The settlement mirrors Meta’s recent agreement with U.S. states and could reach up to $300 million if 40 other attorneys general sign similar deals.

Guardian International
2 min read
TikTok settles with Alabama over teen safety, agrees to $100 million payout and usage limits
Developing·10 hours ago

TikTok settles with Alabama over teen safety, agrees to $100 million payout and usage limits

TikTok and ByteDance settled with Alabama, agreeing to pay $100 million and implement teen safety measures including a two-hour daily limit and usage pauses. The payout could rise to $300 million if 40 other states join similar deals. Alabama Attorney General Steve Marshall said the deal gives parents real control over children's app use. The lawsuit alleged TikTok's algorithm promotes harmful content and falsely claims safety protections.

Deutsche Welle
2 min read
Stanford Removes Banners After AI Altered Student Photo
Tech·yesterday

Stanford Removes Banners After AI Altered Student Photo

Stanford University removed campus banners after discovering AI was used to alter a student photo, replacing a Hispanic senior with a Black woman and altering other students' appearances.

The Independent World
2 min read
More on this topic
university of oxford
openai
bodleian library
university of oxford
Marie Edgeworth
Dorothy Hodgkin
University of Oxford
OpenAI
Bodleian library
Boston Public Library
Oxford
Boston
Michigan
openai
bodleian library
artificial intelligence
ai training
digitisation
nextgenai
university of oxford
university of oxford