Rare Books for AI Training: What Builders Need to Know About Data Sourcing and Governance
npr.org

Rare Books for AI Training: What Builders Need to Know About Data Sourcing and Governance

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRAI labs are purchasing rare offline books, scanning them, and sometimes destroying them to source human-generated text for model training. This practice, revealed by a 404 Media investigation using a tracking device, raises important questions about data provenance, model quality, and governance for AI builders.

AI labs have been quietly buying rare books by the pallet, cutting off their spines, and feeding the pages through industrial scanners to extract text for training large language models. A 404 Media investigation tracked one such shipment to an Amazon AI training warehouse, confirming that the company is using these books to develop its AI products. For builders, this reveals a critical shift in how training data is sourced and why it matters for model quality and governance.

How a tracking device revealed Amazon's book-scanning pipeline

Reporter Emanuel Maiberg placed a tracking device in a rare book order from a bookseller who had noticed unusually large, thematically incoherent purchases. The shipment ended up at Amazon's VGT3 warehouse, which Amazon confirmed is used for AI training data. The books are often niche, self-published, or out-of-print titles that are not available online. To scan them quickly, workers remove the spines and feed pages through high-speed scanners. TechCrunch reports that the originals are sometimes destroyed after scanning. Amazon's statement frames the purchases as helping "develop and improve the products and services our customers use," but the investigation makes clear that AI training is a primary driver.

Why offline books matter for model training quality

The reason labs are turning to physical books is a technical problem called model collapse. When an AI model trains on AI-generated text, its output quality degrades over time. The internet has already been scraped of most human-generated content, so labs need fresh sources of human writing. As Maiberg explained, rare books that never made it online offer a reservoir of clean, human-authored text. This is why orders lack coherent themes: the goal is volume and diversity, not subject expertise. For builders, this means the models you use may be trained on data that was physically extracted from books that are now gone.

What this means for your training data governance

If you are building on top of a foundation model, you should ask your provider about data sourcing. The use of rare books introduces questions about copyright, preservation, and ethical sourcing. Ars Technica notes that the books being destroyed are not valuable first editions but niche titles that may still hold cultural or historical significance. For teams building custom models, documenting sourcing pathways is becoming a governance requirement. Consider whether your training pipeline relies on offline sources and what happens to those sources after scanning. The practice also raises concerns about access to historical texts: once a book is destroyed, its content exists only in a model's weights, not as a retrievable document.

The limits of what we know

This investigation is based on a single tracked shipment, and Amazon's statement is deliberately vague. Not all rare-book purchases may be for AI training; some could support other product development. The legal landscape is also evolving. A court ruling has made it legal for companies to destroy books after scanning, but copyright challenges are ongoing. The scale of the practice is unknown, though reports suggest hundreds of thousands of books have been processed. For now, the key takeaway for builders is that training data sourcing is becoming more complex and less transparent. Understanding where your model's data comes from is no longer just a legal checkbox; it is a quality and governance concern.

FAQs

AI labs need fresh human-generated text to avoid model collapse, where training on AI-generated data degrades quality. Since the internet has been largely scraped, rare books that are not available online provide a clean source of human writing. Investigations show that companies purchase these books in bulk, scan them, and sometimes destroy the originals.

Sources

Latest Tech News