
Destructive Scanning of Rare Books for AI Training Raises a Data Provenance Problem
Published by AINave Editorial • Reviewed by Ramit
The reported destructive scanning of rare books for AI training exposes a trade-off that AI builders should treat as a data governance issue, not only a copyright dispute. Suppliers including ISBNdb are reportedly providing physical books to AI developers, while high-speed scanning can involve removing spines, separating pages, and destroying the original volume. The claimed benefit is cleaner human-written material. The cost may be permanent loss of scarce cultural sources. The reported scanning process and ISBNdb's rationale are described here.
Why pre-2022 books are attractive training data
ISBNdb argues that books published before 2022 are less likely to contain text generated by modern chatbots. It also describes older books as dense, edited, and authoritative compared with parts of the web that may contain synthetic or low-quality material. That is a plausible dataset-quality argument, but it remains a supplier's rationale rather than independent evidence that every pre-2022 book improves model reliability.
For model developers, the distinction matters. A publication date is only one provenance field. Teams still need to know which edition was scanned, whether pages were complete, how optical character recognition errors were handled, and whether the resulting text can be audited or reproduced.
The physical process changes the preservation question
Destructive scanning cuts the spine so pages can pass through automated equipment quickly. That can reduce digitization cost, but it also turns a physical book into a one-way input. For common editions, the loss may be limited. Sellers interviewed in the reporting said some books entering these pipelines may have very few surviving copies after wars, fires, and centuries of handling.
That creates a problem no language model can solve after the fact. A model may retain statistical traces of a text, but that is not equivalent to preserving the edition, its material features, or an independently accessible archival copy. If the scan is private, incomplete, or poorly documented, future researchers may not even be able to verify what the model saw.
What the Anthropic ruling does, and does not, establish
A US court ruling involving Anthropic found that scanning legally purchased books for AI training could qualify as fair use in specific circumstances. The reasoning described in the reporting included the idea that destroying the physical copy during scanning can mean one legal copy replaces another, rather than creating an additional copy. The ruling is reported as context-specific, not as blanket permission for every book, dataset, or jurisdiction.
That legal distinction does not settle the ethics. Ownership may answer whether a company can buy and process a volume. It does not automatically answer whether destroying an ultra-rare cultural artifact is responsible, especially when a preservation-friendly scan is technically possible.
What builders should change in their data workflow
The practical lesson is to make provenance and preservation requirements explicit before accepting a training dataset. Ask suppliers whether they use destructive scanning, how they identify rare items, whether source copies remain available, and what metadata and quality controls accompany the text. NDAs should not prevent a buyer from understanding those risks.
Teams can also separate ordinary books from scarce material: use preservation-friendly digitization for rare works, prefer licensed or institutionally digitized collections where available, and retain records of editions, transformations, exclusions, and OCR confidence. These steps support dataset integrity while reducing the chance that an optimization for acquisition speed creates an irreversible externality.
The unresolved trade-off
Elon Musk has said he asked the SpaceXAI team to preserve rare books in a library and scan them using slower methods, presenting a clear alternative to spine cutting. [That proposal is part of the wider preservation debate](https://www.techradar
Sources
- A cultural crime against humanity? AI firms are destroying ultra-rare books so nobody else can read them
- A cultural crime against humanity? AI firms are destroying...
- A cultural crime against humanity? AI firms are destroying...
- The Great Forgetting: An AI arms race is consuming thousands of rare...
- AI Companies Accused of Destroying Rare Books After... | IBTimes UK
- AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale
- AI Firms Destroying Millions of Rare Books for... | Before It's News






















