Destructive Scanning of Rare Books for AI Training Raises a Data Provenance Problem
techradar.com

Destructive Scanning of Rare Books for AI Training Raises a Data Provenance Problem

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRAI training suppliers are reportedly buying pre-2022 books and using spine-cutting scans to create machine learning datasets. For builders, the issue is larger than copyright: destroying scarce source material can weaken provenance, reproducibility, and trust in training pipelines.

The reported destructive scanning of rare books for AI training exposes a trade-off that AI builders should treat as a data governance issue, not only a copyright dispute. Suppliers including ISBNdb are reportedly providing physical books to AI developers, while high-speed scanning can involve removing spines, separating pages, and destroying the original volume. The claimed benefit is cleaner human-written material. The cost may be permanent loss of scarce cultural sources. The reported scanning process and ISBNdb's rationale are described here.

Why pre-2022 books are attractive training data

ISBNdb argues that books published before 2022 are less likely to contain text generated by modern chatbots. It also describes older books as dense, edited, and authoritative compared with parts of the web that may contain synthetic or low-quality material. That is a plausible dataset-quality argument, but it remains a supplier's rationale rather than independent evidence that every pre-2022 book improves model reliability.

For model developers, the distinction matters. A publication date is only one provenance field. Teams still need to know which edition was scanned, whether pages were complete, how optical character recognition errors were handled, and whether the resulting text can be audited or reproduced.

The physical process changes the preservation question

Destructive scanning cuts the spine so pages can pass through automated equipment quickly. That can reduce digitization cost, but it also turns a physical book into a one-way input. For common editions, the loss may be limited. Sellers interviewed in the reporting said some books entering these pipelines may have very few surviving copies after wars, fires, and centuries of handling.

That creates a problem no language model can solve after the fact. A model may retain statistical traces of a text, but that is not equivalent to preserving the edition, its material features, or an independently accessible archival copy. If the scan is private, incomplete, or poorly documented, future researchers may not even be able to verify what the model saw.

What the Anthropic ruling does, and does not, establish

A US court ruling involving Anthropic found that scanning legally purchased books for AI training could qualify as fair use in specific circumstances. The reasoning described in the reporting included the idea that destroying the physical copy during scanning can mean one legal copy replaces another, rather than creating an additional copy. The ruling is reported as context-specific, not as blanket permission for every book, dataset, or jurisdiction.

That legal distinction does not settle the ethics. Ownership may answer whether a company can buy and process a volume. It does not automatically answer whether destroying an ultra-rare cultural artifact is responsible, especially when a preservation-friendly scan is technically possible.

What builders should change in their data workflow

The practical lesson is to make provenance and preservation requirements explicit before accepting a training dataset. Ask suppliers whether they use destructive scanning, how they identify rare items, whether source copies remain available, and what metadata and quality controls accompany the text. NDAs should not prevent a buyer from understanding those risks.

Teams can also separate ordinary books from scarce material: use preservation-friendly digitization for rare works, prefer licensed or institutionally digitized collections where available, and retain records of editions, transformations, exclusions, and OCR confidence. These steps support dataset integrity while reducing the chance that an optimization for acquisition speed creates an irreversible externality.

The unresolved trade-off

Elon Musk has said he asked the SpaceXAI team to preserve rare books in a library and scan them using slower methods, presenting a clear alternative to spine cutting. [That proposal is part of the wider preservation debate](https://www.techradar

Sources

Latest Tech News