AI labs are buying pre-2022 books as the last slop-free training data
thenextweb.com

AI labs are buying pre-2022 books as the last slop-free training data

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRAI labs are buying pre-2022 print books through broker ISBNdb to avoid AI-generated text contamination. The process destroys books, keeps buyers secret, and raises legal and ethical questions for builders.

AI labs are quietly buying up millions of old printed books to train their models, because these books are the last source of text that no machine has ever touched. The data broker ISBNdb is selling pre-2022 print books in bulk as the only remaining slop-free training data, arguing that books published before the large language model era contain no AI-generated text and offer a verifiable chain of custody. The catch: scanning them at scale means destroying the physical books, and the buyers stay anonymous under strict NDAs.

What happened

ISBNdb, which calls itself the world's largest book database, now sources physical books for AI labs to scan into training data. The company's pitch is blunt: books are "dense, edited, authoritative" and, crucially, printed before 2022, which means they predate the flood of AI-generated content on the open web.

To process large volumes, workers slice off the spines so loose pages can feed through high-speed scanners. The company offers clients strict NDAs and never discloses buyer names. "The optics problem is real," ISBNdb's site reads. "'AI company destroys two million books' is not a headline that generates sympathy."

The legal backdrop matters. A US judge approved Anthropic's $1.5 billion settlement over pirated books and ruled that training on purchased, scanned books counts as fair use, partly because the copying destroyed each print original, so one legal copy replaced another. US distributor Ingram has already warned publishers and offered opt-out mechanisms.

Why AI builders should care

For anyone building or fine-tuning models, data quality is the bottleneck that keeps getting worse. A fast-growing share of online text is now machine-generated, and models risk feeding on their own exhaust, a documented decline called model collapse where each generation ends up a little worse than the last.

Pre-2022 books sidestep that problem entirely. They are fixed, human-authored records that nobody can quietly rewrite. There is a second reason labs want clean paper: data poisoning. Authors have started using tools like Nightshade to lace text with characters that a person reads normally but a model cannot. ISBNdb's own blog cites Anthropic research suggesting that as few as 250 to 500 crafted documents can plant a backdoor in a corpus of trillions of tokens. Pre-2022 books, written before any of these tools existed, sidestep the problem entirely.

Practical implications

If this approach scales, it creates a new data sourcing pipeline that is expensive, destructive, and opaque. Labs pay to buy and destroy physical inventory, then keep the whole operation secret. That makes it hard to audit what data actually went into a model.

Publishers are pushing back. Ingram's opt-out mechanism signals that the industry is moving toward more transparent data-use policies, but the default is still that books get scanned without explicit permission.

For builders, the key takeaway is that clean, verifiable training data has become a competitive advantage worth paying for. If you are training a domain-specific model or fine-tuning on curated text, the same contamination risks apply at smaller scales. A fixed, human-authored corpus from before the AI era is the gold standard, but it comes with legal and ethical complexity.

Caveats

Most of what we know comes from industry reporting and ISBNdb's own marketing materials. Details about which labs are buying, how much they pay, and exact volumes are not disclosed. The legal status of training on scanned, purchased books varies by jurisdiction and is subject to ongoing litigation. The notion of "slop-free" data is framed by proponents and may not fully reflect broader data governance realities. And the spine-removal workflow, while standard for ISBNdb, could be contested on ethical or archival grounds.

FAQs

Slop-free data refers to text that contains no AI-generated content and has a fixed, human-authored provenance. In this context, pre-2022 printed books are pitched as slop-free because they predate large language models and cannot contain synthetic text or backdoor poisoning.

Sources

Latest Tech News