AI Companies Are Buying and Destroying Antique Books for Training Data: What Builders Need to Know
futurism.com

AI Companies Are Buying and Destroying Antique Books for Training Data: What Builders Need to Know

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRAI companies are bulk-purchasing antique books, ingesting their contents for training data, and destroying the physical copies, raising legal and ethical concerns for builders.

AI companies are quietly buying antique books in bulk, scanning their contents to train large language models, and then destroying the physical copies. This practice, which relies on the first-sale doctrine and fair use arguments, has drawn scrutiny from authors, rare booksellers, and the broader AI community. For builders, it raises urgent questions about data provenance, legal risk, and reputational exposure when sourcing training data from physical media.

What happened

According to a settled lawsuit, Anthropic used a hydraulic powered cutting machine to remove pages from physical books it procured from resellers, then scanned them with industrial-grade imaging equipment. A judge found this process to be "transformative" and protected by fair use under the first-sale doctrine, which allows a buyer to do what they want with a purchase without the original copyright holder's permission. However, Anthropic also paid a $1.5 billion settlement to authors for using pirated digital books, highlighting the legal gray area.

The practice has become prevalent enough that intermediaries like ISBNdb now market bulk purchases of 1,000 to one million books per order for AI training. ISBNdb promotes pre-LLM era print books as "structurally clean of modern poisoning tools" and promises to keep buyers anonymous. Small booksellers have reported sudden surges in orders, with one seller noting his inventory of rare and out-of-print books could be destroyed after scanning. Rare booksellers in the Netherlands have also reported bulk orders they suspect are from AI labs.

Why AI builders should care

This controversy directly affects anyone building or deploying large language models. The data sourcing pipeline for training data is becoming a legal and ethical minefield. Using physical books obtained through bulk purchases may seem like a clean workaround to avoid web contamination, but it introduces new risks around consent, copyright, and fair use. The first-sale doctrine and fair use are not settled across all jurisdictions, and the destruction of rare copies adds a reputational dimension that could damage trust in AI products.

For AI teams, the case underscores the importance of data provenance. If your model is trained on data sourced from physical media, you need to understand the chain of custody and whether the original creators or rights holders have consented. Publicized cases like Anthropic's settlement and Meta's accusations create a spotlight that could lead to stricter regulations or litigation.

Practical implications

AI builders should take several concrete steps:

  • Audit data sourcing policies. If your team uses third-party data aggregators, verify their sourcing methods. Intermediaries that promise anonymity may be hiding practices that could expose you to legal or reputational harm.
  • Communicate provenance to stakeholders. Users and investors increasingly care about ethical training data. Being transparent about how data was obtained can build trust and preempt criticism.
  • Scrutinize vendor claims. Vendors like ISBNdb that market "uncorrupted" pre-LLM books may be downplaying the ethical and legal complexities. Request documentation of consent or fair use analysis.
  • Consider alternatives. Licensed datasets, opt-in content partnerships, or synthetic data generation can avoid the risks associated with bulk physical book ingestion.

Caveats

The available evidence is limited and focuses on a developing controversy. Specifics about practices and legality vary by jurisdiction and case details. The extent and scale of bulk-book purchases for AI training remain uncertain, and some booksellers have questioned whether the practice is as widespread as reported. Assertions about specific companies and exact quantities should be treated as reported and contested evidence rather than confirmed universal practice. Legal outcomes may differ in other countries with different copyright and fair use frameworks.

FAQs

The controversy centers on AI companies bulk-purchasing antique and rare books, scanning their contents to train language models, and then destroying the physical copies. This raises questions about consent, copyright, and the ethics of destroying rare cultural artifacts. Reports describe practices like cutting pages with hydraulic machines and intermediaries marketing pre-LLM books as uncontaminated training data.

Sources

Latest Tech News