
Anthropic Project Panama Turned Books Into AI Training Data
Published by AINave Editorial • Reviewed by Ramit
Anthropic’s Project Panama was an internal effort to acquire large quantities of physical books, scan their pages, and use the resulting digital text in developing models behind Claude. The practical lesson for AI builders is straightforward: training data is an infrastructure and governance decision, not merely a model-quality input. The sourcing path can affect legal exposure, auditability, and trust.
Project Panama made data acquisition an industrial process
Unsealed court filings and reporting describe Anthropic spending tens of millions of dollars on used books after hiring former Google executive Tom Turvey in early 2024. Suppliers reportedly included Better World Books and World of Books. Strand Book Store was considered, but the store said it did not supply books to Anthropic.
The books were not preserved after digitization. Workers used hydraulic cutters to remove spines, fed loose pages through high-speed scanners, and recycled the discarded paper. One contractor proposed processing between 500,000 and two million volumes over six months, according to the reported court materials. The project is often called Project Panama, although that terminology comes from reporting and court records rather than a public Anthropic product announcement.
For a model developer, the appeal is clear. Professionally edited books offer dense, human-written material that may be more useful for language quality than large volumes of noisy web text. But a high-quality corpus does not remove the need to document how every part of it was obtained.
The legal distinction is the important part
A June 2025 ruling by US District Judge William Alsup found that some use of books for model training could qualify as fair use because training transforms the material into a new type of work. The ruling did not settle allegations involving illegally obtained digital copies. Anthropic later reached a $1.5 billion settlement over those allegations without admitting liability. Eligible authors and publishers were expected to receive about $3,000 per affected title.
That distinction matters operationally. A company can believe that a training technique is legally transformative while still carrying separate risk from acquiring unauthorized copies. Teams building models or retrieval systems should therefore track source permission, chain of custody, deletion obligations, and the exact system in which data was used.
What builders should take from the controversy
Project Panama is less a lesson about book scanning than about governance under scale. Once a corpus contains millions of items, a vague “we bought it legally” claim is not enough. Builders need title-level or dataset-level provenance, written licenses where possible, controls against ingesting pirated archives, and a process for handling takedown or settlement requirements.
Preservation is another neglected constraint. Destroying a physical copy after scanning may improve throughput, but it can remove rare or out-of-print works from circulation. A preservation-first workflow would retain originals or create independently governed archival copies instead of treating the source object as disposable.
The case also sits inside a wider set of copyright disputes. Reporting described Meta discussions about LibGen in connection with Llama 3, while OpenAI acknowledged downloading LibGen material but said it was removed before ChatGPT’s release. The legal outcomes differ across companies and jurisdictions, so these reports should not be treated as proof that every AI training use has the same status.
For AI product teams, the decision rule is useful: choose the highest-quality dataset whose provenance you can explain to a customer, regulator, author, or court. A marginal gain in model fluency is difficult to defend if the data pipeline cannot show what was acquired, under which rights, and where it went.
Sources
- How Anthropic’s Secret Project Destroyed Millions of Books to Train AI
- Anthropic Knew the Public Would Be Disgusted by How It Was Destroying Physical Books, Secret Documents Reveal
- Why Anthropic Bought and Destroyed Millions of Books to Train AI
- Inside Project Panama, Anthropic's Secret Effort To Scan and Shred the World's Books | IBTimes UK
- “We Don’t Want It to Be Known”: Inside Anthropic’s Secret Plan to Destroy & Scan World Literature
- Anthropic Knew the Public Would Be Disgusted by How It Was Destroying Physical...
- AI labs buy up, rip pages from books to train models
- AI companies are turning old books into training data, 'Fahrenheit 451'-style
- Over 100 authors sue Anthropic for pirating books to train AI
- Inside Project Panama: How Anthropic scanned and destroyed...
- Anthropic Bought Millions of Books, Cut Them Apart... | Stackademic
- AI companies are buying antique books, ingesting their contents to train models, and then destroying them at incredible scale
- 100 authors demand $75M from Anthropic over 'stolen' work to train its systems: lawsuit
- 100 authors sue Anthropic for pirating books, demand $75 million
- Anthropic Project Panama Scans and Destroys Books to Train...






















