Anthropic Project Panama Turned Books Into AI Training Data
greekreporter.com

Anthropic Project Panama Turned Books Into AI Training Data

Tech News
4 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRAnthropic’s reported Project Panama acquired large quantities of physical books, scanned them, and dismantled the copies for AI training. For builders, the story is a warning that data provenance, licensing, and preservation decisions can create product and legal risk long after model training begins.

Anthropic’s Project Panama was an internal effort to acquire large quantities of physical books, scan their pages, and use the resulting digital text in developing models behind Claude. The practical lesson for AI builders is straightforward: training data is an infrastructure and governance decision, not merely a model-quality input. The sourcing path can affect legal exposure, auditability, and trust.

Project Panama made data acquisition an industrial process

Unsealed court filings and reporting describe Anthropic spending tens of millions of dollars on used books after hiring former Google executive Tom Turvey in early 2024. Suppliers reportedly included Better World Books and World of Books. Strand Book Store was considered, but the store said it did not supply books to Anthropic.

The books were not preserved after digitization. Workers used hydraulic cutters to remove spines, fed loose pages through high-speed scanners, and recycled the discarded paper. One contractor proposed processing between 500,000 and two million volumes over six months, according to the reported court materials. The project is often called Project Panama, although that terminology comes from reporting and court records rather than a public Anthropic product announcement.

For a model developer, the appeal is clear. Professionally edited books offer dense, human-written material that may be more useful for language quality than large volumes of noisy web text. But a high-quality corpus does not remove the need to document how every part of it was obtained.

A June 2025 ruling by US District Judge William Alsup found that some use of books for model training could qualify as fair use because training transforms the material into a new type of work. The ruling did not settle allegations involving illegally obtained digital copies. Anthropic later reached a $1.5 billion settlement over those allegations without admitting liability. Eligible authors and publishers were expected to receive about $3,000 per affected title.

That distinction matters operationally. A company can believe that a training technique is legally transformative while still carrying separate risk from acquiring unauthorized copies. Teams building models or retrieval systems should therefore track source permission, chain of custody, deletion obligations, and the exact system in which data was used.

What builders should take from the controversy

Project Panama is less a lesson about book scanning than about governance under scale. Once a corpus contains millions of items, a vague “we bought it legally” claim is not enough. Builders need title-level or dataset-level provenance, written licenses where possible, controls against ingesting pirated archives, and a process for handling takedown or settlement requirements.

Preservation is another neglected constraint. Destroying a physical copy after scanning may improve throughput, but it can remove rare or out-of-print works from circulation. A preservation-first workflow would retain originals or create independently governed archival copies instead of treating the source object as disposable.

The case also sits inside a wider set of copyright disputes. Reporting described Meta discussions about LibGen in connection with Llama 3, while OpenAI acknowledged downloading LibGen material but said it was removed before ChatGPT’s release. The legal outcomes differ across companies and jurisdictions, so these reports should not be treated as proof that every AI training use has the same status.

For AI product teams, the decision rule is useful: choose the highest-quality dataset whose provenance you can explain to a customer, regulator, author, or court. A marginal gain in model fluency is difficult to defend if the data pipeline cannot show what was acquired, under which rights, and where it went.

Sources

Latest Tech News