Inside Project Panama: How Anthropic Scanned and Destroyed Millions of Books for AI Training
In a series of legal filings and unsealed documents made public in early 2026, Anthropic, the AI company behind the Claude chatbot, was revealed to have undertaken an extraordinary — and deeply controversial — data acquisition effort known as Project Panama. According to media reports and court records, this effort involved purchasing, scanning, and then destroying millions of printed books to create a proprietary dataset for training large language models (LLMs) such as Claude.
The revelations have ignited a broad debate over the ethics of AI training datasets, the cultural role of books, the limits of copyright law, and the future of knowledge preservation.
The Mechanics of Project Panama
At its core, Project Panama was a massive industrial digitization project — but with a twist.
Rather than negotiating broad licensing deals for digital text or using public-domain sources, Anthropic reportedly:
- Purchased physical books in bulk, often in tens of thousands at a time.
- Cut the bindings off using hydraulic cutters, enabling bulk scanning on high-speed equipment.
- Scanned every page, creating machine-readable text used in training.
- Recycled the physical paper, meaning that no physical copy of the books remained.
Vendor proposals cited in filings describe plans to process between 500,000 and 2 million books over six months, highlighting the extraordinary scale of the operation — on the order of tens of millions of dollars in book purchases, logistics, and scanning services.
This method stands in contrast to earlier mass digitization efforts such as Google Books, which typically preserved the original volumes even if digital copies were made available.
Anthropic’s Logic — and the AI Arms Race
Insiders reportedly framed Project Panama as a practical and competitive response in the hyper-accelerated AI industry.
According to court documents, Anthropic believed that books represent highly concentrated, carefully edited knowledge — with long-form argumentation and narrative coherence that scraped internet text alone could not provide. They argued that capturing this body of human thought could give AI systems greater reasoning and language abilities.
Executives apparently chose destructive scanning over negotiating licenses at scale because they believed it would be faster and more scalable in a competitive race to dominate AI capability.
That race is real: access to high-quality, well-structured text has become one of the most valuable resources for training cutting-edge models, and access strategies are fiercely competitive among major and emerging AI firms.
Legal and Copyright Controversies
Project Panama surfaced during litigation brought by authors alleging that Anthropic and other AI companies were exploiting copyrighted works without consent.
Some key legal developments include:
- A U.S. federal judge ruled that training AI models on books can qualify as fair use if the use is transformative, but that how data is acquired remains a legally significant question.
- Anthropic settled a separate 2025 lawsuit for $1.5 billion over the earlier use of pirated digital book libraries in training, without admitting wrongdoing.
- Court filings unsealed in 2026 revealed greater detail about Project Panama’s destructive scanning practices and internal discussions about secrecy.
These legal outcomes underline a critical point: while transformative use in training may be defensible under fair use doctrine, the method of acquiring and transforming copyrighted works is still a matter of significant legal contestation.
Authors and Public Backlash
The disclosure of Project Panama triggered backlash from authors, publishers, and cultural commentators.
Creators argued that Anthropic and similar companies were benefiting from the labor and creativity of writers without permission, compensation, or meaningful transparency — effectively turning books into raw material for commercial AI products.
Anthropic’s settlement in 2025 allowed authors to seek compensation, but the broader dissatisfaction speaks to a deeper tension: the imbalance between tech capital’s appetite for training data and the legal and cultural status of stored human knowledge.
Some scholars and commentators have described the project’s destructive nature as a symbolic affront to cultural preservation — destroying physical artifacts of human thought for machine learning. Others point out that similar tactics were used in early digitization projects by libraries and academic institutions, though typically with preservation goals in mind.
Technical Necessity or Cultural Loss?
Advocates within the industry argue that physical scanning — even destructive scanning — is simply a practical necessity when no comprehensive digital corpus with appropriate legal rights exists. High-speed scanning infrastructure has long been used in libraries and archives; the innovation here was its scaling and commercial application.
Critics counter that Project Panama reveals a troubling instrumental view of books and culture: valuing them only for their utility in optimizing AI systems rather than for their historical, aesthetic, or intellectual contributions.
The era of commoditized data has turned cultural artifacts into fodder for algorithmic consumption — raising the question: when AI creators treat books as dispensable inputs, are we witnessing a new form of cultural extraction? If so, what happens to notions of intellectual heritage when physical and often unique forms of human expression are reduced to machine training tokens?
Broader Industry Patterns
Project Panama is not an isolated incident. Court filings in related cases suggest that other AI firms have considered or executed similar strategies, such as:
- Debates within Meta about downloading large shadow libraries of books.
- Past acknowledgements by OpenAI of using large book datasets, although some were subsequently deleted.
These patterns illustrate the systemic nature of data scarcity and acquisition pressures in building state-of-the-art generative models — pressures that pit commercial imperatives against copyright norms and cultural preservation.
The Stakes Ahead
Project Panama has thrust several urgent questions into public view:
- What counts as legitimate access to human knowledge?
- Should AI training environments be subject to more robust cultural and legal oversight?
- Is there a public interest argument for preserving physical copies even as we digitize?
- Who gets to decide how the written record of humanity is absorbed into artificial systems?
As AI advances, these questions will become more consequential — not only for libraries and authors but for societies that still value books as repositories of embodied, situated, and historical intelligence.
Project Panama exposes a fault line where cultural heritage, technological ambition, and legal frameworks collide. How that tension is resolved will shape not just AI futures, but the terrain of knowledge itself.