Back to blog

The Digestion of Books: How AI Companies Buy, Scan and Destroy Physical Libraries

The story that defines why knowledge is ceasing to be a document and becoming a probabilistic weight: Anthropic's Project Panama and the market for physical books used to train models.

Ismael Barea
Ismael Barea
The Digestion of Books: How AI Companies Buy, Scan and Destroy Physical Libraries

This is the story that defines why knowledge is ceasing to be a document and becoming a probabilistic weight.

It is not an algorithm downloading web pages. It is a company buying millions of second-hand books, cutting their spines with a hydraulic guillotine, scanning the loose pages at industrial speed and sending the paper to recycling. No physical copy survives. The book becomes training tokens and you, years later, pay to read its interpretation.

This is not a metaphor. It is what happened with Project Panama and what an entire market is institutionalizing.


The Web Is No Longer a Clean Corpus

For years, AI companies trained their models on the text of the open web. That material has run out and, above all, it has become contaminated.

When a model is trained on text generated by another model, its results degrade in a cycle known as model collapse: first the rare or minority answers in the distribution disappear, then the overall quality of the output worsens, and that second phase is not fixed by fine-tuning. The mechanism was formalized by Shumailov and his colleagues at Oxford, Cambridge and Toronto in a paper published in Nature in 2024.

As if that were not enough, authors who oppose being scanned have started using data poisoning tools: invisible perturbations embedded in text that sabotage the model's training.

74%

of the web pages created contained AI-generated text in early 2025, according to industry estimates.

50

poisoned samples in a batch of 100,000 are enough to degrade a model's results, as demonstrated by Nightshade (Univ. of Chicago, 2024).

500

carefully crafted documents can plant a backdoor in a corpus of trillions of tokens, according to Anthropic research cited by ISBNdb.

In contrast, printed books published before the rise of LLMs offer a structural guarantee: they contain no AI-generated text and they cannot have been modified by poisoning tools. They are the "pure" raw material the web can no longer provide.

Project Panama: The Operation That Guillotined the Books

In January 2026, The Washington Post reconstructed, from more than 4,000 pages of declassified court documents, the operation with which Anthropic obtained book data at industrial scale.

The project, internally codenamed Project Panama, started in early 2024. An internal planning document defined it bluntly: "Project Panama is our effort to destructively scan all the books in the world", and explicitly asked to keep it secret. Anthropic spent tens of millions of dollars buying used books in batches of tens of thousands of copies, from sellers such as Better World Books or World of Books.

The process was industrial and irreversible:

1 · Bulk purchase

Thousands of second-hand copies bought at once, without the seller knowing who is acquiring them.

2 · Hydraulic guillotine

Cutting machines remove the book's spine to leave the pages loose.

3 · Industrial scanning

The pages go through production scanners with a declared OCR accuracy above 99.5%.

4 · Recycling

The paper is destroyed. No physical copy of the book remains.

The vendor proposals evaluated by Anthropic spoke of converting between 500,000 and 2 million volumes in six months, a pace of several thousand books per working day. The paper was sent to recycling: the process was deliberately irreversible and left no physical archive behind.

The scale contrast
Bebelplatz · 1933
20,000
Baghdad · 2003
500,000
Project Panama · 2024-26
500,000
Sarajevo · 1992
1,500,000
Upper scale
2,000,000
Books destroyed. The Bebelplatz burning destroyed 20,000; the fire at Baghdad's National Library in 2003, around 500,000; the destruction of Sarajevo's National Library in 1992, around 1.5 million; Project Panama planned to scan 500,000 to 2 million copies in six months.

500K–2M

books they sought to scan in six months, according to the vendors proposed by Anthropic.

7M+

books downloaded by co-founder Ben Mann from pirate libraries (LibGen, PiLiMi, Books3) in 2021.

$1.5B

settlement by Anthropic with the authors for the pirated portion, in August 2025 (~$3,000 per work).

>99.5%

declared OCR accuracy on the production scanners used for the loose pages.

ISBNdb: The Invisible Library for AIs

The practice is not exclusive to Anthropic. In July 2026, 404 Media's investigation uncovered an entire intermediary market around this demand: old physical books as "clean" training data.

The protagonist was ISBNdb, a company with more than two decades managing a bibliographic database of about 112 million catalogued titles. After the rise of AI, besides selling metadata, it started offering a service of "Printed Books Sourcing for Your AI LLMs Dataset Needs": batches of between 1,000 and one million books per order, with a strict NDA on every engagement and the promise that "your identity, strategy, and targets are never disclosed".

Its sales pitch was as explicit as it was uncomfortable:

ISBNdb · sales pitch

"The world's best AI training data is sitting on a shelf. Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."

ISBNdb · the optics problem

"The optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy."

ISBNdb · the conscience laundering

"Responsible physical sourcing is not book burning. It is the completion of a book's lifecycle: from tree to knowledge to tree again."

The impact was immediate. Booksellers around the world saw massive anonymous orders: one of them went from selling fewer than 20 books a week to hundreds since April 2026; Dutch bookshops received orders for 3,001 titles at once, with a "total disregard for price". Buyers paid premiums for rare and out-of-print books that were then pulped.

Nine days after 404 Media's publication, ISBNdb removed the service page. On July 30, 2026, it denied ever having bought, scanned or sold a single book and called it "a test of market interest". There was no regulator, no court, no institutional notice in between: only journalism. And the legal market that supported it remained intact.

Jul 21, 2026

404 Media publishes the investigation into ISBNdb and the market for books to train models.

Jul 30, 2026

ISBNdb removes the page, denies ever having sold anything and calls it a "market test". The legal market stands.

Why buy and destroy paper instead of licensing books or using existing digital libraries? The answer is a combination of law and economics.

Buying a physical copy is governed by the first sale doctrine (17 U.S.C. § 109): the buyer can do whatever they want with their copy, including destroying it. And in June 2025, federal judge William Alsup ruled in Bartz v. Anthropic that scanning a legally purchased book and using the text to train a model is fair use: the use is transformative and the digital copy replaces the destroyed physical copy, without generating an additional copy.

The ruling, however, drew a clear line: fair use covered the purchased books, but not the more than seven million books downloaded from pirate libraries by Ben Mann. That part cost Anthropic $1.5 billion in a settlement with the authors, the largest on record in the history of U.S. copyright, at about $3,000 per work.

Reading the two decisions together produces a precise incentive: buying second-hand paper and destroying it after scanning costs a fraction of what licensing the content would cost and, in addition, shields against litigation. It is not that AI companies hate books: it is that paper, as a medium, is the perfect legal and economic wrapper for turning it into data.

Why You Should Care

It may seem like a distant story of court documents, but it touches directly on what I do as a knowledge professional and what I already explained in Solving More, Learning Less: while more and more of us delegate the learning process to models, the industry desperately seeks original human knowledge to train them.

Printed public knowledge is being bought, digested and returned under a subscription model paid in tokens. If no physical copy remains to audit what the model answers, the only available version of the past lives on a private server, without Moscow glue or 404 error to make us suspicious. That is exactly the question I explore in The Last Book on the Shelf, and also the other side of the coin I analyzed in If AI multiplies your productivity, who keeps the benefit?: collective memory, like productivity, has an owner.


Official sources


Written by Ismael Barea

AI Engineer at Unit4. Building intelligent software and writing about technology, productivity, and the impact of AI on the developer's daily life.

Ismael Barea

Back to blog