The Last Book on the Shelf: The Privatization of Collective Memory
From censorship that left a scar to AI that destroys paper to synthesize the truth. What happens when knowledge stops being a public document and becomes a probabilistic weight on a private server.
From censorship that left a scar to AI that destroys paper to synthesize the truth. What happens when knowledge stops being a public document and becomes a probabilistic weight on a private server.
The Memory Police
In The Memory Police, by the Japanese writer Yōko Ogawa, the inhabitants of an island suffer the gradual disappearance of everyday things: first the birds, then the photographs, then the maps. What is truly unsettling about the story is not the loss of the physical object, but the fact that, shortly after disappearing, people forget it ever existed. To make sure no one tries to remember, a secret police watches over and erases any trace of the past.
For centuries, changing the story or censoring history required a brutal physical deployment: bonfires, paper guillotines, scissors and an army of officials. When Soviet Russia wanted to erase the disgraced Lavrenti Beria from the Great Soviet Encyclopedia, the publisher sent subscribers replacement pages about the Bering Strait to paste over the original entry. But analog censorship left a scar: the cut of the blade and the glue revealed that something had been hidden there.
Link Rot: A Web That Fades in Real Time
We tend to think of the web as an infinite, immutable archive, but the technical reality is that the internet is falling apart at a staggering pace. According to the Pew Research Center, 38% of the web pages that existed in 2013 have disappeared today. Of all the content generated over the last decade (2013-2023), one in four pages no longer exists. This is what in computing is known as link rot.
In the traditional web, this phenomenon had a clear and honest face: the 404 Page Not Found error.
The 404 was the "visible nothing." A rough but transparent message: it told you that a document, a news story, an investigation or a debate had once been there and someone decided to delete it or let it die. On government sites and news media, more than 20% of the sources cited today return a 404. Digital memory was evaporating, but at least it left the scar of the broken link.
However, the real drama for the artificial intelligence industry happened when link rot collided with the rise of LLMs. Big Tech found itself facing a web caught between two fires:
- The old web disappears under 404 errors.
- The new web is flooded with AI slop: low-quality synthetic content generated by the AIs themselves to rank in search engines.
If you train a language model on today's web, you are training it on a mix of dead links and its own residue. The 404 in the digital world is precisely what triggered the hunt in the analog world: unable to trust a web that erases and contaminates itself, the big tech companies had to go out and buy, guillotine and digest the last immutable repositories of pure human knowledge: paper books published before 2022.
From the Human Web to the Synthetic Web
What is unsettling is not only that the original web is going dark under a rain of 404 errors; what is truly critical is what is filling that void. As the human record disappears through lack of maintenance or server changes, a new web is emerging at industrial speed with AI at the helm.
We have gone from an internet of creators to an internet of generators. Whatever space an abandoned blog or a closed forum leaves free is colonized by a network of automated portals designed to capture impressions and rank in SEO with synthetic text.
This substitution produces an unprecedented phenomenon: the web has stopped being a reflection of human experience and has become a closed circuit of data processed by machines for other machines.
For the tech companies themselves, this has created a catastrophic paradox. If they try to train the next generation of language models on today's web, they are no longer using the collective knowledge of humanity; they are using the digestive residue of the previous AI. The risk of "model collapse" (Model Collapse) from informational inbreeding is so high that the living web no longer serves them.
That is why they needed paper. That is why they needed to guillotine old books. Because in a world where the modern web is written by an algorithm and the old web is erased with a 404, pre-2022 books have become the last uncontaminated oil field on the planet.
To calibrate what is happening now with books, it is worth remembering Bebelplatz. In 1933, the Hitler Youth burned more than 20,000 books in Berlin at a public pyre. Almost a century later, that image remains the quintessential symbol of barbarism, and the condemnation is universal. Today, by contrast, AI companies buy millions of second-hand books, cut their spines with hydraulic guillotines, scan the pages at industrial speed and send the paper to recycling. All with NDAs, trucks and logistics. And it barely sparks any debate. The difference is not in the destruction, but in the visibility: the Nazi burning was a public ritual; industrial digestion is invisible logistics. When ISBNdb, the company that supplies those books, assures us that "responsible physical sourcing is not book burning," it protests too much; its own sales line, "the optics problem is real", reveals that it knows exactly what this is called in history.
An Event That Would Have Been a Revolution
We are facing an atrocious event. If the Bebelplatz burning remains, almost a century later, the universal symbol of barbarism, the industrial digestion of millions of books (which today barely sparks any debate) should set off every alarm. If this had happened in any other era, it would have been a complete revolution: public shock, universal condemnation, banned books. And yet here we are, watching paper being guillotined at industrial scale under the silence of NDAs.
The Library of Alexandria went down in history as the archetype of lost knowledge, the cultural disaster par excellence. Today we are witnessing destruction on an even larger scale, and it barely earns a single opinion column. It is not that collective memory is worth less than before: it is that we no longer recognize it when it is destroyed in silence.
Because it is not just about destroying media. History is being rewritten, or at least concentrated: knowledge ends up in a single point of control where it can be subtly edited, to the taste of whoever manages it. The censorship that left a scar (the Moscow glue, the 404 error) becomes a surgical and invisible edit of the complete record of the past.
Today we are entering a third, far more subtle stage: the era of the invisible nothing.
The Double Control: From Server to Atom
As Julian Assange warned in 2010, the centralization of information on large servers changed the rules of the game. If a media outlet or a server removes a centralized digital file, the document disappears completely from the intellectual record.
However, in recent years we have witnessed a double movement that closes the circle between the digital and physical worlds:
Control of the cloud (servers): when we consult an AI assistant, we do not get a list of links or a 404 error. We get a fluid, assertive, seamless answer. If the original source was deleted, modified or de-indexed before training, the AI will answer just as confidently, omitting the past without leaving any trace of the omission.
Control of matter (physical books): as 404 Media uncovered and I explored in depth in The Digestion of Books, the big AI companies have been massively buying tons of physical books published before the rise of LLMs. They are not after the text of the open web (contaminated with AI slop or synthetic content), but the "purity" of pre-2022 human knowledge.
To process those books at industrial scale, the fast method involves guillotining the spine, separating the pages, scanning them and destroying the physical medium.
From Public Document to Probabilistic Weight
This process hides a profound transformation in the nature of knowledge:
- The paper book was a decentralized, immutable and democratic medium. It was replicated in thousands of bookshelves and libraries. No one could change the printed text in your home with a remote update.
- The AI model absorbs that text, destroys the paper and dissolves the words inside a matrix of billions of mathematical parameters.
- Knowledge stops being a document that can be delimited or quoted and becomes a synthetic interpretation. This produces a privatization of memory: we buy public knowledge on paper, we turn it into training data and we return it to the user under a subscription model paid in tokens.
The Loss of Contrast and the Illusion of Knowing
In Solving More, Learning Less I addressed how the lack of friction when solving problems with AI hinders the construction of our own mental map. If we add to this the destruction or de-indexation of the original sources, the situation becomes critical for the knowledge professional.
If we do not walk the path of understanding and, moreover, the physical or independent sources to audit the model's answer do not exist, how will we know if the AI is hallucinating, applying a bias or rewriting a technical or historical decision?
We risk depending on a "technological priesthood" that guards the only available version of the past on its servers, without Moscow glue or 404 error to make us suspicious.
The Resistance of the Local Archive
Technology increases productivity, but it also rearranges who controls memory. If digital memory is erased through abandonment or by order, and physical memory is digested to train models, the responsibility for the record falls back on the individual.
The Guerrilla Archivist: Marion Stokes
Before the term fake news existed, a librarian from Philadelphia named Marion Stokes understood everything. On November 4, 1979, during the Iran hostage crisis, she pressed "record" on her Betamax and did not stop until her death in 2012: thirty-three years, eight VCRs running in parallel, more than 71,000 tapes and some 400,000 hours of television. She did not do it out of nostalgia: she knew that the networks discarded their own archives to save on storage costs, and that whoever controls the archive controls the memory. "If you don't record the truth, someone else will write it for you," she used to say. Her collection (now digitized at the Internet Archive, with the documentary Recorder: The Marion Stokes Project from 2019) is the physical proof of this post's thesis: the resistance of the local archive, in its purest form, was exercised by a single person.
The real technical resistance in the AI era is not refusing to use these tools, but cultivating two essential habits:
- Custody the source: keep locally, on paper or on decentralized networks, whatever we consider fundamental. Keep the book on the shelf.
- Ask for the source: demand audit and traceability when an answer arrives too clean.
The question is no longer just whether AI will replace part of our work or whether we will solve more while learning less. The question is whether in ten years we will be able to verify whether what a model tells us is real or only the single version it was allowed to keep.
Sources and References
Literature / Fiction:
Web loss and censorship:
- Pew Research Center (2024): When Online Content Disappears. Study on link rot and the 38% of pages missing since 2013.
- Julian Assange - Speech at the Oslo Freedom Forum (2010): the centralization of digital archives on single servers and the invisible disappearance of the historical record.
- Marion Stokes (1979-2012): archivist from Philadelphia; more than 71,000 TV tapes recorded over 33 years and donated to the Internet Archive. Documentary Recorder: The Marion Stokes Project (2019, Matt Wolf).
Buying and destroying books for AI:
- The Washington Post (January 2026): Inside Project Panama: Anthropic destructively scanned millions of books to build Claude. Report based on 4,000 pages of declassified court documents.
- 404 Media (July 2026): AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop. Report on the massive acquisition and digitization/destruction of pre-2022 physical books by AI companies.
- 404 Media (July 2026): Company Offering Printed Books to Train AI Stops After Coverage. The removal of ISBNdb's page and its subsequent denial.
- Shumailov et al. (Nature, 2024): AI models collapse when trained on recursively generated data. Formalization of model collapse through informational inbreeding.
- The Digestion of Books (seed post with the factual detail on Project Panama and ISBNdb).
Internal links (Quantosh.es):
AI Engineer at Unit4. Building intelligent software and writing about technology, productivity, and the impact of AI on the developer's daily life.