The Books Are Gone. The Training Data Stays.
AI companies are reportedly buying, scanning, and destroying physical books to lock up training data before competitors can access it. Here is what that actually means for operators who depend on AI staying honest about what it knows.
The Signal #077 — Dakota’s read on the AI news that actually matters to people running a business.
There is a version of the AI story where the technology democratizes knowledge. More access, lower cost, better answers for more people. That version is getting harder to tell with a straight face.
A guest post published this week on Anna’s Archive, the largest open digital library currently operating, describes something that deserves a slower read than the usual AI news cycle.
What happened
According to the post, several AI companies have been quietly acquiring large quantities of secondhand books through intermediaries, scanning them, and then destroying the physical copies. The goal is to capture training data, specifically text printed before 2022 and therefore “untouched by machines,” before anyone else can.
The most concrete detail in the piece involves Anthropic. The post describes what it calls “Project Panama,” a highly confidential internal program reportedly exposed through a $1.5 billion copyright settlement. According to the piece, Anthropic spent tens of millions of dollars purchasing millions of paper books, scanning them to train its Claude models, and then destroying them all.
The destruction is the part worth sitting with. It is not just about scanning. It is about removing the physical object so competitors cannot access the same source.
Anna’s Archive frames this plainly: after AI companies scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge gets permanently locked on private servers.
Since the beginning of 2025, according to the post, AI-generated content has accounted for more than half of newly published internet content. The archive is now calling for volunteers worldwide to scan and upload books before more are lost, offering recognition, lifetime membership, and in some cases help covering scanning costs for large-scale contributions.
Why it matters for operators
If you are running a business that depends on AI tools for research, drafting, analysis, or customer-facing answers, you have a stake in what those models actually trained on.
Here is the operational problem. The quality and breadth of a model’s training data (the raw text it learned from before you ever typed a prompt) shapes what it knows, what it misses, and where it will confidently hallucinate (make up a plausible-sounding answer to fill a gap). Operators rarely see that layer. You interact with the finished product.
When a law firm uses an AI tool to research case precedent, or a pharmaceutical company uses one to scan historical clinical literature, or a publisher uses one to fact-check references, the assumption underneath all of that is that the model has broad, balanced exposure to what was actually written. If large chunks of pre-2022 print knowledge are now sitting on private servers with no path back to the public domain, that assumption gets shakier.
This is not a theoretical concern about fairness. It is a practical concern about reliability. A model trained on a curated private corpus (a closed collection of text one company controls) will have blind spots that a model trained on a broader, more open corpus would not. You will not be told which gaps exist. You will just get wrong answers presented with the same confidence as correct ones.
What most people get wrong
Most operators hear “AI training data” and mentally file it under “not my problem, that is for the engineers.”
That framing is costing people real money already.
The question of what a model knows is inseparable from the question of whether you can trust its outputs. A SaaS company using AI to auto-generate product documentation, a real estate brokerage using it to summarize market reports, a healthcare administrator using it to surface policy references, all of them are implicitly betting that the model’s underlying knowledge is reasonably complete and reasonably unbiased toward one organization’s interests.
If the most authoritative physical sources for certain topics were scanned by one company and then destroyed, there is no independent way to verify what the model learned from them, whether it learned accurately, or whether competitors’ models are working from an incomplete record. The audit trail is ash.
The Anna’s Archive piece puts it bluntly: while AI companies promise to make human knowledge accessible, some are dismantling the most durable carriers of that knowledge in the process.
The takeaway
Operators do not control what goes into a model’s training run. That is a real constraint. But you do control how much trust you extend to AI outputs in high-stakes decisions, and you do control whether you maintain independent access to primary sources rather than letting AI summaries become the only version of the record you consult.
Knowing how the sausage gets made does not make you a pessimist about AI. It makes you a more careful buyer.
If you want to think through what that looks like practically for your operation, xovionlabs.com is a good place to start.