The Web Is Forgetting. Here Is What That Costs Operators.
Search is degrading, archives are disappearing, and AI systems are accelerating both. Here is what the collapse of the web's collective memory actually means for anyone running a business that depends on information.
The Signal #067 — Dakota’s read on the AI news that actually matters to people running a business.
Most operators assume the information they need is still out there somewhere. You just have to search harder, or ask a better question, or wait for AI to get smarter. That assumption is getting harder to defend.
A detailed piece in The Walrus this week made a case worth sitting with. The argument is not that Google is annoying or that AI is overhyped. It is that the underlying infrastructure storing the web’s knowledge is actively breaking down, and AI systems are speeding up the collapse while pretending to replace what is being lost.
What happened
The Walrus report pulls together several things happening at once, and they are worth naming individually because each one sounds minor in isolation.
Google’s AI summaries are now generating incorrect facts for basic queries. One user in Colorado Springs set up a projector outside and waited for a sunset that, according to Google’s AI answer, had already happened. Small error. But it points to a real pattern: the dominant search engine is now inserting an error-prone AI layer between users and original source pages, making the underlying page harder to find even when it still exists.
Then there is the pollution moving upstream. According to 404 Media, companies are planting content on Reddit specifically to influence the answers generated by AI search tools. The training data and the live query results are both being gamed at the same time.
Meanwhile, the archives that used to serve as the web’s backup memory are under serious strain. FiveThirtyEight, which ran for years as a respected data journalism outlet, had its entire archive deleted by Disney in March 2025 after the remaining staff were laid off. Not archived. Not transferred. Deleted. Wikipedia, which AI systems now scrape directly to generate answers, is seeing dwindling traffic and donations as a result, since users no longer need to click through to the actual page. The resource is becoming the infrastructure of its own slow decline.
The Internet Archive, home to the Wayback Machine and hundreds of billions of web snapshots, is simultaneously dealing with cyberattacks, costly litigation over its digital lending program, and news organizations blocking its crawlers because they fear archived pages could give AI companies indirect access to copyrighted material. The closest thing the web has to a fail-safe backup is buckling.
Link rot was already a slow bleed. Key sections of the United States Constitution briefly disappeared from the Library of Congress website because of a coding error. These are not edge cases.
Why it matters for operators
Here is the practical problem. A lot of business decisions depend on being able to verify something. A contract clause from two years ago. A regulatory guidance document that got updated. A competitor’s pricing page from last quarter. A news story that supports or contradicts a vendor’s claim.
When you send an AI tool to research something and it returns a confident answer, most operators assume that answer is grounded somewhere. The Walrus piece makes clear that grounding is getting shakier. The source pages may no longer exist. The archived versions may be blocked or degraded. The AI summary may be drawing from Reddit threads that were seeded specifically to manipulate AI outputs.
This is not a reason to stop using AI research tools. It is a reason to stop treating their outputs as verified facts. An agency vetting a new media buy, a real estate firm pulling comparable market data, a SaaS company researching a regulatory requirement, any of these workflows break down quietly if the underlying sources are rotting and nobody flags it.
The failure mode is not dramatic. The AI gives you a number. You use the number. The number was wrong three steps back in the chain, and nobody caught it.
What most people get wrong
Most operators frame this as a quality problem with AI models. The model hallucinated. The model was sloppy. Fix the model and the problem goes away.
The Walrus piece reframes it as an infrastructure problem. The model may be doing exactly what it was designed to do, ingesting available sources and generating a plausible summary. The issue is that the sources themselves are degrading. Deleted archives, blocked crawlers, poisoned public forums, and AI summaries that intercept traffic before it ever reaches an original page. That is not a model problem you can patch with a better prompt.
When the corpus (the total body of information an AI draws from) is collapsing, a smarter model just becomes a more confident narrator of bad inputs.
The lesson to carry
AI tools are only as reliable as the information landscape they are drawing from. Right now, that landscape is noisier and thinner than it looks. Operators who build workflows around AI-generated research need a verification step that goes back to primary sources, not just better queries to the same degraded pool.
Treat AI summaries as a starting point, not a citation. Know where your actual source is. And when a page you depend on disappears, notice that it disappeared.
If you want to think through how to build AI workflows with that kind of ground truth built in, xovionlabs.com is a good place to start.