AI Is Devouring the Web: The Internet's Collective Memory Is Vanishing at an Alarming Rate

Generative AI is destroying the economic foundation of web content, threatening the internet's collective memory.
Generative AI is reshaping how information is produced and consumed online, shifting the web from a link economy to an answer economy. As AI chatbots deliver direct answers, original content sites lose traffic and revenue, creators abandon their work, and the diverse knowledge base that underpins the internet erodes. Combined with link rot and the risk of model collapse, humanity's digital collective memory faces an unprecedented crisis.
What We Lose When AI Reshapes the Web
One of the most exhilarating promises when the internet was born was that it could serve as a permanent archive of human knowledge — a collective memory accessible, searchable, and citable by anyone. Yet as generative AI reshapes the production and consumption of online content at an unprecedented pace, that promise is quietly unraveling. A recent article titled As AI eats the web, the internet's collective memory is disappearing sparked intense discussion on Hacker News (146 upvotes, 131 comments), serving as a stark warning about this very phenomenon.
The article's core argument strikes at the heart of the matter: as AI becomes an intermediary layer for information, our connection to original sources is severing, and the original content that carries humanity's collective memory is depreciating, decaying, and disappearing at an alarming rate.

From the "Link Economy" to the "Answer Economy": The Collapse of the Traffic Loop
For the past two decades, the entire web content ecosystem has been built on an implicit contract: content creators provide information, search engines index and distribute it, users click links to visit the original pages, and creators earn advertising revenue or influence in return. It was an imperfect but fundamentally sustainable cycle.
The foundation of this "Link Economy" can be traced back to the PageRank algorithm proposed by Google founders Larry Page and Sergey Brin in 1998. The algorithm's core idea was simple: the more a webpage is linked to by other pages — and the higher the quality of those linking sources — the greater its authority. This mechanism elevated the hyperlink from a mere navigation tool to a standard of value measurement, spawning the entire search engine optimization (SEO) industry and content marketing ecosystem. Websites created high-quality content to attract links and search rankings, then monetized traffic through ad networks like Google AdSense. By some estimates, the global digital advertising market exceeded $680 billion in 2024, with the vast majority built on this "content → links → traffic → ads" loop.
How AI Breaks the Traffic Loop
Now, AI search assistants and chatbots deliver "answers" directly instead of guiding users to click original links. When ChatGPT, Perplexity, or Google's AI Overview summarizes the information you need on the spot, how motivated are you to visit the blogger who painstakingly maintains their website?
Google's AI Overview (formerly Search Generative Experience, or SGE) officially launched in the U.S. market in May 2024. It uses large language models to generate comprehensive answers directly at the top of search result pages. The industry calls this trend "Zero-click Search." According to SparkToro's research, even before AI Overview launched, roughly 65% of Google searches already ended without the user clicking any external link. Perplexity AI takes an even more aggressive approach, positioning itself as an "answer engine" rather than a search engine, synthesizing responses from multiple sources. While it includes citation links, actual click-through rates are far lower than traditional search results.
This triggers a fatal chain reaction:
- Traffic drought: Original content sites see dramatic drops in visits, and ad revenue evaporates with it
- Disappearing creative motivation: When creation goes unrewarded, independent creators and small sites shut down
- Shrinking content sources: The very pool of original content that serves as the foundation for AI training and retrieval is being depleted — by AI itself
This is a self-cannibalizing paradox — AI depends on human-created content to exist, yet it is systematically destroying the economic foundation that sustains that content.
Why the Internet's Collective Memory Is Disappearing
The phrase "collective memory" in the article's title is a concept worth pondering. It refers not just to current information, but to years of accumulated forum posts, personal blogs, technical documentation, and discussion threads from niche communities.
Digital Content Is Far More Fragile Than We Think
Contrary to our intuition, digital content is extraordinarily fragile. The moment a website stops being maintained, its domain expires, or its server shuts down, all the content it hosted vanishes without a trace. By comparison, a printed book — even when out of print — may survive for decades in a library.
In the Hacker News discussion, many users pointed to the worsening problem of "link rot." In 2024, the Pew Research Center published a landmark study finding that as of 2023, approximately 38% of all webpages on the internet were no longer accessible. A classic study by Harvard Law School found that nearly 50% of URLs cited in U.S. Supreme Court decisions had already broken. The situation is equally dire for academic citations — a survey of top journals like Science and Nature found that roughly one-quarter of reference links in papers published more than ten years ago were dead. These figures show that even without AI's impact, preserving digital information already faces enormous challenges. AI accelerating the economic collapse of the content ecosystem only makes the problem worse.
Once-vibrant forums, independent blogs, and specialized communities are shutting down one by one. These places often preserved authentic, in-depth human experience and knowledge that cannot be replicated on major corporate platforms.
The "Pollution" of AI-Generated Content and the Risk of Model Collapse
Even more concerning is that as AI-generated content floods the web, future AI models will inevitably be trained on AI-generated content — what researchers call the risk of "model collapse."
The concept of model collapse was formally introduced in a 2023 paper titled The Curse of Recursion by research teams at the University of Oxford and the University of Cambridge. Their core finding: when AI models are trained on data generated by previous-generation models, after multiple iterations the model's output distribution gradually diverges from the original real-data distribution. Specifically, the model progressively loses the "tails" of the distribution — the uncommon but highly valuable information, minority perspectives, and edge-case but genuine knowledge. Eventually, the model's output degrades into a monotonous "mean," losing both diversity and accuracy. This bears a striking resemblance to inbreeding depression in biology — a shrinking gene pool leads to irreversible decline in population vitality.
Once synthetic content exceeds a certain critical threshold, the quality of the information ecosystem will irreversibly degrade, and authentic human memory will be drowned in an endless sea of machine regurgitation.
Community Debate: The Future of the Open Web
Among the 131 comments on Hacker News, opinions were far from unanimous, reflecting the multifaceted nature of genuine rational discourse.
Pessimists vs. Pragmatists
Some commenters held deeply pessimistic views, arguing that we are witnessing the twilight of the open web, and that the future internet will devolve into walled gardens monopolized by a handful of AI giants.
Other, more pragmatic voices pointed out that transformations in content distribution have happened repeatedly throughout history — from portal sites to search engines, from RSS to social media feeds. Each time, people mourned the death of the old ecosystem, but the web always persisted in new forms. The key question is whether new value distribution mechanisms can be established so that content creators continue to receive fair compensation in the AI era.
An Urgent Call for Digital Archiving
Interestingly, the discussion brought renewed attention to digital preservation organizations like the Internet Archive. Founded by Brewster Kahle in 1996, the Internet Archive is a nonprofit organization dedicated to "universal access to all knowledge," headquartered in a former Christian Science church in San Francisco. Its best-known project, the Wayback Machine, has archived over 866 billion webpage snapshots spanning nearly 30 years. However, the Internet Archive has faced severe legal and financial challenges in recent years: in 2023, it lost a copyright lawsuit brought by publishers and was forced to significantly scale back its digital book lending program; it also suffered a major cyberattack that same year. A more fundamental question remains: given the accelerating pace of content creation and destruction in the AI era, can the Internet Archive's crawling speed and storage capacity keep up with the rate of content disappearing? The organization currently operates on an annual budget of approximately $37 million, funded primarily through donations.
When the commercial content ecosystem cannot guarantee the long-term preservation of information, nonprofit archiving work becomes increasingly vital and precious. Some developers also suggested that individuals should develop the habit of saving local copies of important information rather than relying entirely on the illusion of "eternal cloud storage."
How to Protect Collective Memory in the AI Era
This transformation is irreversible. Rather than simply lamenting, we should focus on strategies for response.
Advice for Content Creators
Content models that rely solely on search traffic carry ever-increasing risk. Building direct audience relationships (email subscriptions, communities, paid memberships) and creating unique value that AI cannot easily replicate (deep original work, authentic experiences, expert judgment) may be a more sustainable path.
Accountability for AI Companies
If AI giants want a long-term supply of high-quality content, they must confront the sustainability of their content sources. Some exploration is already underway, such as content licensing agreements and pay-per-citation mechanisms.
Licensing deals between AI companies and content providers have formed a rapidly growing new market. Between 2023 and 2024, OpenAI signed content licensing agreements with several major media organizations, including the Associated Press (AP), the Financial Times, Axel Springer (parent company of Bild and Politico), and platforms like Reddit, with deal values ranging from millions to hundreds of millions of dollars. Google similarly signed a $60 million annual data licensing agreement with Reddit. However, these deals have sparked widespread controversy: critics point out that only a handful of large media organizations benefit, while the millions of independent blogs, forums, and small websites that truly form the internet's knowledge foundation are completely excluded. Meanwhile, outlets like The New York Times have taken a confrontational approach, suing OpenAI and Microsoft for copyright infringement. The outcome of this legal battle will profoundly shape how value is distributed in the AI-era content ecosystem.
A healthy ecosystem requires AI companies to give back to content creators, rather than extracting value in one direction.
A Call to Every Internet User
Staying alert to information sources, actively supporting high-quality original content, and participating in digital content preservation and archiving — these are contributions ordinary people can make. The internet's collective memory is, at its core, an aggregation of countless individual contributions.
Conclusion
"AI devouring the web" is not alarmism — it is a structural transformation already underway. The convenience that generative AI brings is undeniable, but the cost of that convenience may be the severing of our connection to authentic, diverse, and traceable information sources.
When the day comes that the AI we depend on can only regurgitate content generated by another AI, and the original memories that truly distilled human wisdom and experience have long since silently vanished, what we will have lost is not merely information — it is our capacity to understand the world and trace the truth. Protecting the internet's collective memory may be one of the most important issues of the AI era that we cannot afford to ignore.
Related articles

Apple Watch ECG Detects Atrial Fibrillation, Saves Triathlete's Life: A Real-World Story
Triathlete Connor's heart rate spiked to 219 bpm during a race. His Apple Watch ECG detected AFib, leading to open-heart surgery that fixed a hidden heart condition.

Norcross Maine Forest Fire Maps: A Century-Old Cartographic Legacy and Data Visualization Pioneer
Explore Archie G. Norcross's 1918–1922 Maine forest fire maps—a hand-drawn cartographic masterpiece that pioneered early data visualization and remains valuable for climate research, historical GIS, and AI fire monitoring.

Apogee: A Privacy-First Browser Summarization Extension Rebuilt with Local AI After Mozilla Killed Orbit
After Mozilla killed Orbit, an indie developer rebuilt a fully local AI browser summarization extension called Apogee using Ollama, WebGPU, and Transformers.js—no user data ever leaves your device.