Substack's Bot Epidemic: A Trust Crisis for Subscription-Based Platforms

Substack's AI bot invasion threatens the trust-based business model that subscription platforms depend on.
Substack is facing a growing AI bot problem that threatens its core value proposition. Mass-generated AI content, fake subscription metrics, and spam comments are undermining the trust between creators and readers — the very foundation of the subscription model. This article examines why this threat is especially dangerous for Substack, how it reflects an industry-wide challenge, and what strategies platforms can deploy to fight back.
When Subscription Platforms Meet the AI Flood
Recently, a brief social media post sparked widespread concern across the content creator community: "uh oh, substack's got a bot problem." This seemingly casual tweet actually touches on a core issue that no content platform can afford to ignore — how to maintain authenticity and trust in content ecosystems during the era of generative AI.
Substack, the rapidly growing independent creator subscription platform, has attracted a large base of high-quality writers and loyal paying subscribers with its philosophy of "connecting authors directly to readers, cutting out the middleman." Founded in 2017 by Chris Best, Hamish McKenzie, and Jairaj Sethi, the company's core business model is elegantly simple: creators publish newsletters, readers can subscribe for free or for a fee, and Substack takes a 10% commission on paid subscription revenue. This model bypasses the traditional ad-driven media logic, allowing creators to earn income directly from readers, free from the influence of recommendation algorithms and advertisers on editorial direction. As of 2023, the platform boasts over 35 million active subscription relationships and has attracted numerous prominent journalists and commentators who left traditional media. However, it's precisely this trust-based business model that makes Substack particularly vulnerable to AI bot infiltration.
What Exactly Is Substack's Bot Problem?
In the context of content platforms, a "bot problem" typically encompasses several characteristic scenarios worth examining one by one:
AI-Generated Content Accounts at Scale
With the proliferation of GPT-class large language models, anyone can generate professional-looking long-form articles in minutes. The GPT series of models are built on the Transformer architecture proposed by Google's team in 2017, trained on massive amounts of internet text to learn the statistical patterns of language, enabling them to produce grammatically correct and logically coherent long-form content. The key point is that through API calls, users can batch-generate articles at extremely low cost (just a few cents per thousand tokens), and with automation tools like Zapier and Make, a single person can build a complete content production pipeline in just a few hours.
This has spawned a wave of "AI farm" newsletter accounts — mass-registered, automatically generating content, attempting to infiltrate the ranks of quality creators and siphon off platform recommendation traffic and subscription conversions. This type of content typically lacks original perspectives and merely repackages existing information.
Fake Subscriptions and Engagement Data Manipulation
Bot accounts can artificially inflate subscriber counts, likes, and comment engagement, thereby fabricating a newsletter's apparent "popularity." This exploits the psychological mechanism of "social proof" — a principle systematically articulated by psychologist Robert Cialdini in his classic book Influence: when people face uncertainty, they tend to look at others' behavior to guide their own decisions. When a newsletter displays 100,000 subscribers, new visitors naturally assume "with this many followers, the content must be good," making them more likely to convert into new subscribers.
For creators who rely on social proof to attract genuine readers, this data pollution directly undermines the platform's fair competitive environment. Inflated subscriber numbers don't just deceive ordinary readers — they can also mislead advertisers and brand partners in their assessment of a creator's true influence.
Spam Comments and Scam Funnels
Automated scripts flood popular articles with spam comments, phishing links, and scam messages, disrupting the normal reading experience and potentially compromising platform reputation and user asset security.
Why the Bot Problem Is Especially Lethal for Substack
Substack's value proposition is fundamentally different from traditional social platforms. It doesn't sell traffic — it sells the trust relationship between creators and readers. Readers are willing to pay for a specific author on the premise that the content comes from a real person and offers unique value.
Once the platform is massively diluted by AI bot content, a chain reaction ensues:
- Reader trust collapses: When a paid subscription might deliver AI assembly-line content, users' willingness to pay drops precipitously.
- Quality creators leave: Real authors' work gets buried in a sea of AI content, with their exposure and income taking a hit, ultimately driving them to other platforms.
- Platform brand erosion: Once the perception shifts from "independent creator paradise" to "AI junkyard," recovery becomes extremely difficult.
This is the deeper reason the tweet expressed concern with "uh oh" — it's not just a technical glitch, but a systemic risk that could shake the very foundations of the business model.
An Industry-Wide Problem: AI Content Flooding Goes Far Beyond Substack
It's worth noting that bots and AI content flooding aren't unique to Substack — it's an era-defining challenge for the entire internet content ecosystem.
From Twitter/X's zombie followers to the surge of AI-generated ebooks on Amazon, to the increasingly indistinguishable "expert articles" on Medium and LinkedIn, virtually every platform built on UGC (user-generated content) is racing against the AI flood. The Amazon Kindle Direct Publishing case is particularly telling: since 2023, sellers have been using ChatGPT to "write" entire books in hours and list them for sale, with some even impersonating real authors' names. The science fiction magazine Clarkesworld was forced to temporarily stop accepting new submissions due to the deluge of AI-generated manuscripts. Amazon subsequently limited each account to publishing a maximum of 3 new books per day, though enforcement has been limited in practice.
According to observations from multiple tech media outlets, AI-generated content on some platforms is growing at an exponential rate.
The fundamental contradiction at play is this: AI has driven the marginal cost of content production to near zero, while the cost of content moderation and trust verification remains stubbornly high. The technological barriers are severely asymmetric between attackers and defenders, with the defense naturally at a disadvantage.
Viable Paths for Content Platforms to Combat AI Bots
Facing this challenge, content platforms typically build their defenses across several dimensions:
Strengthened Identity Verification
Raising the barrier for mass bot registration through phone verification, manual review, and creator real-name authentication. Some platforms are beginning to explore "verified human" badge systems. The most radical attempt in this direction is the Worldcoin project (now rebranded as World) — co-founded by OpenAI CEO Sam Altman, it uses iris-scanning devices to create unique digital identity proofs for users, establishing a "Proof of Personhood" mechanism. More moderate approaches adopt progressive verification: phone numbers at the basic tier, government ID at the intermediate tier, and video verification or social graph analysis at the advanced tier. However, all these approaches face difficult trade-offs between privacy protection, accessibility, and cost.
AI Content Detection Technology
Using algorithms to identify characteristic features of AI-generated text, then downranking or flagging suspicious content. Current mainstream detection tools include GPTZero, Originality.ai, and Turnitin's AI detection module, which primarily analyze text "perplexity" and "burstiness" to determine content origin — AI-generated text typically exhibits lower perplexity (i.e., more "predictable") and more uniform sentence variation.
However, it should be noted that current AI detection technology has limited accuracy, and the risk of false positives against real human creators persists. OpenAI shut down its own detection tool in July 2023, admitting it had an accuracy rate of only 26% — a development that itself speaks to the immaturity of this technical approach. AI text that has been manually edited or processed through paraphrasing tools sees detection accuracy drop even further, and non-native English speakers who write in a concise, straightforward style may also be falsely flagged as AI-generated.
Ecosystem Incentive Realignment
A more fundamental approach is adjusting recommendation algorithms and incentive structures to give higher weight to quality original content, while reducing the profit margins for quantity-focused "content farms."
Community Self-Governance and User Reporting
Leveraging the power of real user communities to identify and report suspicious accounts, creating a human-machine collaborative governance model.
Trust Is the Scarcest Resource in the AI Era
The reason this brief tweet warrants deep discussion is that it reveals a reality that's accelerating toward us — in an era when AI can produce content in unlimited quantities, authenticity and trust have paradoxically become the scarcest resources.
For subscription-based platforms like Substack, the ability to effectively curb the bot problem will directly determine whether they can preserve their core value of "genuine human connection." For the content industry as a whole, this is a long-term contest over how to retain human warmth amid the technological flood.
The future winners may not be the platforms with the most content, but the ones that can best prove "the content here is created by real people and worthy of trust."
Key Takeaways
Related articles

Fields Medalist Tim Gowers on the Boundaries and Nature of LLM Mathematical Ability
Fields Medalist Tim Gowers analyzes LLM math capabilities: strong at pattern matching and local reasoning, but fundamentally limited in creative insight and long-range proofs.

Confessions of a Long-Distance Sailor: Systems Thinking and Risk Management in Solo Voyaging
Exploring systems thinking, risk management, and long-term thinking through a long-distance sailor's confessions. The metaphorical parallels between solo sailing and software engineering.

34 Model Iterations in Review: Most Problems Were Actually in the Evaluation Pipeline
After 34 model iterations, an AIOps engineer found most gains came from evaluation bugs. This article details three critical evaluation pitfalls and solutions for MLOps practitioners.