Perplexity Pro Massively Downgraded: From 500 Responses to 6, Paying Users Flee En Masse

Perplexity Pro slashes usage limits by 99%, triggering a user exodus and exposing AI subscription trust issues.
A long-time Perplexity Pro subscriber's Reddit post reveals drastic service degradation: advanced model responses cut from 500 to just 6, image and video quotas nearly eliminated, and an account that vanished for two weeks with no customer support response. The case highlights a fundamental tension in AI subscriptions—companies face massive compute costs but risk destroying user trust through unannounced service cuts, pushing paying customers toward free alternatives like OpenCode and Claude.
A Long-Time User's Angry Confession
Recently, a Reddit post titled "I CAN'T TAKE IT ANYMORE (pro user)" sparked widespread discussion. The poster, a nearly year-long Perplexity Pro subscriber, detailed a series of service degradation issues that ultimately led to cancellation. The post resonated not because of its extremity as an isolated case, but because it reflects a trust crisis prevalent across AI subscription services today.

Perplexity AI is an AI search engine company founded in 2022 by former OpenAI researcher Aravind Srinivas. Its core product combines large language models with real-time web search, positioning itself as an "answer engine" rather than a traditional search engine. Perplexity Pro is its paid subscription tier at roughly $20/month, offering access to multiple advanced models including GPT-4 and Claude, higher usage limits, and premium features like file uploads. Since 2024, Perplexity's valuation has exceeded $9 billion, but as its user base rapidly expands, computational cost pressures have intensified dramatically.
The user began with a preemptive disclaimer: "Before the mods delete my post, I'm not here to shame this company." This somewhat helpless self-protection itself speaks to the weak position paying users occupy when facing service providers. He understands the company needs to support its many employees, but as a loyal user who paid for an entire year, he simply couldn't tolerate the continued decline in service quality.
Perplexity Pro Quota Shrinkage: From "Nearly Unlimited" to "Restrictions Everywhere"
The most impactful part of the post was the user's concrete description of how usage quotas changed. These numbers vividly illustrate the magnitude of Perplexity Pro's service reduction:
Advanced Model Responses Plummeted
Originally, Pro users could get 500 advanced model responses. Now, the user reports getting "about 6" before hitting a paywall urging an upgrade to the Max plan. From 500 to 6—that's nearly a 99% reduction, impossible to explain away as a "normal adjustment" by any measure.
Some context on multi-tier AI subscription pricing is needed here. The Max plan and the $200/month top-tier plan mentioned in the post reflect a tiered strategy commonly adopted by AI companies. Unlike traditional SaaS, AI services have marginal costs that don't approach zero—every advanced model inference consumes expensive GPU compute. Estimates suggest GPT-4-level model inference costs roughly 20-50x more than GPT-3.5, creating a sharp tension between scaling user bases and controlling costs. Traditional SaaS products (like Spotify or Netflix) have core costs in content acquisition and infrastructure maintenance; once systems are built, the incremental cost of serving one more user is negligible. But the economics of AI inference services are completely different: every model call requires real-time GPU memory and compute resources, and the larger the model and longer the context, the more compute consumed. At current market prices, a single NVIDIA H100 GPU costs about $3-4/hour in cloud rental, and one complex GPT-4-level inference might occupy multiple GPUs for several seconds. This means that when millions of users make simultaneous requests, the compute bill grows linearly or even super-linearly. However, cost pressure shouldn't justify unilateral service cuts to users who have already paid.
Image and Video Features Drastically Reduced
Regarding image uploads, the user recalled: "Previously you could upload unlimited images; now I can't even upload 10 before getting blocked." Video functionality changed even more dramatically—"Previously there were nearly unlimited photos per week and about 15 videos; now it's 0 videos and a quota hit after just 2 photos."
This massive feature contraction, for a user who paid annually, essentially amounts to unilaterally reducing the value of a purchased product mid-contract. The user pays the same money but receives progressively less service each month. Multimodal processing (image recognition, video analysis) does indeed consume far more compute than pure text conversations—processing a single image requires first converting it through a vision encoder (like the ViT architecture) into a token sequence, where a medium-resolution image might equate to hundreds or thousands of text tokens. Video analysis requires processing frame-by-frame or by key frames, where a 30-second video might generate computational load equivalent to tens of thousands of tokens. This means one video analysis can cost 50-100x more than an ordinary text conversation. But this precisely demonstrates that the company may have underestimated actual costs during initial pricing, and is now choosing to make already-paying users bear the consequences of that miscalculation.
The Disappearing Account: A Deeper Trust Collapse
If quota shrinkage is a common industry phenomenon, the "disappearing account" event the user described is far more serious.
He recounted: While using the Pro model one day, the system suddenly told him he didn't have Pro access. Upon checking, he discovered his plan had "completely disappeared—nothing there (not even in expired/removed plans)"—despite still retaining his Discord role badge.
More disappointing was the subsequent customer service experience—he contacted Perplexity but "didn't get any response" until his plan was restored two weeks later. During those two weeks, he had already switched to OpenCode and Claude's free version, stating bluntly that "they're much better." Worth noting is that OpenCode is an open-source AI coding assistant terminal tool that runs via command-line interface, allowing users to call various AI models (including those from OpenAI, Anthropic, Google, and other providers) directly using their own API keys, completely bypassing subscription platform quotas. This "Bring Your Own Key" model lets users pay by actual usage, avoiding the "paid but can't use" dilemma inherent in subscriptions while giving users complete control over their toolchain. Claude is a large language model developed by Anthropic; its free version (claude.ai) provides basic conversation capabilities with usage frequency limits but excels in code generation and long-text comprehension. These two alternatives represent two important trends in the AI tools market: first, decentralized solutions from the open-source community that free technical users from dependence on a single platform; second, competitors attracting dissatisfied users through generous free tiers. When a paid service's experience is worse than free alternatives, the subscription value proposition completely collapses.
As compensation, the company gave him 2,000 credits per computer (worth about $20) as an "apology." But the user's experience with the credit system was equally frustrating: a single prompt consumed 1,700 credits, the remaining 300 went toward a Discord bot that was riddled with issues requiring painstaking manual fixes. He then inexplicably ended up at -57 credits, which subsequently reset without explanation.
This credit system is essentially a user-facing quantification of AI inference costs. Under the hood, large language models charge by tokens—the basic unit of text processing. One English word is typically split into 1-3 tokens, while each Chinese character equals roughly 1-2 tokens. Per-million-token pricing varies enormously across models—from under $1 for open-source models like Llama, to about $15 for Claude 3.5 Sonnet, to roughly $30-60 for GPT-4. A complex prompt consuming 1,700 credits might involve long context input (like uploading an entire codebase), multi-step Chain of Thought reasoning, or invocation of the most expensive frontier models. The credit system was designed to let users flexibly choose different models and features, similar to cloud computing's "pay-as-you-go" philosophy, but opaque consumption mechanisms and system bugs (like negative credits) instead increased cognitive burden and distrust. When users can't predict how many credits an operation will consume, the system transforms from "empowering user control" into "manufacturing anxiety and uncertainty." This chaotic credit mechanism further deepened the user's distrust.
Comet Browser Assistant Becomes a Paid-Only Feature
The final straw for this user was the change to the Comet browser feature.
He stated he was "mainly using Comet because I liked its UI, and that assistant that controls the browser was really good." However, he suddenly discovered that this core feature was now "limited to the desktop version of Perplexity" and required upgrading to the $200/month top-tier plan for the full experience.
Comet is Perplexity's AI-powered browser, with its core selling point being a built-in "browser assistant"—an AI Agent capable of directly controlling the browser to execute tasks. This type of functionality falls within the cutting-edge "Computer Use" capability category in the AI industry, similar to Anthropic's Claude Computer Use and OpenAI's Operator. An AI Agent refers to an AI system capable of autonomously planning and executing multi-step tasks—it doesn't just answer questions but can click buttons, fill out forms, and navigate between web pages like a human. Achieving this capability requires the model to possess: visual understanding (recognizing UI elements on screen), spatial reasoning (determining click positions), task planning (decomposing complex goals into step sequences), and error recovery (adjusting strategy when actions fail). Since AI Agents need to continuously interact with environments and perform multi-step reasoning, their compute consumption far exceeds that of ordinary Q&A—a complete browser control task might involve dozens of model calls and screenshot analyses, with each step requiring the model to "observe" the current screen state and "decide" the next action, accumulating token consumption potentially 100x or more that of a normal conversation. This technically explains why the feature was moved to a premium tier. But for users, having a previously available feature locked without any warning is experientially no different from the "free trial then forced payment" playbook.
"I'm not going to pay $200 a month for a decent assistant," he wrote. Now all Pro features have been converted to Max exclusives, and out of fear of hitting limits, he's forced to consider purchasing Sonar 2, which isn't a top-tier model. Sonar is Perplexity's in-house lightweight search-optimized model, fine-tuned from open-source models and designed specifically for fast retrieval and information synthesis. While it excels in response speed and search integration at costs far below GPT-4 or Claude, it shows clear gaps compared to frontier models in complex reasoning, creative writing, code generation, and other tasks requiring deep thinking. It's like an airline repeatedly reducing economy class legroom while telling passengers "upgrade to business class if you want leg space." Ultimately, he chose to cancel his subscription.
The Trust Crisis in AI Subscriptions: A Widespread Industry Concern
This user's closing words deserve deep reflection from everyone in the AI industry: "I think I'm never going to buy an AI plan ever again because nowadays what we pay for can just be taken from us at any point, even if you've paid for a full year of the service."
This statement highlights a core contradiction of the AI subscription economy: users are purchasing a promise, not a fixed product. Against the backdrop of massive compute cost pressures facing AI companies universally, quietly reducing quotas and reassigning feature tiers to control costs has become standard industry practice. According to industry analyst estimates, the average monthly compute cost per active paying user for mainstream AI companies may be 2-5x their subscription fee, meaning most AI subscription services are actually in a "burn cash to acquire customers" phase—they rely on venture capital subsidies to maintain below-cost pricing in order to rapidly accumulate users and market share. This model mirrors the early days of ride-hailing and food delivery platforms. Once investors demand improved unit economics (where revenue per user must cover the cost of serving that user), or when the funding environment tightens, the first things affected are user quotas and feature scope. The frequent pricing strategy adjustments by OpenAI, Anthropic, Google, and others during 2024-2025 are external manifestations of this economic pressure. But the cost of this approach is the gradual erosion of user trust.
Three Lessons for AI Service Providers
From this case, AI service providers should reflect on at least three points:
First, service stability during contract periods. Annual subscribers deserve relatively stable service commitments; significant mid-term reductions are essentially a breach of contract. While most AI services include "service content may change at any time" disclaimers in their user agreements—such clauses are legally known as "unilateral modification clauses" and have existed in the traditional software industry for years—legal compliance doesn't equate to commercial ethical reasonableness. Notably, the EU's Digital Services Act and Digital Markets Act have begun imposing restrictions on platforms' unilateral service changes, and the US FTC issued new rules in 2024 targeting subscription service "dark patterns." As AI subscriptions become widespread, regulators and consumer protection organizations will predictably apply greater scrutiny to such behavior, especially when service reductions reach the extreme 99% magnitude seen in this case.
Second, transparency and communication around changes. The user had no expectation of sudden feature migration (such as the Comet assistant becoming desktop-only) and received no advance notice. This sense of "betrayal" often hurts more than the reduction itself. Best practices should include: 30-day advance notice for major changes, transition periods, and reasonable compensation or refund options for affected annual subscribers. In this regard, some mature SaaS companies (like Notion and Figma) typically offer existing users a "Grandfather Clause" when adjusting pricing or feature tiers—allowing legacy users to continue enjoying original terms for a defined period—an approach that effectively mitigates the sense of deprivation.
Third, customer service responsiveness. An account disappearing for two weeks with no response is a fundamental service failure sufficient to destroy years of accumulated goodwill. For a company valued at nearly $10 billion, this reflects operational capabilities growing far too slowly relative to user base expansion. During AI's rapid growth phase, many companies pour the vast majority of resources into technical R&D and model training while severely underestimating "non-core" functions like customer support and operational assurance. But historical experience shows that technology companies' long-term success often depends on these "boring" but critical foundational capabilities.
Conclusion
It should be noted that this article is based on a single Reddit user's personal experience; specific figures haven't been officially verified and may reflect subjective perceptions and individual differences. However, the numerous sympathetic comments beneath the post—with users noting "all posts now are about the new limits"—indicate this is far from an isolated case.
In today's increasingly competitive AI landscape, model capability is certainly important, but user trust is the true moat for subscription services. When users begin migrating to alternatives like OpenCode and Claude while declaring they'll "never buy an AI plan again," it's a wake-up call for the entire industry. The AI subscription economy stands at a crossroads: continue quietly cutting services in pursuit of short-term cost control, or build more transparent and sustainable business models that maintain user trust? The answer will determine which companies survive this round of elimination.
Related articles

Meta Open-Sources 30B Model Muse Glimmer: A Practical Breakdown of Running a Local Agent on a Single GPU
Meta's Superintelligence Lab open-sources Muse Glimmer, a 30B multimodal Agent model using 4-bit quantization, hybrid attention, and D-Flash speculative decoding to run on a single consumer GPU like the RTX 4090.

OpenAI Open-Sources Codex Security: An In-Depth Review of the AI Security Scanning Tool's Strengths and Limitations
In-depth analysis of OpenAI's open-source Codex Security code scanning tool, comparing it with Snyk, Semgrep, and CodeQL, examining its AI Agent verification, real test data, and current limitations.

Deep Analysis and Defense Guide for Ruby 4.0 Universal RCE Deserialization Gadget Chain
In-depth analysis of Ruby 4.0's universal RCE deserialization gadget chain, covering construction principles, attack surface impact, and Marshal.load security defenses.