Perplexity Quietly Removes Features, Sparking User Backlash: Model Labels Disappear, Pro Downgraded to Flash

Perplexity silently removes model labels and downgrades Pro to Flash, triggering user trust concerns.
Perplexity has quietly removed several features including model identification labels, single message deletion, and replaced Gemini Pro with the cheaper Flash model. Reddit users are raising concerns about declining transparency and service quality. These changes likely reflect cost pressures from API calls, but the lack of communication risks eroding user trust in an increasingly competitive AI search market.
A Community's Collective Frustration
Recently, a Reddit post sparked widespread discussion about Perplexity's product direction. The long-time user didn't mince words in expressing disappointment: "I genuinely want to like Perplexity, and I did for a while... but it's now clear they don't care about their customers at all."
Behind this emotional expression lies a deeper question worth pondering: as an AI search product rapidly iterates, is it sacrificing core user experience in the name of so-called "optimization"?

Perplexity's Technical Architecture and Business Model
To understand the deeper implications of these changes, it's important to first understand Perplexity's technical architecture. As an AI search product, Perplexity is essentially a "model routing + Retrieval-Augmented Generation (RAG)" system. It doesn't develop all underlying large language models in-house. Instead, it calls multiple model providers' services through API interfaces, including OpenAI's GPT series, Anthropic's Claude series, Google's Gemini series, and more. On top of this, Perplexity layers real-time web retrieval, source citations, and conversation management features. This architecture allows it to flexibly offer users multi-model selection, but it also means every query involves real API call costs—which constitute the bulk of its operating expenses.
Three Key Features Perplexity Removed
According to user feedback, Perplexity has recently removed or replaced at least three features, each touching on the core issues of product transparency and user control.
Ability to Delete Individual Messages Removed
Previously, users could delete individual messages within a conversation. While seemingly minor, this feature had practical significance for managing conversation context, cleaning up erroneous queries, and protecting privacy. After its removal, users lost fine-grained control over managing their conversation history.
Model Source Labels ("Prepared by [model]") Disappeared
This is arguably the most controversial change. Previously, Perplexity would display which model generated each response (the "Prepared by [model]" label). For paying users, this label was crucial—it let them clearly know which underlying large language model their query actually invoked, enabling them to judge response quality and verify they were getting their money's worth.
Removing model labels means a decrease in transparency. Users can no longer confirm whether their selected model was actually called, which in a product offering multi-model selection, amounts to shaking the very foundation of trust.
It's worth noting that modern AI products commonly employ "model routing" mechanisms, where the system dynamically decides which model to send a request to based on query complexity, current server load, cost budgets, and other factors. This practice is perfectly reasonable from an engineering standpoint—simple queries don't need the most powerful model, saving significant costs. But it also raises a transparency issue: if a user selects a specific model, is that model actually being used on the backend? Without model labels, users have no way to verify. This is the deeper reason why removing the "Prepared by [model]" tag sparked controversy—it's not just a UI element disappearing, but potentially a deliberate concealment of changes in backend scheduling strategies.
Gemini Pro Silently Replaced with Flash Model
Users pointed out that Perplexity replaced Gemini 3.1 Pro with Gemini 3.7 Flash, rather than keeping both available. It should be noted that the specific version numbers mentioned in the post are from the user's original wording and may contain memory errors. But the core issue stands: Pro series models typically represent stronger reasoning capabilities, while Flash series prioritize speed and cost efficiency.
Within Google's Gemini model family, Pro and Flash represent two fundamentally different design philosophies. The Pro series features larger parameter counts and longer context windows, performing better on complex reasoning, multi-step logical analysis, and long document comprehension tasks. The Flash series is a lightweight version optimized through knowledge distillation and architectural refinement, maintaining basic capabilities while dramatically reducing inference latency and computational costs. According to public pricing, Flash model API call costs are typically one-tenth or even less than Pro models. For queries requiring deep analysis, the quality difference between Pro and Flash can be significant.
The switch from Pro to Flash typically signals a service provider's cost-control trade-off—Flash models cost significantly less to call than Pro models. For users, this may result in decreased response depth and quality.
Product Trade-offs Under Cost Pressure
Viewed against the broader AI industry backdrop, this series of changes isn't hard to understand. As an AI search product integrating multiple large language models, Perplexity faces enormous API call cost pressure. Behind every query is a paid call to model providers like OpenAI, Anthropic, or Google.
Perplexity Pro's subscription is priced at $20 per month, offering unlimited basic queries and limited advanced model queries. However, a single API call to top-tier models like GPT-4 or Claude can cost between $0.01-$0.10 (depending on input/output token counts), and when combined with retrieval service costs, a heavy user's actual service cost per month may far exceed their subscription fee. This is the common dilemma facing all "unlimited" AI subscription services—the faster user growth, the greater the marginal cost pressure. This fundamentally differs from the economies-of-scale logic of traditional SaaS products. Traditional software has marginal service costs approaching zero, but every AI inference consumes real GPU computational resources.
Replacing Pro models with cheaper Flash models directly reduces operating costs; removing model labels may be intended to reduce user confusion and complaints when the backend flexibly routes models (e.g., dynamically switching based on load). These decisions have their commercial logic, but the problem is—these changes were made quietly without adequate communication.
Transparency Is the Foundation of Trust in AI Products
For AI search products, one of the core expectations users have when paying for subscriptions is predictable, verifiable service quality. When model labels are removed and premium models are silently replaced, users lose not just features—they lose trust in the product.
This also reminds all AI product teams: during rapid iteration and cost optimization, "subtracting" features requires far more caution than "adding" them—especially features involving transparency and user control.
User Loyalty in a Competitive Landscape
In 2024-2025, competition in the AI search and AI assistant space is exceptionally fierce. Beyond Perplexity, Google has launched AI Overviews integrated into search results, Microsoft has deeply embedded Copilot into Bing, OpenAI has rolled out ChatGPT search functionality, and Anthropic's Claude continues strengthening its web-connected capabilities. For users, the cost of switching AI tools is essentially zero—there's no data lock-in, no learning curve barrier. This means user retention depends entirely on product experience and trust, and any change perceived as a "downgrade" could cause users to rapidly defect to competitors.
In this market environment, Perplexity's silent removal of features appears particularly risky. Users' tolerance thresholds are dropping while alternative options are multiplying.
Limitations of a Single Source and a Rational Perspective
It should be objectively noted that this discussion is primarily based on a single user's feedback post on Reddit, which contains notably emotional expression. There is currently no official statement from Perplexity regarding these changes, and specific model version information may contain inaccuracies.
Therefore, this should be viewed more as a user experience signal worth monitoring rather than a comprehensive verdict on the product. However, such feedback often represents broader sentiment in communities—when one user complains publicly, there are usually many more silent users harboring similar feelings.
Conclusion: Product Iteration Shouldn't Betray Users
As a star product in the AI search space, Perplexity's every move draws scrutiny. The controversy sparked by these feature removals is essentially the classic tension between product commercialization pressure and user experience.
For users, when choosing AI products, attention should be paid to transparency commitments and communication practices. For product teams, maintaining user trust while controlling costs will be key to sustained growth. After all, in an era where AI tools are highly homogenized and switching costs keep dropping, user loyalty is more fragile than ever.
Related articles

Ify: An AI Solution That Layers on Top of Your Existing Help Desk
Ify is an AI customer service tool that deploys on top of Zendesk, Freshdesk, and other existing help desks — no migration needed. It auto-builds knowledge bases for fast AI support deployment.

Playcall: Open-Source AI Sales Call Analysis Tool — An Affordable Alternative to Gong
Playcall is an open-source AI sales call analysis tool supporting MEDDPICC, BANT, and more. A self-hostable, affordable Gong alternative for SMB sales teams.

BaudBuddy: A Native macOS Serial Terminal with Built-in File Server for Embedded Debugging
BaudBuddy is a native macOS serial terminal for hardware developers, supporting Serial, BLE, Telnet, and RFC 2217, with built-in TFTP/HTTP/FTP file servers for firmware transfers — no account, no tracking, fully local.