Has Claude Been Declining Since Version 4.6? A Two-Year User's In-Depth Critique

A veteran Claude user argues the model peaked at Opus 4.6 and questions whether "more powerful" truly means better.
A Reddit user with over two years of Claude experience argues the model's best days ended with an early version, with each new iteration bringing wordier output, rising costs, and a growing need for custom skills to compensate for poor default behavior. Comparing it to a competing AI that works out of the box with zero configuration, the post sparks a broader debate: do longer reasoning chains and higher benchmark scores actually deliver more value to users? The article presents the user's perspective while noting its inherently subjective and anecdotal nature.
A Veteran User's Disappointment: Has Claude Already Peaked?
A longtime Reddit user recently posted that after more than two years of using Claude, they believe the model peaked at Opus 4.6 — and that every iteration since has only made the experience worse. The post sparked widespread discussion about what AI model "progress" actually means: does a more powerful model always equal a more useful one?
The user's core argument is straightforward: Claude during the Opus 4.6 era wasn't "the sharpest tool in the shed," but it genuinely understood instructions, delivered what was asked, kept costs manageable, and helped them build a large number of projects. Since then, every new version has claimed to bring "a new level of intelligence" — yet in their experience, the actual usability has gone down, not up.

Note: Version names such as "Opus 4.6," "GPT 6," and "fable 5.1" that appear in the original post may reflect the user's misremembering, informal shorthand, or jokes. This article faithfully represents their views without endorsing any specific version numbers as factually accurate.
Three Areas of Regression, According to the User
Declining Output Quality — Language That "Doesn't Sound Human"
The poster complained that the new version's English output has become increasingly "garbage." They also tested other languages and found readability similarly poor. In their view, the model seems to constantly hunt for "caveats and blind spots nobody asked about," trying to appear clever (being a smartass), while failing at the most basic thing: delivering what was actually requested.
This "overthinking" phenomenon is a sentiment many users share about the latest generation of reasoning models. Models now tend to produce longer, more "comprehensive" responses packed with disclaimers and boundary discussions — but for users who just want to get a specific task done quickly, all that extra content is noise, not value.
Rising Costs — "More Powerful" Means More Expensive
The user points out that newer versions justify higher prices by claiming to be "more capable" — but argues that this "capability" might essentially just be a revamped system prompt that keeps the model running longer and generating more unnecessary content. In other words, users are paying more for verbosity.
This highlights a real tension in the industry: when models use "longer chains of thought" as a proxy for capability improvement, token consumption rises accordingly, and real user bills grow with it. The tradeoff between capability and cost is becoming an increasingly important dimension for professional users.
Chain-of-Thought and Its Impact on Token Consumption
Reasoning models (such as Claude 3's extended thinking mode) generate an internal "thinking process" before producing a final answer — and this thinking process is also billed by token. For simple tasks, this means users are paying for the model to essentially "talk to itself," with intermediate steps that often contribute little to the final output. Using Anthropic's pricing structure as an example, output tokens are typically priced at 3–5× the rate of input tokens, so the longer the chain of thought, the more dramatically the bill grows. This design was intended to improve accuracy on complex reasoning tasks — but for everyday structured tasks like code generation or text editing, the gains from extended thinking are often hard to justify against the cost increase. Professional users therefore find themselves making tradeoffs on model versions or configuration parameters.
Skill Overload: Patching Problems That Shouldn't Exist
One of the most vivid complaints in the post: over the past year, AI content creators' videos almost always open with "Hey, I've got a new skill for you..." — skills designed to fix problems that new model versions introduced without anyone asking for them. Users now need extra skills just to teach the model "how to write," "how to design a frontend," or "how to actually be useful."
Worse, these skills themselves consume tokens — and once the context gets compacted, all those instructions vanish, requiring a full reload that burns even more tokens. The user had assumed this was just the unavoidable cost of making AI better — that to get better code, you had to use crude, reductive "caveman skill" prompting.
Context Compaction: How It Works and Why It Hurts
When a conversation approaches the model's context window limit, some AI tools (including Anthropic's Claude.ai and various third-party clients) automatically trigger a "context compaction" mechanism: summarizing earlier conversation content into a shorter digest to free up space for new input. This process can dilute or completely lose previously injected system prompts and custom skill instructions, causing model behavior to drift — carefully tuned output styles or task rules must be reloaded from scratch. For workflows that rely heavily on custom skills, this means spending tokens every time a compaction occurs to "re-awaken" the model's memory, creating a continuously accumulating cost loop. This is the technical root of the "skill consumption → compaction → reload" vicious cycle the poster describes.
The GPT 6 Comparison: The Shock of Out-of-the-Box Usability
What truly shook this user's confidence was the performance of what they call "GPT 6" on their own projects. Their description: zero skill configuration, installed the native Codex from scratch, and the model just worked as expected — no repeated iterations, no need to pair a skill or plugin with every task, and in many cases even lower cost — because "you don't have to hand-hold it through fixing every leftover piece of junk," and it "actually understands how human English is supposed to be written."
This comparison reveals an interesting philosophical divergence in product design: on one side, a model that requires extensive external "skills" and fine-tuning to perform well; on the other, a model whose default behavior already aligns with user expectations right out of the box. For efficiency-focused developers, the appeal of the latter is obvious.
It's worth noting that this is a single user's subjective experience based on their personal projects — a limited sample, and conclusions may vary significantly across different tasks and workflows. But the underlying demand is representative: users want less verbosity, fewer iterations, less skill scaffolding, lower costs, and output that's actually usable.
"Out-of-the-Box" vs. "Tunable": A Fundamental Product Philosophy Divide
Large language models face a fundamental product tradeoff at deployment time: calibrate default behavior to match the intuitive expectations of most users (high "out-of-the-box" usability), or preserve more flexibility for expert users to deeply customize through prompt engineering (high "tunability")? The former typically involves incorporating large amounts of everyday-user preference data during RLHF (Reinforcement Learning from Human Feedback), making the model more concise and direct by default. The latter tends to reserve more "capability headroom" in the default state, with specific behaviors activated through user instructions. Neither strategy is inherently superior — but as a model's primary user base shifts from researchers to efficiency-oriented developers, the weight given to "out-of-the-box" usability tends to rise significantly. The comparison in the post reflects exactly this divide: Claude is perceived by some users as leaning toward the latter, while the competing product is seen as investing more in the former.
Why Still Use Claude at All?
The user candidly admits the only reason they're still on Claude is that OpenAI is still limiting access to the new x20 subscription tier (the premium tier). They don't have much confidence that Claude can claw back ground on what they consider the key dimensions: good writing, less fluff and more action, fewer iterations, less skill dependency, and lower costs. In their view, fully reversing course on all these points "would be very hard for Claude."
A Broader Debate About What "Intelligence" Means
Stripped of the personal venting, this post raises a question worth serious reflection across the entire industry: What actually constitutes model "progress"?
If "smarter" means longer reasoning, more boundary discussions, higher token consumption, and a need for users to "tame" the model with a stack of skills and prompt engineering — how much real-world value does that progress deliver to the average user? By contrast, a model whose default behavior is restrained, precise, and aligned with user intent might be what most people actually want.
That said, we should approach these views with appropriate caution. A single user's experience carries strong subjective bias and is limited to their specific use case. Genuine model capability assessment requires more systematic benchmarking and far larger sample sizes. But feedback from deeply engaged frontline users is precisely what product teams should be listening to most carefully — it serves as a reminder to developers that stacking capabilities should never come at the cost of usability and output quality.
Related articles

AI Agent Developer Job Hunt Guide: Four Hard Standards to Clear Before You Apply
A practical guide for landing AI Agent developer roles: four measurable standards — project runs, problems debuggable, solution explainable, interviews survivable.

Multi-Agent Development Guide: From Monolithic AI to Team Collaboration in Practice
A beginner's guide to multi-agent development covering core advantages, common learning pain points, enterprise tech stacks, and engineering methodology for AI developers.

Agent Skill Routing: Retrieval vs. LLM vs. Two-Stage Architecture Compared
Retrieval or LLM for Agent skill routing? Compare coarse-filter vs. fine-select architectures on latency, accuracy, and cost — with 4 key production considerations.