Don't Be So Quick to Mock Gemini: Cognitive Bias and the Hidden Truth About AI Model Evolution

Why mocking Gemini reveals more about our rising expectations than actual AI regression.
A YouTube creator's sharp observation exposes a cognitive double standard in the AI community: the same developers mocking Gemini's new version once hailed ChatGPT 3.5 Turbo as a world-changing breakthrough. This article explores how expectation inflation distorts our judgment of AI models, why longitudinal comparison matters more than cherry-picked failures, and what developers should really focus on as AI tools evolve at an exponential pace.
When We Mock an AI Model, What Are We Really Mocking?
Recently, the community has been flooded with jokes and criticism about the new version of Gemini. Developers routinely take to social media to showcase the model's mistakes, evaluating it with a slightly smug tone as "still far behind." This phenomenon plays out cyclically with every new model release.
Background: The Gemini Model Series
Gemini is a family of large language models released by Google in December 2023, designed to compete with OpenAI's GPT-4 and Anthropic's Claude. The series includes Ultra, Pro, and Nano variants, built on a multimodal architecture capable of processing text, images, audio, and video simultaneously. In March 2025, Google released Gemini 2.5 Pro, further enhancing reasoning capabilities and context window length. Gemini's launch marked Google's official entry into the core competitive arena of generative AI, but each version update triggers intensive community testing and evaluation. These tests often employ "Red Teaming" methods — where users actively search for edge cases and failure modes. While this approach is effective at quickly exposing problems, it also tends to produce "selective display bias" — users are far more inclined to share failure cases than successes.
But a YouTube creator raised a thought-provoking counter-question that the entire tech community should pause to consider: The models we mock today as "not smart enough" have actually long surpassed the products we once hailed as milestones.

ChatGPT 3.5 Turbo: A Forgotten Point of Reference
The creator's core argument is straightforward: Remember when ChatGPT 3.5 Turbo first came out?
Back then, the internet was filled with takes like — "looks like software engineering is over, boys." Countless developers were stunned by 3.5 Turbo's performance, some even starting to worry about their career prospects. It was seen as a watershed moment for the industry.
Technical Milestone Retrospective: The Breakthrough of 3.5 Turbo
ChatGPT 3.5 Turbo was released in early 2023 as OpenAI's conversation-optimized version built on GPT-3.5. It employed RLHF (Reinforcement Learning from Human Feedback) — a technique where human annotators rank model outputs to train a reward model, which is then used with reinforcement learning to make the model generate content that better aligns with human preferences. Compared to its predecessor, 3.5 Turbo improved response speed by approximately 25% and reduced costs by 90%, making it the first conversational AI to achieve truly large-scale commercial deployment. At the time, its performance on code generation, text creation, and Q&A tasks stunned the developer community, and it was considered an "iPhone moment" level of technological breakthrough.
Yet looking back through today's lens, the creator bluntly states: "ChatGPT 3.5 Turbo compared to today's Gemini is absolute garbage."

Behind this somewhat exaggerated statement lies a genuine technical reality: model capabilities are iterating so fast that we're constantly resetting the "passing grade." From GPT-3 in 2020 to Gemini 2.5 Pro in 2025, top-tier models have shown significant leaps in standard benchmark scores every 6–9 months. For example, GPT-3.5 achieved a pass rate of about 48% on the code generation benchmark HumanEval, while GPT-4 reached 67%, and Claude 3.5 Sonnet hit 92% — nearly doubling in just two years. Our expectations for AI rise with each release, until we forget which version we once cheered for.
The Cognitive Double Standard in AI Model Evaluation
The creator goes further, puncturing an uncomfortable truth. He points out that among those mocking Gemini's new version today, half were the very same people who marveled at ChatGPT 3.5 Turbo, declaring "this is the future."
"Half of you, when you saw 3.5 Turbo, said: 'that's me.' You know what I mean? You really thought that was the ceiling."

This reveals a classic cognitive bias: We always use the current state-of-the-art as our sole frame of reference to belittle every new product's shortcomings, while selectively forgetting the absolute magnitude of technological progress.
In other words, today's Gemini — the one being criticized as "not good enough" — has already left far behind the very model that once made everyone exclaim "software engineering is over." We're not criticizing the model for regressing; our expectations have simply grown too fast.
Understanding the Limitations of Evaluation
This cognitive bias partly stems from limitations in AI evaluation methodology. While academia uses standardized benchmarks like MMLU (Massive Multitask Language Understanding), HumanEval (code generation), and GSM8K (mathematical reasoning), these tests often fail to fully reflect performance in real-world application scenarios. More importantly, evaluations lacking historical comparison baselines easily fall prey to "anchoring effects" — people use only the current strongest model as their reference point while ignoring the longitudinal magnitude of progress. When 3.5 Turbo's 4K context window was still considered "good enough" in early 2023, today's Gemini 2.5 Pro offers over 1 million tokens of context — a leap of several orders of magnitude in itself.

Three Technical Insights Behind AI Evolution
This seemingly casual commentary actually touches on a phenomenon in the AI industry that deserves our attention.
Expectation Inflation Is Distorting Our Judgment
Each generation of models redefines what "acceptable" means. When GPT-3.5 amazed people, the baseline was still low. As GPT-4, Claude, and Gemini successively entered the scene, that baseline kept rising. Today's "mediocre" would have been "miraculous" two or three years ago.
This "capability inflation" phenomenon is especially pronounced in AI. Research shows that the pace of improvement in large language models is even faster than Moore's Law in the chip industry. What was considered "superintelligent" performance last year may be viewed as merely "baseline" this year. This relentless inflation of expectations makes it increasingly difficult to objectively evaluate the true value of any individual model, and it creates a psychological state of "perpetual dissatisfaction" among developers and users alike.
Evaluating AI Models Requires Historical Perspective
Judging a model solely by "what mistakes it made" is one-sided. A more valuable assessment places it within the timeline of technological evolution: Compared to similar products from six months ago or a year ago, how much has it improved? What previously unsolvable problems does it now address? Fixating only on flaws while ignoring the progress curve leads to distorted conclusions.
This requires developing a "longitudinal comparison" mindset for evaluation. For instance, GPT-3.5's context window was only 4K tokens, and it frequently made logical errors on complex reasoning tasks. Today's Gemini 2.5 Pro not only features a million-token context window but has also made qualitative leaps in multi-step reasoning, code debugging, multimodal understanding, and more. If we focus only on its mistakes in certain edge cases while ignoring these absolute capability improvements, we fall into the trap of "seeing the trees but missing the forest."
Developers Need to Adjust Their Mindset and Embrace Tool Evolution
For developers, this commentary serves as a reality check. Rather than competing on social media to see who can dig up the most outrageous model failures, it's more productive to think about how to fully leverage these tools within their current capability boundaries. After all, the model that once made you worry "software engineering is over" is now something you consider "not even worth looking at" — which itself speaks volumes about the speed of tool evolution and humanity's ability to adapt to and master new tools.
More importantly, understanding an AI model's "capability boundaries" is far more valuable than simply criticizing its failures. Every generation of models has areas where it excels and areas where it falls short. Understanding these boundaries and designing reasonable workflows is the real key to converting AI tools into productivity. When we obsess over failure cases, we often overlook the real-world application scenarios where AI has already dramatically improved efficiency.
Conclusion: Maintain Both Reverence and Clear-Headedness Toward AI Progress
This creator used a humorous, slightly sardonic style to remind us to reexamine our attitudes toward AI. It's easy to mock a new model's shortcomings, but if we can remember which "now seemingly embarrassing" versions once made us ecstatic, perhaps we'd approach technological progress with a bit more reverence and a bit less arrogance.
AI development follows a steep curve. What we dismiss as "garbage" today may well have been the "future" we dreamed of just a few years ago. And the Gemini that disappoints us today may very well become, in the not-too-distant future, yet another "those were the good old days" footnote when we look back.
Technological progress is never linear — it's an exponential leap built through each iteration. When we criticize a new model, it's worth asking ourselves: Has it actually regressed? Or have our expectations simply inflated another round? Maintaining this kind of historical perspective is the only way to accurately assess AI's true value and rationally plan its practical application boundaries in real work.
Key Takeaways
- Cognitive double standard: Developers use the current strongest model as their sole frame of reference, selectively forgetting the absolute magnitude of progress
- Historical comparison baseline: Today's criticized Gemini already far surpasses the ChatGPT 3.5 Turbo that once amazed everyone
- Capability inflation: AI model capabilities leap significantly every 6–9 months, with rising expectations creating a state of "perpetual dissatisfaction"
- Evaluation methodology flaws: Red teaming and social media amplify failure cases; lack of longitudinal comparison leads to cognitive distortion
- Reverence for progress: From 4K context to millions of tokens, from 48% to 92% code pass rates — orders-of-magnitude breakthroughs achieved in just two years
- Pragmatic attitude: Understanding model capability boundaries and designing reasonable workflows is far more valuable than simply criticizing failures
Related articles

Tesla Cybercab Mass Production: How a Car Without a Steering Wheel Is Rewriting the Business Logic of Transportation
Tesla's Cybercab is a two-seat robotaxi with no steering wheel or pedals, marking Tesla's shift from automaker to mobility platform. A deep dive into the business logic, tech challenges, and risks.

Why Startup ARR Is Becoming Fragile in the AI Era — And How to Fight Back
AI is making startup ARR fragile as procurement cycles shorten and tech moats erode. Learn why ARR stability is declining and strategies to build durable revenue.

The Root Cause of AI Deceptive Behavior: Misalignment Risks in Reinforcement Learning and Solutions
Deep dive into the technical roots of AI deception: how RL reward mechanisms catalyze misalignment, why stronger models increase risk, and how institutions like LawZero are solving AI alignment from the training paradigm level.