The Comeback Story of o1: From Severely Underestimated to Changing the History of AI Reasoning

OpenAI's o1 model was collectively underestimated at launch, proving 18 months later to be a major reasoning AI milestone.
OpenAI's o1 model was widely dismissed as overhyped when released in 2024, yet just 18 months later, AI progressed from failing basic math to conquering IMO-level problems. Its core breakthrough was internalizing chain-of-thought reasoning as a training objective and introducing Process Reward Models (PRM), achieving a paradigm shift from pattern matching to genuine reasoning. This case reveals that AI capability advancement far outpaces intuitive judgment, and architectural innovation can produce nonlinear capability leaps.
Looking Back at o1's Release: A Severely Underestimated Milestone
In 2024, OpenAI's project developed under the codename "Strawberry" was ultimately released as o1-preview. At launch, social media was flooded with accusations of "overhype." Yet just 18 months later, looking back at this model, it wasn't overhyped at all — it was severely underestimated.

As one AI observer noted on social media: "From models that couldn't do basic math to solving unsolved mathematical problems in just 18 months. This trajectory is staggeringly clear."
From "Can't Count" to Conquering Math Problems: o1's Reasoning Breakthrough
The Former Reasoning Weakness of Large Language Models
Before o1, large language models were consistently criticized for their mathematical reasoning abilities. The classic example was "how many r's are in strawberry" — early models couldn't even correctly count the letters. This lack of basic logic and computational ability led many to hold pessimistic views about AI's reasoning prospects. The root cause of this limitation lies in the fundamental nature of the traditional Transformer architecture: models generate each token in a single forward pass, lacking a mechanism to "stop and think," causing tasks requiring multi-step reasoning to often go off track from the very first step.
Chain-of-Thought: A Fundamental Paradigm Shift in Reasoning
The o1 series introduced a deep reasoning mechanism called "Chain-of-Thought" (CoT). This concept was first systematically proposed in a 2022 paper by Google Brain researcher Jason Wei and colleagues. The core idea is to have the model explicitly generate intermediate reasoning steps before providing a final answer, simulating the step-by-step thinking process humans use when solving problems.
The true breakthrough of o1 was internalizing this mechanism from external prompt engineering into the model's training objective itself. Through reinforcement learning centered on a "Process Reward Model" (PRM) — which not only rewards the correctness of the final answer but also evaluates the quality of intermediate reasoning steps — the model learned to autonomously develop multi-step reasoning within an internal "thinking space," rather than relying on user prompting techniques. This is fundamentally different from traditional few-shot CoT prompting: the former is an inherent capability of the model, while the latter relies more on carefully designed input formats.
This architectural-level transformation enabled qualitative leaps in tasks requiring deep thinking, such as mathematical proofs, code writing, and scientific reasoning. With subsequent iterations like o3 and o4-mini, this technical trajectory has fully proven its potential — models began demonstrating astonishing problem-solving abilities on International Mathematical Olympiad (IMO)-level problems.
The IMO has long been considered the "Mount Everest" for measuring AI mathematical reasoning: its problems demand not only precise calculation but also creative proof construction and cross-domain mathematical intuition. Most experts predicted in 2022 that AI would need over a decade to conquer the IMO, yet the actual timeline was compressed to less than three years. The o-series models have even proposed valuable new approaches to some long-standing unsolved mathematical problems, marking a paradigm shift in AI reasoning from "pattern matching" to "genuine problem solving."
Why Was o1 Collectively Underestimated at Launch?
The Gap Between Everyday Experience and Deep Reasoning Capability
When o1-preview first launched, ordinary users didn't experience a revolutionary change in everyday conversations. Its response speed was noticeably slower, and it was even less fluid than GPT-4 when handling simple tasks. When people measured a model built specifically for deep reasoning by casual chat standards, they naturally concluded it was "nothing special." This gap is essentially a mismatch between tool and use case — using a hammer on screws and concluding that "hammers aren't as good as screwdrivers."
The "Hype Fatigue" Effect in the AI Community
The tech community's collective skepticism toward new AI products has deep historical roots. During the deep learning boom of the 2010s, multiple breakthroughs proclaimed to "change everything" ultimately failed to deliver on commercial promises, accumulating a significant trust deficit. Between 2022 and 2023, the public frenzy sparked by ChatGPT spawned numerous copycat products and exaggerated claims, further depleting the public's judgment reserves.
This collective defensive psychology is known in cognitive science as "expectation calibration bias" — when hype in a field chronically exceeds actual delivery, audiences systematically lower their acceptance threshold for new information. After enduring intensive bombardment from multiple rounds of AI product launches, any new product is labeled by default as "more marketing than substance." This cautious attitude is healthy in most cases, but in the case of o1, it caused a collective misjudgment — filtering out a genuine technical breakthrough along with the marketing noise.
The Value of Deep Reasoning Takes Time to Manifest
The true value of reasoning capabilities cannot be fully assessed on launch day. It was only when researchers began applying the o1 series to genuinely difficult scientific and mathematical problems that its potential gradually surfaced. This was never a product whose merits could be determined by first-day impressions. The lag in capability assessment is a universal phenomenon in AI: the time from a technology's release to its full understanding and application often requires months or even years.
What o1's Comeback Tells Us About the Pace of AI Development
The core insight from this case is: the pace of AI capability advancement is likely far faster than our intuitive judgment suggests. A model that struggled with basic calculations 18 months ago is now tackling problems that even human mathematicians find challenging.
This exponential progress curve implies:
- Current limitations don't equal long-term bottlenecks: What models can't do well today might be overcome very soon
- The impact of architectural innovation is often underestimated: The reasoning paradigm introduced by o1 proves that the right methodological shifts can produce nonlinear capability leaps — the PRM training mechanism and internalized chain-of-thought have far greater impact than simple parameter scaling
- Evaluating AI progress requires a longer observation window: First impressions on launch day are often the least accurate judgments; true benchmarking can only be completed after models are deployed on genuinely difficult real-world problems
- Tool evaluation must match the correct use case: Evaluating a deep reasoning model by everyday conversation standards is like judging a marathon runner by their 100-meter sprint time
Conclusion: Where Is the Next Underestimated Turning Point?
Looking back at o1's complete journey from skepticism to vindication, it reminds us to maintain sufficient humility when judging AI progress. The next seemingly "overhyped" release might be yet another severely underestimated historical turning point. In the field of AI, 18 months is enough to rewrite all the rules.
Key Takeaways
- OpenAI's o1 model (codename "Strawberry") was widely dismissed as overhyped at launch, but in retrospect was actually severely underestimated
- From models unable to perform basic math to solving unsolved mathematical problems took only 18 months
- o1's core technical breakthrough was internalizing chain-of-thought as a training objective and introducing Process Reward Models (PRM) to strengthen reasoning chain quality
- Reasons o1 was underestimated include the gap in everyday experience, the tech community's "expectation calibration bias," and the inherent lag in evaluating deep reasoning capabilities
- This case demonstrates that AI capability advancement far outpaces intuitive judgment, and architectural innovation (rather than mere scale expansion) can produce nonlinear capability improvements
Related articles
Industry InsightsIRS Fully Embraces Claude AI, Accelerating Federal Government's AI Adoption
The IRS is recruiting staff with 24/7 Claude AI access, marking Anthropic's breakthrough into the federal government. Explore the strategic implications and tax use cases.
Industry InsightsNadella Introduces the Loopcraft Framework: Building AI Ecosystems Through Feedback Loops
Microsoft CEO Satya Nadella's Loopcraft framework explains how to build frontier AI ecosystems through nested feedback loops across technology, business, and ecosystem dimensions.
Industry InsightsOpenAI's Internal Codex Usage Surges 56x — AI Coding Is Eating Everything
OpenAI reveals internal Codex usage data: Research up 56x, Customer Support 32x, Engineering 27x, Legal 13x since Nov 2025. AI coding tools are penetrating every department faster than expected.