Why Are AI Labs Sitting on Their Most Powerful Models? Dylan Patel's Deep Dive Into the Logic Behind It

Top AI labs are shelving their best models due to regulation, risking their own competitive flywheel.
Semiconductor analyst Dylan Patel reveals that the world's most powerful AI models were finished training in February but remain unreleased due to safety and regulatory constraints. Ironically, the regulations these labs championed hurt them more than open-source or overseas competitors. With strongest models shelved, revenue per megawatt stagnates, weakening labs' ability to secure scarce compute—threatening the very flywheel that sustains their dominance.
The Most Powerful Models Were Already Trained Back in February
In a recent interview, renowned semiconductor and AI industry analyst Dylan Patel dropped a surprising revelation: the world's most powerful AI models were actually finished training back in February of this year. Yet these models still haven't been released to the public.
Dylan Patel is the founder and chief analyst of SemiAnalysis, a research firm specializing in semiconductors and AI infrastructure. SemiAnalysis is known for its deep technical analysis of AI compute supply chains, chip architectures, and data center economics, and its research reports carry significant influence across Silicon Valley investment circles and the broader tech industry. Patel's analyses are typically grounded in tracking supply chain data, power consumption, chip shipment volumes, and other fundamental indicators—rather than relying solely on public statements from labs—which is why his assessments often reveal the real dynamics beneath the public narrative.
"We've already seen a massive slowdown from AI labs," Patel stated bluntly. This slowdown doesn't stem from a plateau in technical capability, but rather from a series of strategic and regulatory choices. In other words, the leading labs aren't unable to build these models—they're unwilling (or afraid) to release them.
This assessment upends the public's intuitive sense that the AI race delivers "new breakthroughs every quarter." The reality may be that the most cutting-edge results are locked in a vault, and what circulates in the market are merely curated, sub-optimal versions.
Regulatory Lobbying Has Actually Slowed Down the Lobbyists

Patel pointed out a deeply ironic phenomenon: the AI regulations that leading labs have actively championed are actually dragging them down far more than they constrain open-source models or Chinese language models.
The logic isn't hard to follow. The entities bound by strict regulations and safety reviews tend to be the high-profile, closely watched commercial labs that must answer to the public. Open-source communities and overseas teams, by contrast, operate in a relative regulatory gray zone, free to iterate and release faster and more freely.
The open-source AI model ecosystem has experienced explosive growth in recent years. Represented by Meta's Llama series, Mistral, and China's DeepSeek and Qwen, open-source models are rapidly closing the gap with closed-source frontier models. The core characteristic of open-source models is that once the model weights are publicly released, anyone can download, modify, and deploy them—making regulatory enforcement extremely difficult. You simply can't impose retroactive restrictions on open-source weights that have already been widely distributed. By contrast, closed-source labs that serve models via APIs are inherently "regulatorily reachable"—every model update and release is subject to scrutiny from the public and regulatory bodies. This structural asymmetry is precisely the root of the regulatory paradox Patel identified.
This creates a paradox: the leaders who champion "responsible AI" are slowing themselves down through self-imposed constraints, while competitors are rapidly catching up. The "double-edged sword" of regulation is cutting deepest into those who first wielded it.
The List of Shelved "Most Powerful Models"

In the interview, Patel cited several specific cases that outline this pattern of collective shelving:
- OpenAI has not released Astra: Long-anticipated capabilities have been put on hold.
- OpenAI paused training for two weeks: The training process was deliberately halted.
- Anthropic has not released Model 2: This model, referenced in the company's safety evaluations, is widely believed to be the prototype for their next-generation flagship.
Leading labs like Anthropic have established systematic AI safety evaluation frameworks. The most notable is Anthropic's Responsible Scaling Policy (RSP), which classifies models into different ASL (AI Safety Level) tiers based on potential risk levels—similar to the Biosafety Level (BSL) classification system. Each time a model's capabilities reach a new threshold, the lab must complete the corresponding level of safety evaluation before deployment, including testing whether the model can assist with bioweapon development, autonomous replication, cyberattacks, or other dangerous capabilities. The fact that Model 2 was referenced in safety evaluations suggests it may have reached capability boundaries requiring higher-level safety review, thus delaying its release.
What these decisions have in common is that none of them represent an inability to deliver—they represent a deliberate pullback in release cadence. As Patel put it, "They clearly haven't released their best stuff."
For outside observers, this means the models we can test and use may be systematically behind the true state of the art inside these labs. The "strongest" model on public benchmarks may not be the actual ceiling.
The Chain Reaction of Revenue Stagnation

The most commercially insightful part of Patel's analysis is how he connects "shelved models" to the labs' economic models.
He introduced a key metric—revenue per megawatt. This is a core indicator of AI infrastructure economic efficiency. Training and running inference on AI models requires enormous amounts of electricity, and data center capacity is typically measured in megawatts (MW). A typical AI training cluster might consume tens or even hundreds of megawatts of power—equivalent to the electricity consumption of a small city. Revenue per megawatt essentially measures how much commercial return a lab generates per unit of compute investment. When labs don't release their strongest models, other models on the market (including open-source and competitors) become competitive again. The result is that leading labs' revenue per megawatt stagnates—or may even start declining.
"It's not that they're falling behind technically," Patel emphasized. "It's that they're not putting their best stuff out there."
This distinction is crucial. Technical leadership and commercial leadership are not the same thing. When a lab holds the most powerful model but chooses to sit tight, it remains the leader on paper, but at the market revenue level, it may be losing its edge.
Regulation, Compute, and the Breaking Flywheel

Patel pushed the logic chain further to its endpoint: what happens if regulatory factors continue to prevent labs from releasing their strongest models?
The answer is a potentially breaking "flywheel." The "flywheel effect" in AI borrows from the business concept championed by Amazon founder Jeff Bezos: a system where each component reinforces the others, creating a self-accelerating positive cycle. In the AI industry, this flywheel operates as follows:
- Models don't get released → Revenue per megawatt can't climb quickly;
- Revenue growth slows → Labs' ability to pay premiums for compute diminishes;
- Compute purchasing power weakens → They can no longer outbid everyone else for incremental compute capacity.
This flywheel was previously the core mechanism through which leading labs built their moats: deploy the strongest models to generate the highest revenue, then use that revenue to lock down the scarcest and most expensive compute resources, thereby further widening the gap with competitors. Currently, leading labs' compute procurement is already at the multi-billion-dollar scale, with Microsoft, Google, Amazon, and other cloud providers committing over $200 billion in annual capital expenditure. The global supply of high-end AI chips is primarily dominated by NVIDIA, with capacity remaining persistently tight—meaning compute procurement is fundamentally a zero-sum game. Every GPU you secure is a resource your competitor cannot access. Against this backdrop, any deceleration in any part of the flywheel can trigger cascading effects far exceeding expectations.
Once the strongest models are shelved, this positive cycle risks going into reverse. Revenue growth can't keep up, compute bargaining power erodes, and the moat narrows. This is the deeper risk Patel is flagging: regulation doesn't just cause release delays—it could destabilize the entire economic structure that leaders rely on to maintain their advantage.
Conclusion: The Leader's Dilemma
Dylan Patel's analysis offers a different lens for understanding the current AI competitive landscape. Beneath the seemingly calm surface of model release cadences lies a painful balancing act by leading labs between "safety responsibility" and "commercial competition."
They must simultaneously demonstrate restraint and responsibility to regulators and the public, while continuing to lead in the revenue-and-compute flywheel. These two objectives are becoming increasingly difficult to reconcile. When the most powerful models are locked away in a drawer, the real beneficiaries may be the open-source and overseas competitors who don't face the same constraints.
This dilemma also reflects a deeper structural contradiction in AI governance: in a globalized technology race, does unilateral self-restraint inevitably become a unilateral competitive disadvantage? When the cost of being "responsible" is the loss of market share and strategic advantage, how long can leading labs keep walking this tightrope? These questions will likely find their critical answers within the next one to two years.
Related articles

Anthropic Sued: Claude Max 20x Plan Allegedly Delivers Only 6x Usage?
A lawsuit against Anthropic alleges Claude Max's 20x plan delivers only ~6x usage, and the 5x plan just 3.5x. We break down the legal details, community reactions, and the AI subscription transparency crisis.

Cursor Beginner's Guide: A Six-Step Workflow for Managing Changes, Rollbacks, and Validation
New to Cursor and keep breaking things? Learn a six-step dev workflow covering Cursor Rules, Plan mode, Diff review, and Checkpoint rollback to go from guesswork to engineering.

Is Cheap Cursor Reselling Reliable? The Real Risks of Shared Account Pools Exposed
An in-depth analysis of Cursor Pro budget reselling services, exposing the shared account pool model behind so-called legitimate accounts and deep discounts from technical, compliance, and data security perspectives.