Will Chinese Open-Source Models Burst the US AI Bubble? A Deep Dive into the Transmission Chain

Chinese open-source AI models won't create the US AI bubble, but may catalyze its reassessment.
Chinese open-source models like Kimi K3 and DeepSeek are compressing token premiums that US closed-source AI companies depend on, but the real US AI bubble stems from circular capital flows and mismatched return expectations internally. This analysis traces the full chain from performance convergence to price competition to valuation risk, concluding that Chinese models serve as a catalyst—not a cause—of potential market correction.
A Market Shock That Keeps Replaying
On July 16, 2026, Moonshot AI released Kimi K3—the world's first open-weight model with 3 trillion parameters. "Open-weight" means the model's trained parameter files are publicly released, allowing any company or developer to download and deploy inference independently, bypassing the pay-per-use business model of closed-source APIs. This differs slightly from traditional "open source" in software—training data and complete training code are typically not disclosed—but its core commercial impact is the same: it provides users worldwide with a high-performance alternative that doesn't depend on US closed-source vendors.
On the day of the announcement, US tech stocks experienced violent swings; days later, the Trump administration was reported to be considering tighter restrictions on China's frontier AI models. Similar scripts played out during the releases of GLM 5.2 and DeepSeek.
This gave rise to a popular narrative: Will Chinese open-source models burst the US AI bubble? When Chinese open-source models approach the performance of America's top closed-source models at significantly lower prices, are the massive AI infrastructure capital expenditures and sky-high valuations in the US market still justified?
This article dissects the transmission chain from "performance convergence" to "price competition" to "valuation reassessment," identifying exactly which nerves of the US AI industry Chinese open-source models can touch.
Chinese Open-Source Models Haven't Fully Caught Up
First, let's establish a fact: while top Chinese open-source models like GLM 5.2 and Kimi K3 score very high on performance benchmarks—even surpassing US open-source models and top closed-source models in some domains—they have not achieved full parity.
Based on Moonshot AI's published tests and third-party leaderboards, Kimi K3's overall capabilities still trail slightly behind the flagship models of the Gemini and GPT series, though the gap can no longer be neatly summarized as a fixed number of months. In the most complex Agent tasks and scenarios demanding extreme reliability and compliance, top US closed-source models still retain pricing power.
What Chinese open-source models hit first is the vast middle ground of "good enough" task scenarios. For these tasks, users ask: if the job gets done equally well, why should I pay ten times more for a top closed-source model?
Companies Voting with Real Money
This cost sensitivity is already showing up in real production environments. The CEO of Lindy, a US AI startup, publicly stated they switched most of their production hosted Agent traffic from Claude to DeepSeek-related models, cutting inference costs by roughly 90%. After cryptocurrency exchange Coinbase introduced cheaper open-source models, their overall AI bill dropped nearly 50% even as usage actually increased.
On model aggregation platforms like OpenRouter, Chinese open-source models' share of token consumption surged from about 4.5% in the first half of 2025 to over 30% since 2026, with some weeks approaching or exceeding 46%—temporarily overtaking US models. Tokens are the basic unit of measurement for how large language models process text—roughly equivalent to 3/4 of an English word or one Chinese character. AI companies charge users based on the number of input and output tokens, similar to how telecom companies charge by data usage. When Chinese open-source models drive the price per million tokens from several dollars down to a few cents, companies relying on token premiums face enormous margin compression.
From DeepSeek to GLM to Kimi, Chinese models have transformed "low-cost open source" from a one-time shock into a sustained supply.
The Impact Is on Business Models, Not Nationality
Does this mean OpenAI and Anthropic will be massively impacted? It's not that simple. What truly determines how much impact a company absorbs isn't which country it's from, but what it relies on for revenue.

OpenAI's revenue is primarily subscription-based. ChatGPT has over 900 million weekly active users with more than 50 million subscribers paying fixed monthly fees, so token pricing has limited direct impact on subscription revenue. However, its enterprise business already contributes over 40% of revenue, and that segment can't escape price competition either.
Anthropic's revenue structure leans much more heavily toward enterprise and API—annualized revenue was about $9 billion at the end of 2025, surging to tens of billions in 2026, with the company disclosing figures exceeding $47 billion, roughly 80% from API and enterprise contracts billed by usage volume. This directly faces fierce competition from open-source models on pricing.
So Chinese open-source models aren't hitting "US AI companies" as a whole, but rather those companies primarily dependent on the token premium business model. The essence of token premiums is: closed-source models maintain pricing at multiples—sometimes tens of multiples—above marginal cost by monopolizing high-performance capabilities. Once open-source models reach a "good enough" level for equivalent tasks, this premium loses its foundation.
Efficiency Gains Don't Equal Shrinking Compute Demand
Another key question: Chinese open-source models have achieved higher compute efficiency through architectural innovations—MoE (Mixture of Experts) sparse activation dramatically reduces the compute consumed per task. The core principle of MoE is dividing model parameters into multiple "expert" sub-networks, with a gating network dynamically selecting only a small subset to activate during each inference. For example, of Kimi K3's 3 trillion total parameters, only a few hundred billion may actually participate in computation each time. This means the model can have enormous total knowledge capacity while its per-inference computation cost is far below a dense model of equivalent parameter scale—decoupling parameter scale from computation cost. DeepSeek-V3, Mixtral, and other models all employ this architecture.
If more and more companies switch to efficient models, will total global compute demand actually shrink?
Jevons Paradox: Greater Efficiency, Greater Total Consumption
History provides the answer. In 19th-century Britain, economist William Stanley Jevons discovered in his 1865 work The Coal Question that after Watt improved the steam engine, Britain's total coal consumption didn't decline—it soared. The reason: the higher the thermal efficiency of steam engines, the lower the unit cost of burning coal for factories, so more industries adopted steam power to expand production. Each machine consumed less coal, but the total number of machines exploded, pushing total consumption up through greater demand.
This paradox has been repeatedly validated throughout tech history: LED lighting is more than 5x more efficient than incandescent bulbs, yet total global electricity for lighting hasn't decreased, because low costs spawned massive new use cases (decorative lighting, 24-hour commercial lighting, etc.); internet bandwidth costs have fallen by over 99%, yet global data traffic has grown by millions of times.
AI is replaying this script. According to Tirias Research estimates, global token consumption grew from approximately 6.77 quadrillion in 2024 to approximately 20.92 quadrillion in 2025; Goldman Sachs projects that with AI Agent proliferation, token consumption will grow another roughly 24x. The tokens saved through efficiency are devoured by longer reasoning chains and more complex multi-Agent collaboration. Every order-of-magnitude improvement in efficiency unlocks a batch of application scenarios that were previously infeasible due to cost—such as having AI Agents perform reasoning chains of dozens or even hundreds of steps, or having multiple Agents engage in complex collaborative dialogues.
In fact, compute supply remains the core constraint for every lab. Anthropic repeatedly adjusts usage limits through new compute deals while continuing to procure compute at massive scale; Moonshot AI, just days after releasing Kimi K3, suspended new user subscriptions due to surging demand approaching capacity limits. Efficiency is improving, yet demand is hitting supply ceilings.
The True Source of the US AI Bubble: A Complex Web of Capital Circulation
Now for the most critical question: will Chinese models burst the US AI bubble?
A distinction is needed: the preceding discussion concerned industry, while a bubble is a financial phenomenon—it asks whether asset prices can be supported by future profits and cash flows. Where does the US AI bubble come from? The answer is a complex web of capital circulation.

- Amazon and Anthropic: Amazon has cumulatively invested $13 billion, and Anthropic in turn committed to purchasing over $100 billion in compute from Amazon over the next 10 years. The left hand invests $13 billion; the right hand receives $100 billion in long-term orders.
- NVIDIA and xAI: An SPV (Special Purpose Vehicle) plans to raise approximately $20 billion ($7.5 billion equity + $12.5 billion debt), with NVIDIA investing about $2 billion in the equity portion. The SPV uses the money to buy NVIDIA GPUs and lease them to xAI, while the collateral for the $12.5 billion in debt is these very GPUs and leases themselves. An SPV is an independent legal entity established for a specific financing purpose, commonly used to isolate specific assets from the sponsor's overall credit risk. The risk here: if GPUs depreciate faster than expected, or if xAI cannot pay rent, collateral value may not cover the debt, while NVIDIA is simultaneously equipment supplier and investor—creating conflicts of interest and circular dependency.
- Oracle and OpenAI: A 5-year hosting contract totaling approximately $300 billion, yet OpenAI had a net loss of $38.5 billion last year, and rating agencies have publicly warned of default risk.
- CoreWeave: Issues debt backed by GPU assets and Meta's compute leases, while the company's credit rating remains at junk level. Planned capital expenditure for 2026 is $31-35 billion, against revenue of only $12-13 billion. CoreWeave's business model resembles an aircraft leasing company that uses heavy borrowing to purchase planes and leases them to airlines—massively purchasing NVIDIA GPUs and providing them to AI companies via leasing. Its CapEx is nearly 3x revenue, meaning the company is highly dependent on continuous financing to maintain expansion. If accelerating GPU iterations cause older cards to depreciate (similar to early aircraft retirement), or if major clients reduce leases, shrinking collateral values will trigger chain reactions.
Each deal may have sound logic in isolation, but when capital providers, equipment manufacturers, cloud providers, and purchasers are highly interlinked, it becomes much harder for the market to determine whether orders stem from end customers making sustained payments or from financing-backed infrastructure procurement. This creates two risks: independence of demand (is the customer also an investor?) and counterparty credit (will they default when financing tightens?).
This financial mechanism wasn't created by GLM or Kimi, but the price competition from Chinese open-source models may compress AI companies' revenue expectations, serving as a catalyst for the market to reassess the return rates on these long-term contracts. This is the real connection between Chinese open-source models and the US AI bubble.
Two Back-of-the-Envelope Calculations: The Payback Gap and CapEx Growth Rate

First, the payback gap. Bain estimates that by 2030, the world may need approximately $2 trillion in annual AI-related revenue to support the anticipated compute expansion; even accounting for AI-driven cost savings, there could still be an annual revenue gap of approximately $800 billion. JPMorgan's estimate requires approximately $650 billion in new annual revenue to achieve a 10% return. The core question: can today's massive investments ultimately convert into independent, sustained, genuinely collectible end-user revenue? "Independent" means revenue comes from real end users paying for AI value, not from upstream/downstream affiliates confirming revenue with each other; "collectible" means contract amounts actually convert to cash inflows rather than remaining as accounts receivable or long-term commitments.
Second, CapEx growth rate. UBS forecasts hyperscaler capital expenditure growth rates: 76% in 2026 (approximately $673 billion), potentially dropping sharply to 25% in 2027, and narrowing to 6% in 2028. Absolute amounts may still be rising, but for stocks priced on exponential growth, a step-down in growth rate alone can trigger valuation reassessment. This is known in finance as the "second derivative effect"—the market cares not only about growth itself but about the acceleration of growth. When acceleration turns from positive to negative, even if growth continues, the discounting expectations in valuation models adjust dramatically.
So my judgment is: The real powder keg of the US AI bubble is excessively high return expectations; the fuse is the mismatch between CapEx growth rates, financing conditions, and end-user revenue. Both the powder keg and the fuse are inside the US. Chinese open-source models are at most an external spark that might accelerate the reassessment.
The Cautionary Tale of Cisco: Being Right on Technology Doesn't Mean Being Right on Investment
Some may wonder: if underlying demand is genuinely growing, where does the bubble come from? This same confusion was raised 25 years ago, with an almost identical script.

In the 5 years following US telecom deregulation in 1996 (the Telecommunications Act of 1996, which broke the monopoly of giants like AT&T and allowed new entrants to build competing networks), carriers invested over $500 billion in a frenzy of building fiber, switches, and wireless networks—much of it debt-financed. Cisco, selling the equipment, saw revenue surge 850% in 5 years. In March 2000, it reached the world's highest market cap at a 200x P/E ratio. The logic then was identical to today's: traffic will inevitably explode, so no amount of infrastructure is too much.
Then the carrier debt crisis erupted (WorldCom bankruptcy, Global Crossing collapse, etc.), and they collectively slashed CapEx. Cisco's stock plummeted roughly 88% over about two years. But here's the key: internet traffic never stopped growing. Cisco's revenue grew from approximately $19 billion in 2000 to over $60 billion today, yet its stock price didn't surpass its 2000 high until December 2025—taking over 25 years.
The internet didn't fail. What failed was the market paying too high a price for that growth. A correct technology trend, growing industry demand, and investor losses can all happen simultaneously. This is the most counterintuitive feature of bubbles: your judgment about the future can be entirely correct, but if the entry price has already priced in the next decade-plus of growth, investment returns can still be catastrophic.
Of course, institutions like JPMorgan argue this time is different from 2000: the companies bearing the main CapEx—Microsoft, Google, Amazon, Meta—have massive mature businesses and real cash flows, not dependent on debt to survive. But this argument mainly applies to financially healthy cloud giants, not to highly leveraged players like CoreWeave.
Three Key Signals Worth Tracking
If you're following the development of the US AI industry bubble, I recommend focusing on these three signals:
- Cloud giants' CapEx growth rate: Remember those projected numbers—76%, 25%, 6%. If absolute spending is still rising but the growth rate is declining rapidly, the market may reassess growth expectations for the entire supply chain. Specifically, watch the CapEx data disclosed by Microsoft, Google, Amazon, and Meta in quarterly earnings, as well as management guidance for next-quarter capital expenditure. Once multiple giants simultaneously lower guidance, the signal is unmistakable.
- Independence and cash collection of AI revenue: Don't treat long-term contracts as revenue. Is the customer also an investor? Has cash actually been collected? Related-party transactions don't equal fictitious revenue, but the higher the degree of affiliation, the more revenue quality needs verification. Focus on the gap between "operating cash flow" and "revenue" in financial statements—if revenue is growing rapidly but operating cash flow growth lags far behind, it suggests much revenue may be stuck in accounts receivable or non-cash recognition stages.
- GPU depreciation schedules: Cloud giants generally depreciate servers over 5-6 years, but since architecture iteration cycles have shortened (NVIDIA's intervals from A100 to H100 to B200 to next-gen have shrunk to about 18 months), some analysts argue that under 3 years is closer to GPUs' true competitive lifespan. The core debate isn't whether old cards can still function, but whether they can generate returns at sufficiently high utilization rates and prices. If they're still on the books for 6 years but lose competitiveness in 3 years technologically, today's under-depreciation may become tomorrow's asset impairment. The choice of depreciation schedule directly affects current-period profits—extending depreciation flatters the current income statement, but if actual asset value declines faster, it will be exposed at some future point as a one-time impairment loss.
Conclusion: Chinese Open Source Isn't the Creator, but the Catalyst
Returning to the opening question: will China burst the US AI bubble?
- The general capability premium at the model layer will likely be continuously compressed by Chinese open-source models—this is the "yes" part.
- The industry's overall compute demand won't necessarily decline as a result and may continue growing—this is the "no" part.
- As for whether the valuation-layer bubble bursts, it ultimately depends on whether massive capital expenditures can convert into sufficient end-user revenue and free cash flow.
Chinese open source isn't the bubble's creator, but it may force the market to redo the math sooner.
Related articles

Jaithon 3: Analyzing an Experimental Programming Language in Pursuit of Perfect Syntax
An in-depth analysis of Jaithon 3, an experimental language promising "perfect syntax" and high performance, exploring its design philosophy, technical challenges, and community reception.

Gemini 3.5 Pro Rebranded as 3.7 Flash? Decoding the LLM Naming Maze
Reddit users spotted a Gemini 3.5 Pro checkpoint briefly appear on Arena AI before being renamed 3.7 Flash High. We analyze the product strategy and industry naming chaos behind the change.

Sanders Sends Letter to OpenAI and Other AI Giants: Pause Development or Face Legislative Regulation
Senator Sanders sent an open letter to OpenAI, Anthropic, and Meta demanding an immediate AI development pause or face Senate legislation. Analysis of the letter's context and regulatory prospects.