DeepSeek V4.1 Flash Released: Smallest Model Surpasses Flagship Across the Board, Costs Cut to One-Quarter

DeepSeek V4.1 Flash outperforms its own flagship at a fraction of the cost while preparing for a STAR Market IPO.
DeepSeek has released V4.1 Flash, its smallest model that nevertheless surpasses the flagship V4 Pro in performance, speed, and cost. The natively multimodal model merges three previous modes into one unified interface, while cutting KVCache to one-quarter, reducing high-end memory needs by 75%, and slashing storage to one-eighth — making AI Agent deployments far more economical. V4 Pro goes offline September 14, with all traffic routed to Flash. DeepSeek has also retained CITIC Securities for a STAR Market IPO, advancing on technology, pricing, and capital fronts simultaneously.
The Smallest Model That Outperforms the Flagship
DeepSeek has dropped another bombshell — V4.1 Flash is officially live. According to the company, this "smallest" model surpasses their own flagship V4 Pro across every dimension: performance, speed, cost, and total latency. This kind of "David beats Goliath" positioning is rare in the industry. Typically, vendors release larger models to outperform the previous generation, but DeepSeek has managed to leapfrog its flagship with a lighter, more efficient version — a fundamental leap in architectural and engineering efficiency.
For users, this means getting superior inference capabilities at a lower cost than the previous flagship. It also reflects DeepSeek's consistent technical philosophy: don't rely on stacking parameters — instead, push engineering optimization to drive costs to the absolute minimum.

Native Multimodal Design and Three-in-One Mode
V4.1 Flash is natively multimodal, capable of directly understanding image content. More significantly, the product experience has been simplified: the previous three separate modes — "Fast," "Expert," and "Real Image" — have been merged into a single unified entry point.
This consolidation reflects the model's adaptive capabilities under the hood. Users no longer need to manually select a mode; the model automatically routes compute and strategy based on the task type. For everyday users, the interaction barrier is lower; for developers, API calls are more unified, reducing the complexity of engineering integration.
Cost Is the Real Killer Feature: KVCache Reduced to One-Quarter
The most impactful aspect of V4.1 Flash is its dramatic cost structure optimization.
- KVCache slashed to one-quarter of the previous generation: This directly reduces VRAM consumption in long-context and multi-turn conversation scenarios.
- High-end memory requirements drop by 75%: The hardware threshold for deployment is significantly lowered.
- Static storage requirements reduced to one-eighth: Storage costs are compressed to the absolute minimum.

Why This Matters for AI Agent Use Cases
The reduction in KVCache and memory overhead is especially critical for Agent-type tasks. Agents need to maintain long-chain contextual memory, frequently invoke tools, and repeatedly read and write to cache — they are quintessential "cache-hungry" workloads. In the past, this made Agent applications prohibitively expensive to scale.
By cutting cache costs to one-quarter, V4.1 Flash dramatically improves the economics of running long-running, always-on Agents. This could become a pivotal inflection point for the mass adoption of AI Agent applications — only when per-call and long-running costs are low enough will developers dare to deploy Agents in real production environments.
KVCache (Key-Value Cache) is a core technical mechanism in large language model inference. In the Transformer architecture, when processing each token, the model needs to compute attention weights — and the Key and Value vectors from all previous tokens can be cached and reused to avoid redundant computation. This caching mechanism significantly speeds up inference, but at the cost of VRAM usage that grows linearly with context length: the longer the conversation and the more tool calls involved, the larger the KVCache.
In AI Agent scenarios, a single task may involve dozens of rounds of tool calls and accumulated intermediate results, with context windows often reaching tens of thousands or even hundreds of thousands of tokens — making KVCache VRAM consumption a deployment bottleneck. Compressing KVCache overhead to one-quarter means more Agent instances can run concurrently on the same hardware, or longer conversation chains can be supported. For production-grade Agent applications that need to run 24/7, this translates to real, tangible cost savings.
Flagship Goes Offline: Full Billing Switches to Flash Starting September 14
DeepSeek has also provided a concrete timeline: V4 Pro will officially go offline at 12:00 PM on September 14, and all requests will be routed to V4.1 Flash.
This is a remarkably aggressive decision. It means the flagship version will no longer exist as a separate product — all billing will shift to Flash pricing. In other words, users will get capabilities that exceed the old flagship at the price of the "smallest" model.

A Complete Price Restructuring
Taking the flagship offline entirely and routing everything to a lower-cost model is unusual in the industry. Most vendors keep a high-priced flagship to preserve margin, but DeepSeek has chosen to cover all demand at the lowest possible price.
The logic is straightforward: if the new smaller model already outperforms the old flagship on every metric, there's no reason to maintain a slower, more expensive option. Rather than complicating the product lineup, DeepSeek has made a clean break — fully unleashing the price advantage and using extreme value-for-money to capture market share.
The STAR Market (科创板, Shanghai STAR Market) is a dedicated board established by the Shanghai Stock Exchange in 2019, designed to support hard-tech and high-growth companies in raising capital through public listings. It has relatively relaxed profitability requirements, allowing companies with core technological advantages — even if not yet profitable — to apply. AI large model companies typically face enormous upfront R&D costs with commercial pathways still being explored, making STAR Market's listing conditions particularly accommodating for this type of company.
CITIC Securities, as one of China's leading investment banks, taking on DeepSeek's IPO advisory role signals that the capitalization process has entered a substantive phase. A successful listing would give DeepSeek access to continuous public financing channels, enabling it to maintain stronger firepower in compute procurement, R&D investment, and talent competition — without having to rely entirely on shareholder capital injections. This is critical for sustaining technological leadership in the ongoing AI compute arms race.
Capital Markets Moving in Parallel: DeepSeek Prepares for STAR Market IPO
Beyond product developments, DeepSeek has made a significant move on the capital front — the company has retained CITIC Securities to prepare for a STAR Market IPO.

This signals DeepSeek's transition from a pure technology company to a fully-fledged tech enterprise with comprehensive capital market capabilities. Speed, cost, and capital — all three fronts firing simultaneously — form a complete competitive playbook: maintain technological leadership, push prices to the limit, and stockpile financial ammunition.
The Core Dynamic of This LLM Price War
The domestic large model price war has played out across multiple rounds, but DeepSeek's approach this time is particularly noteworthy. This isn't simply a promotional price cut — it's a sustainable cost reduction achieved through fundamental architectural optimization. The costs have genuinely been cut, not propped up by subsidies.
When a company can simultaneously achieve "stronger, faster, and cheaper" while also preparing for an IPO to replenish funds, its position in this competitive landscape is undeniably strong. For the broader industry, this will also pressure other vendors to accelerate their own efficiency improvements — with developers and users ultimately being the biggest beneficiaries.
Conclusion
The release of V4.1 Flash showcases DeepSeek's characteristic engineering-first philosophy — using architectural innovation to achieve extreme cost efficiency. Native multimodal support, unified three-in-one mode, dramatically lower cache costs, flagship retirement, and a STAR Market IPO in the works: every move points toward the same goal — reshaping the market landscape through unbeatable value.
After September 14, all of DeepSeek's compute will run on Flash. This three-pronged offensive — speed, cost, and capital advancing simultaneously — may well become a landmark moment signaling that domestic large model competition has entered a new phase.
Related articles

Invalid Source Material: Unable to Generate a Valid AI/Tech Article
This Twitter source material is an irrelevant marketing tweet with no AI or tech content, making it impossible to generate a valid professional article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material was limited to a single broken tweet with no usable content, making it impossible to produce a complete, high-quality article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material provided was a single vacuous social media tweet with a broken link — insufficient to support writing a complete, factual article.