DeepSeek V4.1 Flash Released: 552B Parameters Outperform V4 Pro Across the Board

DeepSeek's 552B MoE-based V4.1 Flash beats V4 Pro in performance and slashes costs, with full migration on Sept 14.
DeepSeek released V4.1 Flash on September 10th via a rapid 6-tweet announcement. The model uses a 552B-parameter MoE sparse architecture, activating only 8–16B parameters at inference, with KV cache VRAM reduced to one-quarter of the previous generation. It outperforms the former flagship V4 Pro on performance, speed, and cost benchmarks, beats same-tier competitors on agentic coding tasks, and adds native vision understanding. Pricing is lower than the retiring V4 Flash, with a limited-time 50% discount. Starting September 14th, all V4 Pro traffic will migrate to the new model.
Breaking: DeepSeek Drops 6 Tweets to Launch V4.1 Flash
On September 10th, DeepSeek made a bold move, firing off 6 tweets in rapid succession to officially announce the all-new V4.1 Flash model. As one of the smallest models in the mid-size tier, V4.1 Flash is built around three core upgrades: stronger capabilities, faster inference, and higher throughput — with native vision understanding baked in. According to early analysis from Bilibili creators, this model doesn't just outperform the previous flagship V4 Pro in overall benchmarks — it also overtakes competing models at the same tier on several agentic coding benchmarks.
The release timeline is notably aggressive: just two days from internal testing to public launch. This "fast iteration, instant release" cadence has the industry eagerly anticipating the upcoming V4.1 Pro.
V4.1 Flash Architecture Deep Dive: Sparse Design at 552B Parameters
The most striking aspect of V4.1 Flash is its entirely new architecture. The model totals 552B parameters and uses a MoE (Mixture of Experts) sparse activation mechanism, along with a new causal encoder-decoder structure.

The key innovation lies in sparse activation: only 8 billion parameters are activated during input, and 16 billion during output. This means that while the model's total capacity is 552B, the actual compute cost at inference time is kept extremely low. This "large capacity, small activation" philosophy is exactly where modern large models are headed in balancing cost and performance.
Memory efficiency is another major win. According to official data, V4.1 Flash's KV cache VRAM usage is just one-quarter of the previous generation, with SSD storage requirements down to one-eighth. For heavy agentic workloads that involve frequent cache reads and writes, this translates to substantial cost savings — particularly critical for long-running automated agent tasks.
What is MoE? MoE (Mixture of Experts) is the dominant sparse architecture paradigm for large language models today. Unlike traditional dense models that activate all parameters on every forward pass, MoE splits the model into multiple "expert" sub-networks. For each token, a router dynamically selects only a small subset of experts to process it. This allows the model to scale total parameter count dramatically without linearly increasing inference compute — enabling richer knowledge representation during training. GPT-4, Mixtral, and Google's Gemini series all use similar approaches. V4.1 Flash's 552B total parameters map to just 8–16B activated at inference time, an activation ratio of roughly 3–29%. This extreme sparsity maintains throughput while placing high demands on the router's precision.
What is KV Cache? The KV (Key-Value) cache is a critical acceleration mechanism in Transformer inference. During autoregressive generation, each new token requires attention computation against all previous tokens. The KV cache avoids redundant computation by storing previously computed Key and Value matrices in VRAM, at the cost of memory usage that grows linearly with context length. In long-context or multi-turn agentic scenarios, the KV cache often becomes the VRAM bottleneck, directly limiting concurrency and per-inference cost. By reducing KV cache VRAM to one-quarter of the previous generation, V4.1 Flash can support longer context windows or higher concurrent requests on the same hardware — a meaningful practical improvement for automated agents that continuously call tools and maintain long-term memory.
Benchmark Dominance: V4.1 Flash Beats V4 Pro Across Multiple Dimensions
Based on multi-source benchmark data cited by DeepSeek, V4.1 Flash leads the previous flagship V4 Pro across performance, cost, speed, and end-to-end latency.

Particularly noteworthy are the agentic coding benchmarks. In this demanding category — which tests code comprehension and multi-step reasoning — V4.1 Flash not only surpasses V4 Pro but also overtakes competing models at the same tier. This demonstrates that DeepSeek's architectural optimizations have translated into real-world task performance gains, not just raw parameter scaling.
Combined with native vision understanding, V4.1 Flash has effectively evolved from a pure text model into a multimodal foundation — opening up significantly more deployment scenarios in complex agentic applications.
What are Agentic Coding Benchmarks? Agentic coding benchmarks are an emerging evaluation dimension distinct from traditional code generation benchmarks like HumanEval or MBPP, which only test single-turn completion. Agentic benchmarks require models to complete full software engineering tasks across multiple interactive steps: understanding requirements, breaking them down, calling tools, fixing errors, and producing runnable output. Representative benchmarks include SWE-bench (real GitHub issue bug fixes) and extended versions of HumanEval. These benchmarks demand strong long-horizon planning, consistent instruction following, and deep code context understanding — widely considered a more accurate reflection of real-world development workflow performance than single-turn code completion. V4.1 Flash's success in this category indicates genuine improvements in multi-step reasoning chains.
Pricing: V4 Pro Retirement Countdown Begins
Beyond the performance upgrades, the most immediately impactful change for users is the new pricing. DeepSeek has confirmed the following rates per million tokens:
- Cache hit: up to $0.06
- Cache miss: $0.30
- Output: $1.20

DeepSeek is also running a limited-time promotion with 50% off during specific time windows. The overall pricing is already lower than the soon-to-be-retired V4 Flash — more capability at a lower price.
Alongside the new model launch, the retirement timeline for the old flagship has been set: starting at 12:00 PM on September 14th, all V4 Pro traffic will be routed to the new model, billed at the new rates. Tools such as WorkBuddy and OpenCode will also gain full V4.1 Flash support.

Open Source and Large-Scale Deployment Support
On the open-source front, DeepSeek says it will work with the open-source community to improve inference support. For enterprise customers with large-scale deployment needs, the company explicitly stated that deployments at the 2,000-GPU scale can be arranged through direct negotiation — a clear signal of DeepSeek's ambitions in the enterprise market.
Conclusion: Time to Migrate to V4.1 Flash
The launch of V4.1 Flash showcases DeepSeek firing on all cylinders: architectural innovation (MoE sparse activation, causal encoder-decoder), cost optimization (KV cache and storage efficiency), and bold commercial strategy (aggressive pricing cuts, smooth migration path). The two-day turnaround from internal testing to public release also hints that stronger follow-up models like V4.1 Pro are not far behind.
For developers and enterprise users, with the V4 Pro retirement deadline of September 14th approaching, getting familiar with and migrating to V4.1 Flash should be a near-term priority.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.