DeepSeek V4.1 Released: Massive Parameters, Minimal Activation — Smarter, Faster, and More Affordable

DeepSeek V4.1 Fast uses sparse MoE activation for top-tier performance at a fraction of the cost.
DeepSeek V4.1 Fast is now officially available, featuring a 5,520B total parameter MoE architecture that activates only 8B–16B parameters during inference — delivering near-trillion-parameter capability while dramatically cutting hardware and deployment costs. In benchmarks, V4.1 ranks first in terminal, software engineering, and cybersecurity tests, and outperforms its predecessor V4 Pro in 17 of 18 evaluated categories. API pricing has been slashed across the board, with cache-hit input costs dropping by up to 66%. Combined with platform subsidies on WorkBuddy, real-world usage costs roughly one-twentieth of comparable models, underscoring DeepSeek's strategy of expanding adoption through aggressive affordability.
DeepSeek V4.1 Fast Version Makes Its Official Debut
DeepSeek has delivered another major update — the V4.1 Fast version is now officially available. According to an analysis by Bilibili creator Lei Ge, this generation's core highlights can be summed up in three points: stronger capabilities, faster inference, and lower prices. The model footprint has also been continuously optimized, pushing the overall direction toward something smarter and more lightweight.
For developers and enterprise users who closely follow domestic large model progress, the release of DeepSeek V4.1 is undoubtedly a signal worth examining closely. It not only ranks among the top tier in benchmark results, but also makes highly competitive adjustments to deployment costs and usage pricing.
V4.1's New Architecture: Massive Parameters, Minimal Activation
The biggest technical highlight of this release is its entirely new model architecture design. DeepSeek V4.1's total parameter count reaches 5,520B (approximately 5.52 trillion), yet during actual inference, only 8B parameters are activated on input and 16B on output.

This is the classic approach of sparse activation — the MoE (Mixture of Experts) architecture. Lei Ge offered a vivid analogy: "An entire army only needs to send one company to win the battle." In other words, while the model as a whole stores an enormous wealth of knowledge and capability, only a small subset of expert networks is called upon during each inference pass.
The immediate benefit of this design is a dramatically lower barrier for local deployment. Users can run a genuinely smarter large model with far less hardware. For small and medium-sized teams looking to deploy privately but constrained by GPU costs, this is a critical advantage.
DeepSeek V4.1 Benchmark Performance: Solidly in the Top Tier
On the capability side, the official team released comparison data pitting V4.1 against their own flagship V4 Pro, as well as mainstream models including GLM 5.3, Kimi K2, GPT-5.6, and the Claude series.

According to the benchmark chart compiled by Lei Ge, the newly released V4.1 Fast version (highlighted in deep red) delivers standout performance across several key evaluations:
- Terminal 2.1 test: Ranked first among all flagship models compared
- Software Engineering: First
- Cybersecurity: First
- Office Automation and Agent Terminal Exams: First across the board
To be objective, V4.1 still trails some top international models in certain benchmark categories. But viewed in the context of domestic model comparisons, it has firmly established itself in the first tier.
A Comprehensive Leap Over V4 Pro: 17 Out of 18 Benchmarks Surpassed

Even more notable is the comparison with the previous-generation flagship V4 Pro: across 18 evaluated benchmarks, V4.1 Fast outperforms V4 Pro in 17 of them. This generational leap is substantial, demonstrating that the new MoE architecture isn't just a cost-cutting measure — it delivers real, tangible capability gains rather than trading performance for a smaller footprint.
DeepSeek V4.1 API Pricing: Across-the-Board Cuts of Up to 66%
For developers, beyond raw capability, API pricing is often the deciding factor in adoption. V4.1 introduces significant pricing adjustments:
| Billing Item | Old Price (per million tokens) | New Price | Reduction |
|---|---|---|---|
| Input (cache hit) | ¥0.06 | ¥0.02 | ~66% |
| Input (cache miss) | ¥1.50 | ¥1.00 | ~33% |
| Output | ¥4.50 | ¥4.00 | ~11% |
The most dramatic cut is in the cache-hit input scenario, dropping by roughly two-thirds. For applications with large volumes of repeated context or batch processing needs — such as intelligent customer service, document processing, and code analysis — this translates to significantly lower long-term operating costs.
Platform Availability: Exclusive Promotion on WorkBuddy
Worth noting is that DeepSeek V4.1 Fast is now live on the WorkBuddy platform, with an exclusive two-week promotional offer.

According to the platform, WorkBuddy gives each user 100 free credits daily. Running a task with the latest V4.1 model costs only about 2 to 3 credits (at the promotional rate of 0.03x). That works out to roughly 50 free tasks per day for users.
By comparison, running a task with GLM 5.3 on the same platform might cost 30 to 60 credits. In other words, DeepSeek V4.1 on this platform costs approximately one-twentieth of other mainstream models — while delivering a more capable, up-to-date model experience. For users who run AI agent tasks at high frequency daily, that kind of value proposition is genuinely compelling.
Conclusion: Is DeepSeek V4.1 Worth Testing?
The release of DeepSeek V4.1 Fast continues DeepSeek's consistent technical philosophy — finding the balance between performance and cost through architectural innovation. The sparse-activation MoE structure turns "massive parameters, minimal activation" from concept into reality, ensuring high capability ceilings while dramatically reducing the compute required for inference and deployment.
Combined with across-the-board API price cuts and platform-level subsidies, V4.1 signals a clear market strategy: lower the barrier of entry so that more developers and everyday users can access a high-performance model. For readers following the evolution of domestic large models, this is a version well worth getting hands-on with.
Related articles

Open-Source Python SDK: Measuring AI Agent Reliability with SRE Principles
Agent Reliability is an open-source Python SDK that applies SRE's SLO and error budget concepts to AI Agent evaluation, with PASS/FAIL/UNKNOWN states, CI assertions, and zero forced dependencies.

MiniMax RefMod: A Complete Guide to Training-Free Reusable Identity Workflows
MiniMax RefMod offers training-free reusable identity workflows for image, video, and audio generation. Includes Runpod template and tutorial for quick setup.

Invalid Source Material Notice
The source material provided lacks substantive information and is unrelated to AI/tech topics, making it impossible to produce a complete professional article.