DeepSeek V4.1 Flash Hands-On: Flash Dethrones the Pro Flagship

DeepSeek V4.1 Flash outperforms the flagship V4 Pro with native vision, 3× speed, and half the cost.
DeepSeek launched V4.1 Flash with claims of outperforming its own flagship V4 Pro on all key metrics, including native multimodal vision. Hands-on tests across a 3D city generator and an AI diagramming tool confirmed the edge: Flash's built-in visual self-verification, triple the output speed, and superior stability all beat Pro, which lacked vision entirely. Cost-wise, Flash completed two projects for under ¥10 while Pro spent ¥16 on just one. DeepSeek has since announced Pro's retirement with auto-routing to Flash — a move framed as a free upgrade but raising developer concerns about production stability.
Flash Beats Pro Across the Board — DeepSeek Rewrites Its Product Logic
DeepSeek has officially launched the V4.1 Flash model, and the update is far more significant than expected. The new model features an entirely new architecture with native multimodal vision support and is released as open source. More critically, DeepSeek claims that V4.1 Flash comprehensively outperforms the flagship V4 Pro across performance, cost, speed, and total completion time.
This leads to a counterintuitive conclusion: the Flash model, which costs several times less, actually delivers better results than the flagship. With that, the V4 Pro's reason to exist has essentially disappeared — DeepSeek has announced that V4 Pro will soon be retired, with all requests previously routed to Pro automatically redirected to V4.1 Flash. In short, Flash has completely killed off Pro.
Meanwhile, the DeepSeek web interface has been significantly simplified. The previous three modes — "Fast, Expert, and Image" — have been merged into a single unified entry point. Users no longer need to manually switch between models. The AI automatically dispatches tasks based on complexity and activates vision capabilities when it detects an image, resulting in a noticeably smoother experience.
3D Procedural Generator Test: Vision Capability Is the Critical Differentiator
Developer Yupi prepared two different projects for hands-on testing and brought in GLM 5.3 Flash for a side-by-side comparison. The first task was to build a procedural 3D city generator website, then have the AI autonomously verify the result — a test focused on frontend coding, 3D modeling, and visual understanding.
V4.1 Flash hit an output speed of 287 tokens per second, which is remarkably fast. However, tool calls and visual understanding added significant overhead, pushing total completion time to 83 minutes. The finished product looked impressive at first glance — the intricate, hilly cityscape of Chongqing was faithfully recreated, with a full procedural generator control panel on the right side supporting drag-and-scroll perspective navigation.

V4 Pro, by contrast, suffers from one critical weakness: no vision capability. After writing the code, it could only call a sub-agent and rely on another vision-capable model to verify the output. The result was noticeably less detailed modeling of Chongqing, and Shanghai's Oriental Pearl Tower ended up looking like a "candied hawthorn skewer" floating on the water — a clear fail.
GLM 5.3 Flash landed somewhere in between: the atmosphere (lighting, sunset tones) was well done, but there were multiple detail bugs including mesh clipping and inexplicable floating cubes. Overall, V4.1 Flash delivered the highest level of completeness and detail richness, GLM 5.3 Flash was roughly on par, and V4 Pro came in last by a clear margin.

For reference, the world-class GPT-6 Ostra completed a high-fidelity generation with the same prompt in just 18 minutes. That said, Ostra costs several times more, making a direct comparison with Flash somewhat unfair.
Full-Stack Engineering Test: A Double Win on Stability and Speed
The second project was building an AI-powered smart diagramming tool — a test focused on full-stack engineering capability. V4.1 Flash's advantages were even more pronounced in this round.
The interface it produced had a strong tech aesthetic, and it proactively integrated a relevant open-source project while customizing the product name — solid attention to detail. During live testing of the diagramming feature, you could see the AI's thought process in real time; output speed was extremely fast, finishing in around 20 seconds. The nodes were neatly aligned with no misalignment, a clear improvement over V4 Pro's earlier performance. More importantly, the stability was excellent — after many rounds of testing, not a single generation failure occurred.

GLM 5.3 Flash's output also had a polished tech feel and was functionally complete, but it occasionally suffered from connection misalignment and generation failures, making it slightly less reliable. V4 Pro had the most issues: the page launched with bugs, the right-side canvas defaulted to blank with a loading spinner, the AI's reasoning process couldn't stream in real time, and both generation speed and accuracy were mediocre.
The verdict is clear: V4.1 Flash wins this round. Its functionality and interface are on par with GLM 5.3 Flash, yet it finished in half the time with better stability.
Speed and Cost: 3× Faster, Half the Price
Beyond the task tests, an additional speed comparison was run: both DeepSeek models were given the same prompt to write a 10,000-character short novel. V4.1 Flash's output speed was exactly three times that of V4 Pro, with significantly lower time-to-first-token — essentially responding instantly.

The cost comparison is even more striking. V4 Pro alone cost roughly ¥16 just to build the 3D procedural generator. V4.1 Flash completed both projects for under ¥10 total — about ¥5 per project on average.
DeepSeek also updated Flash series pricing at the same time, dropping the cached input price during off-peak hours to ¥0.02 per million tokens. Daily usage costs have dropped dramatically. It's fair to say DeepSeek has once again positioned itself as the go-to model for cost-effective AI coding and everyday productivity.
V4 Pro Retirement Sparks Controversy: Free Upgrade or Business Risk?
Despite V4.1 Flash's impressive showing, the impending retirement of V4 Pro has generated controversy. DeepSeek's intent is to give users a free model upgrade, but many developers have raised concerns: products already integrated with V4 Pro could be disrupted by an automatic model switch, potentially affecting the stability and consistency of live production environments.
This actually highlights a broader issue — when model providers iterate rapidly, how do you balance "technical advancement" against "production environment stability"? For companies that have deeply embedded a large language model into their product, a sudden shift in model behavior introduces unverified risk. It's a reminder for developers to pay close attention to a provider's version management and deprecation policies when selecting a foundational model.
Regardless, the takeaways from this hands-on test are clear: V4.1 Flash has pulled off a genuine upset against its own flagship — delivering lower cost, higher speed, and native multimodal capability — and once again demonstrates DeepSeek's unwavering commitment to the value-for-money path.
Related articles

Invalid Source Material: Unable to Generate a Valid AI/Tech Article
This Twitter source material is an irrelevant marketing tweet with no AI or tech content, making it impossible to generate a valid professional article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material was limited to a single broken tweet with no usable content, making it impossible to produce a complete, high-quality article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material provided was a single vacuous social media tweet with a broken link — insufficient to support writing a complete, factual article.