DeepSeek V4.1 Flash Real-World Test: 200 Million Tokens for Just ¥35

DeepSeek V4.1 Flash handles complex 3D, game, and writing tasks for just ¥35 per 200M tokens.
A Bilibili creator stress-tested DeepSeek V4.1 Flash during closed beta, consuming nearly 200 million tokens at a total cost of roughly ¥35. Tasks included a Qingming Festival 3D scene (91 min), Hello Kitty modeling, a Wipeout-style game, an IKEA assembly webpage, and literary writing. Using Hide inference mode with multi-agent parallel workflows, the model traded time for quality — consistently outperforming expectations while falling short on precision geometry tasks like gears and OpenSCAD watch modeling.
Introduction: Big Surprises from a Budget Model
Recently, a Bilibili content creator ran an in-depth real-world test on DeepSeek V4.1 Flash, a model currently in closed beta. Although the official page still shows it as "in beta," the model is already accessible by calling its model ID directly. The creator completed multiple complex tasks during a busy afternoon peak period, consuming nearly 200 million tokens in total — at a final cost of only around ¥35.
That number alone is striking. 200 million tokens for ¥35 RMB reflects DeepSeek's extreme push on cost efficiency. Officially, DeepSeek V4.1 Flash is described as featuring an entirely new model architecture with stronger multimodal capabilities, faster inference, and lower costs. Based on this hands-on test, the Flash-tier model far exceeded expectations.
Test Configuration: Choosing Hide Inference Mode
The entire test used V4.1 Flash's Hide inference mode rather than Max mode. The creator explained that when Max mode was tried earlier, the model's reasoning process was far too lengthy — the wait times were simply unacceptable. Even with Hide mode, individual task completion times were noticeably longer than GPT and Grok series models.
This reveals something worth noting: while DeepSeek V4.1 Flash is priced extremely affordably, it invests significant reasoning time and iterative cycles to deliver high-quality output. In other words, it trades time for quality — a strategy that compensates for the parameter-scale limitations typical of Flash-tier models. The creator also speculated that the current model is likely larger than its predecessor, V4 Flash.
Along the River During Qingming Festival 3D Scene: A 91-Minute Iteration Journey
One of the most impressive tests involved generating a 3D scene in the style of Along the River During the Qingming Festival (清明上河图). This task used DeepSeek's Hannes tool and invoked roughly 9,000+ lines of local Skill source code. V4.1 Flash's workflow was notably sophisticated: it first read the specifications, then dispatched three sub-agents in parallel, each responsible for textures (Northern Song Dynasty timber-frame architecture), boat and cart structures, and distant trees.

The results showed impressive detail — water rendering, wooden and brick-tile building structures, distinctive windows and doors, the recurring "Zhao Taicheng's Shop" storefront sign, and textures for city walls and bridges all came out quite convincingly. The scene also featured camels, donkeys, horses, and various carts with solid forms and textures. Flaws did exist — some characters clipped through geometry, and the dock boats weren't fully constructed.
What stands out is the process itself: the model's initial scene was riddled with errors. Through reviewing dozens of screenshots and round after round of self-correction, it took 91 minutes to arrive at a convincing final result. This "self-diagnosis + continuous iteration" capability was a recurring highlight throughout the test.
Multi-Task Testing: From 3D Modeling to Game Development
Hello Kitty 3D Modeling
The creator asked the model to build a 3D Hello Kitty scene, including a head, sofa, desk lamp, and office chair. The initial preview showed significant issues with the head, but after 94 minutes of iterative refinement, the metallic finish on the desk and chair improved noticeably. The creator felt the result outperformed similar work they had previously generated with other tools.

Wipeout-Style Obstacle Course Game
When tasked with generating a Wipeout-style (快乐向前冲) obstacle course game, the model produced its own sound effects, designed complete levels, and implemented convincing rope-swing physics. The creator compared its quality to Grok 4.6, calling the two "very, very close." This task took 95 minutes; during that time, after the character repeatedly fell into water at stages three and four, the model proactively made fixes on its own.
IKEA Sofa Assembly Guide Webpage
Given an IKEA sofa assembly PDF, the model was asked to generate an interactive installation guide webpage. The creator gave it high marks: selected text gets a prominent yellow highlight, the three icons at the top are cleverly designed, the layout is clean, and it includes an animation showing the sofa's conversion into a bed. However, there were errors in part details (such as hex plugs) and the mounting position of a semicircular component. The task took 70 minutes, with 54 rounds of image recognition alone.

Unexpected Highlight: Literary Writing from a Flash Model
If the 3D modeling tests demonstrated the model's multimodal and engineering capabilities, the creative writing test completely upended the creator's assumptions about Flash-tier models. The model produced a short story about a person who, on a single afternoon, receives three simultaneous messages: a follow-up medical exam showing normal results, a notice from their landlord about a rent increase, and a "Happy Birthday" text from an ex.

Lines like "It turns out that exhaling in relief can also hurt" and "Three things knock at the door at once: one says you're still alive, one says you have to leave, one says you were once loved" display genuine literary sensitivity and emotional depth. The creator openly admitted, "I've really been underestimating Flash models all along," and invited viewers to rate the piece. This test demonstrated that V4.1 Flash's text generation is already competitive with much larger flagship models.
Weaknesses and Limitations: Mathematical Modeling Still Needs Work
The test wasn't without shortcomings. When tasked with generating 3D gear models — a task involving heavy mathematical computation — results were mediocre across straight, spherical, radial, and tower configurations, taking 75 minutes. The creator noted that most models struggle here, and the best performance they've seen on this type of task came from a different tool.
Additionally, in an OpenSCAD watch modeling task, details like the crown (side roller) and clasp were rendered incorrectly, and the model also added English text that didn't appear in the original reference image. The final rendered video was described as "acceptable but not up to expectations." These examples confirm that Flash-tier models still have meaningful room for improvement in scenarios requiring precise geometric and mathematical computation.
Conclusion: Redefining the Capability Ceiling of "Low-Cost AI"
Overall, DeepSeek V4.1 Flash accomplished a wide range of high-quality tasks — from complex 3D scenes and game development to webpage generation and literary writing — at a cost of roughly ¥35 for nearly 200 million tokens. Its core strengths lie in extreme cost efficiency and powerful self-iterating capability; its weaknesses center on mathematically demanding modeling tasks and longer inference wait times.
For developers and content creators, V4.1 Flash establishes a new reference point: achieving output quality close to flagship models at a fraction of the cost, within an acceptable time budget. When "low cost" no longer means "low quality," the barrier to entry for AI applications continues to fall.
Related articles

Invalid Source Material: Unable to Generate a Valid AI/Tech Article
This Twitter source material is an irrelevant marketing tweet with no AI or tech content, making it impossible to generate a valid professional article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material was limited to a single broken tweet with no usable content, making it impossible to produce a complete, high-quality article.

Insufficient Source Material: Unable to Generate a Valid Article
The source material provided was a single vacuous social media tweet with a broken link — insufficient to support writing a complete, factual article.