DeepSeek V4 Coming Soon: Flash Version May Surpass HY3, GPT-5.6 Reasoning Budget Quietly Cut

DeepSeek V4 Flash nears launch, GPT-5.6 reasoning quietly trimmed, and Seedance 2.5 hits 180s 4K video.
DeepSeek V4 Flash is expected to launch as early as next week, potentially surpassing HY3 and adding native vision for the first time. Meanwhile, OpenAI quietly cut GPT-5.6 Sol's reasoning budget days after launch, drawing user backlash. Anthropic extended Fable 5 access through July 19 ahead of GPT-6, while ByteDance's Seedance 2.5 reportedly generates 180-second 4K videos — a major leap in AI video generation.
Less than 24 hours after the last wave of AI news, the industry is seeing another round of rapid-fire updates. From DeepSeek's next-generation model on the horizon, to OpenAI quietly adjusting GPT-5.6 performance, Anthropic extending Fable 5 access, ByteDance's Seedance 2.5 showcasing impressive video capabilities, and AI protests in San Francisco — this AI summer is accelerating fast. Here's a breakdown of everything worth following.
DeepSeek V4 Is Coming: Flash Version May Surpass HY3
Multiple leaks suggest DeepSeek is preparing to release its next-generation model this month — likely the DeepSeek V4 general availability (GA) release or V4.1, with rumors indicating the Flash version may arrive first. According to credible sources, the GA release of DeepSeek V4 Flash could drop as early as next week, ahead of the Pro variant.
Early reports indicate the new model maintains a similar parameter scale to its predecessor but delivers a clear performance uplift, along with native vision capabilities for the first time. Native vision means the model is trained jointly on image, video, and text data from the pretraining stage — rather than bolting on a vision encoder via a post-hoc "plugin" adapter. The former approach enables deeper semantic alignment in understanding image-text relationships, spatial reasoning, and cross-modal generation; the latter often suffers from modality fragmentation and performance bottlenecks. DeepSeek's previous V2/V3 series focused primarily on text and code, so achieving native vision would represent a fundamental architectural leap from single-modal to multimodal, putting it in direct competition with GPT-4o and Gemini 1.5.
Even more noteworthy: if the Flash version performs as leaked, it could potentially surpass HY3 — the model from Tencent's Hunyuan that currently ranks among the strongest open-source models under 300 billion parameters. The open-source model community typically benchmarks performance by parameter scale: lightweight models (7B, 13B), mid-range models (HY3, DeepSeek V3, in the 100B–300B range), and ultra-large-scale models (such as MiniMax's planned 2.7-trillion-parameter architecture). If DeepSeek V4 Flash can surpass HY3, it would represent a qualitative breakthrough in inference efficiency and parameter utilization — highly valuable for enterprise users deploying on-premise or under constrained compute, while redefining the performance benchmark for its class.

Early Canary Testing Spotted
Some users report that the official GA version of DeepSeek V4 is currently in limited canary testing, open to only a small number of accounts. Early testers note a clear leap in coding and front-end generation capabilities. One notable example: a single-sentence prompt produced a complete Minecraft clone, complete with different ore types and cave systems, with impressive design, mechanics, and overall execution.

Of course, this is a single case and shouldn't be over-interpreted. Quality hasn't yet reached the highest bar set by some recent open-source models, but given that this is the lightweight Flash variant, the result is still remarkable. Additionally, prior leaks suggest DeepSeek isn't just refreshing its existing product line — the company has simultaneously launched development of a much larger frontier flagship model aimed at competing with the 2.7-trillion-parameter Pro model MiniMax is building.
GPT-5.6 Reportedly "Quietly Downgraded": Reasoning Budget Cut
Just four days after GPT-5.6 Sol launched, OpenAI confirmed it had adjusted the model's reasoning budget — a change many developers and users had already noticed as the model feeling "weaker."
The reasoning budget refers to the upper limit on intermediate reasoning tokens that a chain-of-thought (CoT) or "thinking" model is allowed to consume before generating a final answer. OpenAI's o-series and Sol models use dynamic budget allocation to balance answer quality against inference latency: a higher budget allows deeper self-reflection and multi-step verification, but compute cost and response time scale up significantly.
Interestingly, Codex product lead Thibaut initially denied any degradation, claiming "no downgrade, only good things" — but later acknowledged that OpenAI had been experimenting with the model's internal reasoning intensity, referred to internally as "juice values" — essentially a dial controlling thinking depth — and had rolled back those experiments.
Every Reasoning Tier Dropped by One Level
According to user reports, OpenAI reduced Sol's thinking budget to make the model faster and more efficient. The practical effect is that each reasoning tier has been shifted down approximately one level — meaning what "extra high" offered at launch now requires "max" to approximate. The problem: this effectively means the original max tier has "disappeared."
The Terra and Luna models appear unaffected, so the change is specific to Sol and is more pronounced there. OpenAI says the experiments have been rolled back and Sol has not been permanently weakened, but the company continues to fine-tune the reasoning budget and plans further experiments over the coming weeks to manage unexpectedly high usage.
This raises a broader question: Should vendors be allowed to quietly change a model's reasoning capabilities after launch without clearly informing users? In software terms, this is equivalent to silently degrading an API's performance SLA without notice. For developers who have built workflows relying on a specific depth of reasoning, this amounts to a "silent downgrade" that seriously undermines user trust.
Anthropic Extends Fable 5 Access Again
Anthropic has once again extended Fable 5 access through July 19 for all paid plans. Claude Code's weekly rate limits remain at 50% above normal, with users able to apply half their weekly quota to Fable 5, after which they can choose to pay with credits or switch to another model.

The logic here isn't hard to read: with OpenAI expected to launch GPT-6 in the coming weeks, Anthropic has strong incentive to keep users engaged with Fable 5 as long as possible. Meanwhile, the industry has widely noted regression in the Sonnet and Opus model lines, with most resources and compute being heavily redirected toward Fable 5 — which may well be Anthropic's strategic response to the pressure of GPT-6's looming release.
Seedance 2.5: 180-Second 4K Video Generation Turns Heads
ByteDance's Seedance 2.5 is emerging as one of the most anticipated AI video models, with testers already sharing exclusive generation samples. Beyond impressive visual quality, the bigger upgrade is scale — Seedance 2.5 is reportedly capable of generating videos up to 180 seconds long with 4K output, a massive leap over most current video generation models.
Understanding why this matters requires appreciating the core technical barriers in AI video generation. Leading models (Sora, Kling, Runway Gen-3) all face a fundamental quality-length tradeoff: the longer the generated video, the harder it becomes to maintain temporal consistency across frames, with character appearance, lighting, and physical motion prone to "drift" or abrupt discontinuities. 4K output squares the attention computation challenge in the spatial dimension, while 180 seconds means the model must maintain coherent narrative across thousands of frames in the temporal dimension. If Seedance 2.5 can overcome both simultaneously, it likely reflects ByteDance's deep investment in video understanding and distributed training infrastructure.
This is still test-stage material, of course. But if ByteDance can maintain this level of consistency across three continuous minutes of generation, Seedance 2.5 could become the strongest AI video generation model on the market — and signal that AI video is crossing the threshold from "short-clip demos" to "viable for real content production."
San Francisco AI Protests: Can They Actually Stop AI Development?
Hundreds of protesters gathered outside the San Francisco offices of OpenAI, Anthropic, and Google DeepMind, calling for a pause in AI development. Meanwhile, prediction market Polymarket gives AI safety legislation about a 16% chance of passing before year's end.

Attempting to fully halt AI development isn't realistic, and the structural logic runs deep. The core dilemma of AI regulation is the ineffectiveness of unilateral constraint: if one country or company unilaterally slows down, competitors gain a strategic advantage — a classic prisoner's dilemma closely analogous to Cold War nuclear arms control, where effective treaties required multilateral participation and verifiable mechanisms, not unilateral disarmament. Current international AI governance approaches include the EU AI Act's risk-tiered regulatory framework, voluntary safety commitments from the US and UK (such as the Frontier AI Safety Commitments), and the G7 Hiroshima Process. The Polymarket data showing only a 16% chance of US domestic AI safety legislation passing reflects the structural gap between legislative pace and technological progress.
The real conversation should focus on building reasonable safety guardrails and regulatory frameworks, rather than hoping technology can be "paused." Protesters' concerns are understandable on an emotional level, but without a global coordination mechanism, single-point pressure has very limited policy impact. AI is rapidly becoming a geopolitical race, and history shows these races don't stop because of protests.
Closing Thoughts
From DeepSeek V4's imminent arrival and GPT-5.6's quietly adjusted reasoning capabilities, to Seedance 2.5's continued video generation leaps, this dense round of updates again underscores just how intensely competitive the AI landscape has become. Whether it's debates over model transparency or the push and pull between safety and regulation, these are developments every practitioner and user should keep watching. One thing is certain: this AI summer has only just begun to heat up.
Key Takeaways
Related articles

AI Art Prompt Structure Breakdown: Creating a Desert Crystal Pyramid Scene
Breaking down a popular Reddit AI artwork to reveal the five core elements of structured prompts: subject, material, lighting, environment, and atmosphere for AI art scene creation.

$100 Million Deal: AI Gives 50,000 Ukrainian Kamikaze Drones Autonomous Target Lock
A U.S. company struck a $100M deal with Ukraine to deploy AI visual lock-on capabilities on 50,000 cheap kamikaze drones, enabling terminal autonomous guidance to defeat electronic warfare jamming.

The Privacy Boundaries of AI Data Collection: Your Bedroom Is Becoming a Model Training Ground
A humorous tweet about clothes entering AI training data reveals the privacy dilemma of AI data collection. We explore machine unlearning challenges, consent issues, and how users can balance convenience with privacy.