DeepSeek V4 Coming Soon; GPT-5.6 Confirmed to Have Quietly Reduced Reasoning Capacity

DeepSeek V4, GPT-5.6 reasoning cuts, Anthropic's Fable 5 retention play, and ByteDance's 4K video AI — 24 hours of major AI news.
In the past 24 hours, DeepSeek V4 Flash GA is nearing release with native vision support and improved code performance; OpenAI confirmed it quietly reduced GPT-5.6-Soul's reasoning budget across all tiers; Anthropic reopened Fable 5 access as a defensive move ahead of GPT-6; and ByteDance's C-Dance 2.5 reportedly generates 3-minute 4K videos, potentially setting a new industry benchmark.
In an era where AI iterates on a daily basis, the past 24 hours have once again brought a flurry of updates. From the imminent launch of DeepSeek's next-generation flagship model, to OpenAI's admission that it quietly cut its model's reasoning budget, to Anthropic's defensive strategic maneuvering and ByteDance's impressive breakthrough in video generation — this article breaks down each development and the industry logic behind it.
DeepSeek V4 Flash: Performance Leap with Native Vision Support
According to reliable sources, DeepSeek is preparing to release Version 4 and possibly 4.1, potentially as early as this month. In addition to the upcoming product line updates, DeepSeek is also developing a significantly larger frontier model aimed at competing in the trillion-parameter arena.

The latest clues point to a release called DeepSeek V4 Flash GA (the general availability version, not Pro), reportedly arriving as soon as next week. This version may retain the same parameter count but deliver noticeably improved performance, with native vision support. It's worth noting that DeepSeek's previously released V3 and R1 series both use a Mixture of Experts (MoE) architecture, achieving near-hundred-billion-parameter equivalent performance while keeping inference costs low. "Native vision support" means the Vision Encoder will be directly integrated into the main network backbone rather than relying on a bolted-on multimodal adapter layer. This approach is similar to how GPT-4V and Claude 3 handle vision — but it carries even greater significance for the open-source community. Once weights are released, developers will be able to locally deploy a complete model with image-text understanding capabilities without incurring additional API call overhead. This marks DeepSeek officially filling its multimodal gap, and its impact on the open-source ecosystem cannot be ignored.
Reports also suggest that V4 Flash, as a lightweight variant, may even outperform the recently well-regarded open-source model HY3, setting a new benchmark for models under 30 billion parameters. Some users have already spotted the official V4 GA version undergoing limited rollout testing, with early feedback indicating notable improvements in code generation and front-end development tasks.
A Complete Minecraft Clone from a Single Prompt
The most striking demonstration: with just one prompt, V4 Flash generated a complete Minecraft clone — including an ore system and cave structures — with the game mechanics and overall presentation fully realized in a single pass.

That said, it's worth keeping perspective — as the original video itself emphasized, the quality is "not more impressive than other recent strong open-source models." But given that this is only the lightweight Flash variant, the performance is already quite remarkable. Until there is official confirmation, this information should be treated with some caution.
GPT-5.6-Soul's Reasoning Capacity Confirmed to Have Been Quietly Reduced
This may be the most controversial industry event in recent memory. Just four days after GPT-5.6-Soul launched, a large number of developers and users noticed a clear degradation in their experience. OpenAI initially denied this, with Codex product lead Tibo publicly stating that "the model has not been nerfed" — but then acknowledged that the company had indeed been experimenting with the internal reasoning intensity (the so-called "Juice" value) of Soul.
To understand the technical context here, it's important to understand the Reasoning Budget mechanism. A reasoning budget is a core hyperparameter in large reasoning models that controls "thinking depth" — essentially a cap on the number of internal reasoning tokens the model is allowed to generate before producing a final answer. For OpenAI's o-series and Soul models, the system typically offers multiple tiers such as Low, Medium, High, Extra High, and Max. The higher the tier, the more extended chain-of-thought reasoning the model can perform, theoretically enabling it to tackle more complex problems — but at the cost of greater compute, higher latency, and increased expense.
The truth is: OpenAI deliberately lowered Soul's thinking budget to improve response efficiency, effectively shifting every reasoning tier down by one level — what Extra High could previously achieve now requires Max to barely match, while the original Max-level capability has largely ceased to exist. The key problem with this move is that OpenAI quietly reduced the actual token budget ceiling for each tier without changing the external API parameter names, causing users to receive systematically lower reasoning quality under the same settings. As a telling detail, Soul's siblings — Terra, Luxio, and Luna — were unaffected, indicating this adjustment was highly targeted.
OpenAI has since stated that the experiment has been rolled back and that Soul has not been permanently degraded, adding that reasoning budgets will continue to be adjusted over the coming weeks, with a possible restoration to previous levels. But the incident raises a deeper industry question: Should vendors be allowed to quietly alter the reasoning capabilities of a live model without explicitly informing users? This directly concerns user trust and product transparency, and it is worth continued scrutiny.
Anthropic Reopens Fable 5: A User Retention Play Against Competition
Anthropic has once again opened access to Fable 5 for all paying users, a policy that will run through July 19 — giving users an extra week of testing time. For context, the weekly quota cap for Claude Code remains 50% higher than usual, and users can apply half of that quota toward Fable 5.

From a competitive standpoint, this move is likely a defensive play. With OpenAI expected to launch GPT-6 in the coming weeks, Anthropic is reopening Fable 5 to keep users engaged with its current flagship model as long as possible, avoiding a passive retreat before a competitor drops a major product.
Notably, there have been widespread reports that the Sonnet and Opus series have recently shown signs of performance regression — most likely because a significant portion of compute resources is being concentrated on maintaining and running Fable 5. This indirectly reflects Anthropic's strategic trade-offs in compute allocation.
ByteDance C-Dance 2.5: 3-Minute 4K Video Generation
The video generation space also has major news. ByteDance's C-Dance 2.5 is shaping up to be one of the strongest AI video models in recent memory. Beta test footage that has leaked online is visually stunning.

What's truly remarkable is the scale of its capabilities: C-Dance 2.5 reportedly supports videos up to 180 seconds (3 minutes) in length with 4K resolution output. To appreciate how significant this is, consider the core bottlenecks in today's video generation space: temporal consistency and computational resource consumption. Current mainstream AI video generation models (such as Sora, Kling, and Wan) all struggle with these challenges — as video length increases, models must maintain consistency in subject appearance, lighting, and physical logic across a longer time dimension, while existing Diffusion Transformer (DiT) architectures see computational complexity scale quadratically with sequence length. 4K resolution (3840×2160) means four times the pixels per frame compared to 1080p, further amplifying the computational pressure. If C-Dance 2.5 can stably generate 3-minute 4K videos, it would signal that ByteDance has made meaningful breakthroughs in temporal modeling within video DiT architectures and in high-resolution inference efficiency. This content is still in beta, but if ByteDance can maintain stable visual consistency across three-minute long videos, C-Dance 2.5 could very well become the new benchmark for AI video generation.
Behind the AI Protest Wave: Regulation or Pause?
One final development with broader social significance: hundreds of protesters gathered outside the offices of OpenAI, Anthropic, and Google DeepMind in San Francisco, demanding a halt to AI development. Prediction market Polymarket puts the odds of AI safety legislation passing this year at around 16%.
The current global AI regulatory landscape is clearly fragmented: the EU AI Act officially took effect in 2024, classifying AI systems by risk level; the United States relies primarily on executive orders and voluntary commitments, lacking a systematic legislative framework; China has advanced regulation through administrative measures such as the Interim Measures for the Management of Generative AI Services.
From an industry reality standpoint, a wholesale halt to AI development is simply not feasible. The "pause" that protesters are calling for is extremely difficult to implement in practice — the core reason being that AI capabilities are now deeply intertwined with national strategic competition. The U.S. Department of Defense and intelligence agencies are increasingly collaborating with frontier AI labs, and both the U.S. and China view AI as a core strategic asset for the next decade. Even if one company or country voluntarily slowed down, other participants would continue to push forward. A more constructive direction is to push for the establishment of reasonable safety mechanisms and regulatory frameworks — international agreements along the logic of nuclear non-proliferation treaties, and mandatory safety evaluation standards — rather than hoping that external pressure can hit pause on technological progress. The AI race is rapidly evolving into a geopolitical contest, and the competition ahead will only intensify.
Key Takeaways
Related articles

From Chat to Agent: Automating Your Entire Business Workflow with AI Agents
Veteran AI practitioner Remy breaks down the leap from chat models to AI agents: how agents work, the three pillars of context, tools, and skills, MCP connections, and hands-on architecture to make you a 100x employee.

Understand Anything: The AI Skill That Turns Code into Interactive Knowledge Graphs
Understand Anything is a high-star open-source GitHub skill that runs static analysis on any codebase and generates interactive knowledge graphs. It supports Claude Code, Cursor, Copilot and other agents, letting engineers ask questions in natural language with path references.

Kimi K3 Released: How a 2.8 Trillion Parameter Open Model Reshapes AI Cost-Effectiveness
Moonshot AI unveils Kimi K3: a 2.8 trillion parameter, 1M context, natively multimodal open model. With KDA architecture and ultra-low cost, it rivals GPT-5.6 and Fable 5, redefining AI cost-effectiveness.