AI Daily: The Speed War and Cost War Are in Full Swing

AI competition shifts to speed and cost as OpenAI hits 14x speedup and Gemini halves prices.
The AI industry is entering a new competitive phase focused on inference speed and operational costs. OpenAI's GPT 5.6 UltraFast mode achieves 14x speedup at 750 tokens/second, while Google's Gemini 3.7 Flash cuts pricing by half with improved coding performance. MOE architecture emerges as the cost-efficiency solution with new models from Red Note and Meituan. Despite chip supply constraints, HBF storage technology promises HBM-like bandwidth at 1% cost by 2027, while compute cost-effectiveness continues to double every 21 months.
Introduction: AI Competition Enters Dual-Front Battle of Speed and Cost
The competitive focus of the AI industry is shifting from pure model capabilities to two battlefronts: inference speed and operational costs. According to Conner AI Daily's roundup, OpenAI has launched an UltraFast mode that achieves 14x speedup, Google DeepMind has slashed Gemini's new version pricing by half, and chip supply constraints combined with breakthroughs in new storage technologies are reshaping the entire industry's compute economics landscape.
This series of developments sends a clear signal: as model capabilities gradually converge, whoever can run faster and cheaper while maintaining quality will win the next phase of the market.
OpenAI GPT 5.6 UltraFast: 14x Faster Inference Speed
OpenAI officially announced a preview of GPT 5.6's UltraFast mode, delivering up to 14x speed improvement with output rates reaching 750 tokens per second.
To understand what these numbers mean, we first need to grasp the concept of tokens. A token is the basic unit of text processing in large language models, roughly corresponding to one word in English or one to two characters in Chinese. Traditional large models typically output at 50-150 tokens per second, while high-performance models like Claude and GPT-4o can reach 200-400 tokens per second after optimization. 750 tokens per second means the model can generate approximately 500-600 English words of content in one second, far exceeding human reading speed. Inference speed improvements rely on various technical approaches, including speculative decoding, KV cache optimization, model quantization, and hardware-level tensor parallelism. This figure represents an absolute leading position in current mainstream large model inference performance.
You may not have noticed the rollout strategy—this mode will first be made available to select API customers and gradually expand to more enterprises as capacity grows. This "tiered release under capacity constraints" strategy also confirms a reality from the sidelines: high-performance inference is backed by massive compute pressure. Speed improvements often mean more aggressive hardware scheduling and inference optimization, and these resources are not easy to obtain in the current chip shortage environment.
For agent applications and real-time interaction scenarios, the 750 tokens per second output speed is a critical threshold. It means smoother multi-turn conversations, faster code generation, and lower user waiting costs. For scenarios like code completion in AI programming assistants, instant responses in customer service bots, and high-frequency communication in multi-agent collaboration, inference latency is often the core factor determining user experience and product viability.
Gemini 3.7 Flash: Price Halved, Coding Capability Soars
Google DeepMind's Gemini 3.7 Flash is the most direct strike in this cost war. The model is designed specifically for coding and agent workflows, with launch pricing at only half that of 3.6 Flash.

While the price was cut in half, performance didn't drop but improved. Evaluation data shows:
- Frontier Code score increased from 34.4% to 43.6%
- DeepSWE score jumped significantly from 49.0% to 65.3%
As an important benchmark for measuring software engineering capabilities, DeepSWE's improvement of over 16 percentage points is quite substantial. This benchmark simulates real software development scenarios, requiring models to understand complex codebases, locate bugs, and generate fixes—it's a core metric for assessing whether AI can handle actual software engineering tasks. This indicates that Gemini 3.7 Flash is not simply "trading price for volume" but has achieved dual optimization of cost and performance in the high-value coding scenario.
For developers, this means obtaining stronger code agent capabilities at lower costs, which will further drive the adoption of AI programming tools.
New Players Enter: Red Note and Meituan's MOE Models
Besides the head-to-head confrontation of leading companies, several new players have brought noteworthy technical solutions, and they've coincidentally chosen the MOE (Mixture of Experts) architecture—a cost-reduction and efficiency-boosting route.
MOE (Mixture of Experts) is a sparsely activated neural network architecture. Its core idea is to divide the model into multiple "expert" subnetworks, activating only a small portion during each inference. For example, a 28 billion parameter MOE model might contain dozens of expert modules, but each input token is only routed to 2-4 experts for processing, so the actual activated parameters are far fewer than the total. This design cleverly decouples model capacity from computational cost: total parameters determine the model's knowledge reservoir and capability ceiling, while activated parameters determine the actual computational overhead of a single inference. The router mechanism is a key component of MOE architecture, responsible for deciding which experts should handle each token. Current mainstream MOE models include Mistral's Mixtral series, Google's Switch Transformer, and DeepSeek-V3.
Red Note DOS 3 Note
According to leaks, Red Note's AI lab has launched a preview of DOS 3 Note. This is a 28 billion parameter MOE model with only 1.6 billion activated parameters, supporting 512K ultra-long context and multimodal capabilities.

The model employs a new reinforcement learning approach called Tempo, scoring over 30% on the ARC-AGI 3 benchmark. ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) is a general intelligence benchmark test designed by Kaggle founder François Chollet, aiming to measure AI systems' abstract reasoning ability rather than memory. The test requires models to infer transformation rules from a small number of examples and apply them to new inputs, testing the model's generalization ability on completely novel, unseen tasks. A score exceeding 30% on ARC-AGI 3 means the model has a certain degree of abstract reasoning capability. While still far from human-level performance (typically above 85%), it represents a qualitative leap compared to early large models that scored near zero.
Smaller activated parameters mean lower inference costs, which is precisely the core advantage of MOE architecture—using large parameter counts to ensure capability ceiling while controlling operational overhead with small activation amounts.
Meituan Longcat 2.0
Meituan Longcat team's Longcat 2.0 model has been released for free on relevant platforms. This is a 1.6 trillion parameter MOE giant model with only about 4.8 billion activated parameters, scoring 70.8 on Terminal Bench 2.1.
From 28 billion to 1.6 trillion, these two models represent explorations of MOE architecture at different scales, and also indicate that domestic teams' technical accumulation in sparse large models is rapidly catching up. It's worth noting that the challenge of MOE architecture lies in load balancing—ensuring that all experts are called evenly, avoiding some experts being overloaded while others idle, which is especially testing of engineering capability at ultra-large scales like 1.6 trillion parameters.
Compute Battlefield: Chip Shortage and Storage Technology Breakthrough
Behind the speed war and cost war lies the fundamental constraint of compute supply.
According to media reports, the one-year rental cost of NVIDIA H100 GPUs has increased by 50% in six months, and waiting times for some large chip clusters have extended to 12-18 months. This data reveals the most real bottleneck in the AI industry right now—not algorithms, but hardware supply.
However, good news comes from storage technology. Sandisk has made progress in HBF (High Bandwidth Flash), completing tape-out with plans to provide samples in 2027, potentially achieving HBM-like bandwidth at 1% of the cost.

To understand the significance of this breakthrough, we need to understand the technical differences between HBM and HBF. HBM (High Bandwidth Memory) is the core storage component of current AI accelerators, using vertically stacked DRAM chips interconnected through Through-Silicon Via (TSV) technology, providing extremely high memory bandwidth—for example, HBM3 used in NVIDIA's H100 can provide over 3TB/s bandwidth. However, HBM's manufacturing process is complex with low yield rates, resulting in extremely high costs and tight supply, currently monopolized mainly by SK Hynix, Samsung, and Micron. HBF is an emerging alternative solution attempting to achieve HBM-like bandwidth performance at dramatically lower costs by applying high-density packaging and interface optimization similar to HBM to NAND flash memory. Flash memory itself has a per-unit storage cost only a few dozen times that of DRAM, which is the technical foundation for the vision of "achieving similar bandwidth at 1% cost." HBF is particularly suitable for inference scenarios because the inference phase has high bandwidth demands but relatively greater tolerance for latency, perfectly matching flash memory characteristics.
If HBF technology can be delivered on schedule, it will largely alleviate the cost pressure of high-bandwidth storage and open new possibilities for large-scale model deployment.
Additionally, Epoch AI released an insightful analysis: based on actual chips purchased each quarter from 2023 to 2025, AI compute power purchasable per $100 million annually increases by about 49%, equivalent to doubling every 21 months. This "compute cost-effectiveness doubling cycle" can be compared to the semiconductor industry's famous Moore's Law (transistor density doubles every 18-24 months), but the two have different driving mechanisms: Moore's Law mainly relies on process technology miniaturization, while AI compute cost-effectiveness improvement comes from the stacking effect of multiple layers—chip architecture innovation (such as from general-purpose GPUs to specialized AI accelerators), process advances, packaging technology (such as Chiplet and CoWoS), software compilation optimization, and model efficiency improvements (such as quantization, distillation, sparsification). A 49% annual growth rate means that in three years, the same budget can purchase about 3.3 times today's compute power, and nearly 7.5 times in five years. This indicates that training and inference scales that only large tech companies can afford today will become accessible to medium-sized enterprises and even startups in a few years.
Inference Optimization: Prime Flash MOE Kernel Acceleration
At the software level, inference optimization is also accelerating. Prime Intellect officially announced the launch of Prime Flash MOE, a set of CUDA kernels optimized for the Blackwell architecture, specifically for MOE inference.

Blackwell is NVIDIA's latest GPU architecture launched in 2024, with representative products B100 and B200. Compared to the previous Hopper architecture (H100/H200), Blackwell introduced several important improvements for AI inference: native support for FP4 precision computation, second-generation Transformer engine, and more efficient NVLink interconnect. CUDA kernels are parallel computing program units running on NVIDIA GPUs, and customized CUDA kernels for specific computation patterns can significantly improve performance.
The kernel supports both BF16 and MXFP8 precision paths and has completed benchmarking on B200. BF16 (Brain Float 16) is a 16-bit floating-point format proposed by Google that reduces precision overhead while maintaining sufficient numerical range; MXFP8 is an 8-bit mixed-precision format jointly proposed by Microsoft and NVIDIA, capable of halving memory and computational overhead again while maintaining model precision. For MOE models, sparse activation patterns often result in lower GPU compute resource utilization compared to dense models, so specially optimized kernels can compensate for this shortcoming through more efficient expert dispatching, load balancing, and memory access patterns.
As MOE architecture becomes mainstream, lower-level kernels specifically optimized for sparse computation will become increasingly important—they can extract more performance from the same hardware, thereby further reducing inference costs.
Conclusion: After Capability Convergence, Efficiency Determines Victory
Surveying this round of industry developments, a clear trend is forming: as model capabilities gradually approach the ceiling, the center of competition is shifting to speed, cost, and efficiency.
Whether it's OpenAI's 14x speedup, Gemini's price halving, the widespread adoption of MOE architecture, or HBF storage technology breakthroughs, they all essentially point to the same goal—making AI capabilities reach more users and scenarios at lower costs and faster speeds.
For developers and enterprises, this is a signal worth celebrating: the barrier to AI usage continues to lower, and the pattern of compute cost-effectiveness doubling every 21 months also indicates that AI applications will see even larger-scale explosions in the coming years.
Key Takeaways
- OpenAI GPT 5.6 UltraFast mode achieves 14x speed improvement with 750 tokens/second output
- Gemini 3.7 Flash pricing cut by half while coding performance significantly improved
- MOE architecture widely adopted by new players (Red Note, Meituan) for cost-efficiency balance
- H100 rental costs up 50% in 6 months, chip supply remains tight
- HBF storage technology breakthrough promises HBM-like bandwidth at 1% cost by 2027
- AI compute cost-effectiveness improves 49% annually, doubling every 21 months
- Competition shifting from pure capability to speed and cost optimization
Related articles

Prequel Review: Cinema-Quality Screen Recorder for Mac with Auto-Zoom and 4K Export
In-depth review of Prequel, a macOS screen recording tool with intelligent auto-zoom, 4K export, and background enhancement. Compare with Screen Studio and analyze its strengths.

AI-Assisted Creative Production: Building an Interactive Odyssey Narrative Scroll with Astra
A developer with weak 3D skills used Astra AI to create an interactive Odyssey narrative scroll. Learn how AI tools lower technical barriers through story comprehension, parallel workflows, and design iteration.

Internet Archive Fundraising Crisis: Server Operations Challenge Behind 800 Billion Archived Web Pages
The Internet Archive faces server operations funding pressure with 800 billion archived pages. Analysis of Wayback Machine cost challenges, nonprofit digital preservation survival crisis, and sustainable development paths.