GLM-5.3 Open-Source Tops Charts, Gemini 3.7 Flash Released, DeepSeek V4 Pro Withdrawn

GLM-5.3 leads open-source, Gemini 3.7 Flash ships cheaper and stronger, DeepSeek V4 Pro gets pulled within hours.
August 14 saw a flurry of AI launches: Zhipu's GLM-5.3 built on a 743B base delivers 50% better coding and promises open weights; Google's Gemini 3.7 Flash hits 65.3% on DeepSWE at half the previous price; DeepSeek V4 Pro 0813 was quietly released then withdrawn in under 24 hours; OpenAI unveiled UltraFast API (750 tok/s via Cerebras) and a macOS Computer History memory feature. MiniMax launched Music 3.0, and China's domestic chip and compute ecosystem advanced significantly.
A New Milestone for Open-Source LLMs: Zhipu's GLM-5.3 Released
August 14th was an extraordinarily eventful day in the AI space, with multiple leading companies launching new products in rapid succession—alongside the dramatic withdrawal of a flagship model. First up is Zhipu's GLM-5.3.
According to reports from the AI Compass channel on Bilibili, GLM-5.3 is built on a massive 743B base model with extensive post-training, targeting two core scenarios: coding and agents, while also setting a new benchmark for open-source models in cybersecurity. The 743B (743 billion parameters) base is one of the largest parameter-scale foundation models in the open-source ecosystem. While parameter count directly impacts a model's knowledge capacity and reasoning ability, simply stacking parameters isn't a silver bullet—post-training is equally critical. Post-training typically includes supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and alignment training for specific scenarios. Zhipu's deep post-training focused on coding and agents means the model has undergone extensive specialized optimization in code generation, multi-step reasoning, and tool calling. In practice, its coding performance evaluation shows approximately a 50% improvement over the previous GLM-5.2—a significant iterative leap that directly reflects the depth of post-training optimization.
Across multiple public benchmarks, GLM-5.3 consistently ranks at the top among open-source models. Even more commendable is Zhipu's commitment to releasing all weights within two weeks after completing security hardening. This "security first, then openness" strategy addresses industry concerns about LLM safety while maintaining the open stance characteristic of China's open-source community. For developers, a model with dramatically improved coding capabilities that's also fully open-source is undoubtedly an extremely attractive option.
Google's Rapid Iteration: Gemini 3.7 Flash Makes a Major Leap in Coding
Google DeepMind's update cadence is equally impressive. Just three weeks after the previous version, Gemini 3.7 Flash has officially arrived, positioned as the strongest work model for coding and agent scenarios.
This generation of Flash shows major improvements on coding benchmarks: Frontier Code scores reach 43.6%, while DeepSWE hits 65.3%. Frontier Code focuses on frontier-level code generation and reasoning tasks, covering high-difficulty scenarios like algorithm design and system architecture, effectively differentiating models' ceilings on complex programming tasks. DeepSWE simulates real software engineering workflows, including bug fixes, code reviews, cross-file refactoring, and other practical development scenarios that closely mirror engineers' daily work. Gemini 3.7 Flash achieving 65.3% on DeepSWE means it can provide correct or high-quality solutions for nearly two-thirds of real software engineering tasks—a level that already has practical value. Even more critical is the pricing strategy—Gemini 3.7 Flash costs only half of 3.6 Flash, demonstrating Google's aggressive approach to cost efficiency.
The model is already available on mainstream development tools including Cursor, Devin, and GitHub Copilot, meaning developers can immediately experience its capabilities in their actual workflows. The combination of "faster, cheaper, stronger" is the consistent product positioning of the Flash series and Google's core weapon in facing intense competition.
DeepSeek V4 Pro Flagship Withdrawn: An Episode Lasting Less Than 24 Hours
The most dramatic event of this period is undoubtedly the unexpected withdrawal of DeepSeek V4 Pro 0813.

According to reports, DeepSeek's V4 Pro 0813 release notice on its official website and open platform announcements have all been withdrawn, less than 24 hours after the quiet release on the evening of August 12th. Intriguingly, however, the related information for DeepSeek V4 Pro 0813 remains visible in the official API documentation.
This kind of "release then withdraw" operation is uncommon in the industry and may involve model quality control, release schedule adjustments, or other strategic considerations. Regardless of the reason, this episode reflects the extreme tension in current LLM competition—companies want to seize the first-mover advantage in releases while also needing to carefully balance quality and stability. For the outside world, we may need to wait for DeepSeek's subsequent official announcement.
OpenAI's Two-Pronged Attack: UltraFast API and Computer History Memory
OpenAI has also been highly active in this round of competition, launching two significant features in one go.
First is the new UltraFast API service tier. This service runs GPT-5.64 with speeds up to 14x faster, powered by Cerebras, outputting up to 750 tokens per second. Cerebras Systems is a company specializing in AI acceleration hardware, with its core product WSE (Wafer Scale Engine) being the world's largest single-chip processor, integrating trillions of transistors on an entire wafer. Compared to traditional GPU solutions, Cerebras's core advantage in inference scenarios lies in ultra-low latency and ultra-high throughput, as its architecture eliminates inter-chip communication bottlenecks. OpenAI chose Cerebras for compute support precisely for this architectural advantage—750 tokens per second output speed is roughly 10x or more than conventional API services, meeting the needs of real-time conversations, multi-turn agent calls, and other latency-sensitive application scenarios. This tier primarily targets latency-sensitive agent applications and high-concurrency scenarios, currently in early access.

Second is the optional Computer History feature. This feature captures recent computer activity on macOS as a local timeline and transforms it into memory usable by ChatGPT and Codex. From a technical perspective, this falls under "environment-aware AI"—it captures users' screen activity, app switching, and operation records at the macOS system level, building a localized timeline data stream, then structuring this information for injection into ChatGPT's context window. This is fundamentally different from traditional chat memory: the latter only remembers conversation content, while Computer History remembers the user's behavioral trajectory across their entire computer. With it, ChatGPT Work can restore context across selected apps and websites, understanding scenarios like "what document were you just viewing in your browser" or "which code section did you modify in your IDE," enabling AI to provide more coherent and precise assistance based on the user's recent computer operations. You might not have noticed that this feature is off by default and requires users to actively enable it—a reasonable design consideration for privacy protection, as screen activity data is far more sensitive than ordinary chat records.
Dual Evolution in Music Generation and Developer Tools
Beyond the large models themselves, the supporting application ecosystem is also evolving rapidly.
MiniMax released Music 3.0, a next-generation open-weight, production-grade music generation model. It can compose, arrange, perform, and produce an entire song in one pass based on creative concepts and optional lyrics, with single tracks up to 5 minutes long. The model targets creators and production scenarios, with weights available for download, further lowering the barrier to AI music creation.

On the developer tools front, Cursor launched Bios for Cloud Agents, which continuously prepares ready-to-use development environment replicas in the background, allowing cloud agents to start working without waiting for configuration. The company claims up to 10x faster environment startup, approximately 3x faster first-token generation, and maintained operation even if dependency updates fail. This feature comes at no additional cost and will become the default option starting August 17th. This type of infrastructure optimization is a key piece in improving agent practicality—when an AI coding assistant's environment startup time shrinks from minutes to seconds, the developer experience undergoes a qualitative transformation.
China's Compute Infrastructure Accelerates: Token Pricing Standardization and Chip Adaptation
On the industry side, Chinese companies' strategic moves are equally noteworthy.
China Mobile announced at its interim results briefing that it will launch a standardized nationwide Token pricing plan for individual and enterprise customers. Tokens are the basic unit of measurement for how large language models process text, with one Chinese character corresponding to approximately 1.5-2 tokens. Currently, industry LLM API pricing is generally based on per-million-token billing units, but there are massive differences between vendors in pricing structures, billing granularity, and package configurations, creating difficulties for enterprise procurement and cost accounting. China Mobile's standardized pricing plan is similar to how the telecom industry early on unified call charges and data into standard packages—essentially commoditizing AI inference capabilities. In the first half of the year, the company completed multi-tier Token pricing design with pilot programs across multiple provinces, while building a unified AI model platform with the intention of incorporating LLM calls into a standardized billing system. This marks AI capabilities gradually moving toward "utility-style" infrastructure operation, not only lowering the barrier for SMEs to use large models but also establishing the billing foundation for large-scale AI adoption.

Chip maker Cambricon also delivered impressive half-year results: H1 revenue of 5.996 billion yuan, up 108.13% year-over-year; net profit of 2.311 billion yuan, up 122.61% year-over-year. Chairman Chen Tianshi revealed that the company has completed inference adaptation for five mainstream Chinese LLMs—GLM, DeepSeek, Qwen, Kimi, and MiniMax—with the sixth-generation chip architecture currently in development. It's worth noting that LLM inference adaptation is far from simply "swapping in a different chip." Mainstream LLM training and inference frameworks (such as PyTorch, vLLM, TensorRT-LLM, etc.) have long been built around the NVIDIA CUDA ecosystem, with operator libraries, parallelism strategies, and memory management all deeply coupled. For domestic chips to achieve adaptation, they need to redevelop or port at multiple levels including compilers, operator optimization, and distributed inference frameworks. Cambricon's completion of inference adaptation for five mainstream models means its software stack has achieved general capabilities covering mainstream architectures including MoE (Mixture of Experts) and Dense models—an important indicator of the maturity of China's domestic AI chip ecosystem. The deep collaboration between domestic chips and domestic LLMs is building a more self-reliant and controllable AI technology stack.
AI Model Capability Rankings: A Snapshot of the Current Competitive Landscape
Finally, let's look at the current AI model capability rankings. Claude Opus 5 leads the intelligence leaderboard at 63 points and dominates the agent rankings; GPT-5.6 leads the coding competition; Claude 3.8 Max holds steady at third on the agent rankings, proving that open-source is closing in on closed-source camps. The newcomer Gemini 3.7 Flash enters at ninth on the intelligence chart with 56 points, highlighting its speed advantage.
Overall, the dense cluster of releases on August 14th clearly outlines three main threads of current AI competition: open-source models continuing to approach closed-source top-tier performance, coding and agents becoming the core battleground, and compute power and cost efficiency determining commercial success. Whether it's Zhipu, Google, OpenAI, or domestic chip manufacturers, all are accelerating on their respective tracks. This race is far from over and deserves our continued attention.
Related articles

Claude Autonomously Designs Proteins with 35% Success Rate, Far Exceeding Human Expert Performance
Anthropic's Claude achieves 35% wet-lab success rate in autonomous protein design, far surpassing the 10-15% human expert average, signaling AI's move toward real scientific productivity.

Perplexity Discover's Multilingual Support Suddenly Disappears — Why Are International Users Upset?
Perplexity Discover's multilingual news feature suddenly dropped non-English support, frustrating international users. We analyze possible causes and the broader challenges of AI product internationalization.

GitHub Daily · August 20: Mojo Tops the Charts & The Local-First Open Source Rebellion
GitHub Trending Aug 20: Mojo tops charts for AI compute stack ambitions, OpenLogi surges 1225 stars with local-first philosophy, and privacy rebellion dominates.