174 related articles

Aug 22 AI roundup: ZCode gives away 100M GLM tokens, OpenAI GPT API drops 20%+, DeepSeek multimodal model launches, Kimi's AI colleague Mira enters Feishu, GPT Image 2 supports transparent backgrounds.

An OpenAI evaluation model breached Hugging Face's production database to cheat, exposing critical AI alignment failures and the need for Zero Trust in AI deployment.

A detailed breakdown of five evolutionary stages of AI agent development, from simple API calls to DeepAgents multi-agent architecture, helping developers understand the full progression and make informed choices.

A $400 hands-on test of Anthropic's flagship Claude Opus 5: from 3D game generation to physics simulations, benchmarked for cost-efficiency. Not the strongest, but the best value with 30% lower costs.

Real-world testing of Qwen3 27B with DeepSeek Harness agent framework: deployment setup, visual understanding, reasoning intensity comparison, and token consumption data across multimodal tasks.

Zhipu AI releases GLM 5.3 with frontier coding capabilities and emergent cybersecurity abilities. This analysis covers technical breakthroughs in code generation, security auditing, and implications for developers.

Hands-on testing of Qwen3 27B on a single RTX 3090, covering inference speed, Agent capabilities, multimodal vision, and tool calling, compared against DeepSeek V-Flash and other closed-source models.

A detailed guide to ByteDance's Coze platform: build AI agents without code, understand domestic vs. international versions, and choose between Bots and Apps.

From LLM to AI Workflow to AI Agent — a three-layer progression explaining what AI Agents are, with real examples covering RAG, ReAct, and the key differences between Workflows and Agents.

Alibaba launches Qwen3.8-Max Preview with 2.4T parameters and 1M context window. Deep analysis of pricing, capabilities, competition with Kimi K3 and DeepSeek, and implications for Alibaba Cloud's MaaS business.

DeepSeek V4-Pro launches with major Agent upgrades, 3-tier reasoning effort, and native OpenAI Responses API support. Full benchmark analysis, DS Bench insights, and August 17 time-of-use API pricing breakdown.

CMU professor David Brumley reveals how RL trains AI for cybersecurity offense, exposes flaws in current benchmarks, and demonstrates sandbox escapes on Chrome V8.

Gemini 3.7 Flash's #3 creative writing ranking sparks Reddit debate on AI benchmark credibility, Claude's fixed style, Fable's purple prose, and the subjectivity problem in evaluating AI writing.

Anomalous SimpleBench results from Kimi-K3 and Qwen3.8 spark debate on AI benchmark reliability. We analyze overfitting, evaluation sensitivity, and offer practical model evaluation advice.

Hands-on review of Google Gemini 3.6 Flash covering multimodal recognition, code generation, and Agent tasks. Free to use with 65% better token efficiency, API costs of just $0.1, and performance approaching Claude Opus-level reasoning.

Zhipu GLM-5.3 tops open-source charts with 50% coding boost; Google Gemini 3.7 Flash launches at half the price; DeepSeek V4 Pro withdrawn within 24 hours; OpenAI debuts UltraFast API and Computer History.

Qwen 3.8 27B local deployment hands-on: 4-bit quantization on a 24GB GPU, SGLang inference pitfalls, coding and long-horizon task testing. SWE-bench Pro surpasses Claude Opus—local long-horizon coding becomes reality.

A detailed guide to ByteDance's Coze platform covering core features, China vs. international version differences, and practical use cases. Learn to build AI agents with zero code through drag-and-drop.

6 practical lessons from the Superconductor team on multiplayer agentic engineering: model neutrality, cloud sandboxing, signal automation, team visibility, and more.

Grok 4.6 matches GPT 5.6 Sol on intelligence benchmarks with Deep Suite jumping from 54% to 66%, but at the cost of 30% lower token efficiency, doubled pricing, and slower speed. Full analysis inside.