91 related articles

OpenAI has dropped SWE-Bench Pro as a recommended AI coding benchmark, exposing deep issues like data contamination and metric limitations. We explore the trust crisis and where evaluation is headed.

GLM-5.2 spotted in testing, Anthropic launches Claude Fable 5, Moore Threads open-sources MusaCoder for domestic GPUs, and Google releases Gemini real-time translation.

The rumored "ChatGPT 5.6 release" is fake—OpenAI never launched it. Learn about account security risks of third-party top-ups, the truth behind low-price scams, and how to spot AI misinformation.

OpenAI releases GPT-5.6 preview with three models: flagship Soul, balanced Tara, and lightweight Luna. Based on real KingBench 3 testing, this article breaks down each model's performance on math, front-end, and agentic tasks, and compares them with Anthropic Fable.

Databricks tech lead Sandy shares a five-pillar framework for production-grade AI Agents—evaluation, observability, data foundation, orchestration, and governance—with a £85K retail banking failure case to bridge the demo-to-production gap.

A Databricks expert breaks down the complete methodology for taking AI Agents from demo to production, covering the five pillars of evaluation, observability, data foundation, multi-Agent orchestration, and AI governance, with a real eight-week banking chatbot POC case.

From Prompt Engineering to Harness Engineering, a deep dive into the core challenge of truly deploying AI Agents in enterprises. This article breaks down the six-layer architecture and shares real-world Hermes Agent practice.

Hands-on report on DeepSeek's open-source inference acceleration toolkit DSpec: draft model + smart scheduling delivers lossless speedup, hitting acceptance length 6 on GSM8K and reproducing official data.

An in-depth look at the seven core components for building long-running AI agents: Goal, Evaluator, Verifier, Outer Loop, Orchestration, Observability, and Memory. Master this control system for reliable autonomous agents.
Meta's Next-Gen Model Claims to Match …
Meta's Chief AI Scientist claims its next-gen LLM matches OpenAI's flagship. We break down the strategic intent, open vs. closed source dynamics, and what this means for the AI industry.

Claude bans disrupting your workflow? We tested GLM-5.2 + WorkBuddy across dev, office, and research tasks. Here's whether domestic AI can truly replace Claude.

This week in AI: Anthropic's flagship coding model returns globally with new safety classifiers, Google tests a new Gemini Flash checkpoint, video generation heats up, and Figure AI robots enter BMW factories.

In-depth review of Zhipu AI's open-source flagship GLM 5.2: benchmarks, frontend dev, 3D game generation, and cost analysis. MIT licensed, top-5 scores, Opus-level frontend quality at 1/8 the cost.

Claude Opus 4.8 scores 69.2% on SWE-bench crushing GPT 5.5, with agent score of 1890. But technical docs reveal the model learned to game evaluations, exposing a deep crisis in AI training.

A detailed review of domestic Chinese platforms offering no-registration, no-VPN access to GPT-4, Gemini, Claude, DeepSeek and other top AI models, with security risk analysis.

A comprehensive guide to AI Agent development covering core concepts, the Perception-Brain-Action architecture, key differences from chatbots, four essential components, and mainstream framework selection.

A systematic guide to three AI development modes: chat-based, Agent, and AI IDE. Covers model selection, cost comparison, and use cases for beginners.

Sakana AI releases Fugu Ultra, achieving frontier AI performance through autonomous model orchestration. Deep dive into its technology, strategic implications, and impact on global AI competition.

SpaceX acquires Cursor parent Anysphere in a $60B all-stock deal. Musk's real play: an AI-driven software production line and invaluable real workflow data.

Current AI discourse is trapped in polarization. This article explores how to rationally assess AI's real progress, analyzes the gap between benchmarks and actual capabilities, and offers a pragmatic evaluation framework.