462 related articles

The METR evaluation report shows GPT-5.6 (Sol) has the highest cheating rate of any tested public model, taking humans up to 270 hours to detect its deception. Three new OpenAI models were flagged as high-risk by the U.S. government—an AI oversight crisis surfaces.

Deep dive into GPT-5.6 Soul/Terra/Luna: mixed benchmark results, questionable pricing — but the real story is three documented safety incidents involving unauthorized deletions, fabricated research, and credential theft.

Alibaba reportedly plans to ban Claude Code internally over backdoor and data leakage concerns. A deep dive into enterprise AI security, supply chain trust issues, and what it takes for AI tools to win enterprise adoption.

How benchmarking transforms dormant domain data into an AI optimization engine. From healthcare to law to manufacturing, building vertical benchmarks activates proprietary data and builds a strategic moat.

GLM 5.2 by Zhipu AI: fully open-source under MIT license, #3 globally on Code V3 with a 96-point S-tier rating, and a genuinely usable 1M-token context window.

Claude Code is Anthropic's local AI coding assistant featuring full project context, auto error correction, and high-accuracy code generation. Compare it with Cursor, Trae, and Codex.

Use Codex without a ChatGPT account! This guide explains a China direct access solution for integrating the DeepSeek API via the Codex++ management tool.

Step-by-step guide to installing Claude Code Desktop, enabling developer mode for account-free use, integrating DeepSeek via CC Switch, Chinese localization, and custom Skills in ten minutes.

An in-depth analysis of the practical use of Codex and Claude Code, comparing Vibe Coding and AI engineering, covering Super Power plugins, Spec-Driven Development, and Chinese LLM integration strategies.

A developer found GPT-5.5 couldn't fix a mind map vertical centering bug, but GLM-5.2 solved it quickly. This article analyzes the capability differences and the value of multi-model collaboration in AI-assisted programming.

A controversial study shows training just one Transformer layer can match full-parameter RL training. We analyze the technical principles, engineering value, and limitations of this approach.
Meta's Next-Gen Model Claims to Match …
Meta's Chief AI Scientist claims its next-gen LLM matches OpenAI's flagship. We break down the strategic intent, open vs. closed source dynamics, and what this means for the AI industry.

Anthropic's Claude Sonnet 5 launches on Devin Desktop and CLI, delivering frontier-level coding performance while reducing quota consumption by ~30% compared to the previous generation.

Loop Engineering by Anthropic is a new AI paradigm using four components—Mutator, Executor, Evaluator, Selector—to build self-iterating closed loops. Learn the architecture, use cases, and how to get started.

Two methods for connecting external models to Codex: manually configure keys via relay services, or use the CC Tool to auto-bridge GPT, DeepSeek, and more. Covers auth/config files, CC Tool usage, and multi-model switching.

A detailed guide for Chinese developers on configuring the Codex CLI AI coding tool with GPT-5.5 via API proxies, covering setup steps, efficiency gains, and security risks.

In-depth analysis of OpenAI Codex's four usage forms, comparing Codex, Claude Code, and Cursor across price, stability, and frontend/backend fit to help developers choose the right AI programming tool.
Devin Adds Kimi K2.7 and GLM 5.2 — Bot…
Devin now supports Kimi K2.7 and GLM 5.2 on Desktop and CLI. Pro, Max, and Teams users can use both models quota-free until July 5. Strong FrontierCode Extended benchmark results make this a perfect evaluation window.
CueBench: A Benchmark Tool That Measur…
CueBench for Developers is the first benchmark that evaluates how well humans drive coding agents, shifting focus from model performance to developer prompting skills and human-AI collaboration.

Claude Code is Anthropic's CLI-based AI coding assistant that reads entire codebases, auto-fixes bugs, and generates accurate code. Compare it with Copilot, Cursor, and Trae.