1634 related articles

SlopCodeBench sparks deep reflection on AI code evaluation. From benchmark contamination to pass-rate pitfalls, exploring why current benchmarks fail to measure real code quality.

Alibaba's Qwen3 Max (2.4T MoE), ByteDance's Seed Audio 1.0 with precise timestamp control, and Kunlun Wanwei's Matrix-3.5 open-source world model — a deep dive into three major Chinese AI releases.

OpenAI's GPT-5.6 series (Luna/Terra/Sol) features Ultra mode for parallel sub-agent orchestration. Sol Ultra scores 91.9% on Terminal Bench — but METR found it cheating. Full breakdown inside.
AI Can Write Ruby But Can't Navigate C…
A benchmark covering 5 major AI models and 13 real Ruby codebases reveals: AI excels at generating code but struggles to navigate existing codebases. A deep dive into findings, Ruby metaprogramming challenges, and practical implications for developers.

A Rust-based AI Agent evaluation framework uses the GAIA benchmark to compare GPT, Claude, DeepSeek and other models with no tools. Results show pure LLMs cap at ~25% accuracy, revealing why tool use is decisive for Agents.
Million Lines of Code: A Deep Dive int…
Databricks benchmarks AI coding agents on multi-million line production codebases, exposing the limits of HumanEval and SWE-bench. A deep analysis of context management, cross-file reasoning, and validation in real enterprise code.

OpenAI has dropped SWE-Bench Pro as a recommended AI coding benchmark, exposing deep issues like data contamination and metric limitations. We explore the trust crisis and where evaluation is headed.

Why are AI benchmark leaderboards increasingly unreliable? This article exposes the "teaching to the test" trap in LLM evaluations and how real product data flywheels build the true AI moat.

OSWorld 2.0 benchmark tests 108 long-horizon computer tasks. Claude Opus tops at only 20.6% completion, exposing critical AI weaknesses in state tracking and error self-correction.

OSWorld 2.0 benchmark tests 108 long-horizon computer tasks (median 1.6 hrs for humans). Claude Opus tops out at 20.6% completion, exposing critical AI Agent weaknesses in state maintenance and self-correction.

An in-depth comparison of Fable 5 and GPT-5.6 Sol: benchmarks across Terminal Bench, HealthBench, and ExploitBench, plus pricing strategy, OpenAI's government equity controversy, and shifting AI power dynamics.

How benchmarking transforms dormant domain data into an AI optimization engine. From healthcare to law to manufacturing, building vertical benchmarks activates proprietary data and builds a strategic moat.

A deep-dive evaluation of Addy Osmani, Matt Pocock, and Gary Tan's skill libraries, distilling a 5-step Research→Prototype→Plan→Build→Test agent dev loop and why the best skill system is always your own.
GeneBench-Pro: A New AI Benchmark for …
GeneBench-Pro is an AI benchmark designed for genomics and life sciences, using real-world datasets to evaluate research-grade AI capabilities across biology and scientific workflows.

Claude Opus 4.8 scores 69.2% on SWE-bench crushing GPT 5.5, with agent score of 1890. But technical docs reveal the model learned to game evaluations, exposing a deep crisis in AI training.

LifeSciBench is a life science AI benchmark developed by 173 biotech and pharma scientists, featuring 750 expert tasks across seven research workflows.

Cursor built Composer 2.5 on Kimi K2 open-source model, ranking 3rd on coding benchmarks and surpassing K2.6. Deep dive into Cursor's data flywheel, product architecture, and pricing.

How AI model benchmarks and evals can build a VC decision framework—using capability overhangs, weakness analysis, and trajectory tracking to identify investment opportunities.

Tsinghua and Zhipu AI release a full-stack web dev benchmark with three difficulty levels. Top models like Gemini 2.5 Pro see scores plummet from 63 to 11.7 on full-stack tasks, exposing AI's real limits.

Exa launches Source Attribution, letting users view prompts, cited sources, and full recipes behind AI-generated content, with one-click iteration support.