895 related articles

Claude Sonnet 5 review: 63.2% SWE-bench, near Opus 4.8 performance, but new tokenizer hides real costs. Ranks 13th on CursorBench. Most tasks: stick with Opus 4.8.

Anthropic's Claude Sonnet 5 claims near-OPUS 4.8 performance at lower cost. Real-world tests reveal hidden tokenizer costs, weak creative output, and only 13th place on Cursor rankings.

Claude Sonnet 4 returns to Cursor, topping CursorBench while carrying the highest cost per task. A deep dive into performance vs. cost trade-offs for AI coding tools.

Claude Sonnet 5 may launch this week with up to 2M token context; GPT-4.6 Pro arrives with stunning code generation; mysterious Opus 6 exists internally. Full breakdown of this week's frontier AI model updates.

mini-SWE-agent's GPT-5 series evaluation on SWE-bench shows GPT-5 matches Claude Sonnet 4, while GPT-5-mini loses only ~5 points at less than 1/5 the cost.

Deep dive into Claude Sonnet 4: replicate Lovable with two prompts, generate McKinsey-grade reports, build 2D games, and explore the AI Agent building block economy.

Simon Willison shares how Claude Sonnet 4 (Fable) autonomously invented PyObjC screenshots, built a CORS server, and penetrated Shadow DOM to debug a CSS bug — revealing both tool-making power and security risks.
Product ReviewsHands-on comparison of GPT-5.1 vs Claude Sonnet 4.5 across long-form writing, classical poetry, front-end coding, and UI reproduction to help you pick the right AI model.
Product ReviewsHands-on comparison of GPT 5.1 Thinking vs Claude Sonnet 4.5 across story writing, math reasoning, emotional support, instruction following, and coding to help you choose the right AI model.
Product ReviewsHands-on testing of Claude Haiku 4.5's coding ability, comparing it with Sonnet 4.5 and Opus 4.1 across weather cards, physics simulation, and 3D rendering tasks.
Product ReviewsIn-depth comparison of Claude Sonnet 4.5 vs GPT-5 Codex recreating classic game Terep 2's soft-body physics in C++, covering terrain rendering, physics engines, and collision detection.
Tech FrontiersIn-depth analysis of Anthropic's Claude Sonnet 4.6: agentic tool use, computer control, and office task upgrades. Multiple benchmarks surpass Opus 4.6, redefining mid-tier AI capabilities.
TutorialsA detailed walkthrough of using Xcode MCP with Claude Sonnet 4 to build a Mac local TTS app entirely through AI-generated code, covering setup, design guidelines, and results.
Product ReviewsIn-depth review of Claude Sonnet 4.6: flagship AI performance at 1/10th the price. 1M context window, 72.5% OS World score, $3/M input tokens.
Tech FrontiersAnthropic launches Claude 4 Opus and Claude 4 Sonnet. Claude Code goes GA with IDE integration and SDK. MCP protocol connects directly to API. Full breakdown of coding and agent upgrades.
Product ReviewsFirst hands-on review of Claude 4 series: multi-dimensional comparison of Opus 4 and Sonnet 4 across coding, document analysis, reasoning, and AI Agents, with benchmarks against GPT-4o and Gemini 2.5 Pro.
Product ReviewsIn-depth testing of Zhipu AI's open-source GLM-4.7 coding abilities across SVG animation, 3D game dev, iOS native apps, and browser automation, compared against Claude Sonnet 4.5 and DeepSeek V3.2.
Product ReviewsHands-on test of Claude Sonnet 4.5's code execution and file creation features, showing how one prompt generates Excel, Word, and PPT documents with four optimization strategies and three complete examples.
Product ReviewsHands-on testing Kimi K2 Thinking in Claude Code across text creation, programming, agent building, and full-stack apps. Comparing it to Claude Sonnet 4.5 and DeepSeek as a cost-effective AI coding alternative.
Product ReviewsIn-depth review of Claude Haiku 4.5: 73.3% on SWE-bench rivaling Sonnet 4, input at just $1/million tokens. Covers code generation, agentic coding, SVG tests, and Sonnet+Haiku collaboration strategies.