99 related articles
TutorialsComplete guide to downloading, installing, and configuring ByteDance's Trae AI coding tool. Free built-in LLMs like DeepSeek for AI-assisted Python programming.
TutorialsCan't use Claude Code in China? This hands-on guide tests Cursor and Windsurf as alternatives for accessing the Opus 4 model, both supporting Alipay payment with no overseas phone number required.
TutorialsA beginner's guide to physical AI robot development covering the complete tech stack from GPU hardware, Linux, Python, deep learning, computer vision to ROS2, with a clear learning roadmap.
Product ReviewsCompare Trae CN, Cursor, VS Code plugins, and Claude Code across cost, ease of use, and flexibility. Get free and low-cost AI coding setup recommendations for students.
Product ReviewsRovo Agent is Atlassian's AI coding CLI tool offering 20M free Claude 4 Sonnet tokens daily, ranked #1 on SWE-bench. Learn about its adaptive memory system, installation, and hands-on experience.
Product ReviewsDeep dive into ByteDance's open-source Trae Agent — the free AI coding CLI tool topping SWE-bench. Covers installation, features, and comparisons with Claude Code and Gemini CLI.
Product ReviewsHands-on review of Alibaba's open-source Qwen Code CLI command-line coding tool, covering installation, API configuration, development experience, and comparison with Claude Code. Powered by a 480B parameter MoE model.
Tech FrontiersAnthropic launches Claude Code Web with browser and mobile coding, plus Haiku 4.5 at 1/3 the cost of flagship models. Parallel tasks, multi-model collaboration, and seamless cloud-to-local workflows.
Product ReviewsIn-depth review of Claude Haiku 4.5: 73.3% on SWE-bench rivaling Sonnet 4, input at just $1/million tokens. Covers code generation, agentic coding, SVG tests, and Sonnet+Haiku collaboration strategies.
Tech FrontiersAnthropic releases Claude Haiku 4.5, a distilled version of Sonnet 4.5 with near-flagship coding performance, double the speed, and one-third the cost. Scores 73.3 on SWE-Bench, ideal for developers seeking cost-efficiency.
Tech FrontiersSWE-bench opens evaluation environments, task sets, trajectories, and training recipes, dramatically lowering the barrier to AI coding agent development.
Tech FrontiersSWE-agent Multimodal officially released with image viewing and web browser debugging capabilities for automated frontend visual bug detection and fixes, plus the new SWE-bench Multimodal benchmark.
Tech FrontiersSWE-bench launches its official blog for in-depth content on AI coding evaluation, AI Agents, and toolchains—signaling a new phase of maturity and standardization in AI programming benchmarks.
Tech FrontiersQwen team leads open-source models on SWE-bench, demonstrating strong software engineering capabilities. This article analyzes SWE-bench standards, Qwen's progress, and the value of open-source AI coding tools.
Deep DivesAnthropic's Advisor Strategy lets Sonnet execute tasks while Opus serves as advisor, cutting costs 12% while boosting SWE-Bench by 2.7 points. A new multi-model AI Agent paradigm explained.
TutorialsA deep dive into the MLflow open-source AI engineering platform, covering experiment tracking, LLM evaluation, model deployment, and monitoring to help teams efficiently manage the ML lifecycle.
Product ReviewsComprehensive comparison of 80+ AI coding agent tools, with SWE-Bench benchmark rankings covering Devin, Cursor, Claude Code, GitHub Copilot and more, plus pricing analysis to help developers choose.
Product ReviewsDeep dive into the crafta-bench open-source project, a benchmark tool designed for Cursor Background Agents. Explore AI coding Agent evaluation dimensions, industry trends, and practical implications.