1548 related articles

Complete guide to OpenCode AI coding tool: two installation methods, model configuration, Agent types, custom commands, MCP extensions, Agent SQL, with practical examples.

How to choose local vision language models on M4 Pro 64GB? Compare Qwen2.5-VL, Llama 3.2 Vision, and more, with tool recommendations for Ollama, LM Studio, and MLX.

Explore how multi-agent simulations let AI agents autonomously build civilizations. From Stanford's AI Town to civilization-scale simulations, discover memory mechanisms, emergent behavior, and implications for social science and AI safety.

Ksyon is a fully local AI robot project using the lightweight vision-language model Moondream for environmental perception, combined with lifelike head movements and a sarcastic personality for engaging human-robot interaction.

Deep dive into Google DeepMind's Gemini Robotics 2 model, analyzing its VLA architecture breakthroughs, generalization gains, dexterous manipulation advances, and the future of embodied AI.

Deep dive into the enterprise voice agent Sandwich Architecture (STT+LLM+TTS): design principles, architecture comparison, 0.3s latency optimization, and real-world pitfalls in barge-in and tool calling.

Google's SKILL.state method replaces full conversation history with structured state, cutting Agent token usage from 1.1M to 65K (94% reduction) in 100-step benchmarks while maintaining accuracy.

Exploring the practical value of synthetic tactile datasets for robotic grasping. Analyzing the data scarcity challenge, stick-slip physics modeling, 500K-row simulation data, and Sim-to-Real gap solutions.

Explore seven waves of retrieval system evolution: from BM25 lexical search to generative agentic retrieval, revealing how these technologies stack and work together in production.

A systematic overview of the evolution from AI, machine learning, deep learning, and Transformer to LLMs, covering generative AI principles, model selection, and the future of AI Agents.

Cursor's CEO reveals OpenAI models carry only ~5% of user traffic, with 95%+ going to Anthropic Claude and competitors. A deep dive into AI coding tool model preferences and industry implications.

Why do we always underestimate how fast AI models evolve? From linear thinking bias to exponential growth realities, and what it means for developers, investors, and users.

Reddit user reports ChatGPT voice mode cloning their voice. Analysis of OpenAI's disclosed unauthorized voice generation risk, technical causes, and safety guardrail limitations.

OpenAI will cut off model access to Cursor on November 12 following SpaceX's acquisition. Learn the reasons, developer impact, and multi-model coping strategies.

Tencent Hunyuan's WorldClaw generates explorable, editable 3D worlds from text. Deep dive into its multi-model Agent architecture, AI-native game engines, AI pharma funding, and data strategy shifts.

Explore how AI identifies counterfeit cosmetics through computer vision packaging inspection, spectral analysis, and multimodal detection, plus real-world challenges and blockchain-integrated anti-counterfeiting ecosystems.

Real-world testing of Gemini Flash vs Pro across three projects: racing game, subscription app, and luxury website. Flash is 3x faster and cheaper, but Pro remains essential for production accuracy.

Alpamayo 2 Super is an open-source reasoning model for autonomous driving with commercial deployment support. Explore its reasoning capabilities, robotics backbone architecture, and OpenMDW-1.1 license.

Hands-on comparison of DeepSeek V4 Pro vs. OpenAI Codex recreating Don't Starve from scratch. DeepSeek excels at planning but gameplay breaks down; Codex delivers complete features. A deep dive into how model capability and engineering environment interact.

An open-source game behavior capture tool that synchronously records gameplay video and keyboard/mouse input with frame-level alignment, providing structured datasets for imitation learning and world model research.