59 related articles

Gemini Nano's on-device AI model currently has limited language support, with no official timeline for RTL languages like Hebrew and Arabic. This article explores the technical bottlenecks, commercial priorities, and future outlook.

Conversational AI shines in the lab but fails in real conversations. This article analyzes voice assistants' core weaknesses—model architecture or overly "clean" data? Covering ASR, VAD, and end-to-end systems engineering.

RoughCut is a fully automated AI editing tool generated by Codex, supporting talking-head, unboxing, and commentary modes with a semi-automated publishing system.

A deep dive into the physical AI companion device "Amis": combining personalized character design, emotional dialogue, and daily assistant features to explore how AI hardware fills modern emotional needs.

OpenAI unveils the GPT-Live voice model family, with full-duplex interaction enabling AI to listen and speak simultaneously and delegate complex reasoning to GPT-5.5. GPQA benchmark jumps from 45% to 80%.

OpenAI launches GPT Live, a voice AI model family supporting full-duplex real-time conversation, deep task delegation, multimodal interaction, and instant translation, with reasoning near GPT-5 level.

An in-depth hands-on test of GPT's real-time voice feature, covering Cantonese and Sichuanese dialect recognition, emotional tone switching, complex role-play, and cross-voice contextual memory—objectively presenting the true level and remaining gaps of AI voice interaction.

OpenAI's GPT Live full-duplex voice model, Grok 4.5 coding model with Cursor, and ByteDance's Seedream 5.0 Pro image generation launched together. A deep dive into three AI releases.

MosiAI open-sources MOSS-Transcribe-Diarize-0.9B: unified speech transcription and speaker diarization, 128K context for 90-min audio, hotword boosting, SGLang Day-0 support, edge-deployable.

Zhipu's GLM-5.2 is now fully available to Coding Plan users with long-context support and dual reasoning tiers. Meanwhile, Anthropic faces U.S. export controls and OpenAI is under multi-state investigation.

The gap between AI power users and everyone else isn't about prompt tricks — it's about understanding LLMs, multimodal models, workflows, and agents. Build your complete AI mental model here.

AI dream interpretation and personality analysis are trending on social media, but can AI really understand you? This article unpacks the technical limits and hidden risks—from the Barnum Effect to LLMs.

A clear, in-depth guide to how AI Agents work: the paradigm shift from traditional programs, the perception-decision-action loop, and the four pillars—LLMs, tool calling, memory, and RAG.

An exclusive look at the AI Engineer Summit dress rehearsals, decoding the paradigm shift from research to production. A deep dive into AI Engineer challenges, RAG, agent systems, and AI engineering as a distinct discipline.

Tencent Hunyuan and Tsinghua jointly release DiscoBench, the first benchmark evaluating search agents' dynamic ambiguity clarification. Covering 463 ambiguity instances across 11 domains, it reveals real weaknesses of mainstream LLMs.

A political news story about British satirical candidate 'Count Binface' sparked debate in the tech community: why does AI struggle to understand sarcasm, contrast humor, and cultural context? An in-depth analysis of LLM limitations.

Researchers found a hidden authentication backdoor in multiple Tenda router firmware versions, letting attackers bypass passwords to gain admin access. Learn the technical principles, impact, and protection tips.

A tweet about an AI character "attending" a July 4th party reveals how virtual and real-world social boundaries are dissolving. A deep dive into AI socialization, anthropomorphism, and attention scarcity.

How does AI Agent automate TV subtitle workflows end-to-end? This demo covers audio extraction, ASR, timestamp sync, and line optimization — GLM5 processes a 26-min video in just 10 minutes.

AI-powered smart home devices are reshaping household security. This article analyzes sociotechnical threat models—from prompt injection and data breaches to domestic abuse—and explores responsible AI home design principles.