96 related articles

Meeting recordings, mixed languages, and background noise causing speech-to-text to drop words or produce gibberish? This article dives deep into ASR hallucination causes and offers practical solutions.

KRAFTON partners with NVIDIA ACE to build PUBG Ally, a next-gen AI teammate with voice recognition, LLM reasoning, and real-time tactical decisions—pioneering the evolution from NPC to CPC.

LangChain releases four major updates: OpenWiki for auto-generating codebase docs, voice agent tutorials, Harbor evaluation integration, and deepagents programmable sub-agents.

WisprGemma is an open-source, browser-local voice input tool built on WebGPU and Transformers.js. One Gemma model handles speech recognition and text polish — your voice never leaves your device.

OpenAI announces GPT-5.6 Sol Ultra coming to Codex and its most powerful real-time voice model yet, GPT Realtime 2.1; Tencent's Toast lands on iOS; Anthropic finds brain-like structures in Claude.

MosiAI open-sources MOSS-Transcribe-Diarize-0.9B: unified speech transcription and speaker diarization, 128K context for 90-min audio, hotword boosting, SGLang Day-0 support, edge-deployable.

Tencent Hunyuan HY3 official version is open-sourced under Apache 2.0, priced as low as 1 yuan per million input tokens, with major gains in agents, reasoning, coding, and long context. On the same day, Meituan open-sourced its trillion-parameter LongCat 2.0.

Want to break into LLM development but not sure where to start? This guide breaks the core skills into four progressive layers — from basic knowledge to RAG, fine-tuning, Agents, and multimodal — so you can align with real enterprise needs and land the job.

Agent Draw is an AI whiteboard built on TLDraw that lets you speak or type to have an AI agent draw flowcharts and diagrams in real time. A deep dive into its tech, design, and use cases.

Xiaomi XiaoAI 10.1-inch Smart Control Panel features AI LLM Q&A, WeChat calling, and whole-home Mi IoT control. Priced at 839 yuan, ~679 yuan after national subsidy. An in-depth review of AI capabilities, screen experience, and smart home integration.

Google Gemini's latest Drops update: real-time voice-to-image generation lowers creative barriers, plus new small business AI tools. A deep dive into multimodal AI strategy.

A deep dive into Claude-real-video: how keyframe extraction, image captioning, and ASR convert video into structured LLM-readable input for model-agnostic video understanding.

How does AI Agent automate TV subtitle workflows end-to-end? This demo covers audio extraction, ASR, timestamp sync, and line optimization — GLM5 processes a 26-min video in just 10 minutes.

AI customer service is a core tool for digital transformation. This guide covers its value, use cases, and implementation logic, including efficiency gains, cost reduction, and data-driven optimization.

Build a content creator competitor radar with Vibe Coding in 3 hours: auto-fetch videos, local FunASR transcription, structured viral video analysis, and Feishu dashboard sync.

A roundup of 12 trending open-source AI agent projects on GitHub, covering video generation, agent frameworks, skill packs, code engines, security scanning, and voice processing.

XiaoWu is a fully local AI voice input method powered by on-device LLMs for accurate offline speech recognition, smart punctuation, and minimalist interaction — no internet required.

A systematic breakdown of the three core AI Agent modules (Control, Perception, Action), with deep analysis of AutoGPT, BabyAGI, HuggingGPT, LlamaIndex architectures and Chain-of-Thought reasoning.

CLI-WeChat-Bridge is an open-source tool that bridges AI CLI tools like Codex, Claude Code, and Open Code to WeChat, supporting voice input, file transfer, and multi-CLI parallel switching.

Deep dive into OpenLLMVTuber, a 10K-star open-source AI virtual character framework integrating ASR, LLM, TTS, and Live2D with voice interruption, visual perception, and modular architecture.