1136 related articles

Speech To Markdown is a free macOS/iOS app that converts voice to structured Markdown notes using local LLMs. Fully offline, no API keys needed, with global hotkey dictation.

Peekinduck uses two AI voice agents—Demo Duck and Guide Goose—sharing customer memory to unify pre-sales demos, onboarding, and post-sales support for B2B SaaS teams.

Echologue is a privacy-first AI voice journal that processes data locally with end-to-end encryption. This analysis examines its product design, technical architecture, and indie developer philosophy.

Microsoft's open-source voice AI project VibeVoice rapidly gained 50K+ GitHub Stars, focusing on emotional expression and natural prosody. A deep dive into its technology, strategy, and applications.

GitHub Trending July 29: Microsoft's VibeVoice leads voice AI open-source wave, MoonshotAI's FlashKDA CUDA kernel surges 25%, and open-source alternatives rise.

A deep comparison of Pipecat Flows and Vapi Squad for voice AI agent architecture — covering latency, accuracy, multi-agent handoffs, and when to use each.

OpenAI's upgraded voice assistant can speak dialects, do real-time simultaneous interpretation, teach English, and even get flustered. Here's what changed.

Build a production AI voice agent with Claude Code + Telnyx single-stack — no code needed, live phone number in 5 minutes. Covers 5 business scenarios including appointment booking, lead qualification, and support triage.

OpenAI's GPT-Live full-duplex voice model enables natural simultaneous conversation with a reasoning delegation architecture pairing real-time dialogue with GPT-5.5 deep reasoning. Now live for 150M users.
Self-Hosted Voice AI Assistant: Bringi…
Explore a self-hosted voice AI assistant built for Asterisk and FreePBX: keep data on-premises, integrate with existing PBX, replace legacy IVR, and deploy local voice intelligence affordably.

OpenAI's GPT-Live voice model family brings full-duplex interaction, task delegation, GPT-5-level intelligence, real-time translation, and image understanding to voice AI.

OpenAI's new voice model delivers near-zero-latency bidirectional conversation with real-time multilingual simultaneous interpretation across Cantonese, Spanish, and English.

An open-source project adds 43 game voice packs to Claude Code, covering 500+ lines from StarCraft, Red Alert, and Portal, triggered on task completion, authorization prompts, and errors.

OpenAI's GPT-Live voice model tackles the cocktail party problem through Background Robustness — enabling precise speaker focus in noisy, multi-person environments with natural multi-turn dialogue.

AI zero-shot voice cloning needs just 3 seconds of audio to impersonate anyone. Learn the 3 tiers of voice fraud evolution and practical defenses like family code words and video verification.

Can AI be conscious? Exploring GPT-4o voice model technology and philosophy — the Problem of Other Minds, Descartes, and what makes consciousness real.
Voice Cloned in Three Seconds: Why AI …
Just 3 seconds of audio lets AI clone your voice for fraud. Learn the tech behind AI voice scams, why defenses fail, and practical tips like code words to protect yourself.

OpenAI's GPT Live One powers a new ChatGPT voice mode with full duplex conversation, real-time web reasoning, and live translation. Here's a deep dive into all three breakthroughs.

OpenAI launches ChatGPT Voice powered by GPT Live One, featuring full-duplex real-time conversation, multi-task reasoning, and live translation. A deep dive into its capabilities and what it means for the future of voice AI.

OpenAI's ChatGPT Voice with GPT-Live 1 achieves true full-duplex voice conversation — supporting interruptions, real-time reasoning, web search, and live translation.