AI Voice Workshop + AI Audio Workshop: One Person Can Complete the Entire Production Pipeline for Radio Drama-Quality Audiobooks

AI Voice Workshop + AI Audio Workshop automate the entire audiobook production pipeline from voice acting to post-production.
AI Voice Workshop and AI Audio Workshop Studio Edition form a complete production system covering the entire audiobook creation workflow. AI Voice Workshop solves the 'who reads' problem with text analysis, character recognition, voice matching, emotion control, and batch synthesis. AI Audio Workshop solves 'how to make it sound good' through an AI Agent-based post-production director system that automatically matches BGM, ambient sounds, and effects to achieve radio drama-level quality. Together, they enable one person to accomplish what previously required an entire team, dramatically lowering the barrier to audiobook production.
From Voice Acting to Post-Production: AI Agents Reshape the Audiobook Production Pipeline
After six months, AI Voice Workshop (AI声工坊) and AI Audio Workshop (AI音工坊) Studio Edition have officially released their updates. This isn't a simple TTS dubbing tool—it's a complete production system covering the entire audiobook creation workflow: from text analysis, character recognition, and voice matching to emotion control, batch synthesis, post-production mixing, and everything in between, all within a single workstation.
TTS (Text-to-Speech) technology has undergone a long evolution from early concatenative synthesis and parametric synthesis to today's deep learning-based synthesis. Early TTS systems sounded mechanical and lacked emotion, while modern neural network-based TTS models (such as Tacotron, VITS, GPT-SoVITS, etc.) can generate natural speech that closely resembles real human voices, with support for emotion control and voice cloning. The audiobook domain demands far more from TTS than ordinary voice assistants—it requires stable long-duration output, seamless multi-character switching, and naturally flowing emotional dynamics. This is precisely why TTS tools specifically designed for audiobook scenarios have unique and irreplaceable value.
For audiobook professionals and enthusiasts, this means one person can now accomplish the workload that previously required an entire team. Traditional audiobook production is a highly labor-intensive industry—for a 300,000-character novel, professional voice actors typically need 40-60 hours of dry recording, and the post-production team requires equal or even more time for noise reduction, editing, scoring, and mixing. A complete audiobook production team usually includes planning editors, voice actors (hosts), post-production audio engineers, quality review personnel, and more, with production cycles ranging from weeks to months and costs from thousands to tens of thousands of yuan. This also explains why a large number of web novels on the market cannot be turned into audiobooks—production capacity bottlenecks severely constrain content supply. Now, the maturation of the AI toolchain is breaking through this bottleneck. More critically, the learning curve is minimal—virtually zero barrier to entry.
AI Voice Workshop: Solving the "Who Reads" Problem
Full-Process One-Stop Audiobook Workstation
The positioning of AI Voice Workshop Studio Edition is crystal clear: it's a full-process workstation designed for audiobook production. Specifically, it covers the following core stages:
- Text analysis and chapter splitting: After importing a novel TXT file, the program automatically splits it by chapter headings
- Character recognition and voice matching: Automatically identifies characters in the text and matches each one with an appropriate AI voice
- Emotion control and batch synthesis: Supports both natural language control and reference voice emotion modes
- Post-processing and project export: Built-in post-processing plugins like VTS, supports script export

Actual Audiobook Production Workflow
The entire production process can be summarized in six steps:
Step 1: Prepare the original novel text as a TXT file, ensuring chapter titles are properly formatted. Step 2: Create a new project card in the audiobook management interface. Step 3: Import the text for automatic chapter splitting—the program converts the raw text into parseable structured text.
Step 4: Automatically match voices to characters. The system intelligently analyzes character traits and matches appropriate voices. Since the narrator role is particularly important, users are advised to manually assign it. Step 5: Configure the emotion control mode—select "natural language control" for character dialogue to achieve rich emotional expression, and select "reference voice emotion mode" for narration to maintain stable output. Step 6: Click generate, wait for batch TTS conversion to complete, then export the audio.
It's worth explaining the technical differences between these two emotion control modes in depth. The "natural language control" mode guides the model to generate speech with corresponding emotions through text descriptions (such as "said sadly" or "shouted angrily"), relying on the large language model's deep understanding of emotional semantics. This is ideal for the rich and varied emotional expressions in character dialogue. The "reference voice emotion mode" provides a reference audio clip as a style anchor, letting the model mimic that audio's speed, intonation, and emotional characteristics during synthesis—ideal for narration scenes that need to maintain long-term consistency. Each mode has its pros and cons: the former is flexible but may produce unstable emotional fluctuations, while the latter delivers stable output but its expressiveness is limited by the quality of the reference sample.
One detail worth noting: precise adjustment of emotion modes is the key to producing premium audiobooks—this requires users to continuously explore and accumulate experience through actual use.
Rich Auxiliary Features
Beyond the core voice-over workflow, AI Voice Workshop also provides several practical features: script export (with customizable styles), remote recording assistance (supporting CV remote collaboration with user authorization mechanisms to prevent unauthorized operations), and more.

AI Audio Workshop: Solving the "How to Make It Sound Good" Problem
AI Post-Production Director — The Core Highlight of Agent Architecture
If AI Voice Workshop solves "who reads," then AI Audio Workshop solves "how to make it sound good." This is a one-stop workstation specifically designed for audiobook post-production, covering audio transcription, intelligent analysis, sound effect matching, emotion rendering, multi-track mixing, and batch export.
The biggest highlight of AI Audio Workshop is its built-in AI Post-Production Director system, designed with an Agent architecture:
- Chief Director Agent: Analyzes the entire book's genre and emotional tone, formulating the overall post-production plan
- Executive Director Agent: Automatically matches BGM, ambient sounds, and various sound effects for specific scenes
AI Agent is one of the core paradigms in current AI application development. Unlike traditional single-call LLM usage, Agent architecture allows AI systems to autonomously plan tasks, invoke tools, execute iteratively, and adjust strategies based on feedback. In AI Audio Workshop, the Chief Director Agent and Executive Director Agent form a hierarchical multi-Agent collaboration system: the Chief Director handles global decisions—similar to a human director reading through a script and setting the overall style direction, determining the work's emotional tone, pacing, and sound effect style; the Executive Director handles scene-by-scene implementation—similar to a sound designer selecting and arranging effects for specific passages, handling audio events at every emotional turning point. The core advantage of this architecture is its ability to handle complex contextual dependencies, maintaining logical coherence in sound effect arrangements across scenes, rather than simply processing each text segment in isolation.
This means even without years of post-production experience, the system can automatically understand plot progression and emotional changes, helping you arrange radio drama-level sound effects.

Five-Category Sound Effect Library System
AI Audio Workshop supports five independent sound effect library categories:
| Effect Type | Usage Description |
|---|---|
| BGM | Background music |
| SFX Short Effects | Instantaneous effects like actions, impacts |
| Ambient Effects | Scene atmosphere creation |
| Transition Effects | Chapter/scene transitions |
| Range Effects | Sustained effects for specific areas |
This five-category classification actually corresponds to the standard audio layering logic in professional radio drama and film/TV post-production. BGM establishes emotional tone and narrative pacing, SFX short effects provide action feedback and rhythmic accents (like door sounds, footsteps, weapon clashes), ambient effects build spatial awareness and immersion (like rain, marketplace bustle, forest birdsong), transition effects handle narrative rhythm connections to avoid abrupt jump cuts, and range effects provide a sustained atmosphere layer for specific passages (like the sound of distant fighting on a battlefield). In professional audio post-production, the proper layering and volume balancing of these elements—i.e., mixing—is the key to determining audio quality. In traditional production, sound designers need to individually select, trim, and time-align materials from massive asset libraries, and the AI Agent's intervention automates this tedious process.
Users can build or import custom sound effect libraries through the built-in library management program. The quality of your sound effect library directly determines the quality of the final product, which also means the system's ceiling depends on the user's own accumulation of resources.
Complete Audiobook Post-Production in Three Steps
The post-production workflow has been compressed to three steps: select sound effect libraries, load dry audio files, and start the pipeline task. The program supports batch concurrent processing with multiple tasks running simultaneously. After processing is complete, users can view the allocation of each audio track through the built-in console—all AI-deployed audio events are displayed in a structured track format for easy secondary adjustments.
Real Case Results: Full Coverage Across Multiple Genres
The video demonstrates several complete cases covering mainstream audiobook genres:
- Fantasy/Cultivation: Intense atmosphere rendering for blood energy bursts and battle scenes
- Youth/Campus: Delicate emotional expression for graduation farewells, complemented by warm ambient sounds
- Children's Stories: The adventure story of Firefly Ah Xing, with lively tone and upbeat pacing
- Suspense/Crime: Oppressive atmosphere in interrogation room scenes, with high-tension character dialogue
- Historical/Cultural: The story of Han Xin, with natural and fluid narration style

From the actual results, emotion control and sound effect matching across different genres have achieved quite impressive quality. The interrogation room atmosphere in suspense works and the emotional rendering of graduation scenes in youth/campus works are particularly noteworthy, approaching the level of professional radio dramas. These cases also validate the Agent architecture's generalization ability across different narrative styles—the same system can adaptively match entirely different post-production styles for different genres without any manual mode switching.
Versions and Costs: Nearly Negligible Usage Fees
Studio Edition vs. Personal Edition Comparison
Both programs offer Studio and Personal editions:
- Studio Edition: Built-in LLM channels, full feature set, suitable for studios and professional creators, with only minimal LLM API call fees
- Personal Edition: No built-in LLM, requires users to connect their own model provider APIs (compatible with OpenAI-style interfaces), suitable for hobbyists with some technical ability
The Studio Edition's fee structure is designed to cover server and LLM call costs, making it nearly negligible for studios or individual creators. While the Personal Edition lacks some commercial auxiliary features, careful adjustment can still produce excellent results. The Personal Edition's compatibility with OpenAI-style interfaces means users can connect to virtually all mainstream LLM services including OpenAI, Claude, and domestic Chinese providers (such as Zhipu, Tongyi, DeepSeek, etc.), offering extremely high flexibility.
Voice Copyright Reminder
The developers specifically emphasize: the built-in test voices are sourced from various novel audiobooks, and the developers themselves hold no copyright over them—commercial use of these voices is not supported. For commercial purposes, be sure to import your own properly licensed voice resources.
Summary: One Person, Replacing an Entire Audiobook Production Team
The combination of AI Voice Workshop + AI Audio Workshop truly achieves full-process automation of audiobook production from voice acting to post-production. This isn't a toy-level demo—it's a production-ready toolchain. The introduction of AI Agent architecture transforms post-production from "requires years of experience" to "done in three steps," representing a revolutionary improvement in production efficiency for the entire audiobook industry.
Of course, no matter how powerful the tools are, they still need human guidance. Fine-tuning emotion modes, accumulating sound effect libraries, selecting character voices—these details determine whether the final product is merely "listenable" or truly "enjoyable." But at the very least, the barrier has been dramatically lowered, allowing creators to devote more energy to the content itself rather than tedious technical processes. From a broader perspective, the emergence of such tools may give rise to a new generation of individual audiobook creators—people who don't need professional recording studios or teams, just passion for content and an ear for sound, to produce high-quality audio works. The supply side of audiobook content may be on the verge of a structural transformation.
Related articles
Product ReviewsThe Programmer's Desk Setup Guide: Building a Workspace That Feels Like Home
Discover how programmers build productive, comfortable workspaces. From multi-monitor setups to ergonomic design, explore the desk philosophy that drives focus and flow.
Product ReviewsQoder vs Cursor Real-World Comparison: Which $20/Month AI IDE Is Better?
Hands-on comparison of Qoder vs Cursor AI IDEs: Agent autonomy, human interaction count, and architecture decisions. Qoder needed only 2 interactions vs Cursor's 8.
Product ReviewsCursor Cloud Agent Demo: Eliminating Bottlenecks Across the Entire Software Development Lifecycle
Deep analysis of Cursor's Cloud Agent demo showing how cloud VMs, automated test artifacts, and a full-chain control plane systematically eliminate human bottlenecks across the software development lifecycle.