AI Voice Clone Audiobook Dubbing: Solo Multi-Cast Pipeline Earning over 10,000 RMB Monthly
Workflow: Every morning, use a Python script to automatically split novel or publication text by chapter, tagging each segment wit
Key Fields
FIELD STAMPS🔧 Workflow
Every morning, use a Python script to automatically split novel or publication text by chapter, tagging each segment with character names and emotion labels. Preprocessing a 300,000-character novel takes about 0.5 to 1 hour. At noon, use the ElevenLabs Creator plan or local XTTS-v2 open-source model to generate 3 to 5 character voice audio files respectively. Batch-calling the speech synthesis API takes about 1.5 to 2 hours to process an entire book. In the afternoon, use Audacity for noise reduction and audio format conversion, manually reviewing and listening to flag stuttering or incorrect pausing segments, and re-generating corrections. This takes about 2 to 3 hours. In the evening, upload the finished audio to platforms like Ximalaya and Van Shang Yousheng, fill in tags and categories, and set a revenue share ratio of about 0.5 hours. On weekends, invest an additional 2 to 3 hours reviewing playback data across categories, retiring single books with under 500 plays, and listing new trending categories.
🛠 Setup Requirements
Hardware requirements: An ordinary computer (16GB RAM with a dedicated graphics card preferred) and stable network. Local deployment of XTTS-v2 requires an NVIDIA graphics card (8GB VRAM minimum). Software proficiency: Master basic Python operations, be able to run the AI-NovelSpeaker-V2 project script by qzw881130 on GitHub, and complete text splitting and batch speech synthesis API calls. Audio processing: Install Audacity for post-processing noise reduction and audio format conversion, with a learning curve of about 1 to 2 weeks to independently complete the entire workflow. Scaled operations: Once familiar, solidify the workflow into batch processing scripts so a single person can simultaneously operate 5 to 10 audiobooks across different categories. Upgrade path: Once production is stable, gradually migrate to the ElevenLabs paid plan to enhance sound quality, starting at $22 per month, bringing sound quality close to professional single-cast levels.
🧰 Toolchain
- 🔧 ElevenLabs
- 🔧 XTTS-v2
- 🔧 Audacity
- 🔧 Van Shang Yousheng Platform
- 🔧 AI-NovelSpeaker-V2
- 🔧 Ximalaya Creator Center
💰 Revenue
Monthly income of about 8,000 to 15,000 RMB, calculated based on platform playback revenue sharing for single audiobooks, with Ximalaya's share ratio at around 50% to 70%. A single book with 100,000 monthly plays generates about 2,000 to 4,000 RMB in revenue. The Van Shang Yousheng platform added new revenue-sharing channels after its public beta in March 2026. One person can simultaneously operate 5 to 10 multi-category audiobooks (urban romance, suspense/mystery, wuxia/xianxia, etc.). Production cost for a 300,000-character novel is about 240 RMB (Van Shang Yousheng charges under 8 RMB per 10,000 characters), and revenue share is collected monthly based on play counts after listing. Long-tail category audiobooks can still generate passive income 3 to 6 months after listing, with total income growing linearly alongside the number of listed categories.
💸 Cost
ElevenLabs Creator subscription is about $22 per month (approx. 160 RMB), supporting the generation of about 100,000 characters of audio monthly. The open-source XTTS-v2 solution incurs only electricity and GPU depreciation costs, with electricity for a 300,000-character novel costing about 20 to 30 RMB. Van Shang Yousheng charges under 8 RMB per 10,000 characters, making production costs for a 300,000-character novel about 240 RMB. Audacity is free to use, and Python and GitHub open-source scripts are zero cost. Running XTTS-v2 on a cloud GPU server costs about 200 to 400 RMB monthly, suitable for creators without a local graphics card.
⏱ Time Investment
About 4 to 6 hours daily, including text preprocessing (0.5 to 1 hour), AI batch generation (1.5 to 2 hours), manual review and correction (2 to 3 hours), and listing management (0.5 hours). An additional 2 to 3 hours spent on weekends reviewing data and selecting products. Adding 2 to 3 new audiobooks per month, while existing inventory only requires monitoring revenue share data. Total monthly work hours for a single person are about 120 to 150 hours to operate 8 to 10 audiobooks, which can be compressed to 3 hours daily after workflow semi-automation.
🚀 Getting Started
Step 1: Choose the zero-cost open-source XTTS-v2 solution, download the AI-NovelSpeaker-V2 project from GitHub for local deployment, and run through the complete workflow of a 50,000 to 100,000-character short audiobook. Step 2: Upload the finished product to the Ximalaya Creator Center or Van Shang Yousheng platform to verify if the revenue share link is working smoothly, and observe first-week playback data. Step 3: Adjust categories and voice selection based on platform data feedback. Increase investment in single books with over 1,000 plays, and decisively eliminate those with under 500 plays. Step 4: Once monthly income stabilizes above 3,000 RMB, consider migrating to the ElevenLabs paid plan to improve sound quality and expand categories, gradually semi-automating the workflow to reduce manual review time.
🔑 Keys to Success
- ✅ Multi-character voice clone quality is the core barrier; a single book requires at least 3 to 5 different timbres to distinguish characters.
- ✅ The manual review step cannot be omitted; AI speech stuttering and phrasing errors require manual tagging and regeneration.
- ✅ Choose audio platforms with traffic-based revenue sharing for listing rather than merely selling finished products to achieve continuous passive income.
- ✅ Operating multiple categories simultaneously disperses the risk of single-category traffic fluctuations; urban romance and suspense mystery are high-playback categories.
- ✅ Text sources must be legally authorized; public domain books and self-copyrighted content are safe choices.
⚠️ 风险
- ⚠️ Copyright audits on audio platforms are tightening; ensure text sources are legally authorized or face the risk of takedowns.
- ⚠️ AI voice cloning involves voiceprint copyright; using celebrity or other people's voices carries legal risks.
- ⚠️ Platform revenue-sharing policies may adjust, leading to unstable production costs and earnings ratios.
- ⚠️ Open-source model sound quality still falls short of professional single-casts, leaving the high-end audiobook market share limited.
- ⚠️ Competition in AI audio content platforms is intensifying, and overcapacity in the second half of 2026 may lead to lower revenue-share unit prices.
📌 Real Cases
- 📌 The Van Shang Yousheng platform officially launched its public beta in March 2026. Developed by the original team behind Lanren Tingshu, it introduced fully automated AI multi-cast audiobook creation features, with production costs under 8 RMB per 10,000 characters and a 300,000-character novel costing about 240 RMB to produce. Individual creators without recording studio experience can list products and monetize via revenue sharing. Hundreds of creators onboarded during the first week of public beta.
- 📌 Zhimai Pai AI reported that individual practitioners using the XTTS-v2 open-source voice clone solution deployed the AI-NovelSpeaker-V2 project locally to produce audiobooks at zero subscription cost. A single person operating multi-category listings simultaneously on platforms like Ximalaya earned about 10,000 to 15,000 RMB monthly. Core operations involved text splitting plus batch speech synthesis plus manual review and correction.
- 📌 FlowPix reported real data on making money with AI dubbing software, where an individual creator used the ElevenLabs Creator plan at $22 monthly to produce audiobooks and listed them on Ximalaya for revenue sharing based on plays. A single month with 100,000 plays yielded around 2,000 to 4,000 RMB, and simultaneously operating 5 books achieved a monthly income of 10,000 RMB.
- 📌 The book2podcast project (GitHub user SPA3K) automatically transforms entire books into three-person podcast shows, launching the 'Three People, One Book' channel on Xiaoyuzhou. It completed full-process automation from documents to dialogue scripts to speech synthesis, validating a compoundable production model from text to multi-cast audio.