VibeVoiceFusion is a full-stack, multi-speaker voice generation web system featuring LoRA fine-tuning, batch generation, and VRAM optimization. Based on Microsoft's VibeVoice (AR + diffusion architecture)
Python
490
292 commits
updated Oct 5, 2026
A Complete Web Application for Multi-Speaker Voice Generation
Built on Microsoft's VibeVoice Model
Features • Demo Samples • Get Started • Documentation • Community • Contributing
VibeVoiceFusion is a web application for multi-speaker speech generation with voice cloning and speech recognition, built on Microsoft's VibeVoice models. It brings text-to-speech, file transcription and live microphone transcription together in one bilingual (English/Chinese) web interface. It also covers LoRA fine-tuning and an OpenAI-compatible API, and it runs on consumer GPUs with 10GB+ of VRAM.
/v1/audio/speech, /v1/audio/transcriptions and the Realtime WebSocket /v1/realtime, so existing OpenAI SDKs work as-isListen to voice generation samples created with VibeVoiceFusion. Click the links below to download and play:
🎧 Pandora's Box Story (BFloat16 Model)
Generated with bfloat16 precision model - Full quality, 14GB VRAM
🎧 Pandora's Box Story (Float8 Model)
Generated with float8 quantization - Optimized for 7GB VRAM with comparable quality
🎭 东邪西毒 - 西游版 (Journey to the West Version)
Multi-speaker dialog with distinct voice characteristics for each character
Build docker image
# Clone the repository
git clone https://github.com/zhao-kun/vibevoicefusion.git
cd vibevoicefusion
# Build and the docker image
docker compose build vibevoice
After build successfully, run command:
docker run -d \
--name vibevoicefusion \
--gpus all \
-p 9527:9527 \
-v $(pwd)/workspace:/workspace/zhao-kun/vibevoice/workspace \
zhaokundev/vibevoicefusion:latest
Access the application at http://localhost:9527
The Docker image is available on Docker Hub, and you can launch VibeVoiceFusion using the following command.
docker pull zhaokundev/vibevoicefusion
docker run -d \
--name vibevoicefusion \
--gpus all \
-p 9527:9527 \
-v $(pwd)/workspace:/workspace/zhao-kun/vibevoice/workspace \
zhaokundev/vibevoicefusion:latest
Build Time: 18-28 minutes | Image Size: ~12-15GB
1. Install Backend Dependencies
# Clone the repository
git clone https://github.com/zhao-kun/vibevoice.git
cd vibevoice
# Install Python package
pip install -e .
2. Download Pre-trained Model
Download from HuggingFace (choose one):
Place files in ./models/vibevoice/
3. Install Frontend Dependencies (for development)
cd frontend
npm install
4. Build Frontend (for production)
cd frontend
npm run build
cp -r out/* ../backend/dist/
Production Mode (single server):
# Start backend server (serves both API and frontend)
python backend/run.py
# Access at http://localhost:9527
Development Mode (separate servers):
# Terminal 1: Start backend API
python backend/run.py # http://localhost:9527
# Terminal 2: Start frontend dev server
cd frontend
npm run dev # http://localhost:3000
For quick testing without setting up projects:
Speaker 1: Hello) or narration (plain text)Key Points:
This guide walks you through the complete process of creating multi-speaker voice generation from start to finish.
Start by creating a new project or selecting an existing one. Projects help organize your voice generation work with metadata and descriptions.
Create and manage projects from the home page
Actions:
The project will be automatically selected and you'll be navigated to the Speaker Role page.
Upload reference voice samples for each speaker. The system supports various audio formats (WAV, MP3, M4A, FLAC, WebM).
Upload and manage voice samples for each speaker
Actions:
Tips:
Create a dialog session and write the multi-speaker conversation. The dialog editor supports drag-and-drop reordering and real-time preview.
Multi-speaker dialog editor with visual and text modes
Actions:
Dialog Format (Text Mode):
Speaker 1: Welcome to our podcast!
Speaker 2: Thanks for having me. It's great to be here.
Speaker 1: Let's dive into today's topic.
Narration Mode:
For single-speaker content like audiobooks, articles, or podcasts, use Narration Mode:
Speaker N: prefixesThis is the first paragraph of your narration.
This is the second paragraph. No speaker formatting needed.
The narrator voice you selected will read all the text.
Features:
Configure generation parameters and start the voice synthesis process. Monitor real-time progress and manage generation history.
Generation interface with parameters, live progress, and history
Actions:
float8_e4m3fn (recommended): 7GB VRAM, faster loadingbfloat16: 14GB VRAM, full precisionReal-Time Monitoring:
Generation interface with parameters, live progress, and history
Generation History:
For CLI-based generation without the web UI:
python demo/local_file_inference.py \
--model_file ./models/vibevoice/vibevoice7b_float8_e4m3fn.safetensors \
--txt_path demo/text_examples/1p_pandora_box.txt \
--speaker_names zh-007 \
--output_dir ./outputs \
--dtype float8_e4m3fn \
--cfg_scale 1.3 \
--seed 42
CLI Arguments:
--model_file: Path to model .safetensors file--config: Path to config.json (optional)--txt_path: Input text file with speaker-labeled dialog--speaker_names: Speaker name(s) for voice file mapping--output_dir: Output directory for generated audio--device: cuda, mps, or cpu (auto-detected)--dtype: float8_e4m3fn or bfloat16--cfg_scale: Classifier-Free Guidance scale (default: 1.3)--seed: Random seed for reproducibilitySee Demo Model Tools for all command-line tools: TTS, Realtime 0.5B, ASR, and model/voice-preset conversion.
Environment variables (optional):
export WORKSPACE_DIR=/path/to/workspace # Default: ./workspace
export FLASK_DEBUG=False # Production mode
| Configuration | GPU Layers | VRAM Usage | Speed | Target Hardware |
|---|---|---|---|---|
| No offloading | 28 | 11-14GB | 1.0x | RTX 4090, A100, 3090 |
| Balanced | 12 | 6-8GB | 0.70x | RTX 4070, 3080 16GB |
| Aggressive | 8 | 5-7GB | 0.55x | RTX 3060 12GB |
| Extreme | 4 | 4-5GB | 0.40x | RTX 3080 10GB |
Float8 quantization is only supported on NVIDIA RTX 40 and 50 series GPUs. See docs/offloading.md for offloading details.
Development API URL (frontend/.env.local):
NEXT_PUBLIC_API_URL=http://localhost:9527/api/v1
vibevoice/
├── backend/ # Flask API server
│ ├── api/ # REST API endpoints
│ │ ├── projects.py # Project CRUD
│ │ ├── speakers.py # Speaker management
│ │ ├── dialog_sessions.py # Dialog CRUD
│ │ ├── generation.py # Voice generation
│ │ ├── dataset.py # Dataset management
│ │ └── training.py # LoRA training
│ ├── services/ # Business logic layer
│ ├── models/ # Data models
│ ├── task_manager/ # Background task queue
│ ├── inference/ # Inference engine
│ ├── training/ # Training engine & state management
│ ├── i18n/ # Backend translations
│ └── dist/ # Frontend static files (production)
├── frontend/ # Next.js web application
│ ├── app/ # Next.js pages
│ │ ├── page.tsx # Home/Project selector
│ │ ├── quick-generate/ # Quick generation (no project)
│ │ ├── speaker-role/ # Speaker management
│ │ ├── voice-editor/ # Dialog editor
│ │ ├── generate-voice/ # Generation page
│ │ ├── dataset/ # Dataset management
│ │ └── fine-tuning/ # LoRA training page
│ ├── components/ # React components
│ ├── lib/ # Context providers & utilities
│ │ ├── ProjectContext.tsx
│ │ ├── SessionContext.tsx
│ │ ├── SpeakerRoleContext.tsx
│ │ ├── GenerationContext.tsx
│ │ ├── TrainingContext.tsx
│ │ ├── GlobalTaskContext.tsx
│ │ ├── i18n/ # Frontend translations
│ │ └── api.ts # API client
│ └── types/ # TypeScript type definitions
└── vibevoice/ # Core inference library
├── modular/ # Model implementations
│ ├── custom_offloading_utils.py # Layer offloading
│ └── adaptive_offload.py # Auto VRAM config
├── processor/ # Input processing
└── schedule/ # Diffusion scheduling
For complete API documentation including request/response examples, see docs/APIs.md.
workspace/
├── projects.json # All projects metadata
├── _quick-generate/ # Quick generation storage
│ ├── voices/ # Uploaded voice samples
│ ├── outputs/ # Generated audio files
│ └── history.json # Generation history
└── {project-id}/
├── voices/
│ ├── speakers.json # Speaker metadata
│ └── {uuid}.wav # Voice files
├── scripts/
│ ├── sessions.json # Session metadata
│ └── {uuid}.txt # Dialog text files
├── output/
│ ├── generation.json # Generation metadata
│ └── {request_id}.wav # Generated audio files
├── datasets/
│ ├── datasets.json # Dataset metadata
│ └── {dataset-id}/
│ ├── datasets.jsonl # Dataset items (one JSON per line)
│ ├── audio/ # Audio files
│ └── voice_prompts/ # Voice prompt files
└── training/
├── training_history.json # Training job metadata
└── lora_output/
└── {lora-name}/
├── model_epoch_*.safetensors # Checkpoint files
└── model_final.safetensors # Final model
RTX 4090 (24GB VRAM):
| Configuration | VRAM | Generation Time | RTF | Quality |
|---|---|---|---|---|
| BFloat16, No offload | 14GB | 15s (50s audio) | 0.30x | Excellent |
| Float8, No offload | 7GB | 16s (50s audio) | 0.32x | Excellent |
RTX 3060 12GB:
| Configuration | VRAM | Generation Time | RTF | Quality |
|---|---|---|---|---|
| Float8, Balanced | 7GB | 30s (50s audio) | 0.60x | Excellent |
| Float8, Aggressive | 6GB | 40s (50s audio) | 0.80x | Good |
RTF (Real-Time Factor) < 1.0 means faster than real-time
Share your projects and experiences:
Important: This project is for research and development purposes only.
DO:
DO NOT:
By using this software, you agree to use it ethically and responsibly.
We welcome contributions from the community! Here's how you can help:
# Backend tests (when available)
pytest tests/
# Frontend tests (when available)
cd frontend
npm test
# Manual testing
# 1. Create project
# 2. Add speakers
# 3. Create dialog
# 4. Generate voice
# 5. Verify output quality
This project follows the same license terms as the original Microsoft VibeVoice repository. Please refer to the LICENSE file for details.
If you use this implementation in your research, please cite both this project and the original VibeVoice paper:
@software{vibevoice_webapp_2024,
title={VibeVoice: Complete Web Application for Multi-Speaker Voice Generation},
author={Zhao, Kun},
year={2024},
url={https://github.com/zhao-kun/vibevoice}
}
@article{vibevoice2024,
title={VibeVoice: Unified Autoregressive and Diffusion for Speech Generation},
author={Microsoft Research},
year={2024}
}
# Try Float8 model
--dtype float8_e4m3fn
# Enable layer offloading in web UI
# Or use CLI with manual configuration
# Adjust CFG scale (try 1.0 - 2.0)
--cfg_scale 1.5
# Use higher precision model
--dtype bfloat16
# Change port in backend/run.py
app.run(host='0.0.0.0', port=9528)
cd frontend
rm -rf node_modules .next
npm install
npm run build
Made by the VibeVoice Community
VibeVoiceFusion is a full-stack, multi-speaker voice generation web system featuring LoRA fine-tuning, batch generation, and VRAM optimization. Based on Microsoft's VibeVoice (AR + diffusion architecture)
Python
490
292 commits
updated Oct 5, 2026
A Complete Web Application for Multi-Speaker Voice Generation
Built on Microsoft's VibeVoice Model
Features • Demo Samples • Get Started • Documentation • Community • Contributing
VibeVoiceFusion is a web application for multi-speaker speech generation with voice cloning and speech recognition, built on Microsoft's VibeVoice models. It brings text-to-speech, file transcription and live microphone transcription together in one bilingual (English/Chinese) web interface. It also covers LoRA fine-tuning and an OpenAI-compatible API, and it runs on consumer GPUs with 10GB+ of VRAM.
/v1/audio/speech, /v1/audio/transcriptions and the Realtime WebSocket /v1/realtime, so existing OpenAI SDKs work as-isListen to voice generation samples created with VibeVoiceFusion. Click the links below to download and play:
🎧 Pandora's Box Story (BFloat16 Model)
Generated with bfloat16 precision model - Full quality, 14GB VRAM
🎧 Pandora's Box Story (Float8 Model)
Generated with float8 quantization - Optimized for 7GB VRAM with comparable quality
🎭 东邪西毒 - 西游版 (Journey to the West Version)
Multi-speaker dialog with distinct voice characteristics for each character
Build docker image
# Clone the repository
git clone https://github.com/zhao-kun/vibevoicefusion.git
cd vibevoicefusion
# Build and the docker image
docker compose build vibevoice
After build successfully, run command:
docker run -d \
--name vibevoicefusion \
--gpus all \
-p 9527:9527 \
-v $(pwd)/workspace:/workspace/zhao-kun/vibevoice/workspace \
zhaokundev/vibevoicefusion:latest
Access the application at http://localhost:9527
The Docker image is available on Docker Hub, and you can launch VibeVoiceFusion using the following command.
docker pull zhaokundev/vibevoicefusion
docker run -d \
--name vibevoicefusion \
--gpus all \
-p 9527:9527 \
-v $(pwd)/workspace:/workspace/zhao-kun/vibevoice/workspace \
zhaokundev/vibevoicefusion:latest
Build Time: 18-28 minutes | Image Size: ~12-15GB
1. Install Backend Dependencies
# Clone the repository
git clone https://github.com/zhao-kun/vibevoice.git
cd vibevoice
# Install Python package
pip install -e .
2. Download Pre-trained Model
Download from HuggingFace (choose one):
Place files in ./models/vibevoice/
3. Install Frontend Dependencies (for development)
cd frontend
npm install
4. Build Frontend (for production)
cd frontend
npm run build
cp -r out/* ../backend/dist/
Production Mode (single server):
# Start backend server (serves both API and frontend)
python backend/run.py
# Access at http://localhost:9527
Development Mode (separate servers):
# Terminal 1: Start backend API
python backend/run.py # http://localhost:9527
# Terminal 2: Start frontend dev server
cd frontend
npm run dev # http://localhost:3000
For quick testing without setting up projects:
Speaker 1: Hello) or narration (plain text)Key Points:
This guide walks you through the complete process of creating multi-speaker voice generation from start to finish.
Start by creating a new project or selecting an existing one. Projects help organize your voice generation work with metadata and descriptions.
Create and manage projects from the home page
Actions:
The project will be automatically selected and you'll be navigated to the Speaker Role page.
Upload reference voice samples for each speaker. The system supports various audio formats (WAV, MP3, M4A, FLAC, WebM).
Upload and manage voice samples for each speaker
Actions:
Tips:
Create a dialog session and write the multi-speaker conversation. The dialog editor supports drag-and-drop reordering and real-time preview.
Multi-speaker dialog editor with visual and text modes
Actions:
Dialog Format (Text Mode):
Speaker 1: Welcome to our podcast!
Speaker 2: Thanks for having me. It's great to be here.
Speaker 1: Let's dive into today's topic.
Narration Mode:
For single-speaker content like audiobooks, articles, or podcasts, use Narration Mode:
Speaker N: prefixesThis is the first paragraph of your narration.
This is the second paragraph. No speaker formatting needed.
The narrator voice you selected will read all the text.
Features:
Configure generation parameters and start the voice synthesis process. Monitor real-time progress and manage generation history.
Generation interface with parameters, live progress, and history
Actions:
float8_e4m3fn (recommended): 7GB VRAM, faster loadingbfloat16: 14GB VRAM, full precisionReal-Time Monitoring:
Generation interface with parameters, live progress, and history
Generation History:
For CLI-based generation without the web UI:
python demo/local_file_inference.py \
--model_file ./models/vibevoice/vibevoice7b_float8_e4m3fn.safetensors \
--txt_path demo/text_examples/1p_pandora_box.txt \
--speaker_names zh-007 \
--output_dir ./outputs \
--dtype float8_e4m3fn \
--cfg_scale 1.3 \
--seed 42
CLI Arguments:
--model_file: Path to model .safetensors file--config: Path to config.json (optional)--txt_path: Input text file with speaker-labeled dialog--speaker_names: Speaker name(s) for voice file mapping--output_dir: Output directory for generated audio--device: cuda, mps, or cpu (auto-detected)--dtype: float8_e4m3fn or bfloat16--cfg_scale: Classifier-Free Guidance scale (default: 1.3)--seed: Random seed for reproducibilitySee Demo Model Tools for all command-line tools: TTS, Realtime 0.5B, ASR, and model/voice-preset conversion.
Environment variables (optional):
export WORKSPACE_DIR=/path/to/workspace # Default: ./workspace
export FLASK_DEBUG=False # Production mode
| Configuration | GPU Layers | VRAM Usage | Speed | Target Hardware |
|---|---|---|---|---|
| No offloading | 28 | 11-14GB | 1.0x | RTX 4090, A100, 3090 |
| Balanced | 12 | 6-8GB | 0.70x | RTX 4070, 3080 16GB |
| Aggressive | 8 | 5-7GB | 0.55x | RTX 3060 12GB |
| Extreme | 4 | 4-5GB | 0.40x | RTX 3080 10GB |
Float8 quantization is only supported on NVIDIA RTX 40 and 50 series GPUs. See docs/offloading.md for offloading details.
Development API URL (frontend/.env.local):
NEXT_PUBLIC_API_URL=http://localhost:9527/api/v1
vibevoice/
├── backend/ # Flask API server
│ ├── api/ # REST API endpoints
│ │ ├── projects.py # Project CRUD
│ │ ├── speakers.py # Speaker management
│ │ ├── dialog_sessions.py # Dialog CRUD
│ │ ├── generation.py # Voice generation
│ │ ├── dataset.py # Dataset management
│ │ └── training.py # LoRA training
│ ├── services/ # Business logic layer
│ ├── models/ # Data models
│ ├── task_manager/ # Background task queue
│ ├── inference/ # Inference engine
│ ├── training/ # Training engine & state management
│ ├── i18n/ # Backend translations
│ └── dist/ # Frontend static files (production)
├── frontend/ # Next.js web application
│ ├── app/ # Next.js pages
│ │ ├── page.tsx # Home/Project selector
│ │ ├── quick-generate/ # Quick generation (no project)
│ │ ├── speaker-role/ # Speaker management
│ │ ├── voice-editor/ # Dialog editor
│ │ ├── generate-voice/ # Generation page
│ │ ├── dataset/ # Dataset management
│ │ └── fine-tuning/ # LoRA training page
│ ├── components/ # React components
│ ├── lib/ # Context providers & utilities
│ │ ├── ProjectContext.tsx
│ │ ├── SessionContext.tsx
│ │ ├── SpeakerRoleContext.tsx
│ │ ├── GenerationContext.tsx
│ │ ├── TrainingContext.tsx
│ │ ├── GlobalTaskContext.tsx
│ │ ├── i18n/ # Frontend translations
│ │ └── api.ts # API client
│ └── types/ # TypeScript type definitions
└── vibevoice/ # Core inference library
├── modular/ # Model implementations
│ ├── custom_offloading_utils.py # Layer offloading
│ └── adaptive_offload.py # Auto VRAM config
├── processor/ # Input processing
└── schedule/ # Diffusion scheduling
For complete API documentation including request/response examples, see docs/APIs.md.
workspace/
├── projects.json # All projects metadata
├── _quick-generate/ # Quick generation storage
│ ├── voices/ # Uploaded voice samples
│ ├── outputs/ # Generated audio files
│ └── history.json # Generation history
└── {project-id}/
├── voices/
│ ├── speakers.json # Speaker metadata
│ └── {uuid}.wav # Voice files
├── scripts/
│ ├── sessions.json # Session metadata
│ └── {uuid}.txt # Dialog text files
├── output/
│ ├── generation.json # Generation metadata
│ └── {request_id}.wav # Generated audio files
├── datasets/
│ ├── datasets.json # Dataset metadata
│ └── {dataset-id}/
│ ├── datasets.jsonl # Dataset items (one JSON per line)
│ ├── audio/ # Audio files
│ └── voice_prompts/ # Voice prompt files
└── training/
├── training_history.json # Training job metadata
└── lora_output/
└── {lora-name}/
├── model_epoch_*.safetensors # Checkpoint files
└── model_final.safetensors # Final model
RTX 4090 (24GB VRAM):
| Configuration | VRAM | Generation Time | RTF | Quality |
|---|---|---|---|---|
| BFloat16, No offload | 14GB | 15s (50s audio) | 0.30x | Excellent |
| Float8, No offload | 7GB | 16s (50s audio) | 0.32x | Excellent |
RTX 3060 12GB:
| Configuration | VRAM | Generation Time | RTF | Quality |
|---|---|---|---|---|
| Float8, Balanced | 7GB | 30s (50s audio) | 0.60x | Excellent |
| Float8, Aggressive | 6GB | 40s (50s audio) | 0.80x | Good |
RTF (Real-Time Factor) < 1.0 means faster than real-time
Share your projects and experiences:
Important: This project is for research and development purposes only.
DO:
DO NOT:
By using this software, you agree to use it ethically and responsibly.
We welcome contributions from the community! Here's how you can help:
# Backend tests (when available)
pytest tests/
# Frontend tests (when available)
cd frontend
npm test
# Manual testing
# 1. Create project
# 2. Add speakers
# 3. Create dialog
# 4. Generate voice
# 5. Verify output quality
This project follows the same license terms as the original Microsoft VibeVoice repository. Please refer to the LICENSE file for details.
If you use this implementation in your research, please cite both this project and the original VibeVoice paper:
@software{vibevoice_webapp_2024,
title={VibeVoice: Complete Web Application for Multi-Speaker Voice Generation},
author={Zhao, Kun},
year={2024},
url={https://github.com/zhao-kun/vibevoice}
}
@article{vibevoice2024,
title={VibeVoice: Unified Autoregressive and Diffusion for Speech Generation},
author={Microsoft Research},
year={2024}
}
# Try Float8 model
--dtype float8_e4m3fn
# Enable layer offloading in web UI
# Or use CLI with manual configuration
# Adjust CFG scale (try 1.0 - 2.0)
--cfg_scale 1.5
# Use higher precision model
--dtype bfloat16
# Change port in backend/run.py
app.run(host='0.0.0.0', port=9528)
cd frontend
rm -rf node_modules .next
npm install
npm run build
Made by the VibeVoice Community