allssai/voxcpm-kazakh-tts

最强哈萨克语TTS:Multilingual TTS Engine – Kazakh-Enhanced Edition

Python

10

9 commits

updated Feb 24, 2026

See the code

README

VoxCPM Kazakh TTS - Multilingual Text-to-Speech Engine

License Python PyTorch

A multilingual text-to-speech system based on VoxCPM 1.5, with LoRA fine-tuning specifically optimized for Kazakh language.

Quick Start • Documentation • 中文文档


✨ Key Features

  • 🌍 Multilingual Support: Kazakh, Chinese, English, and mixed text
  • 🎭 Zero-Shot Voice Cloning: Clone any voice with 3-10 seconds of reference audio
  • ⚡ Real-Time Synthesis: Fast high-quality speech generation
  • 🎯 Kazakh Optimization: Enhanced Kazakh language quality through LoRA fine-tuning
  • 🎨 Bilingual Interface: Full support for Chinese and Kazakh
  • 📊 Real-Time Progress: Display inference process and detailed logs
  • 🔧 Flexible Configuration: Adjustable speed, pitch, inference steps, and more
  • 🎵 Voice Management: Complete CRUD operations for voice presets

🚀 Quick Start

Prerequisites

  • Python 3.8+
  • 8GB+ RAM
  • 4GB+ Disk Space
  • (Recommended) NVIDIA GPU with CUDA

Three-Step Deployment

1. Clone the Repository

git clone https://github.com/allssai/voxcpm-kazakh-tts
cd voxcpm-kazakh-tts

2. Download LoRA Model

The LoRA model is hosted on HuggingFace and needs to be downloaded separately:

# Method 1: Using huggingface-cli (Recommended)
pip install huggingface_hub
huggingface-cli download ErnarBahat/VoxCPM-KazakhTTS-Lora --local-dir ./lora

# Method 2: Manual Download
# Visit https://huggingface.co/ErnarBahat/VoxCPM-KazakhTTS-Lora
# Download all files to ./lora/ directory

3. Install and Launch

Linux/macOS:

chmod +x install.sh
./install.sh    # Auto-install dependencies
./start.sh      # Launch application

Windows:

install.bat     # Auto-install dependencies
start.bat       # Launch application

Docker:

docker-compose up -d

4. Access the Application

Open your browser and visit: http://localhost:7860

First Launch Note:

  • The first run will automatically download the VoxCPM base model (~1.5 GB)
  • Download time depends on network speed, typically 5-15 minutes
  • The model will be cached locally, subsequent launches take only 10 seconds

For detailed instructions, see INSTALL.md

📖 Usage

Open your browser and visit http://localhost:7860 to access the web interface.

For detailed usage instructions, see the documentation.

🎯 Technical Details

⚠️ Common Issues

  • First launch is slow: The first run downloads the base model (~1.5 GB), taking 5-15 minutes. Subsequent launches take only 10 seconds.
  • Port already in use: Change the port in web_app.py or stop the process using port 7860.
  • CUDA not available: The system will automatically use CPU mode (slower but functional).

For more troubleshooting, see INSTALL.md.

📊 System Requirements

Minimum: Python 3.8+, 8GB RAM, 4GB Disk

Recommended: Python 3.10+, 16GB RAM, NVIDIA GPU (6GB+ VRAM), 10GB Disk

📄 License

This project is based on VoxCPM 1.5 and follows the corresponding open-source license.

🙏 Acknowledgments

📞 Support

For questions or suggestions, please check the project documentation:


Enjoy multilingual speech synthesis! 🎉

allssai/voxcpm-kazakh-tts

最强哈萨克语TTS:Multilingual TTS Engine – Kazakh-Enhanced Edition

Python

10

9 commits

updated Feb 24, 2026

See the code

README

VoxCPM Kazakh TTS - Multilingual Text-to-Speech Engine

License Python PyTorch

A multilingual text-to-speech system based on VoxCPM 1.5, with LoRA fine-tuning specifically optimized for Kazakh language.

Quick Start • Documentation • 中文文档


✨ Key Features

  • 🌍 Multilingual Support: Kazakh, Chinese, English, and mixed text
  • 🎭 Zero-Shot Voice Cloning: Clone any voice with 3-10 seconds of reference audio
  • ⚡ Real-Time Synthesis: Fast high-quality speech generation
  • 🎯 Kazakh Optimization: Enhanced Kazakh language quality through LoRA fine-tuning
  • 🎨 Bilingual Interface: Full support for Chinese and Kazakh
  • 📊 Real-Time Progress: Display inference process and detailed logs
  • 🔧 Flexible Configuration: Adjustable speed, pitch, inference steps, and more
  • 🎵 Voice Management: Complete CRUD operations for voice presets

🚀 Quick Start

Prerequisites

  • Python 3.8+
  • 8GB+ RAM
  • 4GB+ Disk Space
  • (Recommended) NVIDIA GPU with CUDA

Three-Step Deployment

1. Clone the Repository

git clone https://github.com/allssai/voxcpm-kazakh-tts
cd voxcpm-kazakh-tts

2. Download LoRA Model

The LoRA model is hosted on HuggingFace and needs to be downloaded separately:

# Method 1: Using huggingface-cli (Recommended)
pip install huggingface_hub
huggingface-cli download ErnarBahat/VoxCPM-KazakhTTS-Lora --local-dir ./lora

# Method 2: Manual Download
# Visit https://huggingface.co/ErnarBahat/VoxCPM-KazakhTTS-Lora
# Download all files to ./lora/ directory

3. Install and Launch

Linux/macOS:

chmod +x install.sh
./install.sh    # Auto-install dependencies
./start.sh      # Launch application

Windows:

install.bat     # Auto-install dependencies
start.bat       # Launch application

Docker:

docker-compose up -d

4. Access the Application

Open your browser and visit: http://localhost:7860

First Launch Note:

  • The first run will automatically download the VoxCPM base model (~1.5 GB)
  • Download time depends on network speed, typically 5-15 minutes
  • The model will be cached locally, subsequent launches take only 10 seconds

For detailed instructions, see INSTALL.md

📖 Usage

Open your browser and visit http://localhost:7860 to access the web interface.

For detailed usage instructions, see the documentation.

🎯 Technical Details

⚠️ Common Issues

  • First launch is slow: The first run downloads the base model (~1.5 GB), taking 5-15 minutes. Subsequent launches take only 10 seconds.
  • Port already in use: Change the port in web_app.py or stop the process using port 7860.
  • CUDA not available: The system will automatically use CPU mode (slower but functional).

For more troubleshooting, see INSTALL.md.

📊 System Requirements

Minimum: Python 3.8+, 8GB RAM, 4GB Disk

Recommended: Python 3.10+, 16GB RAM, NVIDIA GPU (6GB+ VRAM), 10GB Disk

📄 License

This project is based on VoxCPM 1.5 and follows the corresponding open-source license.

🙏 Acknowledgments

📞 Support

For questions or suggestions, please check the project documentation:


Enjoy multilingual speech synthesis! 🎉

Languages

Python

98.6%