allssai/Spark-TTS-Kazakh

Production-ready application for Kazakh text-to-speech synthesis with voice cloning capabilities. Built on fine-tuned Spark-TTS model with REST API, web interface, and automatic script conversion.

Python

12

20 commits

updated Feb 17, 2026

See the code

README

Kazakh Spark-TTS Tool

Fine-tuned Spark-TTS for Kazakh text-to-speech with voice cloning, supporting Cyrillic and Tote Zhazu (Arabic) scripts.

Key Features

  • 🎯 High-Quality Kazakh TTS: Natural and fluent speech synthesis
  • 🎤 Voice Cloning: Clone any voice with 3-10 seconds of reference audio
  • 📝 Dual Script Support: Cyrillic and Tote Zhazu scripts
  • ⚡ Fast Inference: Optimized for real-time generation

The GitHub repository includes:

  • ✅ FastAPI REST API server
  • ✅ Web-based user interface
  • ✅ Complete documentation and examples
  • ✅ Easy installation and deployment

image

Quick start

# Clone repository
git clone https://github.com/allssai/Spark-TTS-Kazakh.git
cd Spark-TTS-Kazakh

# Model weights
Hugging Face: `ErnarBahat/Spark-TTS-Kazakh`

```bash
huggingface-cli download ErnarBahat/Spark-TTS-Kazakh --local-dir ./pretrained_models/Kazakh-Spark-Final

# or
git lfs install
git clone https://huggingface.co/ErnarBahat/Spark-TTS-Kazakh pretrained_models/Kazakh-Spark-Final

# Install dependencies
pip install -r requirement.txt

# Run application
python app.py

Server: http://localhost:8002

Usage

REST API

import requests

# Text-to-speech
requests.post("http://localhost:8002/tts", json={"text": "Сәлеметсіз бе!", "script": "cyrillic"})

# Voice cloning
files = {"audio": open("reference.wav", "rb")}
data = {"text": "Сәлеметсіз бе!"}
requests.post("http://localhost:8002/clone", files=files, data=data)

CLI

python infer.py --text "Сәлеметсіз бе!" --output output.wav
python infer.py --text "Сәлеметсіз бе!" --reference voice.wav --output cloned.wav

Project structure

├── app.py                 # FastAPI server
├── infer.py              # Inference script
├── cli/                  # SparkTTS core modules
├── sparktts/             # SparkTTS library
├── static/               # Web UI
├── src/                  # Source code
├── config_axolotl/       # Training configs
├── example/              # Example files
└── examples/             # Code examples

License

Apache License 2.0 — see LICENSE.

Citation

@misc{kazakh_spark_tts_2026,
  title={Spark-TTS-Kazakh: Fine-tuned Kazakh Text-to-Speech Model},
  author={Ernar Bahat},
  year={2026},
  publisher={GitHub},
  howpublished={\url{https://github.com/allssai/Spark-TTS-Kazakh}},
  note={Model: \url{https://huggingface.co/ErnarBahat/Spark-TTS-Kazakh}}
}

allssai/Spark-TTS-Kazakh

Production-ready application for Kazakh text-to-speech synthesis with voice cloning capabilities. Built on fine-tuned Spark-TTS model with REST API, web interface, and automatic script conversion.

Python

12

20 commits

updated Feb 17, 2026

See the code

README

Kazakh Spark-TTS Tool

Fine-tuned Spark-TTS for Kazakh text-to-speech with voice cloning, supporting Cyrillic and Tote Zhazu (Arabic) scripts.

Key Features

  • 🎯 High-Quality Kazakh TTS: Natural and fluent speech synthesis
  • 🎤 Voice Cloning: Clone any voice with 3-10 seconds of reference audio
  • 📝 Dual Script Support: Cyrillic and Tote Zhazu scripts
  • ⚡ Fast Inference: Optimized for real-time generation

The GitHub repository includes:

  • ✅ FastAPI REST API server
  • ✅ Web-based user interface
  • ✅ Complete documentation and examples
  • ✅ Easy installation and deployment

image

Quick start

# Clone repository
git clone https://github.com/allssai/Spark-TTS-Kazakh.git
cd Spark-TTS-Kazakh

# Model weights
Hugging Face: `ErnarBahat/Spark-TTS-Kazakh`

```bash
huggingface-cli download ErnarBahat/Spark-TTS-Kazakh --local-dir ./pretrained_models/Kazakh-Spark-Final

# or
git lfs install
git clone https://huggingface.co/ErnarBahat/Spark-TTS-Kazakh pretrained_models/Kazakh-Spark-Final

# Install dependencies
pip install -r requirement.txt

# Run application
python app.py

Server: http://localhost:8002

Usage

REST API

import requests

# Text-to-speech
requests.post("http://localhost:8002/tts", json={"text": "Сәлеметсіз бе!", "script": "cyrillic"})

# Voice cloning
files = {"audio": open("reference.wav", "rb")}
data = {"text": "Сәлеметсіз бе!"}
requests.post("http://localhost:8002/clone", files=files, data=data)

CLI

python infer.py --text "Сәлеметсіз бе!" --output output.wav
python infer.py --text "Сәлеметсіз бе!" --reference voice.wav --output cloned.wav

Project structure

├── app.py                 # FastAPI server
├── infer.py              # Inference script
├── cli/                  # SparkTTS core modules
├── sparktts/             # SparkTTS library
├── static/               # Web UI
├── src/                  # Source code
├── config_axolotl/       # Training configs
├── example/              # Example files
└── examples/             # Code examples

License

Apache License 2.0 — see LICENSE.

Citation

@misc{kazakh_spark_tts_2026,
  title={Spark-TTS-Kazakh: Fine-tuned Kazakh Text-to-Speech Model},
  author={Ernar Bahat},
  year={2026},
  publisher={GitHub},
  howpublished={\url{https://github.com/allssai/Spark-TTS-Kazakh}},
  note={Model: \url{https://huggingface.co/ErnarBahat/Spark-TTS-Kazakh}}
}

Languages

Python

88.9%

JavaScript

3.9%

CSS

3.4%

HTML

2.3%

Shell

1.6%