AnuarSv/IndexTTS2-Kazakh

Model

IndexTTS2 Kazakh Finetune

3

16 commits

2 linked in READMEs

updated Sep 24, 2026

See the code

README

IndexTTS2 Kazakh Finetune

Nuraidar Mambetaly · Anuar Sultanbekov · Altair Balakhazy

Fine-tuned GPT component of IndexTTS2 on Kazakh language data.

Training Details

Validated on Tesla V100 16GB. Inference uses approximately 10GB VRAM.

Audio Sample

Text: Менің досымның жақсы және жаман әдеттері бар. Жақсы әдетт ол маған көмектеседі. Ол мені әрқашан қолдайтын болады. Ал жаман әдетті кейде ол менің затымды рұқсатсыз алып кетуі мүмкін. Кейде ол мені ренжітеді. Бірақ артынан менден кешірім сұрайды.

Listen to sample

Training Curves

Training curves

Limitations

The training data contains a single speaker throughout all 4.8GB. Voice cloning does not generalize to other voices at this stage. The model produces Kazakh speech but defaults to the training speaker regardless of the reference audio provided.

This will improve with a multi-speaker dataset. Contributions and experiments are welcome.

Installation

Download the base model first:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="IndexTeam/IndexTTS-2",
    local_dir="checkpoints"
)

Then replace the GPT checkpoint and tokenizer with the Kazakh versions:

# Download this repo
hf download TilLabs/IndexTTS2-Kazakh --local-dir kz_files

# Replace
cp kz_files/gpt.pth checkpoints/gpt.pth
cp kz_files/bpe.model checkpoints/bpe.model
cp kz_files/config.yaml checkpoints/config.yaml

Clone the training branch and run:

git clone -b training_v2 https://github.com/JarodMica/index-tts.git
cd index-tts
uv sync --all-extras
uv run webui.py --model_dir checkpoints --fp16

Config Change

The base model uses number_text_tokens: 12000. This finetune uses 2000 to match the Kazakh BPE vocabulary. The config.yaml in this repo already has this applied.

What's in This Repo

FileDescription
gpt.pthFine-tuned GPT checkpoint (Kazakh)
bpe.modelSentencePiece BPE tokenizer trained on Kazakh text
config.yamlModel config with number_text_tokens: 2000

Files not included (s2mel.pth, feat1.pt, feat2.pt, qwen0.6bemo4-merge) are part of the base model and should be downloaded from IndexTeam/IndexTTS-2.

indextts
kazakh
text-to-speech

AnuarSv/IndexTTS2-Kazakh

Model

IndexTTS2 Kazakh Finetune

3

16 commits

2 linked in READMEs

updated Sep 24, 2026

See the code

README

IndexTTS2 Kazakh Finetune

Nuraidar Mambetaly · Anuar Sultanbekov · Altair Balakhazy

Fine-tuned GPT component of IndexTTS2 on Kazakh language data.

Training Details

Validated on Tesla V100 16GB. Inference uses approximately 10GB VRAM.

Audio Sample

Text: Менің досымның жақсы және жаман әдеттері бар. Жақсы әдетт ол маған көмектеседі. Ол мені әрқашан қолдайтын болады. Ал жаман әдетті кейде ол менің затымды рұқсатсыз алып кетуі мүмкін. Кейде ол мені ренжітеді. Бірақ артынан менден кешірім сұрайды.

Listen to sample

Training Curves

Training curves

Limitations

The training data contains a single speaker throughout all 4.8GB. Voice cloning does not generalize to other voices at this stage. The model produces Kazakh speech but defaults to the training speaker regardless of the reference audio provided.

This will improve with a multi-speaker dataset. Contributions and experiments are welcome.

Installation

Download the base model first:

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="IndexTeam/IndexTTS-2",
    local_dir="checkpoints"
)

Then replace the GPT checkpoint and tokenizer with the Kazakh versions:

# Download this repo
hf download TilLabs/IndexTTS2-Kazakh --local-dir kz_files

# Replace
cp kz_files/gpt.pth checkpoints/gpt.pth
cp kz_files/bpe.model checkpoints/bpe.model
cp kz_files/config.yaml checkpoints/config.yaml

Clone the training branch and run:

git clone -b training_v2 https://github.com/JarodMica/index-tts.git
cd index-tts
uv sync --all-extras
uv run webui.py --model_dir checkpoints --fp16

Config Change

The base model uses number_text_tokens: 12000. This finetune uses 2000 to match the Kazakh BPE vocabulary. The config.yaml in this repo already has this applied.

What's in This Repo

FileDescription
gpt.pthFine-tuned GPT checkpoint (Kazakh)
bpe.modelSentencePiece BPE tokenizer trained on Kazakh text
config.yamlModel config with number_text_tokens: 2000

Files not included (s2mel.pth, feat1.pt, feat2.pt, qwen0.6bemo4-merge) are part of the base model and should be downloaded from IndexTeam/IndexTTS-2.

indextts
kazakh
text-to-speech