The reference training stack: the data pipeline and the three training stages — pretrain, LoRA fine-tune, and DPO — that produce the model served by gepard-inference.
GEPARD — Generative, Prosody-aware, Autoregressive text-to-speech for Realtime Dialogues — is a decoder-only TTS model built to be served by a stock LLM engine (vLLM) without custom CUDA kernels in the decode loop. A standard full-attention Qwen3.5 backbone predicts discrete audio codes; an FSQ-based NVIDIA NanoCodec turns them into a 22.05 kHz waveform. Everything non-standard — zero-shot voice cloning, text-repetition, classifier-free guidance — is kept out of the autoregressive loop or distilled into the weights, so generation stays single-pass.
This repository is the training side of Gepard. It takes a tokenized speech corpus and runs the full lifecycle end to end:
Everything is driven by one unified Hydra config tree (conf/) with a single entry file per phase. The design, architecture, data pipeline, and every training stage are documented in docs/MODEL_GUIDE.md. Adding a language the model has never seen has its own recipe: docs/NEW_LANGUAGE_GUIDE.md. Full experimental detail is in the technical report (gepard_techreport.pdf).
The model these stages produce is:
Requires a CUDA GPU and Python 3.12. The setup script builds a local venv/ with a CUDA-matched PyTorch, the NeMo codec, and the Gepard training package.
# 1. (once per machine, optional) system packages — nvcc, python3.12 headers, git-lfs
make system-deps
# 2. create venv/ and install the training stack
make setup
# 3. authenticate Hugging Face (dataset + checkpoint pull/push)
make login
Two optional environments layer on top: make setup_dpo (NeMo codec + Whisper + WER, for the DPO data pipeline) and make setup_inference (to run the demo notebook / smoke-test a trained checkpoint).
Point at a source corpus and run the stages. The one fully public, pre-tokenized example dataset — already encoded with the default codec — is nineninesix/emolia_filtered_nano_codec_21_dataset (a filtered slice of Emilia); uncomment its block in conf/data/full_corpus.yaml to build on open data end to end.
make dataset # build the tokenized training corpus (prepare stage)
make train # pretrain the base model (4-GPU FSDP)
make finetune # LoRA short-phrase fine-tune (single GPU)
DPO (optional, four stages — see MODEL_GUIDE §7 / §12.5):
make setup_dpo
make dpo-sample-sharded SHARDS=4 # rollouts (run `make dpo-progress` in a 2nd terminal)
make dpo-score && make dpo-pairs # reward + preference pairs
make dpo-train # LoRA + stop-head DPO
Publish a trained checkpoint:
make merge CHECKPOINT=checkpoints/checkpoint-N # fold LoRA adapters (for fine-tune (sft) only)
make upload REPO=you/model CHECKPOINT=<merged-dir> # bf16 + auto-generated model card
Every stage composes its config from conf/ and validates it before any GPU time — see MODEL_GUIDE §8 (config system) and §12 (training). One-off variations layer on an entry via experiment= presets (conf/experiment/, §12.8).
Demo notebook. notebooks/inference_demo.ipynb is meant for quick in-project checks — loading a just-trained checkpoint and listening to it without leaving the repo. For pure inference (no training dependencies), the gepard-inference repo ships an equivalent Colab.
make help # show all commands
make system-deps # install apt deps (nvcc, python3.12, git-lfs) — sudo, once
make setup # build venv/ and install the training stack
make login # authenticate Hugging Face
make dataset # build the tokenized training corpus
make train # pretrain (4-GPU FSDP) │ make resume CHECKPOINT=…
make finetune # LoRA fine-tune (single GPU) │ make finetune-resume CHECKPOINT=…
make dpo-dataset # DPO stages 1–3 (sample→score→pairs)
make dpo-train # DPO stage 4 (LoRA + stop head)
make merge # fold LoRA adapters into a servable checkpoint
make upload # push a checkpoint + model card to the Hub
make test # config + golden-baseline tests
If you use this work in your research, please cite:
@software{gepard_2026,
author = {Abdurazakov Ulanbek, Pavlov Denis, and Bakashov Nursultan},
title = {Gepard: Real-Time Decoder-Only TTS Native to vLLM},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/nineninesix/gepard-1.0}},
note = {Open-source, vLLM-native autoregressive TTS}
}
Gepard builds on and is inspired by a great deal of open work. The main pieces:
@misc{qwen3,
title={Qwen3 Technical Report},
author={Qwen Team},
year={2025},
eprint={2505.09388},
archivePrefix={arXiv}
}
@inproceedings{kwon2023vllm,
title={Efficient Memory Management for Large Language Model Serving with PagedAttention},
author={Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph E and Zhang, Hao and Stoica, Ion},
booktitle={Proceedings of the 29th Symposium on Operating Systems Principles (SOSP)},
pages={611--626},
year={2023},
eprint={2309.06180},
archivePrefix={arXiv}
}
@article{nvidia2025nanocodec,
title={NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference},
author={Casanova, Edresson and Neekhara, Paarth and Langman, Ryan and Hussain, Shehzeen and Ghosh, Subhankar and Yang, Xuesong and Juki{\'c}, Ante and Li, Jason and Ginsburg, Boris},
journal={arXiv preprint arXiv:2508.05835},
year={2025}
}
@article{he2024emilia,
title={Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation},
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Jin, Zhizheng and others},
journal={arXiv preprint arXiv:2407.05361},
year={2024}
}
@article{mentzer2023fsq,
title={Finite Scalar Quantization: VQ-VAE Made Simple},
author={Mentzer, Fabian and Agustsson, Eirikur and Tschannen, Michael and Malireddy, Srikanth and Alshina, Elena},
journal={arXiv preprint arXiv:2309.15505},
year={2023}
}
@article{nvidia2024magpie,
title={Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment},
author={Neekhara, Paarth and Hussain, Shehzeen and Ghosh, Subhankar and Li, Jason and Valle, Rafael and Badlani, Rohan and Ginsburg, Boris},
journal={arXiv preprint arXiv:2406.17957},
year={2024}
}
@article{ho2022cfg,
title={Classifier-Free Diffusion Guidance},
author={Ho, Jonathan and Salimans, Tim},
journal={arXiv preprint arXiv:2207.12598},
year={2022}
}
@article{rafailov2023dpo,
title={Direct Preference Optimization: Your Language Model is Secretly a Reward Model},
author={Rafailov, Rafael and Sharma, Archit and Mitchell, Eric and Ermon, Stefano and Manning, Christopher D and Finn, Chelsea},
journal={arXiv preprint arXiv:2305.18290},
year={2023}
}
@article{meng2024simpo,
title={SimPO: Simple Preference Optimization with a Reference-Free Reward},
author={Meng, Yu and Xia, Mengzhou and Chen, Danqi},
journal={arXiv preprint arXiv:2405.14734},
year={2024}
}
@article{voicestar2025,
title={VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation},
author={Peng, Puyuan and Li, Shang-Wen and Mohamed, Abdelrahman and Harwath, David},
journal={arXiv preprint arXiv:2505.19462},
year={2025}
}
Apache 2.0 — see the LICENSE file for details.
Gepard loads the NVIDIA NeMo NanoCodec (nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps) at runtime. That model is not covered by Apache 2.0 — it is governed by the NVIDIA Open Model License Agreement. See the NOTICE file for third-party attribution.
6 commits
Jupyter Notebook
58.3%
Python
37.9%
Shell
2.4%
Makefile
1.4%
The reference training stack: the data pipeline and the three training stages — pretrain, LoRA fine-tune, and DPO — that produce the model served by gepard-inference.
GEPARD — Generative, Prosody-aware, Autoregressive text-to-speech for Realtime Dialogues — is a decoder-only TTS model built to be served by a stock LLM engine (vLLM) without custom CUDA kernels in the decode loop. A standard full-attention Qwen3.5 backbone predicts discrete audio codes; an FSQ-based NVIDIA NanoCodec turns them into a 22.05 kHz waveform. Everything non-standard — zero-shot voice cloning, text-repetition, classifier-free guidance — is kept out of the autoregressive loop or distilled into the weights, so generation stays single-pass.
This repository is the training side of Gepard. It takes a tokenized speech corpus and runs the full lifecycle end to end:
Everything is driven by one unified Hydra config tree (conf/) with a single entry file per phase. The design, architecture, data pipeline, and every training stage are documented in docs/MODEL_GUIDE.md. Adding a language the model has never seen has its own recipe: docs/NEW_LANGUAGE_GUIDE.md. Full experimental detail is in the technical report (gepard_techreport.pdf).
The model these stages produce is:
Requires a CUDA GPU and Python 3.12. The setup script builds a local venv/ with a CUDA-matched PyTorch, the NeMo codec, and the Gepard training package.
# 1. (once per machine, optional) system packages — nvcc, python3.12 headers, git-lfs
make system-deps
# 2. create venv/ and install the training stack
make setup
# 3. authenticate Hugging Face (dataset + checkpoint pull/push)
make login
Two optional environments layer on top: make setup_dpo (NeMo codec + Whisper + WER, for the DPO data pipeline) and make setup_inference (to run the demo notebook / smoke-test a trained checkpoint).
Point at a source corpus and run the stages. The one fully public, pre-tokenized example dataset — already encoded with the default codec — is nineninesix/emolia_filtered_nano_codec_21_dataset (a filtered slice of Emilia); uncomment its block in conf/data/full_corpus.yaml to build on open data end to end.
make dataset # build the tokenized training corpus (prepare stage)
make train # pretrain the base model (4-GPU FSDP)
make finetune # LoRA short-phrase fine-tune (single GPU)
DPO (optional, four stages — see MODEL_GUIDE §7 / §12.5):
make setup_dpo
make dpo-sample-sharded SHARDS=4 # rollouts (run `make dpo-progress` in a 2nd terminal)
make dpo-score && make dpo-pairs # reward + preference pairs
make dpo-train # LoRA + stop-head DPO
Publish a trained checkpoint:
make merge CHECKPOINT=checkpoints/checkpoint-N # fold LoRA adapters (for fine-tune (sft) only)
make upload REPO=you/model CHECKPOINT=<merged-dir> # bf16 + auto-generated model card
Every stage composes its config from conf/ and validates it before any GPU time — see MODEL_GUIDE §8 (config system) and §12 (training). One-off variations layer on an entry via experiment= presets (conf/experiment/, §12.8).
Demo notebook. notebooks/inference_demo.ipynb is meant for quick in-project checks — loading a just-trained checkpoint and listening to it without leaving the repo. For pure inference (no training dependencies), the gepard-inference repo ships an equivalent Colab.
make help # show all commands
make system-deps # install apt deps (nvcc, python3.12, git-lfs) — sudo, once
make setup # build venv/ and install the training stack
make login # authenticate Hugging Face
make dataset # build the tokenized training corpus
make train # pretrain (4-GPU FSDP) │ make resume CHECKPOINT=…
make finetune # LoRA fine-tune (single GPU) │ make finetune-resume CHECKPOINT=…
make dpo-dataset # DPO stages 1–3 (sample→score→pairs)
make dpo-train # DPO stage 4 (LoRA + stop head)
make merge # fold LoRA adapters into a servable checkpoint
make upload # push a checkpoint + model card to the Hub
make test # config + golden-baseline tests
If you use this work in your research, please cite:
@software{gepard_2026,
author = {Abdurazakov Ulanbek, Pavlov Denis, and Bakashov Nursultan},
title = {Gepard: Real-Time Decoder-Only TTS Native to vLLM},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/nineninesix/gepard-1.0}},
note = {Open-source, vLLM-native autoregressive TTS}
}
Gepard builds on and is inspired by a great deal of open work. The main pieces:
@misc{qwen3,
title={Qwen3 Technical Report},
author={Qwen Team},
year={2025},
eprint={2505.09388},
archivePrefix={arXiv}
}
@inproceedings{kwon2023vllm,
title={Efficient Memory Management for Large Language Model Serving with PagedAttention},
author={Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph E and Zhang, Hao and Stoica, Ion},
booktitle={Proceedings of the 29th Symposium on Operating Systems Principles (SOSP)},
pages={611--626},
year={2023},
eprint={2309.06180},
archivePrefix={arXiv}
}
@article{nvidia2025nanocodec,
title={NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference},
author={Casanova, Edresson and Neekhara, Paarth and Langman, Ryan and Hussain, Shehzeen and Ghosh, Subhankar and Yang, Xuesong and Juki{\'c}, Ante and Li, Jason and Ginsburg, Boris},
journal={arXiv preprint arXiv:2508.05835},
year={2025}
}
@article{he2024emilia,
title={Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation},
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Jin, Zhizheng and others},
journal={arXiv preprint arXiv:2407.05361},
year={2024}
}
@article{mentzer2023fsq,
title={Finite Scalar Quantization: VQ-VAE Made Simple},
author={Mentzer, Fabian and Agustsson, Eirikur and Tschannen, Michael and Malireddy, Srikanth and Alshina, Elena},
journal={arXiv preprint arXiv:2309.15505},
year={2023}
}
@article{nvidia2024magpie,
title={Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment},
author={Neekhara, Paarth and Hussain, Shehzeen and Ghosh, Subhankar and Li, Jason and Valle, Rafael and Badlani, Rohan and Ginsburg, Boris},
journal={arXiv preprint arXiv:2406.17957},
year={2024}
}
@article{ho2022cfg,
title={Classifier-Free Diffusion Guidance},
author={Ho, Jonathan and Salimans, Tim},
journal={arXiv preprint arXiv:2207.12598},
year={2022}
}
@article{rafailov2023dpo,
title={Direct Preference Optimization: Your Language Model is Secretly a Reward Model},
author={Rafailov, Rafael and Sharma, Archit and Mitchell, Eric and Ermon, Stefano and Manning, Christopher D and Finn, Chelsea},
journal={arXiv preprint arXiv:2305.18290},
year={2023}
}
@article{meng2024simpo,
title={SimPO: Simple Preference Optimization with a Reference-Free Reward},
author={Meng, Yu and Xia, Mengzhou and Chen, Danqi},
journal={arXiv preprint arXiv:2405.14734},
year={2024}
}
@article{voicestar2025,
title={VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation},
author={Peng, Puyuan and Li, Shang-Wen and Mohamed, Abdelrahman and Harwath, David},
journal={arXiv preprint arXiv:2505.19462},
year={2025}
}
Apache 2.0 — see the LICENSE file for details.
Gepard loads the NVIDIA NeMo NanoCodec (nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps) at runtime. That model is not covered by Apache 2.0 — it is governed by the NVIDIA Open Model License Agreement. See the NOTICE file for third-party attribution.
6 commits
Jupyter Notebook
58.3%
Python
37.9%
Shell
2.4%
Makefile
1.4%