dignome/kitten_tts2

Model

Kitten TTS 2 for audio.cpp

0

4 commits

2 linked in READMEs

updated Oct 3, 2026

See the code

README

Kitten TTS 2 for audio.cpp

Powered by Stellon Labs.

Native audio.cpp GGUF conversion of KittenML/kitten-tts-2, with the full S3 meanflow decoder, speaker encoders, and prepared voice presets included in one file.

Use dignome/audio.cpp-custom, branch kittentts2, or an audio.cpp build that includes the same port. The published implementation is commit a96e66c6. The GGUF uses audio.cpp's model packaging and is intended for this loader; generic GGUF support in another application does not imply compatibility.

Download

The GGUF includes all model components, tokenizer, configuration, model specification, 48 prepared voice entries, and license/attribution files. Separate reference WAVs, voice JSON files, or model downloads are not needed for preset synthesis or custom voice cloning.

The language-model projections use Q8_0, with F16 embeddings/tied output head. The speaker projection retains its original BF16 storage, and the S3 decoder and speaker encoders retain F32 weights. This is a mixed-precision package, not an all-Q8 conversion.

Features

  • Native C++/GGML inference on CPU and NVIDIA CUDA; no Python, LibTorch, TorchScript, or ONNX Runtime dependency at inference time.
  • All 47 upstream presets, plus the additional PreparedBruno entry.
  • Named English voices such as Bruno, Bella, Jasper, and Luna.
  • Language presets: Arabic, Chinese, French, German, Hindi, Italian, Portuguese, Russian, and Spanish.
  • Custom voice cloning from a 1–30 second reference recording and its transcript.
  • Chunked long-form synthesis and mono 24 kHz output.

Ten languages have been validated: English and the nine languages listed above. Upstream advertises additional languages, but those are not validated by this port. Select a matching language preset or supply a reference recording and transcript in the target language; there is no separate language switch.

Build and run

See the port documentation for full instructions.

git clone --branch kittentts2 https://github.com/dignome/audio.cpp-custom.git
cd audio.cpp-custom
cmake -S . -B build-kitten2 -DAUDIOCPP_MODEL_SET=custom -DAUDIOCPP_MODELS=kitten_tts2 -DENGINE_ENABLE_CUDA=OFF
cmake --build build-kitten2 --config Release --target audiocpp_cli audiocpp_server --parallel 8

For NVIDIA, configure with -DENGINE_ENABLE_CUDA=ON and use --backend cuda. Visual Studio executables are under build-kitten2/bin/Release/; single-configuration builds use build-kitten2/bin/. The examples below assume the executable is on your PATH and the GGUF is in the current directory.

audiocpp_cli --family kitten_tts2 --model kitten-tts2-native-q8-multilingual.gguf --backend cpu --threads 8 --task tts --voice-id Bruno --text "Hello from native Kitten TTS 2." --out hello.wav

German preset:

audiocpp_cli --family kitten_tts2 --model kitten-tts2-native-q8-multilingual.gguf --backend cpu --task tts --voice-id German --text "Guten Tag! Dies ist ein Beispiel." --out german.wav

Voice cloning:

audiocpp_cli --family kitten_tts2 --model kitten-tts2-native-q8-multilingual.gguf --backend cpu --threads 8 --task clon --voice-ref reference.wav --reference-text "The exact words spoken in the reference recording." --text "The new words to say in this voice." --out cloned.wav

For server configuration, the model family is kitten_tts2. A model entry may use "id": "kitten-tts2", "family": "kitten_tts2", "task": "tts", "mode": "offline", and "path" pointing to the downloaded GGUF. API requests use the configured model ID; for example, "model": "kitten-tts2" and "voice": "Bruno".

Validation and limitations

CPU and NVIDIA CUDA synthesis and cloning were validated on Windows. The target branch was freshly built and tested on CPU; the earlier CUDA validation used the same model and CUDA implementation before transfer to this branch. See the validation record for exact environments and results.

The UI estimates 7 GB VRAM for this package, including headroom. This is guidance, not a guaranteed minimum or maximum for every request. CUDA was tested on a 16 GB RTX 4060 Ti; AMD and other GPU backends have not been validated.

This is an experimental port. Streaming, automatic reference transcription, the upstream English text normalizer, and the smaller student decoders are not supported here. Write numbers and abbreviations as spoken words when needed. The full default decoder is included.

Sources and license

  • Kitten TTS 2: KittenML/kitten-tts-2, snapshot baa41e5d2c5f64be0095365672a7858542261271.
  • S3 decoder: ResembleAI/chatterbox-turbo, revision 749d1c1a46eb10492095d68fbcf55691ccf137cd, s3gen_meanflow.safetensors.
  • Speaker encoder: the speaker/model.safetensors distributed with the Kitten snapshot, attributed upstream to pyannote/embedding.

Kitten model materials retain the Stellon Labs Community License in LICENSE. Commercial use requires registration under that agreement, and exceeding its revenue or funding limits requires a separate license. The bundled decoder and speaker components retain their respective MIT terms in native/LICENSE and speaker/LICENSE. See NOTICE for attribution and conversion details.

Conversion and native integration published by dignome. Bulk of the work was started with Sol 6.1 and finished with Astra.

audio.cpp
gguf
multilingual
text-to-speech
voice-cloning

dignome/kitten_tts2

Model

Kitten TTS 2 for audio.cpp

0

4 commits

2 linked in READMEs

updated Oct 3, 2026

See the code

README

Kitten TTS 2 for audio.cpp

Powered by Stellon Labs.

Native audio.cpp GGUF conversion of KittenML/kitten-tts-2, with the full S3 meanflow decoder, speaker encoders, and prepared voice presets included in one file.

Use dignome/audio.cpp-custom, branch kittentts2, or an audio.cpp build that includes the same port. The published implementation is commit a96e66c6. The GGUF uses audio.cpp's model packaging and is intended for this loader; generic GGUF support in another application does not imply compatibility.

Download

The GGUF includes all model components, tokenizer, configuration, model specification, 48 prepared voice entries, and license/attribution files. Separate reference WAVs, voice JSON files, or model downloads are not needed for preset synthesis or custom voice cloning.

The language-model projections use Q8_0, with F16 embeddings/tied output head. The speaker projection retains its original BF16 storage, and the S3 decoder and speaker encoders retain F32 weights. This is a mixed-precision package, not an all-Q8 conversion.

Features

  • Native C++/GGML inference on CPU and NVIDIA CUDA; no Python, LibTorch, TorchScript, or ONNX Runtime dependency at inference time.
  • All 47 upstream presets, plus the additional PreparedBruno entry.
  • Named English voices such as Bruno, Bella, Jasper, and Luna.
  • Language presets: Arabic, Chinese, French, German, Hindi, Italian, Portuguese, Russian, and Spanish.
  • Custom voice cloning from a 1–30 second reference recording and its transcript.
  • Chunked long-form synthesis and mono 24 kHz output.

Ten languages have been validated: English and the nine languages listed above. Upstream advertises additional languages, but those are not validated by this port. Select a matching language preset or supply a reference recording and transcript in the target language; there is no separate language switch.

Build and run

See the port documentation for full instructions.

git clone --branch kittentts2 https://github.com/dignome/audio.cpp-custom.git
cd audio.cpp-custom
cmake -S . -B build-kitten2 -DAUDIOCPP_MODEL_SET=custom -DAUDIOCPP_MODELS=kitten_tts2 -DENGINE_ENABLE_CUDA=OFF
cmake --build build-kitten2 --config Release --target audiocpp_cli audiocpp_server --parallel 8

For NVIDIA, configure with -DENGINE_ENABLE_CUDA=ON and use --backend cuda. Visual Studio executables are under build-kitten2/bin/Release/; single-configuration builds use build-kitten2/bin/. The examples below assume the executable is on your PATH and the GGUF is in the current directory.

audiocpp_cli --family kitten_tts2 --model kitten-tts2-native-q8-multilingual.gguf --backend cpu --threads 8 --task tts --voice-id Bruno --text "Hello from native Kitten TTS 2." --out hello.wav

German preset:

audiocpp_cli --family kitten_tts2 --model kitten-tts2-native-q8-multilingual.gguf --backend cpu --task tts --voice-id German --text "Guten Tag! Dies ist ein Beispiel." --out german.wav

Voice cloning:

audiocpp_cli --family kitten_tts2 --model kitten-tts2-native-q8-multilingual.gguf --backend cpu --threads 8 --task clon --voice-ref reference.wav --reference-text "The exact words spoken in the reference recording." --text "The new words to say in this voice." --out cloned.wav

For server configuration, the model family is kitten_tts2. A model entry may use "id": "kitten-tts2", "family": "kitten_tts2", "task": "tts", "mode": "offline", and "path" pointing to the downloaded GGUF. API requests use the configured model ID; for example, "model": "kitten-tts2" and "voice": "Bruno".

Validation and limitations

CPU and NVIDIA CUDA synthesis and cloning were validated on Windows. The target branch was freshly built and tested on CPU; the earlier CUDA validation used the same model and CUDA implementation before transfer to this branch. See the validation record for exact environments and results.

The UI estimates 7 GB VRAM for this package, including headroom. This is guidance, not a guaranteed minimum or maximum for every request. CUDA was tested on a 16 GB RTX 4060 Ti; AMD and other GPU backends have not been validated.

This is an experimental port. Streaming, automatic reference transcription, the upstream English text normalizer, and the smaller student decoders are not supported here. Write numbers and abbreviations as spoken words when needed. The full default decoder is included.

Sources and license

  • Kitten TTS 2: KittenML/kitten-tts-2, snapshot baa41e5d2c5f64be0095365672a7858542261271.
  • S3 decoder: ResembleAI/chatterbox-turbo, revision 749d1c1a46eb10492095d68fbcf55691ccf137cd, s3gen_meanflow.safetensors.
  • Speaker encoder: the speaker/model.safetensors distributed with the Kitten snapshot, attributed upstream to pyannote/embedding.

Kitten model materials retain the Stellon Labs Community License in LICENSE. Commercial use requires registration under that agreement, and exceeding its revenue or funding limits requires a separate license. The bundled decoder and speaker components retain their respective MIT terms in native/LICENSE and speaker/LICENSE. See NOTICE for attribution and conversion details.

Conversion and native integration published by dignome. Bulk of the work was started with Sol 6.1 and finished with Astra.

audio.cpp
gguf
multilingual
text-to-speech
voice-cloning