OPPOer/CuteTTS

Model

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

11

3 commits

1 linked in READMEs

updated Aug 24, 2026

See the code

README

EN | 中文

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

GitHub paper

CuteTTS logo

  • A lightweight (~230M-parameter) continuous autoregressive TTS model that runs efficiently on GPUs, CPUs, and Apple silicon.
  • Ultra-low latency: ~40 ms to the first audio chunk and a throughput of ~9× real time on an NVIDIA RTX 4090.
  • Excellent speech quality and voice cloning performance.
  • Web demo, Python API, and CLI.
  • Multilingual support: English, Chinese, French, German, and Spanish.

CuteTTS architecture
CuteTTS performance

Zero-shot voice-cloning performance

ModelParams.LibriSpeech test-clean WER (%) ↓LibriSpeech test-clean SIM ↑Seed-TTS EN WER (%) ↓Seed-TTS EN SIM ↑Seed-TTS ZH WER (%) ↓Seed-TTS ZH SIM ↑
MOSS‑TTS8B1.9867.71.8470.91.3777.0
Qwen3‑TTS1.7B2.3570.31.6671.40.9177.0
FireRedTTS‑21.5B4.3264.21.9566.51.1473.6
MOSS‑TTS‑Nano0.1B4.1048.44.6249.93.1364.3
F5‑TTS0.3B2.4266.01.8367.01.5676.0
ZipVoice0.1B2.0567.41.7069.71.4075.1
IndexTTS21.5B2.4770.02.2270.61.0276.5
CosyVoice 30.5B1.9969.72.0271.81.1678.0
VoxCPM22B3.0174.01.8475.30.9779.5
VibeVoice1.5B3.0468.91.1674.4
DiTAR0.6B2.3967.01.6973.51.0275.3
VibeVoice‑Realtime0.5B2.0069.52.0563.3
Pocket TTS0.1B1.5949.11.6350.7
CuteTTS0.2B2.1678.92.0476.51.4177.8
CuteTTS‑distill0.2B2.4176.82.0374.21.4775.6
cutetts
safetensors
text-to-speech
voice-cloning

Contributors

MinMinLiang

3 commits

OPPOer/CuteTTS

Model

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

11

3 commits

1 linked in READMEs

updated Aug 24, 2026

See the code

README

EN | 中文

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

GitHub paper

CuteTTS logo

  • A lightweight (~230M-parameter) continuous autoregressive TTS model that runs efficiently on GPUs, CPUs, and Apple silicon.
  • Ultra-low latency: ~40 ms to the first audio chunk and a throughput of ~9× real time on an NVIDIA RTX 4090.
  • Excellent speech quality and voice cloning performance.
  • Web demo, Python API, and CLI.
  • Multilingual support: English, Chinese, French, German, and Spanish.

CuteTTS architecture
CuteTTS performance

Zero-shot voice-cloning performance

ModelParams.LibriSpeech test-clean WER (%) ↓LibriSpeech test-clean SIM ↑Seed-TTS EN WER (%) ↓Seed-TTS EN SIM ↑Seed-TTS ZH WER (%) ↓Seed-TTS ZH SIM ↑
MOSS‑TTS8B1.9867.71.8470.91.3777.0
Qwen3‑TTS1.7B2.3570.31.6671.40.9177.0
FireRedTTS‑21.5B4.3264.21.9566.51.1473.6
MOSS‑TTS‑Nano0.1B4.1048.44.6249.93.1364.3
F5‑TTS0.3B2.4266.01.8367.01.5676.0
ZipVoice0.1B2.0567.41.7069.71.4075.1
IndexTTS21.5B2.4770.02.2270.61.0276.5
CosyVoice 30.5B1.9969.72.0271.81.1678.0
VoxCPM22B3.0174.01.8475.30.9779.5
VibeVoice1.5B3.0468.91.1674.4
DiTAR0.6B2.3967.01.6973.51.0275.3
VibeVoice‑Realtime0.5B2.0069.52.0563.3
Pocket TTS0.1B1.5949.11.6350.7
CuteTTS0.2B2.1678.92.0476.51.4177.8
CuteTTS‑distill0.2B2.4176.82.0374.21.4775.6
cutetts
safetensors
text-to-speech
voice-cloning

Contributors

MinMinLiang

3 commits