5uck1ess/tts-bench

Speed and samples benchmark: for all types of text to speech (TTS) models on Windows/Linux/Mac.

324

stars

402

commits

Python

primary language

Sep 5, 2026

updated

README

tts-bench

Bench for local text-to-speech (TTS) models. Three lenses, on whatever hardware you put it on:

  • Speed — cold + warm TTFA (time to first audio), RTFx (realtime speed; higher = faster than realtime), memory, on CPU / CUDA / Apple Silicon
  • Listen — every model on every prompt, default voice + voice cloning, with inline audio players, so you can pick a model by ear
  • Scores — objective metrics per model: UTMOS (naturalness), WER (intelligibility), SIM (cloning fidelity), scored over the bench prompts via seed-tts-eval-style ASR + speaker-verification. Sortable, with a Default/Cloning toggle.

An objective quality score (NAQ) was prototyped but isn't part of the bench — the v2 features didn't track subjective ranking closely enough to publish, so it was pulled and is being redesigned separately. The bench measures speed; quality is by-ear via the Listen lens.


▶ Demos

5uck1ess.github.io/tts-bench — listen to every model, no install. Three lenses:

  • Listen — one consolidated gallery with an inline <audio> player for every model on every prompt, in default voice and voice cloning (each clone sits next to the reference it's imitating). Browse by prompt (compare all models on one sentence) or by model (audition one model across prompts); only one clip plays at a time. Audio is rig-independent, so each sample is sourced once from the highest-fidelity rig and tagged with where it came from. Quality, prosody, and artifacts are obvious in 5 seconds — benchmark tables can't show that.
  • Speed — per-rig leaderboards (Ryzen 9 9950X3D + RTX 5090, Apple M4, Ryzen + RTX 3090) with cold/warm TTFA, RTFx, and memory, sortable. Pick the box you actually own.
  • Scores — objective metrics per model (UTMOS naturalness, WER intelligibility, SIM cloning fidelity), with an interactive top-15 chart, sortable tables, and a Default/Cloning toggle. Human votes remain the preference ground truth; these are objective backstops.

Full per-rig reports (every model × prompt × device, plus by-prompt samples) are linked from the Archive.


🗳 Vote

Quality is subjective, so the ground truth is your ears. The companion TTS Voting Arena is a public, blind A/B listening test — two clips, no model names shown, pick the one that sounds better. No login, ~5 seconds a vote.

  • Default voice — which model sounds more natural?
  • Cloning — which clone better matches the reference voice?

Votes feed a live human-preference Elo leaderboard right there on the arena. This is where the "best sounding" and cloning calls above come from — every vote sharpens the ranking.

→ Vote now at the TTS Arena


Quick start

Requires uv and Python 3.11. ~10-15 min install. Disk for the full set is large: ~39 GB of per-model venvs in the repo, plus ~125 GB of model weights downloaded to your Hugging Face cache (~/.cache/huggingface, not the repo) — ~165 GB all-in. Individual models are far smaller, so installing a subset costs a fraction of that.

# Windows — everything, or just the models you want
.\install.ps1
.\install.ps1 kokoro,piper,miso
python bench.py
# macOS / Linux — everything, or just the models you want
./install.sh
./install.sh kokoro piper miso
python bench.py

Pass model names to install only those (names = the venvs/<name> slugs, which match the tables below — lowercase, e.g. kokoro, f5tts, chatterbox, miso). A few share one install: neutts covers NeuTTS Air + Nano, chatterbox both ChatterBox variants, vibevoice the 0.5B/1.5B, moss_tts both MOSS checkpoints, inflect both Inflect v2 sizes, qwentts the 1.7B Base + both 0.6B checkpoints, fish is Fish Speech 1.5. Add scoring (plus scoring_sim on Linux) for the objective-metrics venv. bench.py only runs models whose venv exists, so a partial install benches cleanly — install more models later by re-running with new names.

Interactive feel-test: python speak.py kokoro. One-shot A/B comparison: python compare.py "your phrase". See docs/architecture.md for the runner protocol and how to add a model.


TLDR (June 2026)

Fastest:

  • CPU (Ryzen 9 9950X3D, Windows): Piper — 107ms warm TTFA, 59× RTFx
  • CUDA (RTX 5090): Kokoro — 67ms warm TTFA, 104× RTFx
  • CPU + MPS (Apple M4, 16 GB): Piper — 208ms warm TTFA, 32× RTFx

Best sounding: No objective ranking right now — the NAQ score is paused pending redesign. Open the Demos site and use the Listen lens.

Best cloning — blind A/B votes (these measure voice-match preference, not intelligibility):

    1. OmniVoice — top on voice/accent match (24-1-3), but it can garble or drop words; a timbre-focused A/B vote doesn't penalize that, so read this as "best voice match," not "best overall clone." Audition it first — objective WER (the new Scores lens) is meant to catch exactly this gap.
    1. Echo-TTS — near-tied #1 (21-1-6), clean 44.1 kHz
    1. IndexTTS-2 — third (16-2-5), accent held

→ full per-rig results · → full cloning ranking


Models tracked (71)

Predefined voices

ModelParamsReleasedPredefinedCloningMultilingualSRExpressiveLicense
Audio8 TTS 0.6B (compiled)601MAug 2026✓ (11)44.1kApache 2.0
Audio8 TTS Preview 0.1B170MAug 2026✓ (11)44.1kAudio8 Community v1.0
Audio8 TTS Preview 0.6B601MJul 2026✓ (11)44.1kApache 2.0
Inflect-Micro v29.36MJul 2026✓ (1)— (en)24k2 knobsApache 2.0
Inflect-Nano v23.96MJul 2026✓ (1)— (en)24k2 knobsApache 2.0
KittenTTS Nano 0.1<100MAug 202524kApache 2.0
Kokoro82MDec 202424kApache 2.0
LFM2.5-Audio 1.5B1.5BDec 2025✓ (4)— (en)24kLFM Open v1.0
LuxTTS123MJan 202622.05kMIT
Magpie-TTS357MDec 2025✓ (9)22.05kemotion voices*NVIDIA OML
Maya13BOct 2025✓ (voice desc)24ktags + descApache 2.0
MeloTTS~52MFeb 2024✓ (4)44.1kMIT
Orpheus TTS3BMar 2025✓ (8)— (en)24ktagsApache 2.0
OuteTTS 1.0 1B~1BApr 2025✓ (12)44.1kCC-BY-NC-SA 4.0 + Llama 3.2
Parler-TTS Mini v1878MJun 2024✓ (voice desc)44.1kdesc*Apache 2.0
Piper~15MJan 202322.05kGPL-3.0
Qwen3-TTS 0.6B CustomVoice0.6BJan 2026✓ (9)24kApache 2.0
sanoTTS Amy1.46MJul 2026✓ (1)— (en)22.05kknobGPL-3.0
sanoTTS Heart-Nano294KSep 2026✓ (1)— (en)24kknobGPL-3.0
Scylla's Band~103MJul 2026✓ (10)✓ (4)24k6 knobsApache 2.0
Soprano 1.1 80M80MJan 202632kApache 2.0
Supertonic 399MMay 2026✓ (31)24ktagsMIT + OpenRAIL-M
Vaniq-Edge8.91MAug 2026✓ (1)— (en)24kknobMIT
VibeVoice Realtime 0.5B0.5BDec 202524kMIT
Voxtral 4B TTS4BNov 2025✓ (20)24kCC-BY-NC 4.0

Zero-shot cloning

ModelParamsReleasedPredefinedCloningMultilingualSRExpressiveLicense
Breeze TTS 23.47BAug 2026✓ (zh+en)24ktags + descBreezeBlue Research (NC)
ChatterBox1.2BApr 202524kknobMIT
ChatterBox Turbo744MDec 202524ktags*MIT
Coqui XTTS-v2750MOct 2023✓ (17)24kCPML (non-commercial)
CosyVoice 3 0.5B0.5BDec 202524kdescApache 2.0
Dia 1.6B-06261.6BJun 202544.1ktagsApache 2.0
dots.tts (soar)2BJun 2026✓ (24)48kApache 2.0
DramaBox3.3BApr 2026— (en)48kdescLTX-2 Community (NC)
Echo-TTS~2.8BDec 202544.1ktagsCC-BY-NC-SA 4.0
F5-TTS v1330MOct 2024✓ (zh+en)24kCC-BY-NC
Fish Speech 1.5~500MNov 202444.1kCC-BY-NC-SA 4.0
Fish Speech S2-Pro4BMar 2026✓ (80+)44.1ktagsResearch (non-commercial)
Higgs Audio v3 TTS4BJun 2026✓ (100)24ktagsResearch (NC)
IndexTTS-21.5BJun 2025✓ (zh+en)24kemo-ref + desc + knobApache 2.0
LongCat-AudioDiT 1B1.42BMar 2026✓ (zh+en)24kMIT
LongCat-AudioDiT 3.5B3.83BMar 2026✓ (zh+en)24kMIT
Mars5-TTS1.2BJun 202424kAGPL-3.0
MetaVoice-1B1.2BFeb 202448kApache 2.0
MioTTS 0.1B0.1BFeb 2026✓ (en+ja)44.1kFalcon-LLM
MioTTS 0.6B0.6BFeb 2026✓ (en+ja)44.1kApache 2.0
MiraTTS0.5BDec 202548kknobMIT
Miso TTS 8B8.2BMay 2026— (en)24kModified MIT
MOSS-TTS v1.08B (Qwen3)Feb 2026✓ (20)24kApache 2.0
MOSS-TTS v1.58B (Qwen3)May 2026✓ (31)24ktags (pause)Apache 2.0
MOSS-TTS-Nano100MApr 2026✓ (zh+en)48kApache 2.0
NeuTTS Air748MSep 202524kApache 2.0
NeuTTS Nano229MDec 2025✓ (4)24kApache 2.0
OmniVoice~1BMar 2026✓ (600+)24ktags*Apache 2.0 code / CC-BY-NC weights
OpenVoice v2~100MApr 202422.05kknobMIT
Pocket-TTS100MJan 2026✓ (6)24kApache 2.0
Qwen3-TTS 0.6B Base0.6BJan 202624kApache 2.0
Qwen3-TTS 1.7B Base1.7BJan 202624kApache 2.0
Qwen3-TTS 1.7B (CUDA-graph)1.7BJan 202624kMIT
Sesame CSM-1B1BMar 202524kApache 2.0
Sopro V2 Turbo120MAug 2026✓ (4)24kApache 2.0
Sopro V2 Turbo (streaming)120MAug 2026✓ (4)24kApache 2.0
Step-Audio-EditX3BOct 202524ktags + descApache 2.0
StyleTTS 2~148MJun 202324kknobMIT
VibeVoice 1.5B1.5BAug 202524kMIT
VibeVoice 7B7BSep 202524kMIT
VoxCPM22BApr 2026✓ (30)48kdescApache 2.0
WavTTS0.67BMay 2026✓ (zh+en)16kMIT code / CC-BY-NC 4.0 weights
ZipVoice123MJun 2025✓ (zh+en)24kApache 2.0
Zonos v0.11.6BFeb 202544.1kemo-ref + knobApache 2.0
Zonos28B (MoE, ~900M active)Jun 202644.1kknobApache 2.0

Expressive column — what explicit emotion/delivery control the model offers: tags = inline cues in the text itself ((laughs), [sigh], <laugh>); desc = natural-language style/emotion instructions; knob = numeric or preset parameter (exaggeration, style enum, pitch/speed); emo-ref = emotion conditioned on a separate reference clip or emotion vector; = none (for cloning models, expression simply follows the reference clip). * = caveat applies. Exact syntax, sources, and caveats per model: docs/expressive-control.md. Note the bench feeds every model the same plain prompts for fairness, so these features are not exercised in any score.

Full per-model gotchas + license details: docs/known-issues.md. Models considered but excluded: docs/considered.md.

Predefined vs Cloning. Predefined models have fixed/selectable speaker voices baked into the weights — they speak with no reference needed. Cloning (zero-shot) models have no voice of their own: they synthesize whatever voice you hand them as a reference clip at inference. Given no reference, a pure zero-shot model falls back to a bundled sample (this bench uses chris_hemsworth_15s.wav), so its "default voice" is just a clone of that clip. A few models do both (e.g. Voxtral has 20 presets and cloning).

Newly added, not yet benched. qwentts_06b_custom (Qwen3-TTS 0.6B CustomVoice, 9 preset timbres) is registered and smoke-tested but carries no speed or objective numbers yet — the bench + UTMOS/WER/SIM scoring pass runs on the Linux rig. The 0.6B Base checkpoint (qwentts_06b) is registered too, but like both 1.7B Base rows it is reference-only — Base has no model-native preset voice, so its no-reference run clones the house chris_hemsworth_15s.wav and it appears under Cloning, never Default. Its cloning lens is disabled by an intermittent decode runaway (identical input yielded 12.9 s of audio on one call and 655.3 s on the next) — see docs/known-issues.md.

Rig availability: Voxtral is Mac (MLX, preset-voice only) + Linux (vLLM, cloning); Fish S2-Pro / MetaVoice / Step-Audio-EditX / Higgs Audio v3 / dots.tts / Zonos2 / Orpheus / CosyVoice 3 are Linux-only (CUDA) — Higgs v3 is the one server-backed model (it runs via a Docker sgl-omni HTTP server, not an in-process load), dots.tts and Zonos2 require Linux-only dependencies/compiled CUDA kernels, and Orpheus (vLLM) + CosyVoice 3 (cu121 / torch 2.3.1) have no Blackwell-compatible Windows path; Echo-TTS and DramaBox are Windows + Linux (CUDA-only, no CPU/MPS; DramaBox needs ~18 GB VRAM). The rest run on Windows + Linux CUDA, most on CPU/MPS too. Per-rig speed + samples on the Demos site.


Voice cloning

48 of the 71 tracked models can clone a voice from a reference clip. Three reference formats supported (wav only / wav + transcript / HF-gated wav). Drop a reference into reference/, then python bench.py --reference reference/myvoice.wav.

Reference-format docs + the blind-vote cloning ranking (28 cloning models, human-preference A/B, frozen at 397 votes; the live arena board now has 738 cloning votes): docs/cloning.md.


Test hardware

MachineUsed for
Windows desktop (Ryzen 9 9950X3D / 128 GB / RTX 5090 32 GB)Windows CPU + CUDA bench rows
Linux workstation (Ryzen 9 5900XT / 64 GB / RTX 3090 24 GB, Ubuntu Server 24.04)Linux CPU + CUDA; the only rig that runs Fish-Speech S2 natively
Mac (Apple M4 / 16 GB / M4 GPU)Mac CPU + MPS bench rows

If you reproduce on different hardware, file an issue or PR with your results and we'll add a column.


Docs


License

MIT for the bench code in this repo. Each TTS model has its own license — see docs/known-issues.md for the full per-model table.


Support

If this bench saved you a weekend of writing your own:

Buy me a coffee at ko-fi.com

Contributors

5uck1ess

399 commits

estebanstifli

3 commits

5uck1ess/tts-bench

Speed and samples benchmark: for all types of text to speech (TTS) models on Windows/Linux/Mac.

324

stars

402

commits

Python

primary language

Sep 5, 2026

updated

README

tts-bench

Bench for local text-to-speech (TTS) models. Three lenses, on whatever hardware you put it on:

  • Speed — cold + warm TTFA (time to first audio), RTFx (realtime speed; higher = faster than realtime), memory, on CPU / CUDA / Apple Silicon
  • Listen — every model on every prompt, default voice + voice cloning, with inline audio players, so you can pick a model by ear
  • Scores — objective metrics per model: UTMOS (naturalness), WER (intelligibility), SIM (cloning fidelity), scored over the bench prompts via seed-tts-eval-style ASR + speaker-verification. Sortable, with a Default/Cloning toggle.

An objective quality score (NAQ) was prototyped but isn't part of the bench — the v2 features didn't track subjective ranking closely enough to publish, so it was pulled and is being redesigned separately. The bench measures speed; quality is by-ear via the Listen lens.


▶ Demos

5uck1ess.github.io/tts-bench — listen to every model, no install. Three lenses:

  • Listen — one consolidated gallery with an inline <audio> player for every model on every prompt, in default voice and voice cloning (each clone sits next to the reference it's imitating). Browse by prompt (compare all models on one sentence) or by model (audition one model across prompts); only one clip plays at a time. Audio is rig-independent, so each sample is sourced once from the highest-fidelity rig and tagged with where it came from. Quality, prosody, and artifacts are obvious in 5 seconds — benchmark tables can't show that.
  • Speed — per-rig leaderboards (Ryzen 9 9950X3D + RTX 5090, Apple M4, Ryzen + RTX 3090) with cold/warm TTFA, RTFx, and memory, sortable. Pick the box you actually own.
  • Scores — objective metrics per model (UTMOS naturalness, WER intelligibility, SIM cloning fidelity), with an interactive top-15 chart, sortable tables, and a Default/Cloning toggle. Human votes remain the preference ground truth; these are objective backstops.

Full per-rig reports (every model × prompt × device, plus by-prompt samples) are linked from the Archive.


🗳 Vote

Quality is subjective, so the ground truth is your ears. The companion TTS Voting Arena is a public, blind A/B listening test — two clips, no model names shown, pick the one that sounds better. No login, ~5 seconds a vote.

  • Default voice — which model sounds more natural?
  • Cloning — which clone better matches the reference voice?

Votes feed a live human-preference Elo leaderboard right there on the arena. This is where the "best sounding" and cloning calls above come from — every vote sharpens the ranking.

→ Vote now at the TTS Arena


Quick start

Requires uv and Python 3.11. ~10-15 min install. Disk for the full set is large: ~39 GB of per-model venvs in the repo, plus ~125 GB of model weights downloaded to your Hugging Face cache (~/.cache/huggingface, not the repo) — ~165 GB all-in. Individual models are far smaller, so installing a subset costs a fraction of that.

# Windows — everything, or just the models you want
.\install.ps1
.\install.ps1 kokoro,piper,miso
python bench.py
# macOS / Linux — everything, or just the models you want
./install.sh
./install.sh kokoro piper miso
python bench.py

Pass model names to install only those (names = the venvs/<name> slugs, which match the tables below — lowercase, e.g. kokoro, f5tts, chatterbox, miso). A few share one install: neutts covers NeuTTS Air + Nano, chatterbox both ChatterBox variants, vibevoice the 0.5B/1.5B, moss_tts both MOSS checkpoints, inflect both Inflect v2 sizes, qwentts the 1.7B Base + both 0.6B checkpoints, fish is Fish Speech 1.5. Add scoring (plus scoring_sim on Linux) for the objective-metrics venv. bench.py only runs models whose venv exists, so a partial install benches cleanly — install more models later by re-running with new names.

Interactive feel-test: python speak.py kokoro. One-shot A/B comparison: python compare.py "your phrase". See docs/architecture.md for the runner protocol and how to add a model.


TLDR (June 2026)

Fastest:

  • CPU (Ryzen 9 9950X3D, Windows): Piper — 107ms warm TTFA, 59× RTFx
  • CUDA (RTX 5090): Kokoro — 67ms warm TTFA, 104× RTFx
  • CPU + MPS (Apple M4, 16 GB): Piper — 208ms warm TTFA, 32× RTFx

Best sounding: No objective ranking right now — the NAQ score is paused pending redesign. Open the Demos site and use the Listen lens.

Best cloning — blind A/B votes (these measure voice-match preference, not intelligibility):

    1. OmniVoice — top on voice/accent match (24-1-3), but it can garble or drop words; a timbre-focused A/B vote doesn't penalize that, so read this as "best voice match," not "best overall clone." Audition it first — objective WER (the new Scores lens) is meant to catch exactly this gap.
    1. Echo-TTS — near-tied #1 (21-1-6), clean 44.1 kHz
    1. IndexTTS-2 — third (16-2-5), accent held

→ full per-rig results · → full cloning ranking


Models tracked (71)

Predefined voices

ModelParamsReleasedPredefinedCloningMultilingualSRExpressiveLicense
Audio8 TTS 0.6B (compiled)601MAug 2026✓ (11)44.1kApache 2.0
Audio8 TTS Preview 0.1B170MAug 2026✓ (11)44.1kAudio8 Community v1.0
Audio8 TTS Preview 0.6B601MJul 2026✓ (11)44.1kApache 2.0
Inflect-Micro v29.36MJul 2026✓ (1)— (en)24k2 knobsApache 2.0
Inflect-Nano v23.96MJul 2026✓ (1)— (en)24k2 knobsApache 2.0
KittenTTS Nano 0.1<100MAug 202524kApache 2.0
Kokoro82MDec 202424kApache 2.0
LFM2.5-Audio 1.5B1.5BDec 2025✓ (4)— (en)24kLFM Open v1.0
LuxTTS123MJan 202622.05kMIT
Magpie-TTS357MDec 2025✓ (9)22.05kemotion voices*NVIDIA OML
Maya13BOct 2025✓ (voice desc)24ktags + descApache 2.0
MeloTTS~52MFeb 2024✓ (4)44.1kMIT
Orpheus TTS3BMar 2025✓ (8)— (en)24ktagsApache 2.0
OuteTTS 1.0 1B~1BApr 2025✓ (12)44.1kCC-BY-NC-SA 4.0 + Llama 3.2
Parler-TTS Mini v1878MJun 2024✓ (voice desc)44.1kdesc*Apache 2.0
Piper~15MJan 202322.05kGPL-3.0
Qwen3-TTS 0.6B CustomVoice0.6BJan 2026✓ (9)24kApache 2.0
sanoTTS Amy1.46MJul 2026✓ (1)— (en)22.05kknobGPL-3.0
sanoTTS Heart-Nano294KSep 2026✓ (1)— (en)24kknobGPL-3.0
Scylla's Band~103MJul 2026✓ (10)✓ (4)24k6 knobsApache 2.0
Soprano 1.1 80M80MJan 202632kApache 2.0
Supertonic 399MMay 2026✓ (31)24ktagsMIT + OpenRAIL-M
Vaniq-Edge8.91MAug 2026✓ (1)— (en)24kknobMIT
VibeVoice Realtime 0.5B0.5BDec 202524kMIT
Voxtral 4B TTS4BNov 2025✓ (20)24kCC-BY-NC 4.0

Zero-shot cloning

ModelParamsReleasedPredefinedCloningMultilingualSRExpressiveLicense
Breeze TTS 23.47BAug 2026✓ (zh+en)24ktags + descBreezeBlue Research (NC)
ChatterBox1.2BApr 202524kknobMIT
ChatterBox Turbo744MDec 202524ktags*MIT
Coqui XTTS-v2750MOct 2023✓ (17)24kCPML (non-commercial)
CosyVoice 3 0.5B0.5BDec 202524kdescApache 2.0
Dia 1.6B-06261.6BJun 202544.1ktagsApache 2.0
dots.tts (soar)2BJun 2026✓ (24)48kApache 2.0
DramaBox3.3BApr 2026— (en)48kdescLTX-2 Community (NC)
Echo-TTS~2.8BDec 202544.1ktagsCC-BY-NC-SA 4.0
F5-TTS v1330MOct 2024✓ (zh+en)24kCC-BY-NC
Fish Speech 1.5~500MNov 202444.1kCC-BY-NC-SA 4.0
Fish Speech S2-Pro4BMar 2026✓ (80+)44.1ktagsResearch (non-commercial)
Higgs Audio v3 TTS4BJun 2026✓ (100)24ktagsResearch (NC)
IndexTTS-21.5BJun 2025✓ (zh+en)24kemo-ref + desc + knobApache 2.0
LongCat-AudioDiT 1B1.42BMar 2026✓ (zh+en)24kMIT
LongCat-AudioDiT 3.5B3.83BMar 2026✓ (zh+en)24kMIT
Mars5-TTS1.2BJun 202424kAGPL-3.0
MetaVoice-1B1.2BFeb 202448kApache 2.0
MioTTS 0.1B0.1BFeb 2026✓ (en+ja)44.1kFalcon-LLM
MioTTS 0.6B0.6BFeb 2026✓ (en+ja)44.1kApache 2.0
MiraTTS0.5BDec 202548kknobMIT
Miso TTS 8B8.2BMay 2026— (en)24kModified MIT
MOSS-TTS v1.08B (Qwen3)Feb 2026✓ (20)24kApache 2.0
MOSS-TTS v1.58B (Qwen3)May 2026✓ (31)24ktags (pause)Apache 2.0
MOSS-TTS-Nano100MApr 2026✓ (zh+en)48kApache 2.0
NeuTTS Air748MSep 202524kApache 2.0
NeuTTS Nano229MDec 2025✓ (4)24kApache 2.0
OmniVoice~1BMar 2026✓ (600+)24ktags*Apache 2.0 code / CC-BY-NC weights
OpenVoice v2~100MApr 202422.05kknobMIT
Pocket-TTS100MJan 2026✓ (6)24kApache 2.0
Qwen3-TTS 0.6B Base0.6BJan 202624kApache 2.0
Qwen3-TTS 1.7B Base1.7BJan 202624kApache 2.0
Qwen3-TTS 1.7B (CUDA-graph)1.7BJan 202624kMIT
Sesame CSM-1B1BMar 202524kApache 2.0
Sopro V2 Turbo120MAug 2026✓ (4)24kApache 2.0
Sopro V2 Turbo (streaming)120MAug 2026✓ (4)24kApache 2.0
Step-Audio-EditX3BOct 202524ktags + descApache 2.0
StyleTTS 2~148MJun 202324kknobMIT
VibeVoice 1.5B1.5BAug 202524kMIT
VibeVoice 7B7BSep 202524kMIT
VoxCPM22BApr 2026✓ (30)48kdescApache 2.0
WavTTS0.67BMay 2026✓ (zh+en)16kMIT code / CC-BY-NC 4.0 weights
ZipVoice123MJun 2025✓ (zh+en)24kApache 2.0
Zonos v0.11.6BFeb 202544.1kemo-ref + knobApache 2.0
Zonos28B (MoE, ~900M active)Jun 202644.1kknobApache 2.0

Expressive column — what explicit emotion/delivery control the model offers: tags = inline cues in the text itself ((laughs), [sigh], <laugh>); desc = natural-language style/emotion instructions; knob = numeric or preset parameter (exaggeration, style enum, pitch/speed); emo-ref = emotion conditioned on a separate reference clip or emotion vector; = none (for cloning models, expression simply follows the reference clip). * = caveat applies. Exact syntax, sources, and caveats per model: docs/expressive-control.md. Note the bench feeds every model the same plain prompts for fairness, so these features are not exercised in any score.

Full per-model gotchas + license details: docs/known-issues.md. Models considered but excluded: docs/considered.md.

Predefined vs Cloning. Predefined models have fixed/selectable speaker voices baked into the weights — they speak with no reference needed. Cloning (zero-shot) models have no voice of their own: they synthesize whatever voice you hand them as a reference clip at inference. Given no reference, a pure zero-shot model falls back to a bundled sample (this bench uses chris_hemsworth_15s.wav), so its "default voice" is just a clone of that clip. A few models do both (e.g. Voxtral has 20 presets and cloning).

Newly added, not yet benched. qwentts_06b_custom (Qwen3-TTS 0.6B CustomVoice, 9 preset timbres) is registered and smoke-tested but carries no speed or objective numbers yet — the bench + UTMOS/WER/SIM scoring pass runs on the Linux rig. The 0.6B Base checkpoint (qwentts_06b) is registered too, but like both 1.7B Base rows it is reference-only — Base has no model-native preset voice, so its no-reference run clones the house chris_hemsworth_15s.wav and it appears under Cloning, never Default. Its cloning lens is disabled by an intermittent decode runaway (identical input yielded 12.9 s of audio on one call and 655.3 s on the next) — see docs/known-issues.md.

Rig availability: Voxtral is Mac (MLX, preset-voice only) + Linux (vLLM, cloning); Fish S2-Pro / MetaVoice / Step-Audio-EditX / Higgs Audio v3 / dots.tts / Zonos2 / Orpheus / CosyVoice 3 are Linux-only (CUDA) — Higgs v3 is the one server-backed model (it runs via a Docker sgl-omni HTTP server, not an in-process load), dots.tts and Zonos2 require Linux-only dependencies/compiled CUDA kernels, and Orpheus (vLLM) + CosyVoice 3 (cu121 / torch 2.3.1) have no Blackwell-compatible Windows path; Echo-TTS and DramaBox are Windows + Linux (CUDA-only, no CPU/MPS; DramaBox needs ~18 GB VRAM). The rest run on Windows + Linux CUDA, most on CPU/MPS too. Per-rig speed + samples on the Demos site.


Voice cloning

48 of the 71 tracked models can clone a voice from a reference clip. Three reference formats supported (wav only / wav + transcript / HF-gated wav). Drop a reference into reference/, then python bench.py --reference reference/myvoice.wav.

Reference-format docs + the blind-vote cloning ranking (28 cloning models, human-preference A/B, frozen at 397 votes; the live arena board now has 738 cloning votes): docs/cloning.md.


Test hardware

MachineUsed for
Windows desktop (Ryzen 9 9950X3D / 128 GB / RTX 5090 32 GB)Windows CPU + CUDA bench rows
Linux workstation (Ryzen 9 5900XT / 64 GB / RTX 3090 24 GB, Ubuntu Server 24.04)Linux CPU + CUDA; the only rig that runs Fish-Speech S2 natively
Mac (Apple M4 / 16 GB / M4 GPU)Mac CPU + MPS bench rows

If you reproduce on different hardware, file an issue or PR with your results and we'll add a column.


Docs


License

MIT for the bench code in this repo. Each TTS model has its own license — see docs/known-issues.md for the full per-model table.


Support

If this bench saved you a weekend of writing your own:

Buy me a coffee at ko-fi.com

Contributors

5uck1ess

399 commits

estebanstifli

3 commits

Languages

Python

82.4%

Shell

8.7%

PowerShell

7.2%

HTML

1.7%