List of open-source TTS, voice cloning, and music generation models
494
37 commits
updated Sep 10, 2026
A curated list of open-source Text-to-Speech (TTS) and voice cloning models. Models are sorted by release date (newest first).
| Model | Voice Cloning | ASR | Languages | Streaming | License |
|---|---|---|---|---|---|
| AuK | ✅ | ❌ | — | — | ![MIT][license-mit] |
| AuK-Flash | ✅ | ❌ | — | — | ![MIT][license-mit] |
| rumik-oss 1 | ❌ | ❌ | 22 Indic languages + English | — | ![Other][license-other] |
| Irodori-TTS-v4.1-Anime | — | ❌ | Japanese | — | ![MIT][license-mit] |
| ICE-012 Audio | ✅ | ❌ | 590 | ✅ | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| TontaubeV1 | ✅ | ❌ | 7 | ✅ | ![Other][license-other] |
| Breeze TTS 2 | ✅ | ❌ | 2 | ✅ | ![Other][license-other] |
| Sopro v2 Turbo | ✅ | ❌ | 4 | ✅ | ![Apache 2.0][license-apache-2.0] |
| CuteTTS | ✅ | ❌ | 5 | ✅ | ![Apache 2.0][license-apache-2.0] |
| Rynsan TTS | — | ❌ | Khasi, Garo, Pnar, English, Hindi | — | ![CC BY 4.0][license-cc-by-4.0] |
| Audio8 TTS Preview 0.1B | ✅ | ❌ | 8 | — | ![Other][license-other] |
| Kiseki-TTS | ❌ | ✅ | Japanese | — | ![MIT][license-mit] |
| FireRedTTS3 | ✅ | ❌ | 24 | — | ![Apache 2.0][license-apache-2.0] |
| Audio8-TTS-Preview-0.6b | ✅ | ❌ | Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish | ❌ | ![Apache 2.0][license-apache-2.0] |
| NeuTTS-2E | ❌ | ❌ | English | ✅ | ![Other][license-other] |
| Scylla's Band | ❌ | ❌ | en_us, en_gb, es, it | ✅ | ![Apache 2.0][license-apache-2.0] |
| sanoTTS | ❌ | ❌ | English, Nepali, Hindi, Vietnamese, Indonesian, Chinese | ❌ | ![Other][license-other] |
| FreyaTTS | ❌ | ❌ | Turkish | ❌ | ![Apache 2.0][license-apache-2.0] |
| Inflect-Nano-v2 | ❌ | ❌ | English | ❌ | ![Apache 2.0][license-apache-2.0] |
| Gepard | ✅ | ❌ | English, Spanish, Portuguese, Dutch | ✅ | ![Apache 2.0][license-apache-2.0] |
| Higgs Audio v3 TTS | ✅ | ❌ | 102 | ✅ | ![Research Only][license-research-only] |
| dots.tts | ✅ | ❌ | Multilingual | ❌ | ![Apache 2.0][license-apache-2.0] |
| Confucius4-TTS | ✅ | ❌ | 14 | ❌ | ![Apache 2.0][license-apache-2.0] |
| WavTTS | ✅ | ❌ | English, Chinese | ❌ | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| MOSS-TTS | ✅ | ❌ | 31 | ✅ | ![Apache 2.0][license-apache-2.0] |
| VoxFlash-TTS | ✅ | ❌ | Chinese, English | ❌ | ![Apache 2.0][license-apache-2.0] |
| Miso TTS | ✅ | ❌ | English | ❌ | ![MIT][license-mit] |
| Raon-OpenTTS-1B | ✅ | ❌ | English | — | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| OronTTS | ✅ | ❌ | Mongolian, Kazakh | ❌ | ![MIT][license-mit] |
| Supertonic 3 | ✅ | ❌ | 31 | ✅ | ![OpenRAIL-M][license-openrail-m] |
| Scenema Audio | ✅ | ❌ | English, German, French, Spanish, Italian, Portuguese, Japanese, Chinese, Korean, Russian, Arabic, Hindi, Swahili | ❌ | ![Other][license-other] |
| Dramabox | ✅ | ❌ | English | ❌ | ![Other][license-other] |
| Sarashina2.2-TTS | ✅ | ❌ | Japanese, English | ❌ | ![Research Only][license-research-only] |
| LongCat-AudioDiT | ✅ | ❌ | Chinese, English | ❌ | ![MIT][license-mit] |
| SILMA TTS | ✅ | ❌ | Arabic, English | — | ![Apache 2.0][license-apache-2.0] |
| Fish Audio S2 Pro | ✅ | ❌ | 80+ | ✅ | ![Research Only][license-research-only] |
| LongCat-Next | ✅ | ✅ | Chinese, English | ✅ | ![MIT][license-mit] |
| Voxtral-4B-TTS | ✅ | ❌ | 9 | ✅ | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Blue (Light Blue) TTS | ✅ | ❌ | Hebrew, English, Spanish, Italian, German | ❌ | ![MIT][license-mit] |
| KittenTTS | ✅ | ❌ | English, Multiple | ✅ | ![Apache 2.0][license-apache-2.0] |
| Ming-omni-tts | ✅ | ❌ | Chinese, English | ❌ | ![Apache 2.0][license-apache-2.0] |
| SoulX-Singer | ✅ | ❌ | Mandarin, English, Cantonese | ✅ | ![Apache 2.0][license-apache-2.0] |
| SoproTTS | ✅ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Qwen3-TTS | ✅ | ❌ | 10 | ✅ | ![Apache 2.0][license-apache-2.0] |
| TADA | ❌ | ❌ | English | ❌ | ![Other][license-other] |
| Irodori-TTS-500M-v2 | ✅ | ❌ | Japanese | ❌ | ![MIT][license-mit] |
| KugelAudio | ✅ | ❌ | 23 European languages | ✅ | ![MIT][license-mit] |
| LEMAS-TTS | ✅ | ❌ | 10 | ❌ | ![Apache 2.0][license-apache-2.0] |
| MioTTS-2.6B | ✅ | ❌ | English, Japanese | ✅ | ![LFM][license-lfm] |
| MOSS-TTS-Nano | ✅ | ❌ | 20 | ✅ | ![Apache 2.0][license-apache-2.0] |
| NeuTTS | ✅ | ❌ | English, Spanish, German, French | ✅ | ![Apache 2.0][license-apache-2.0] |
| OmniVoice | ✅ | ❌ | 600+ | ❌ | ![Apache 2.0][license-apache-2.0] |
| T5Gemma-TTS | ✅ | ❌ | English, Chinese, Japanese | ❌ | ![MIT][license-mit] |
| TinyTTS | ❌ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| VoxCPM2 | ✅ | ❌ | 30 | ✅ | ![Apache 2.0][license-apache-2.0] |
| Soprano | ❌ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| GLM-TTS | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Echo-TTS | ✅ | ❌ | English | ❌ | ![MIT][license-mit] |
| VibeVoice-Realtime | ✅ | ❌ | Multilingual | ✅ | ![MIT][license-mit] |
| Fun-CosyVoice 3.0 | ✅ | ❌ | 9 + 18+ Chinese dialects | ✅ | ![Apache 2.0][license-apache-2.0] |
| LFM2-Audio-1.5B | ✅ | ✅ | English | ✅ | ![LFM][license-lfm] |
| Marvis-TTS | ✅ | ❌ | English, French, German | ✅ | ![Apache 2.0][license-apache-2.0] |
| IndexTTS2 | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Maya1 | ✅ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Step-Audio-EditX | ✅ | ❌ | Mandarin, English, Sichuanese, Cantonese, Japanese, Korean | ✅ | ![Apache 2.0][license-apache-2.0] |
| KaniTTS | ❌ | ❌ | English, German, Chinese, Korean, Arabic, Spanish | ✅ | ![LFM][license-lfm] |
| VibeVoice-Finetuning | ❌ | ❌ | — | ❌ | ![MIT][license-mit] |
| VoxCPM | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| FireRedTTS2 | ✅ | ❌ | EN, ZH, JP, KO, FR, DE, RU | ✅ | ![Apache 2.0][license-apache-2.0] |
| Audio Flamingo 3 (AF3) / Audio Flamingo Next | ❌ | ✅ | Multi-lingual | ✅ | ![Apache 2.0][license-apache-2.0] |
| ZipVoice | — | — | Chinese, English | — | ![Apache 2.0][license-apache-2.0] |
| Fish Speech | ✅ | ❌ | 8 | ✅ | ![Apache 2.0][license-apache-2.0] |
| Chatterbox | ✅ | ❌ | 23+ | ❌ | ![MIT][license-mit] |
| Orpheus-TTS | ✅ | ❌ | Multilingual | ✅ | ![Apache 2.0][license-apache-2.0] |
| MegaTTS3 | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Spark-TTS | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Step-Audio | ✅ | ✅ | Chinese, English, Japanese | ✅ | ![Apache 2.0][license-apache-2.0] |
| Kokoro-82M | ✅ | ❌ | 8 | ✅ | ![Apache 2.0][license-apache-2.0] |
| KokoClone | ✅ | ❌ | 7 | ✅ | ![Apache 2.0][license-apache-2.0] |
| LuxTTS | ✅ | ❌ | - | ✅ | ![Apache 2.0][license-apache-2.0] |
| MiMo-Audio | ✅ | ✅ | Multi-lingual | ✅ | ![Apache 2.0][license-apache-2.0] |
| SoulX-Podcast | ✅ | ❌ | Mandarin, English, Cantonese, Sichuanese, Henanese | ✅ | ![Apache 2.0][license-apache-2.0] |
| VieNeu-TTS | ✅ | ❌ | Vietnamese | ✅ | ![Apache 2.0][license-apache-2.0] |
| Dia | ✅ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| MeloTTS | ❌ | ❌ | English, Spanish, French, Chinese, Japanese, Korean | ❌ | ![MIT][license-mit] |
| Kimi-Audio | ✅ | ✅ | Multi-lingual | ✅ | ![MIT][license-mit] ![Apache 2.0][license-apache-2.0] |
| eSpeak-NG | ❌ | ❌ | 100+ | ✅ | ![Other][license-other] |
Description: AuK is a 1.5B foundation model from Tencent for speech generation and editing, trained on millions of hours of diverse audio data. Through a single natural-language instruction interface it unifies an unusually broad task set: zero-shot TTS (speak text in the reference voice) and instruct TTS (voice from a description alone, no reference), content editing (rewrite what is said; even lyric editing that preserves melody and voice), acoustic editing (pitch by semitones, speed, volume), paralinguistic editing (emotion, timbre, de-accent, nonverbal sounds, whisper conversion), and enhancement & separation (denoise/dereverberate, speech separation, music/vocal separation, target-speaker extraction). Architecture: a diffusion transformer with layer-fusion weights, conditioned by a Qwen2.5-Omni-3B MLLM encoder and a separate VAE (loaded at runtime). Day-0 SGLang-Omni serving support, Gradio and ComfyUI integrations, and a task Cookbook are provided. Released under MIT.
Release Date: September 9, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| License | ![MIT][license-mit] |
| Parameters | 1.5B |
| Architecture | diffusion transformer + layer fusion, Qwen2.5-Omni-3B MLLM encoder, separate VAE |
| Variants | AuK (this, base) + AuK-Flash (distilled, 4-step inference) |
| Editing | content, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion |
| Enhancement Separation | speech enhancement, speech separation, music separation, target speaker extraction |
| Deployment | SGLang-Omni (day-0), Gradio, ComfyUI |
Features: Unifies generation and the full editing/enhancement/separation spectrum in one instruction-following model — most systems pick one lane (TTS, or editing, or separation); AuK does zero-shot + instruct TTS, lyric rewriting with melody preservation, emotion/timbre/de-accent/whisper paralinguistic edits, and source separation through the same natural-language interface. A diffusion transformer with layer fusion, distilled into a 4-step AuK-Flash variant for fast inference.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| ComfyUI-AuK | ComfyUI node | ComfyUI-AuK |
· · · · · · · · · · · · · ·
Description: AuK-Flash is the distilled variant of AuK, Tencent's 1.5B foundation model for speech generation and editing, optimized for fast 4-step inference. It exposes the same natural-language instruction interface as the base model: zero-shot TTS (reference voice) and instruct TTS (voice description, no reference), content and lyric editing, pitch/speed/volume acoustic edits, emotion/timbre/de-accent/nonverbal/whisper paralinguistic edits, plus speech enhancement and speech/music/target-speaker separation. Architecture matches the base: diffusion transformer with layer-fusion weights, Qwen2.5-Omni-3B MLLM encoder, and a separate runtime-loaded VAE. Released under MIT.
Release Date: September 9, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| License | ![MIT][license-mit] |
| Parameters | 1.5B |
| Architecture | diffusion transformer + layer fusion (distilled to 4 inference steps), Qwen2.5-Omni-3B MLLM encoder, separate VAE |
| Base Model | tencent/AuK |
| Editing | content, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion |
| Enhancement Separation | speech enhancement, speech separation, music separation, target speaker extraction |
| Deployment | SGLang-Omni, Gradio, ComfyUI |
Features: Distills the AuK foundation model's diffusion transformer down to 4 inference steps, making the full generate-and-edit capability set (including enhancement and separation) practical for interactive use — traded against the base model's maximum quality, with the two variants loadable side-by-side from the same codebase.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: rumik-oss 1 is a 3B multilingual text-to-speech model from rumik ai, trained on fewer than 70,000 hours of speech while performing competitively with existing TTS models. It covers 22 Indic languages in their native scripts and romanized forms plus English, supporting both single-language and code-switched synthesis. Delivery is conditioned via <description="..."> tags (tone, accent, pace), with inline vocalization control (<laugh>, <chuckle>, <sigh>). It extends CohereLabs/tiny-aya-fire with discrete speech tokens from the mimi codec: following the flattened codec-token formulation used in llama-mimi, text conditioning and audio generation share a single autoregressive sequence, predicting eight codebook tokens per frame before advancing, with the frozen mimi decoder reconstructing the 24 kHz waveform. The model ships with 4 fixed voices (Ira, Aisha, Siya, Zoya) that perform equally well across all 22 languages; there is no zero-shot voice cloning. Licensed under Cohere's CC-BY-NC-4.0 with acceptable-use addendum (research and non-commercial use only).
Release Date: September 6, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | 22 Indic languages + English (native scripts and romanized; code-switching supported) |
| License | ![Other][license-other] |
| Parameters | 3B (3,381,533,697 BF16) |
| Architecture | CohereLabs/tiny-aya-fire backbone + flattened mimi codec tokens (8 codebooks/frame, llama-mimi formulation) |
| Audio Codec | kyutai/mimi (frozen decoder), 24 kHz output |
| Pronunciation | ✅ |
| Highlights | description-conditioned delivery (<description> tags), inline vocalizations (<laugh>/<chuckle>/<sigh>) |
| Variants | rumik-oss-1 (post-trained), rumik-oss-1-base (speaker-conditioned pre-post-training) |
Features: Brings competitive multilingual TTS to 22 Indic languages with under 70k training hours, using a flattened mimi codec-token formulation (single autoregressive sequence for text conditioning + audio) on the tiny-aya-fire backbone. Code-switched synthesis, description-conditioned delivery, and inline vocalization tags are first-class capabilities, and its 4 voices perform equally well across all 22 languages — unusual, as most TTS voices are language-specific.
Links: ![HuggingFace][link-huggingface] ![Blog][link-blog] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Irodori-TTS-v4.1-Anime is a Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small using anime-style speech data. Because the base model's annotation pipeline is not publicly documented, the fine-tuning data was annotated independently — so caption conditioning and emoji controls may behave differently from the base model. The full-precision checkpoint (0.8B params, F32) ships at the repository root, with quantized variants (int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic) in subdirectories. It follows the base model's MIT License and ethical restrictions; inference uses the original Irodori-TTS repository.
Release Date: September 4, 2026
| Feature | Value |
|---|---|
| Voice Cloning | — |
| Asr | ❌ |
| Languages | Japanese |
| License | ![MIT][license-mit] |
| Parameters | ~0.8B (766,052,385 F32) |
| Architecture | Irodori-TTS (Aratako) fine-tune; caption-conditioned with emoji controls |
| Base Model | Aratako/Irodori-TTS-v4.1-Small |
| Variants | int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic |
Features: A community fine-tune that ports the Irodori-TTS line into the anime-voice domain using an independently built annotation pipeline (since the base model's is undocumented), and ships the result with five ready-made quantization variants (int4/int8/fp8) for efficient inference.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: ICE-012 Audio is a multilingual text-to-speech model from DarkPs (a FanuonAI organization) with streaming output and reference-based voice cloning. Its defining trait is breadth of language coverage — 590 language names/variants are accepted (name or 2–3-letter ID, with a language-agnostic fallback), including 13 Arabic dialects ("Lahgtna" variants) alongside the full ISO list. It introduces an active acoustic adapter — conditioning codec embeddings before the backbone and refining hidden states after it. Voice is controllable along six axes: gender (male/female), age (child → elderly), pitch (5 levels), accent (10 English accents), style (e.g. whisper), and speed (0.5–2.0×), plus an --auto-voice mode where the model picks a voice automatically. The checkpoint is ~714M parameters (F16) and runs via transformers with trust_remote_code=True. Released under CC BY-NC 4.0.
Release Date: August 29, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 590 names/variants (incl. 13 Arabic Lahgtna dialects; language-agnostic fallback) |
| Streaming | ✅ |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Parameters | ~714M (714,409,993 F16) |
| Architecture | causal LM with active acoustic adapter (conditions codec embeddings pre-backbone, refines hidden states post-backbone) |
| Voice Controls | gender, age, pitch, accent, style, speed; auto-voice mode |
Features: The active acoustic adapter wraps the backbone on both sides — conditioning codec embeddings before it and refining hidden states after — while a six-axis voice-control space (gender/age/pitch/accent/style/speed plus auto-voice) and 590-language coverage make it one of the broadest single-checkpoint TTS releases for dialect and minority-language synthesis.
Links: ![HuggingFace][link-huggingface] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: TontaubeV1 is a multilingual text-to-speech model from TontaubeAI (craitech) designed for expressive voice cloning, long-form generation, and low-latency streaming. Its release contains four causal codebook predictors: CB0 generates semantic audio and duration from text, while progressively smaller CB1–CB3 add acoustic detail. CB0 uses a Qwen3-1.7B-derived transformer trunk and CB1–CB3 progressively shallower Qwen3-0.6B-derived trunks, each with a two-layer audio-token head. The four output streams are decoded with DualCodec, and the inference path uses VibeVoice's acoustic encoder/decoder for continuous reconstruction and streaming. It ships bundled synthetic voices plus zero-shot cloning from up to 60 s of reference audio, with public speaking styles audiobook, conversational, and agentic. Released under the Tontaube Community Model License 1.0, which is explicitly not open-source.
Release Date: August 26, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 7 (English, German primary; Spanish, French, Italian, Dutch, Portuguese secondary) |
| Streaming | ✅ |
| License | ![Other][license-other] |
| Parameters | ~2.87B (2,873,962,498; CB0 1.83B + CB1 449M + CB2 327M + CB3 269M) |
| Architecture | 4-stage Qwen3-derived codebook cascade (CB0–CB3) + DualCodec + VibeVoice decode |
| Styles | audiobook, conversational, agentic |
Features: The four-stage codebook cascade (CB0 semantic+duration → CB1–CB3 progressive acoustic refinement) lets a single multilingual model deliver expressive, long-form, low-latency speech with strong zero-shot cloning. On the 1,088 English zero-shot Seed-TTS examples it posts 1.66% mean utterance-level WER (measured with Whisper large-v3 at semantic temperature 0.6), and the RTX-5090 streaming path reaches ~200 ms to first encoded audio.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Breeze TTS 2 is an open-weight text-to-speech model from BreezeBlue / RESONIA built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following supports reference-free voice design (create a voice from a text description) and reference-guided voice direction (clone a voice while steering tone, emotion, pace, delivery), alongside standard reference-audio voice cloning. Ultra-low-latency streaming reaches 0.32 RTF (≈3.1× real time with the warmed-up fast path) and under 40 ms time-to-first-audio on an NVIDIA H100, emitting 24 kHz PCM. Source code is Apache-2.0; model weights are governed by the BreezeBlue Research and Non-Commercial License (commercial use needs written authorization from RESONIA).
Release Date: August 25, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 2 (English, Chinese) |
| Streaming | ✅ |
| License | ![Other][license-other] |
| Parameters | 3B (3,466,363,713) |
| Architecture | seq2seq backbone + depth decoder + codec with CUDA-graph fast path (no named backbone disclosed) |
| Highlights | #1 open-weight on Artificial Analysis TTS leaderboard; vocal events inline (laugh/cough) |
Features: Pairs natural-language voice control with real-time streaming: a single model handles reference-free voice design (no reference audio needed) and voice direction (clone + steer prosody), and ships a CUDA-graph fast path that hits sub-40 ms TTFA at ~3.1× real time on H100 — open-weight quality that the authors claim exceeds frontier proprietary TTS.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| ComfyUI-Breeze-TTS-2 | ComfyUI node | ComfyUI-Breeze-TTS-2 |
· · · · · · · · · · · · · ·
Description: Sopro (Portuguese for "breath") is a lightweight voice-cloning text-to-speech family. This repo ships sopro-v2-turbo, a 120M-parameter open model that streams and runs comfortably on a laptop CPU or in the browser (ONNX runtime), reaching SOTA-level intelligibility against much larger systems. It supports zero-shot voice cloning from 5–20 s of reference audio, four languages (English, European Portuguese, French, German), and a streaming path with ~300 ms time-to-first-audio on a laptop CPU (0.24 RTF offline / 0.21 RTF streaming on an M3 CPU, 0.07 RTF on H100). Released under Apache-2.0.
Release Date: August 25, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 4 (English, European Portuguese, French, German) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 120M (121,574,193) |
| Deployment | in-browser ONNX runtime; int8 AR weights on CPU; causal vocoder |
| Architecture | autoregressive TTS + chunked-attention streaming path + causal vocoder (F5-TTS/CosyVoice/Vocos lineage acknowledged) |
Features: Packs SOTA-level intelligibility into a 120M footprint that runs in the browser or on a laptop CPU, with a chunked-attention + causal-vocoder streaming path (~300 ms TTFA) — making zero-shot multilingual voice cloning practical for on-device and edge deployment rather than GPU-only serving.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: CuteTTS is a lightweight (~230M-parameter) continuous autoregressive TTS model from OPPO that models continuous latents rather than discrete codec tokens, running efficiently on GPUs, CPUs, and Apple silicon. It delivers ultra-low latency — ~40 ms to the first audio chunk and ~9× real-time throughput on an RTX 4090 — with strong speech quality and zero-shot voice cloning (best-in-comparison 78.9 SIM on LibriSpeech test-clean). Multilingual support covers English, Chinese, French, German, and Spanish. A distilled variant (CuteTTS-distill) trades slight quality for further efficiency. Ships with a web demo, Python API, and CLI.
Release Date: August 24, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 5 (English, Chinese, French, German, Spanish) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | ~230M |
| Architecture | continuous autoregressive modeling of latents + speaker encoder + audio VAE (discrete-codec-free design) |
| Variants | CuteTTS, CuteTTS-distill |
Features: Autoregressively models continuous latent audio representations instead of discrete codec tokens, eliminating codebook-related artifacts and quantization loss at only ~230M parameters. Combined with a lightweight speaker encoder and audio VAE, this yields best-of-class speaker similarity among compared open models and ~40 ms first-chunk latency while remaining practical for CPU/Apple-silicon inference.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Rynsan TTS is a multilingual text-to-speech model that extends k2-fsa/OmniVoice to support Khasi, Garo, and Pnar — languages of Meghalaya, India that have historically had limited representation in modern speech technology. Rather than building a system from scratch, Rynsan retains the multilingual capabilities of the base model while adding speech data for these low-resource Khasic languages. Developed under the Tynrai AI initiative, its broader goal is accessible speech technology for the diverse languages and dialects of Meghalaya, supporting their preservation and use in voice-based applications. A live demo is available at ri.tynrai.in/demo. The repository is gated (manual access approval).
Release Date: August 22, 2026
| Feature | Value |
|---|---|
| Asr | ❌ |
| Languages | 5+ (English, Hindi + extension languages Khasi kha, Garo grt, Pnar pbv; base OmniVoice supports more) |
| License | ![CC BY 4.0][license-cc-by-4.0] |
| Parameters | ~0.61B |
| Architecture | OmniVoice (k2-fsa) multilingual TTS, extended fine-tune |
| Base Model | k2-fsa/OmniVoice |
| Developer | Toiar / Tynrai AI |
Features: Extends a modern multilingual TTS foundation to three substantially under-resourced Khasic languages — a rare production-oriented entry for indigenous-language speech tech, aimed at accessibility, education, and language preservation rather than benchmark leadership.
Links: ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Audio8 TTS Preview 0.1B is the smallest release in the Audio8 TTS family ("the smallest zero-shot TTS worth running"): a ~170M-parameter generative model plus a separate ~120M-parameter codec decoder, making the complete audio generation stack much smaller than most modern multilingual TTS systems. It supports speech generation and zero-shot voice cloning (reference audio + matching transcript). Primary languages are Chinese and English, with German, Spanish, French, Italian, Japanese, and Korean as experimental/multilingual-evaluation targets. Released under the custom Audio8 Community License v1.0: non-commercial use is free, and commercial use is free only for entities with annual revenue under US$2M.
Release Date: August 19, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 8 (Chinese + English primary; de/es/fr/it/ja/ko experimental) |
| License | ![Other][license-other] |
| Parameters | ~0.17B main model (+ ~120M codec decoder) |
| Architecture | Audio8 Falcon H1 DualAR — slow AR (semantic tokens) + fast AR (codec codebooks), 10 codebooks × 4096 entries |
| Audio Codec | bundled 44.1 kHz neural codec (~21.5 frames/s) |
| Context | up to 2,048 packed text/audio positions |
| Variants | 0.1B (this), 0.6B |
Features: Packs practical zero-shot cloning into a ~170M-parameter model using an Falcon-H1-derived DualAR design (slow AR predicts semantic tokens per frame; fast AR predicts the frame's 10 codebooks conditioned on the slow hidden state). On Seed-TTS it posts EN WER 1.662 at only ~0.17B — within reach of 4B+ systems — and ships with its own 44.1 kHz codec so no external codec checkpoint is needed.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Kiseki-TTS is a small, fast Japanese text-to-speech model from telecomadm1145 built on top of Qwen/Qwen3-TTS-Tokenizer-12Hz. It generates discrete neural audio codec tokens at 12.5 Hz (4–6× fewer autoregressive steps than 50–75 Hz codecs) and decodes them to waveform with the Qwen3 TTS codec. The acoustic decoder is a linear-time Mamba2 SSM rather than a self-attention stack, so generation cost is constant per frame — memory does not grow with utterance length and there is no KV cache to manage. Because TTS and ASR were trained jointly in a single multi-task run, the same checkpoint also performs ASR (Japanese speech → text, reading only codec layer 0). It is a single-domain voice (ASMR-style Japanese training data) with no speaker conditioning or voice cloning.
Release Date: August 15, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ✅ |
| Languages | Japanese only (ja) |
| License | ![MIT][license-mit] |
| Parameters | ~0.41B (0.33B backbone + 78M audio branch) |
| Architecture | Transformer encoder (12 layers, bidirectional self-attention) + cross-attention → Mamba2 SSM decoder (6 layers, no causal self-attention) |
| Audio Codec | Qwen3-TTS-Tokenizer-12Hz (12.5 Hz, 16 quantizer layers) |
| Base Model | Kiseki-1.1-0.3B (seq2seq translation model) |
| Training Data | telecomadm1145/asmr_archive_qwentts_encoded |
Features: The decoder deliberately omits causal self-attention — temporal context is carried entirely by the Mamba2 recurrent state while text conditioning enters through cross-attention whose K/V are computed once during prefill. This yields O(1) state per frame (a fixed SSM tensor plus a 3-frame conv window) instead of an O(T) KV cache, so long-form synthesis degrades gracefully past the ~41 s training ceiling instead of hitting a memory cliff. Combined with the 12.5 Hz codec and a shared multi-token-prediction head that resolves all 16 codebook layers in one trunk pass, the model is both compute- and memory-bandwidth-bound rather than quadratic in length.
Links: ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: FireRedTTS3 is a unified speech generation and editing system from the FireRed Team built on semantically enriched continuous speech representations. It ships in two variants: FireRedTTS3-Base (zero-shot voice cloning across 24 languages and 21 Chinese dialects) and FireRedTTS3-Instruct (natural-language voice design and combined semantic + acoustic speech editing in one model). Beyond cloning, it supports instruction-based voice design (no reference audio needed) and editing operations such as insertion / deletion / substitution (semantic) and speed / pitch / volume changes (acoustic).
Release Date: August 5, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 24 (plus 21 Chinese dialects) |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | Qwen3 backbone + patch-level diffusion autoregressive (DiTAR) + RedAE codec + CAM++ speaker encoder |
| Variants | Base (cloning), Instruct (cloning + voice design + editing) |
Features: Represents speech with semantically enriched continuous (non-quantized) representations, enabling a single system to do zero-shot cloning, text-driven voice design, and fine-grained semantic + acoustic editing. On Seed-TTS-eval it reaches an average WER/CER of 3.04% with 78.8% speaker similarity; MiniMax-MLS-Test average SIM 84.8%.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| FireRedTTS3-ComfyUI | ComfyUI node | FireRedTTS3-ComfyUI |
· · · · · · · · · · · · · ·
Description: Audio8 TTS Preview 0.6B is a 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning. It uses a DualAR architecture inspired by Fish Audio S2 Pro: a slow AR transformer predicts one semantic token per audio frame, and a fast AR transformer predicts the frame's codec codebooks conditioned on the slow hidden state and preceding codebooks. The bundled 44.1 kHz neural audio codec handles both reference-audio encoding and waveform decoding — no additional codec checkpoint is required. The model supports 11 recommended languages (Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish) with zero-shot voice cloning from a reference audio clip + matching transcript.
Release Date: July 28, 2026
| Feature | Value |
|---|---|
| Parameters | 601,159,424 (0.6B, excluding the codec) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish (11) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | DualAR (slow AR + fast AR), inspired by Fish Audio S2 Pro |
| Slow Ar | 24 layers, width 896, 14 attention heads, 2 KV heads |
| Fast Ar | 4 layers, width 896, 14 attention heads, 2 KV heads |
| Acoustic Tokens | 10 codebooks, 4,096 entries per codebook |
| Codec | 44.1 kHz, 2,048 samples per model frame (~21.5 frames/s), bundled (no external codec needed) |
| Context Length | up to 2,048 packed text/audio positions |
| Sample Rate | 44,100 Hz |
| Inference | transformers with trust_remote_code=True; CUDA-capable GPU recommended |
| Dependencies | torch>=2.5.0, torchaudio>=2.5.0, transformers>=4.57.0,<5, soundfile, safetensors |
| Preview Status | language coverage intentionally limited; broader multilingual + Chinese dialect support planned |
| Library Name | transformers (custom_code) |
| Pipeline Tag | text-to-speech |
| Createdat | 2026-07-28T07:53:00Z |
Features: The DualAR architecture is the technical centerpiece: rather than a single autoregressive decoder predicting all codebook levels sequentially (the standard codec-LLM TTS pattern), Audio8 splits the work into a slow AR that predicts one semantic token per audio frame and a fast AR that predicts the frame's remaining codec codebooks conditioned on the slow hidden state. This separation lets the semantic-level reasoning happen at the slow AR's 24-layer depth while the acoustic codebook prediction stays lightweight at 4 layers — reducing the total compute per frame without sacrificing semantic quality. The bundled 44.1 kHz codec (no external codec checkpoint needed) and the 10-codebook / 4,096-entry acoustic token design give the model self-contained high-fidelity output at a compact 0.6B scale, making it one of the smallest multilingual zero-shot-cloning TTS systems shipping at 44.1 kHz.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: NeuTTS-2E is a super-fast, highly realistic, on-device emotional text-to-speech model from Neuphonic. It is the next generation after NeuTTS Air / Nano (which continue to ship for multilingual + zero-shot-cloning contexts) — narrowed in scope to an English-only alpha focused on:
Release Date: July 21, 2026
| Feature | Value |
|---|---|
| Parameters | 0.2B (compact LM backbone + codec) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English (English-only alpha) |
| Streaming | ✅ |
| License | ![Other][license-other] |
| Backbone | compact LM backbone tuned for emotional TTS token generation |
| Codec | efficient codec (compact, paired with the LM) |
| Speakers | 4 fixed (emily, paul, sophie, steven) |
| Emotions | 6 + neutral (angry, disgusted, fearful, happy, sad, surprised, neutral) |
| Emotion Control Mode | single-argument selection (no composable multi-axis axes like Scylla's Band) |
| Input Format | text only — no phonemizer, no system dependencies |
| On Device | yes (laptop-class CPU real-time / better-than-real-time) |
| Distribution Formats | safetensors (torch), Q4 GGUF, Q8 GGUF |
| Formats In Collection | neuphonic/neutts-2e (safetensors), neuphonic/neutts-2e-q4-gguf (smallest footprint), neuphonic/neutts-2e-q8-gguf (mid-tier compression) |
| Gguf Features | imatrix, conversational, endpoints_compatible |
| Pipeline Tag | text-to-speech |
| Library Name | (HF tag does not declare transformers / safetensors stem beyond safetensors itself) |
| Downloads | 194 / 241 / 216 (torch / q4 / q8) |
| Intended Use | embedded voice agents, on-device assistants, toys, privacy-sensitive applications |
| Comparison With Air Nano | Air/Nano continue to ship for zero-shot cloning + multilingual contexts; 2E is the next-gen focused English emotional variant |
| Safety Note | model is alpha; legitimate project landing is neuphonic.com (not neutts.com) |
Features: The technical center of NeuTTS-2E is maximum speed per parameter
at on-device budgets — the 0.2B LM + codec pair delivers
real-time-or-better on laptop-class CPUs while exposing
discrete categorical emotion control (angry / disgusted /
fearful / happy / sad / surprised / neutral) plus a
fixed four-speaker cast for consistency in agent / toy /
accessibility voice personas. Two design choices distinguish it from
the surrounding TTS field:
First, the categorical emotion surface is single-axis and
discrete (one emotion per call), not the continuous multi-axis
composable vector surface used by models like Scylla's Band
([neurotica base + 6-axis continuous strengths]). The project's
positioning — production-grade agents + toys + accessibility —
benefits from a one-argument API where emotion="happy" is the
explicit operational state. The release locks emotional mode at
generation time, which simplifies downstream filtering / guardrails.
Second, the distribution-shape design (one model, three
deployment formats) is a deliberate on-device-first posture: the
safetensors torch build for max-quality GPU/server; Q8 GGUF for
mid-tier compression; Q4 GGUF for the small-footprint embedded
target. All three are direct llama.cpp-compatible drops of the
same model — no retraining-per-format — letting users pick size vs
quality at deployment time without changing the production API.
The combined CPU-first + GGUF-first design pattern is the
opposite of the cloud-first TTS systems in this list — and is what
makes 2E suitable for embedded voice agents, toys, and
privacy-sensitive applications where audio + text must remain
on-device.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Collection][link-collection]
· · · · · · · · · · · · · ·
Description: Scylla's Band is a multilingual, multi-voice, expressive TTS model from Spybyscript, designed specifically for local and self-hosted inference through ONNX Runtime (with an experimental LiteRT backend for explicit native / mobile use). The architecture is a continuous-latent TTS family:
Release Date: July 19, 2026
| Feature | Value |
|---|---|
| Parameters | not stated (architecture: 4-layer duration predictor (192 hidden) + 12-layer rectified-flow acoustic generator (512 hidden, AdaLN, QK norm)) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | en_us, en_gb, es, it (4 public text-input languages) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Sample Rate | 24,000 Hz |
| Managed Voices | 10 (ariadne, felix, gwen, ink, max, orpheus, rex, scylla, stone, tuesday) |
| Voice Default Locale | en_us for most; ink / orpheus / tuesday default to en_gb |
| Voice Style Dim | 128 (style features) + 32 (prosody features) |
| Affect Axes | 6 (calm, joy, anger, sadness, sarcasm, questioning — all continuous in [0, 1]) |
| Affect Overlay Axes | sarcasm, questioning (mixable with any core delivery) |
| Affect Cfg Scope | duration + acoustic-flow prediction (preserves voice / reference) |
| Encoders Default | ONNX Runtime (Python CLI / Python API / Android sample / libscyllasband native) |
| Encoders Experimental | LiteRT (experimental / explicit-selection) |
| Cli Quality Default | 8-step Heun sampling |
| Graph Budgets | 512 G2P text tokens / 512 phone frames / 640 latent frames |
| Latent Target Buckets | 256 / 384 / 512 / 640 (smallest-fit selection) |
| Vocoder | Scylla's Band acoustic adapter + frozen charactr/vocos-mel-24khz |
| Hop Lengths | 256 (waveform) / 512 (latent) |
| Text Frontend | phrase-level multilingual G2P (74-phone vocabulary) |
| Span Context | 3 segments over up to 768 phones with 512-dim context state |
| Prefix Context | up to 24 acoustic latent frames from preceding chunk |
| Long Form Features | boundary metadata + punctuation pause floors + prefix-latent carryover + span context |
| Group Speak Input | [voice], [voice:language], [voice:language:axis=value,...] annotations |
| Bundle Contract | 1.0.0 / scyllasband-duration-flow |
| Intended Use | single-voice speech synthesis (10 voices); en/es/it; long-form narration; multi-voice dialogue from tagged text; continuous affect control; ONNX desktop/server; ONNX + LiteRT native/mobile |
| Not Intended | arbitrary-speaker cloning / impersonation / fraud / deceptive speech |
| Distributions | training data, trainer checkpoints, and export tooling not distributed |
| Cli Commands | download, validate-bundle, list-voices, normalize-text, speak, group-speak, stream, plan |
| Library Name | onnxruntime (tags include onnx, tflite, litert, duration-flow) |
Features: Three design decisions distinguish Scylla's Band in the multilingual
TTS class. First, decoupling duration and acoustic flow as
separate rectified-flow stages — duration is a 192-hidden, 4-layer
predictor operating on a 512-phone window, acoustic latents a
512-hidden, 12-layer AdaLN / QK-norm generator at 24-dim. This split
lets affect-CFG act on both stages independently while retaining
voice / reference conditioning, supporting the 6-axis continuous
composability. Second, 6 affect axes (with sarcasm and
questioning as overlays mixed with any core delivery) instead of
mutually-exclusive discrete emotion classes — calm=0.5, joy=0.5 is
a valid input, and axes stay in [0, 1] so multi-axis states are
expressible without combinatorial blow-up. Third, the ONNX-first
runtime design with libscyllasband native + an experimental
LiteRT backend sits at a budget most neural TTS systems don't
target — the 8-step Heun default and 512/512/640 fixed graph budget
keep the model usable on CPU and mobile, and the inference-only
release surface (training data + checkpoints not distributed) is the
complement of the latency / mobile inference focus.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: sanoTTS is the smallest known neural text-to-speech family. The name sano (सानो) is Nepali for "small". Each voice weighs 294k to 2.3M parameters — smaller than the smallest voice in prior families (TinyTTS at 1.62M; Inflect Nano at 4.63M; Kokoro at 82M) and the family fits in under 4 MB per voice with zero runtime dependencies (the espeak-ng phonemizer is bundled). Voices run real-time on a ~$3 ESP32-S3 microcontroller (output through a GPIO into an LM386 and a speaker) and live in the browser via WebAssembly — no server, no upload, no NPU. The full neural stack is duration → acoustic → decoder, quantized to int8, with the espeak-ng phonemizer included. 11 voices across 6 languages ship: English, Nepali, Hindi, Vietnamese, Indonesian, and Chinese — including the 294k heart-nano voice (337 KB) and the mel-based heart / heart-nano pair that predicts a 100-band spectrogram rendered by a noise-fed ConvNeXt + iSTFT decoder at 24 kHz. The inference runtime was relicensed MIT (September 2026); the project as a whole remains GPLv3 via espeak-ng. The project page at ampixa.github.io/sanoTTS hosts a live browser synthesis demo for every voice.
Release Date: July 13, 2026
| Feature | Value |
|---|---|
| Parameters | 294k–2.3M per voice (smallest = the 294k "heart-nano" voice, 337 KB) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English, Nepali, Hindi, Vietnamese, Indonesian, Chinese (6 languages, 11 voices) |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Architecture | full neural stack — duration model → acoustic model → decoder |
| Quantization | int8 (W8/A12, corr 0.9995+; piperlite portable C99) |
| Runtime Microcontroller | ESP32-S3 (real-time RTF 0.41, GPIO → LM386 → speaker) |
| Runtime Browser | WebAssembly (no server, no upload, no NPU) |
| Runtime Footprint | under 4 MB per voice, zero dependencies |
| Voices | 11 (English: amy / kristin / hfc / amy-1p1m / amy-1p8m / robot / heart / heart-nano; one voice each for NE / VI / ID / ZH + shared lang voices) |
| Phonemizer | espeak-ng (bundled) |
| License Split | inference runtime MIT; project overall GPL-3.0 (copyleft from espeak-ng) |
| Library Name | sanotts |
| Training Method | distillation (per voice) |
Features: The hard constraint — be the smallest neural TTS family known,
real-time on a $3 microcontroller — drives the entire stack.
Conventional sub-100M TTS systems are too large for an ESP32's flash
and RAM. sanoTTS keeps the full duration → acoustic → decoder
neural pipeline (no espeak-NG-only fallback, no concatenative
hybrid), quantizes everything to int8, and bundles the phonemizer
so the whole voice ships in under 4 MB with zero runtime dependencies.
The newest heart / heart-nano voices switch to a mel-based recipe
(100-band spectrogram + noise-fed ConvNeXt + iSTFT decoder at 24 kHz),
bringing the smallest voice down to 294k parameters / 337 KB —
a per-voice footprint 100× smaller than Kokoro and 2× smaller than
TinyTTS while still leading SCOREQ / UTMOS in the sub-15M class —
and the demo synthesizes every voice live in the browser via
WASM, so the smallest-known neural TTS is also the only one that
runs unattended on a $3 chip and a $0 web page.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: FreyaTTS is a 183M-parameter Turkish text-to-speech model. It is tokenizer-free at the character level — 92 symbols in its Turkish vocabulary — so there is no phonemizer or G2P step in either training or inference. Speech is generated with a non-autoregressive conditional flow-matching DiT in the frozen AudioVAE2 latent space (25 Hz, 64-dim latents, 16 kHz encode / 48 kHz decode). Training runs from scratch on Turkish speech: a pretraining stage followed by SFT stage 1/2 for voice lock and short-utterance coverage. Output is 48 kHz mono. On the project's Freya-TR-Eval benchmark the model reports WER 8.0% / CER 3.0%, ranking 3rd of 7 among open sub-1B Turkish TTS systems — a deliberate single-target-speaker, no-cloning design choice for a focused foundation release. The evaluation dataset is freya-tr-eval.
Release Date: July 7, 2026
| Feature | Value |
|---|---|
| Parameters | 183.2M |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | Turkish (tr) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | conditional flow-matching diffusion transformer (DiT), non-autoregressive, 32-step Euler ODE, no CFG |
| Tokenizer | character-level (92 Turkish symbols; no phonemizer, no G2P) |
| Latent Space | frozen AudioVAE2 (Apache-2.0, openbmb/VoxCPM2), 64-dim at 25 Hz |
| Codec Io | 16 kHz encode / 48 kHz decode |
| Sample Rate | 48,000 Hz |
| Training | from scratch on Turkish speech; pretraining + SFT stage 1/2 (voice lock + short-utterance coverage) |
| Evaluation | Freya-TR-Eval — WER 8.0% / CER 3.0%, 3rd of 7 open sub-1B Turkish TTS |
| Library Name | freyatts |
Features: Two design choices are worth flagging. First, tokenizer-free character-level Turkish: by training directly on the 92-symbol Turkish alphabet with no phonemizer or G2P grapheme-to-phoneme step, the model removes a dependency that is fragile for agglutinative Turkish morphology and that often degrades quality when ported to low-resource Turkic relatives. Second, non-autoregressive conditional flow-matching in a frozen AudioVAE2 latent space: the 25 Hz / 64-dim bottleneck keeps the DiT small (183M) while inheriting a separately-trained audio codec's representation, letting a focused single-language-non-multilingual release ship at a fraction of the parameter budget of multilingual foundation TTS systems. The deliberate "no cloning, single target speaker" choice is a scope-lowering move that lets the foundation release put all its capacity into Turkish speech quality rather than spread it across zero-shot speaker adaptation.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Inflect-Nano-v2 is a complete local text-to-waveform speech synthesis model with 3,966,721 deployable parameters — under 4M total. It is a VITS-architecture fixed-voice English TTS designed for CPU or CUDA inference with deterministic seeds, long-text handling, and 24 kHz mono output. The full FP32 checkpoint is 15.97 MB, making it one of the smallest complete neural TTS systems that produces natural-sounding speech without a separate vocoder or phonemizer dependency. The model ships with a public adaptation toolkit for preparing data, auditing train/validation splits, adapting a fixed voice or language, resuming training, evaluating checkpoints, and exporting PyTorch or ONNX packages. A sibling Inflect-Micro-v2 (9.36M parameters) prioritizes quality below 10M; Nano prioritizes footprint below 4M. Both share one public API.
Release Date: June 25, 2026
| Feature | Value |
|---|---|
| Parameters | 3,966,721 (3.97M deployable) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | VITS (end-to-end text-to-waveform) |
| Sample Rate | 24,000 Hz |
| Footprint | 15.97 MB FP32 |
| Inference | CPU or CUDA; PyTorch + ONNX export |
| Determinism | deterministic seeds for reproducible generation |
| Long Text | automatic text splitting and handling |
| Input Format | text (no phonemizer or system dependencies) |
| Adaptation Toolkit | data prep, split auditing, voice/language adaptation, training resume, checkpoint eval, PyTorch/ONNX export |
| Sibling Model | Inflect-Micro-v2 (9.36M params, quality-prioritized below 10M) |
| Api | one public API across Micro and Nano sizes |
| Library | pytorch |
| Metrics | WER |
| Inference False On Hf | yes (no hosted HF inference endpoint; local-only) |
Features: Inflect-Nano-v2's defining constraint is completeness under 4M parameters: the entire text-to-waveform pipeline — no separate vocoder, no phonemizer, no system dependencies — fits in 3.97M deployable parameters and a 15.97 MB FP32 checkpoint. This is smaller than even sanoTTS's smallest voice (745k) when measured by complete-pipeline footprint, though sanoTTS ships per-voice weights rather than a single fixed-voice checkpoint. The VITS end-to-end architecture is the enabler: by folding the acoustic model and vocoder into a single jointly-trained network, Inflect avoids the multi-stage pipeline overhead that makes most neural TTS systems larger. The public adaptation toolkit extends the fixed-voice design into a customizable platform — users can prepare data, adapt a voice or language, resume training, and export PyTorch or ONNX packages — making the 4M-parameter footprint a starting point for domain-specific TTS rather than a dead-end fixed-voice release.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: GEnerative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue. Gepard is built for low-latency, high-throughput streaming conversation: the model starts speaking the moment text begins arriving, generating audio piece by piece instead of waiting for a full sentence. It is a single decoder-only autoregressive language model built on Qwen3.5 (14 layers, hidden 1024, 8 heads) with ≈556M total parameters (backbone + audio interface + voice-cloning compressor). Audio is produced through NVIDIA NeMo NanoCodec — Finite Scalar Quantization at 22.05 kHz, 21.5 frames/s, 1.89 kbps — with the full 32-channel FSQ frame sampled in one step. Reports ~25× real time on a single RTX 5090 with first-audio-chunk latency around 50 ms; a 96 GB Blackwell card serves up to 256 concurrent conversations. CFG refinement is baked into the weights so quality gain comes at no extra two-pass cost at inference, though the two-pass mode is still selectable as a quality dial.
Release Date: June 22, 2026
| Feature | Value |
|---|---|
| Parameters | ~556M (555,694,169; Qwen3.5 backbone + audio interface + voice-cloning compressor) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | English (US/UK), Spanish (es-MX), Portuguese (pt-BR), Dutch (NL) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Audio Codec | NVIDIA NeMo NanoCodec (FSQ, 22.05 kHz, 21.5 fps, 1.89 kbps; NVIDIA Open Model License) |
| Sample Rate | 22,050 Hz |
| Backbone | Qwen3.5 full-attention transformer (14 layers, hidden 1024, 8 heads; ~500M params) |
| Inference | vLLM |
| Throughput | 256 conversations on one 96 GB Blackwell (RTX Pro 6000) GPU |
| Benchmark | Seed-TTS-eval leader on perceived quality (NISQA-MOS 4.25, NOI 4.16, COL 4.16, DIS 4.51) trading some WER/SIM |
Features: A prosody-aware autoregressive single-pass frame generator: the whole 32-channel FSQ audio frame is sampled in one step (no depth transformer), and CFG quality refinement is baked into the weights rather than incurred at inference as a two-pass cost — so the publicly reported TTFA of ~50 ms and 25× real time on a single RTX 5090 represent the quality-on path, not a cheap-fast preview. Voice cloning is decoupled into a separate up-front compressor, which means cloning is "free" at run-time once the reference clip is encoded — a structural choice that supports serving hundreds of conversations per GPU. A stop-head weight update (2026-08-06) fixed premature stopping at sentence boundaries and lifted the effective duration ceiling; on Seed-TTS-eval Gepard leads the compared systems on perceived quality (NISQA-MOS 4.25) while trading some speaker similarity and WER for its streaming-first design.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo] ![Paper][link-paper] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Boson AI's flagship conversational TTS: an ~4B autoregressive decoder over interleaved text and audio tokens from the Higgs Tokenizer (8 codebooks at 25 fps / 24 kHz). Built for voice chat rather than narration, it covers 102 languages with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.
Release Date: June 4, 2026
| Feature | Value |
|---|---|
| Parameters | 4B (BF16, 36 layers, hidden=2560, GQA 32/8) |
| Architecture | Autoregressive decoder (Qwen3-style) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 102 (85 with WER/CER <5, 17 between 5-10) |
| Streaming | ✅ |
| Audio Output | 24 kHz |
| License | ![Research Only][license-research-only] |
Features: Interleaved text/audio token modelling with a delay-pattern multi-codebook embedding/head: a single autoregressive stack emits both modalities and supports inline <|category:value|> control tags (emotion/style/sfx/prosody) inserted at any point in the target text.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive TTS system from Rednote-HiLab. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE, with no discrete tokens anywhere in the pipeline. It achieves the best average performance on Seed-TTS-Eval (WER 0.94 / 1.30 / 6.60 on zh / en / zh-hard) and the highest speaker similarity on a 24-language MiniMax multilingual benchmark, with broad cross-lingual voice cloning.
Release Date: June 3, 2026
| Feature | Value |
|---|---|
| Parameters | 2B (semantic encoder + LLM + AR flow-matching acoustic head) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Multilingual (24+ languages; zh / en focus) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Sample Rate | 48 kHz |
| Tokenizer | 48 kHz AudioVAE (continuous, no discrete tokens) |
Features: A fully continuous autoregressive pipeline that keeps generation in waveform-latent space end-to-end (no discrete-code phase), pairing an LLM-side semantic encoder with an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE — yielding SOTA seed-TTS-Eval scores and the strongest speaker-similarity number (83.9 avg) on the 24-language MiniMax multilingual benchmark.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Confucius4-TTS is an LLM-based text-to-speech system from NetEase Youdao designed for multilingual and cross-lingual synthesis. It uses a speech encoder + LLM (Text2Semantic) + flow-matching Semantic2Acoustic architecture that allows zero-shot voice cloning without a required reference transcript and explicit cross-lingual voice transfer with unaccented output across languages. Covers Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese with code-switching and emotion transfer.
Release Date: June 2, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | 14 (zh, en, ja, ko, de, fr, es, id, vi, th, pt, it, ru, ms) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | speech encoder + LLM (T2S) + flow-matching head (S2A) |
Features: Cross-lingual voice transfer without accent drift: the same reference voice stays consistent when the speaker switches languages — backed by a speech encoder + LLM backbone pipeline (T2S) with a flow-matching acoustic decoder (S2A) and training that bundles 14 languages with code-switched, emotion-preserving decoding.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: WavTTS is an end-to-end zero-shot TTS framework that synthesizes speech directly in the raw waveform space — explicitly skipping the intermediate mel-spectrogram, VAE-latent, or codec-token representations that most modern TTS stacks use. It is built on a flow-matching diffusion transformer (DiT) with waveform patchification, multi-scale mel-spectrogram supervision, and an optimized noise schedule. Forked from F5-TTS at the codebase level but replaces the whole acoustic pipeline.
Release Date: May 28, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English, Chinese |
| Streaming | ❌ |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] ![MIT][license-mit] |
| Sample Rate | 16 kHz |
| Training Data | Emilia |
| Architecture | Flow-matching DiT, raw waveform patchification, multi-scale mel supervision |
| Training Steps | 1.2M |
Features: Skip every intermediate waveform representation (no mel, no VAE, no codec tokens): a flow-matching DiT produces raw-waveform patches directly, supervised at multiple mel scales and an optimized noise schedule — yielding high-quality zero-shot TTS at 16 kHz from a single end-to-end stack.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: MOSS-TTS is a production-grade Text-to-Speech foundation model developed by the OpenMOSS Team and MOSI.AI. The current public v1.5 release preserves the original 1.0 capabilities — zero-shot voice cloning, long-form speech generation, token-level duration control, Pinyin/IPA pronunciation supervision, multilingual synthesis, and code-switching — and extends multilingual continued training from 20 languages to 31 languages including Cantonese, Dutch, Finnish, Hindi, Macedonian, Malay, Romanian, Swahili, Tagalog, Thai, and Vietnamese. v1.5 improves speaker similarity, reduces cloning variance on long-reference / short-text scenarios, follows punctuation-driven prosody more reliably, and adds explicit inline pause markers (e.g., [pause 3.2s]).
Release Date: May 25, 2026
| Feature | Value |
|---|---|
| Parameters | 8B (Delay), 1.7B (Local) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 31 (extended from v1.0's 20) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Max Duration | 1 hour |
| Pause Control | yes (inline markers like [pause 3.2s]) |
| Lang Tag Control | yes (set language= in user message) |
Features: v1.5 widens MOSS-TTS from 20 → 31 languages with stronger per-language multilingual synthesis (control via a language tag in the user message), more stable cloning under long-reference / short-text conditions, punctuation-driven prosody that holds up across long sentences, and explicit inline pause tokens ([pause 3.2s]) for scripted narration control.
Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: VoxFlash-TTS is a zero-shot voice-cloning text-to-speech engine built around extreme latent compression. The VAE encodes 24 kHz waveforms into a 9 frames/s latent space — roughly 8× more compressed than EnCodec (75 fps) and 2.4× more than Stable Audio (21.5 fps). Generating 10 s of audio therefore requires the diffusion model to produce just 90 latent vectors rather than hundreds or thousands of tokens, with downstream quadratic savings in attention cost. A ConvNeXtV2-based phoneme encoder followed by a novel coarse-alignment algorithm (cheaper than cross-attention) maps text into the latent sequence; a modern diffusion head then iteratively refines speech latents that the lightweight VAE decoder renders back to waveforms. The architecture targets low-latency, low-resource deployment — consumer-grade GPUs and edge devices — with Chinese and English zero-shot cloning. The project card lists inference: false on HF (no hosted inference endpoint), but the project page at voxflash.github.io carries the abstract, demo examples, and ablations.
Release Date: May 22, 2026
| Feature | Value |
|---|---|
| Parameters | not stated (ConvNeXtV2 phoneme encoder + diffusion head + lightweight VAE decoder) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Chinese, English |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Audio Codec | VoxFlash VAE (9 Hz / 9 fps latent, 24 kHz input) |
| Compression Ratio | ~8× tighter than EnCodec (75 fps), ~2.4× tighter than Stable Audio (21.5 fps) |
| Phoneme Encoder | ConvNeXtV2 + coarse-alignment algorithm (no cross-attention) |
| Diffusion Head | modern multi-step iterative refinement |
| Decoder | lightweight VAE decoder |
| Sample Rate | 24,000 Hz |
| Inference | local CUDA ≥ 12.3.2; no HF hosted endpoint |
| Training Dataset | seed-tts-eval |
| Metrics | word_error_rate, speaker_similarity |
Features: The central technical move is compressing the audio latent space to 9 frames/s instead of the conventional 75 fps (EnCodec) or 21.5 fps (Stable Audio). This is not a quantization tweak — it is a temporal-decimation architectural choice that shrinks the sequence length the diffusion model has to traverse, and because attention cost scales quadratically with sequence length the end-to-end compute drops by orders of magnitude. Combined with a coarse-alignment phoneme-to-latent map that avoids cross-attention entirely (using a ConvNeXtV2 phoneme encoder instead), VoxFlash hits millisecond-level inference latency on consumer-grade and edge hardware for zero-shot Chinese + English cloning, where conventional latent-diffusion TTS systems are too slow for real-time edge deployment.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Miso TTS 8B is a text-to-speech model from Miso Labs built on the Sesame Conversational Speech Model (CSM) architecture. A large Llama-3.2-style backbone consumes text/audio-frame embeddings and predicts codebook 0 of the Mimi audio token stream, while a smaller 300M autoregressive audio decoder predicts codebooks 1–31 in codebook depth. The model is designed for high-quality conversational speech and voice continuation from a short prompt audio clip.
Release Date: May 21, 2026
| Feature | Value |
|---|---|
| Parameters | 8B (backbone llama-3.2-style) + 300M (audio decoder) = 8.3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Architecture | Sesame-style CSM (two transformer stack: backbone + audio decoder) |
| Audio Tokenizer | Mimi (32 codebooks, vocab 2051, max seq 2048) |
| Library | pytorch |
Features: A two-transformer Sesame-style CSM (Llama 8B backbone consumes text + audio frames and produces backbone codebook-0 prediction; a 300M audio decoder autoregresses over codebook depth via Mimi's 32-codebook stack) — letting the larger backbone spend capacity on linguistic / speaker conditioning while a leaner decoder handles fine-grained codebook-by-codebook generation.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Raon-OpenTTS is an open-data, open-weight zero-shot TTS system from KRAFTON that performs on par with state-of-the-art closed-data models. This is the 1B variant (1048M parameters). Both model weights and training data are public: Raon-OpenTTS-Core is 510.1K hours of English speech, quality-filtered from the 615K-hour public Raon-OpenTTS-Pool using combined DNSMOS, WER, and VAD rank-based filtering. It ranks 1st or 2nd in WER and SIM among recent zero-shot TTS models on Seed-TTS-Eval and CV3-Eval, and achieves the best average WER/SIM on Raon-OpenTTS-Eval across Clean, Noisy, Wild, and Expressive regimes. A smaller Raon-OpenTTS-0.3B variant is also available.
Release Date: May 21, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English only (trained on 11 English speech datasets) |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Parameters | 1048M |
| Architecture | DiT (Diffusion Transformer) based on F5-TTS with flow matching; dim=1408, depth=28, heads=24 |
| Audio Output | 80-ch mel-spectrogram at 16 kHz, HiFi-GAN vocoder (LibriTTS) |
| Training Data | Raon-OpenTTS-Core (510.1K hours), 520K updates on 48× B200 |
Features: Demonstrates that fully open data + open weights can match proprietary SOTA: on Seed-TTS-Eval it reaches 1.78 WER / 0.749 SIM (vs Qwen3-TTS 1.46/0.715 at 1.7B), and best overall robustness (WER 2.81 / SIM 0.695) across four acoustic regimes on its own Raon-OpenTTS-Eval benchmark. The pipeline pairs large-scale rank-based data curation (DNSMOS + WER + VAD filtering of a 615K-hour pool) with an efficient F5-TTS-derived DiT.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Dataset][link-dataset]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| ComfyUI-Raon-OpenTTS | ComfyUI node | ComfyUI-Raon-OpenTTS |
· · · · · · · · · · · · · ·
Description: OronTTS is a non-autoregressive text-to-speech model from btsee, an F5-TTS fork specialized for Mongolian (Khalkha Cyrillic) and Kazakh (Cyrillic). It uses Flow Matching + Diffusion Transformer + Vocos (dim 1024, depth 22, 16 heads, vocab 65, 24 kHz sample rate), trained on the btsee/mbspeech_mn corpus (3,846 Mongolian speech samples) and outputs zero-shot synthesis from a short reference audio + language tag.
Release Date: May 16, 2026
| Feature | Value |
|---|---|
| Parameters | (not stated) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Mongolian (Khalkha Cyrillic), Kazakh (Cyrillic) |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Architecture | F5-TTS (OT-CFM + DiT + Vocos) |
| Dim | 1024 |
| Depth | 22 |
| Heads | 16 |
| Vocab Size | 65 |
| Sample Rate | 24000 Hz |
| Mel Bins | 100 |
| Training Data | btsee/mbspeech_mn (3,846 Mongolian speech samples) |
Features: F5-TTS re-purposed for low-resource Cyrillic languages (Mongolian + Kazakh) — non-autoregressive flow-matching DiT over a tight 65-word vocab. Trained on a small (~3.8k sample) Mongolian corpus; the architecture is small enough that Khalkha Cyrillic and Kazakh Cyrillic share the same checkpoint via the lang tag at inference time.
Links: ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Supertonic 3 is the third-generation open-weight release from Supertone. It is a lightweight, on-device text-to-speech system that runs with ONNX Runtime entirely on the user's machine (no network, no API call) and ships as a Python SDK (pip install supertonic). Compared with the Supertonic 2 base (5 languages, 66 M params), v3 expands to 31 languages and adds expression tags (<laugh>, <breath>, <sigh>), more stable reading on long utterances, and higher speaker similarity across the core language set.
Release Date: May 6, 2026
| Feature | Value |
|---|---|
| Parameters | (not stated on card; Supertonic 2 baseline 66 M — likely similar or smaller weight class) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | 31 (expanded from Supertonic 2's 5) |
| Streaming | ✅ |
| License | ![OpenRAIL-M][license-openrail-m] |
| On Device | yes (ONNX Runtime, no cloud call) |
| Expression Tags | yes (<laugh>, <breath>, <sigh>) |
Features: A more compact on-device multilingual TTS: ONNX-Runtime inference everywhere, 31 languages from a single small open-weight encoder, and discrete expression tags that the decoder interprets inline — without a separate speaker-emotion control path or a cloud-rendered audio round-trip.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi]
· · · · · · · · · · · · · ·
Description: Scenema Audio is a zero-shot expressive voice cloning and speech generation model from ScenemaAI. It is built on an audio diffusion transformer extracted from the audio branch of Lightricks' LTX 2.3 (a 22B audiovisual model) — keeping the in-the-wild acoustic quality the bigger model learned while specializing for speech output. Generation is prompt-driven: a <speak> tag carries a voice description, gender, optional scene (ambient audio around the voice), and language; an <action> tag shifts emotional state mid-generation. Action tags cover rage, grief, joy, fear, exhaustion; voice prompt can describe timbre/pitch/breathiness/rasp/resonance plus character archetypes ("Tony Soprano having a breakdown"). Supports zero-shot voice cloning from 10-20 seconds of reference audio with some emotional variability, automatic long-form narration by splitting text and maintaining voice continuity, and 13 multilingual built-ins.
Release Date: April 26, 2026
| Feature | Value |
|---|---|
| Parameters | (audio diffusion transformer of LTX 2.3, weights ~9.8 GB bf16 / ~4.9 GB INT8 + ~6.7 GB pipeline) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | 13 (en, de, fr, es, it, pt, ja, zh, ko, ru, ar, hi, sw) |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Parent Model | Lightricks LTX-2.3 (audio branch) |
| Prompt Format | <speak voice=… gender=… scene=… language=…> XML with <action> tag for shifting emotion |
| Long Form Narration | yes (auto-splits text while preserving voice continuity) |
| Quantized | yes (INT8 weights at ~4.9 GB, identical quality) |
Features: A standalone audio diffusion transformer extracted from a much bigger multimodal source: the model inherits how people actually sound in real scenes (angry, laughing, whispering, crying, exhausted, terrified) and exposes that capacity through a <speak> + <action> prompt interface — emotional state shifts within a single generation, instead of being a token-level or speaker-level conditioning problem.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Dramabox is Resemble AI's expressive TTS, distributed under the LTX-2 Community License. It is an IC-LoRA fine-tune of the LTX-2.3 3.3B audio-only branch (Diffusion Transformer + flow matching), conditioned on Gemma 3 12B text embeddings. Generation is prompt-driven: speaker identity, emotion, delivery, laughs, sighs, breaths, pauses, and transitions are all expressed inside a natural-language description, with an optional 10-second voice reference that clones the target timbre.
Release Date: April 17, 2026
| Feature | Value |
|---|---|
| Parameters | 3.3B (LTX-2.3 audio backbone, IC-LoRA fine-tune) + 12B Gemma 3 text encoder (conditioning only) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Base Model | Lightricks/LTX-2.3 (audio branch) |
| Architecture | DiT + flow matching, IC-LoRA fine-tune, Gemma 3 12B text embeddings |
| Inference Time | ~2.5 s / generation (warm server) |
Features: IC-LoRA fine-tune of LTX-2.3's audio branch leaves the heavy text-understanding work to Gemma 3 12B and lets the DiT do the expressive rendering — so what's normally multimodal-stage orchestration collapses into a single prompt-driven TTS where speaker identity, emotion, and delivery are encoded in the prompt itself, and the timbre comes from a 10-second voice reference when present.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Sarashina2.2-TTS is a Japanese-centric text-to-speech system from SB Intuitions built on a large language model. It supports both Japanese and English, delivers strong pronunciation accuracy on Japanese text through large-scale end-to-end training, and reproduces a speaker's voice, speaking style, and acoustic characteristics from a short reference clip (zero-shot). Training data is sourced exclusively from legitimately acquired, properly licensed speech archives per the Sarashina Model NonCommercial License Agreement v2.0 (released April 24, 2026).
Release Date: April 16, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | Japanese, English |
| Streaming | ❌ |
| License | ![Research Only][license-research-only] |
| Base Model | sbintuitions/sarashina2.2-0.5b-instruct-v0.1 |
| Cross Lingual | yes (Japanese ↔ English, code switching) |
Features: Japanese-optimized TTS fine-tuned on responsibly-licensed Japanese training corpora with explicit cross-lingual code-switching to English in a single utterance; reference audio carries speaking style and speaker identity together, so the same prompt yields narration, broadcast, conversation, or customer-service delivery without separate style conditioning.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: State-of-the-art diffusion-based TTS model operating directly in waveform latent space. Developed by Meituan's LongCat team, it requires only a Waveform VAE and Diffusion backbone, effectively mitigating compounding errors.
Release Date: March 30, 2026
| Feature | Value |
|---|---|
| Parameters | 1B / 3.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | Chinese, English |
| Streaming | ❌ |
| Audio Output | 24000 Hz |
| License | ![MIT][license-mit] |
Features: Adaptive Projection Guidance (APG) replaces traditional classifier-free guidance for elevated generation quality. Outperforms Seed-TTS on zero-shot voice cloning benchmarks.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: SILMA TTS v1 is a high-performance, 150M-parameter bilingual (Arabic/English) TTS model developed by SILMA AI. Built on the F5-TTS diffusion architecture, it was pretrained from scratch using tens of thousands of hours of high-quality public and proprietary data. It supports instant voice cloning with less than 8 seconds of reference audio (the reference transcript can also be left empty — it is transcribed on the fly), full support for Arabic Tashkeel (diacritics, auto-enriched via CATT when absent), NeMo-based text normalization, and an RTF around 0.12 on an RTX 4090. Released under a commercial-friendly license: code MIT, model weights Apache-2.0. The model is 100% compatible with F5-TTS v1.1.7 tooling for inference and fine-tuning.
Release Date: March 13, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 2 (Arabic MSA/Fusha + English) |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 150M |
| Architecture | F5-TTS Diffusion Transformer with flow matching (pretrained from scratch, F5-TTS v1.1.7-compatible) |
| Pronunciation | ✅ |
| Cost | RTF ≈ 0.12 (RTX 4090) |
Features: Brings native-level Arabic synthesis to a 150M footprint: one of the smallest open F5-TTS-family models pretrained from scratch rather than fine-tuned, with first-class Arabic handling (Tashkeel-aware pronunciation via CATT enrichment, NeMo text normalization) alongside English, plus instant zero-shot cloning under fully permissive licensing (Apache-2.0 weights / MIT code).
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Fish Audio S2 Pro is a leading text-to-speech model with fine-grained inline control of prosody and emotion. It combines reinforcement learning alignment with a dual-autoregressive architecture for high-quality speech synthesis.
Release Date: March 10, 2026
| Feature | Value |
|---|---|
| Parameters | ~10 GB (BF16) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 80+ (Tier 1: En, Zh, Jp) |
| Streaming | ✅ |
| License | ![Research Only][license-research-only] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Native multimodal foundation model by Meituan LongCat Team processing text, vision, and audio under a single autoregressive objective. Industrial-strength model with strong speech synthesis and voice cloning.
Release Date: March 2026
| Feature | Value |
|---|---|
| Parameters | 3B (MoE A3B) |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| Audio Output | 24 kHz |
| License | ![MIT][license-mit] |
Features: Discrete Native Autoregression Paradigm (DiNA) unifying modalities in shared discrete token space. Combines visual understanding, generation, and audio processing in single model.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Frontier, open-weights text-to-speech model developed by Mistral AI. Designed to be fast, instantly adaptable, and produces lifelike speech with natural prosody and emotional range.
Release Date: March 2026
| Feature | Value |
|---|---|
| Parameters | 4B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | 9 (En, Fr, Es, De, It, Pt, Nl, Ar, Hi) |
| Streaming | ✅ |
| Audio Output | 24 kHz |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
Links: ![HuggingFace][link-huggingface] ![Demo][link-demo] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: BlueTTS (project page: lightbluetts.com) is a multilingual text-to-speech library. Built around slim ONNX graphs that run on ONNX Runtime with first-class CPU support and optional accelerators — OpenVINO (Intel), CUDA ORT (NVIDIA), TensorRT, and ONNX Runtime stock CPU. Targets five languages — Hebrew, English, Spanish, Italian, German — including inline mixed-language with XML-style tags in the text prompt. Inference is deliverable as a PyPI package (blue-onnx), a Rust crate, or directly from the pinned ONNX graphs on the HF Hub; the v2 release ships a slimmed opset-17 ONNX bundle (notmax123/blue-onnx-v2) that's intended for both FP32 production and the experimental INT8 weight-only fallback.
Release Date: February 27, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | Hebrew, English, Spanish, Italian, German |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Runtime | ONNX Runtime (stock CPU; OpenVINO / CUDA / TensorRT optional) |
| Speed | "fastest open-source TTS" (per project description) |
| Graph Format | ONNX opset 17 (slim, full-precision; experimental weight-only INT8 fallback) |
| Distribution | PyPI + HuggingFace + Rust |
Features: A CPU-first multilingual TTS that ships both slimmed ONNX graphs and a Python package where the same code path runs on stock CPU ONNX Runtime by default — and optionally accelerates on OpenVINO / CUDA ORT / TensorRT — so deployment doesn't gate on GPU availability. Languages include Hebrew (with explicit G2P normalization) — a comparatively rare open-source TTS target — plus standard European languages, all from MIT-licensed weights distributed via both Hugging Face and PyPI.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![PyPI][link-pypi] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: KittenTTS is an open-source realistic text-to-speech model designed for lightweight deployment. It is a state-of-the-art TTS model under 25MB with just 15 million parameters, running without GPU on any device.
Release Date: February 24, 2026 (v0.8.1)
| Feature | Value |
|---|---|
| Parameters | 15M-80M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | English, Multiple |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Ming-omni-tts is a high-performance unified audio generation model in the Ming 2.0 series. It uses a custom 12.5 Hz continuous tokenizer and a Patch-by-Patch compression strategy that drives the LLM inference frame rate down to 3.1 Hz, enabling fine-grained control over speech rate, pitch, volume, emotion, and dialect (notably Cantonese at ~93 % accuracy). It supports 100+ premium built-in voices plus zero-shot voice design from natural-language prompts and is the first autoregressive model that jointly generates speech, ambient sound, and music in a single channel.
Release Date: February 11, 2026
| Feature | Value |
|---|---|
| Parameters | 16.8B (3B active, MoE; A3B) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | Chinese, English, Cantonese |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Patch-by-Patch compression drives the inference frame rate to 3.1 Hz, drastically cutting LLM-side latency for podcast-style audio while preserving naturalness. A custom 12.5 Hz continuous tokenizer plus a DiT head jointly produce speech, ambient sound, and music in a single output channel — an "in-the-scene" listening experience rather than TTS-on-top-of-a-track.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: SoulX-Singer is a high-fidelity, zero-shot singing voice synthesis model for generating realistic singing voices for unseen singers without fine-tuning.
Release Date: February 6, 2026
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Mandarin, English, Cantonese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: SoproTTS is a lightweight English text-to-speech model with zero-shot voice cloning. It uses dilated convolutions (WaveNet-style) and lightweight cross-attention layers instead of the common Transformer architecture.
Release Date: February 4, 2026 (v1.5)
| Feature | Value |
|---|---|
| Parameters | 135M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Rtf | 0.05 (CPU M3) |
| Training-Cost | ~$100 |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Qwen3-TTS is an open-source series of Text-to-Speech models developed by Alibaba Cloud. Supports stable, expressive, and streaming speech generation with free-form voice design.
Release Date: January 22, 2026
| Feature | Value |
|---|---|
| Parameters | 0.6B-1.7B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 10 (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: TADA is a unified speech-language model from Hume AI built around a Text-Acoustic Dual-Alignment tokenizer: for every text/subword token there is exactly one corresponding speech vector, so the audio stream stays 1:1 aligned with text. As a TTS model, each autoregressive step covers one text token and dynamically determines the duration and prosody for that token, breaking the fixed-frames-per-second constraint that drives most modern TTS backbones. As a speech-language model, it generates a text token and the speech for the preceding token in the same dual step.
Release Date: January 12, 2026
| Feature | Value |
|---|---|
| Parameters | 1B (Llama 3.2 1B base) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Base Model | meta-llama/Llama-3.2-1B |
| Tokenization | 1:1 text–acoustic dual alignment (one speech vector per text token) |
| Dynamic Duration | yes (each autoregressive step covers one text token, duration is determined per-token) |
Features: A dual-alignment speech–text tokenizer that decouples autoregression from a fixed audio frame rate: each text token owns exactly one speech vector, and the model synthesizes the whole segment for that token in one step, regardless of how long the spoken form is — eliminating transcript hallucination and the latency overhead of constant-frame-rate codecs while staying as compact as Llama 3.2 1B.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi] ![Paper][link-paper] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: Japanese Text-to-Speech model based on Rectified Flow Diffusion Transformer. Features emoji-based style and sound effect control by embedding emojis in input text for expressive speech generation.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 500M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | Japanese |
| Streaming | ❌ |
| Audio Output | 48kHz waveform |
| License | ![MIT][license-mit] |
Features: Key Feature: Emoji annotation control - insert specific emojis into text to control speaking styles, emotions, and sound effects.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Open-source TTS for European languages with 7B parameters. Outperformed ElevenLabs in human preference testing.
Release Date: Early 2026
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 23 European languages |
| Streaming | ✅ |
| License | ![MIT][license-mit] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Part of the LEMAS (Large-scale Extensible Multilingual Audio Suite) project. Zero-shot multilingual TTS with 0.3B parameters supporting 10 languages with word-level precise editing capabilities.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 0.3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 10 (zh/en/de/fr/es/pt/it/ru/id/vi) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Special-Feature | Word-level editing (LEMAS-Edit) |
Features: Built on 150,000+ hours of multilingual speech data with word-level timestamps. Includes LEMAS-Edit for precise word-level speech editing via masked token infilling.
Links: ![Website][link-website] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Lightweight, high-speed LLM-based TTS model for English and Japanese with minimal resource usage.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 2.6B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | English, Japanese |
| Streaming | ✅ |
| License | ![LFM][license-lfm] |
| Rtf | 0.135-0.145 |
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Ultra-lightweight open-source multilingual speech generation model with only 0.1B parameters. Designed for realtime speech generation that runs directly on CPU without GPU.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 0.1B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | 20 |
| Streaming | ✅ |
| Audio Output | 48 kHz Stereo |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Pure autoregressive architecture with MOSS-Audio-Tokenizer-Nano. Compresses audio to 12.5 Hz token stream using RVQ with 16 codebooks. Runs on 4-core CPU.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: NeuTTS is a collection of open-source on-device TTS models with instant voice cloning. Built off LLM backbones with GGUF format quantizations for efficient on-device deployment.
Release Date: Early 2026
| Feature | Value |
|---|---|
| Parameters | 360M (Air), 120M (Nano) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | English, Spanish, German, French |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| On-Device | yes (GGUF quantizations) |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Massive multilingual zero-shot TTS model scaling to 600+ languages. Uses diffusion language model-style discrete non-autoregressive architecture with single-stage text-to-acoustic mapping.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 600+ |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Training-Data | 581k hours |
Features: Simplified single-stage architecture vs conventional two-stage pipelines. Full-codebook random masking strategy with LLM initialization for superior intelligibility. Noise-robust prompt processing.
Links: ![Website][link-website] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Multilingual TTS model with voice cloning and duration control, built on the T5Gemma encoder-decoder LLM architecture. Supports batch generation for multiple audio variations.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 2B-2B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | English, Chinese, Japanese |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Vram | 7.6-10.6 GB |
Features: PM-RoPE positional encoding with XCodec2 audio codec. Low-VRAM options with CPU offloading. Batch inference efficiency with single encoder pass.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: The smallest English TTS model with only 1.6 million parameters. End-to-end neural network achieving ~53x real-time synthesis speed on CPU via ONNX optimization.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | ~3.4 MB (ONNX FP16) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Ultra-compact architecture optimized for CPU-only deployment. Multi-platform support via Python and Node.js APIs. Works on laptops, edge devices, and embedded systems.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: OpenBMB's next-generation tokenizer-free diffusion autoregressive TTS model with 2 billion parameters. Supports 30 languages with automatic detection, voice design from text descriptions, and high-fidelity voice cloning.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 2B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 30 (+ 9 Chinese dialects) |
| Streaming | ✅ |
| Audio Output | 48 kHz |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Tokenizer-free design with LocEnc → TSLM → RALM → LocDiT pipeline. Built-in super-resolution via AudioVAE V2 for 48kHz output.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Soprano is an ultra-lightweight, on-device text-to-speech (TTS) model designed for expressive, high-fidelity speech synthesis at unprecedented speed. The 1.1 release ships an 80M-parameter backbone that achieves up to 20× real-time generation on CPU and 2000× real-time on GPU, with lossless streaming (<250 ms latency on CPU, <15 ms on GPU), <1 GB memory usage at inference, and infinite generation length (automatic text splitting). Output sample rate is 32 kHz, with widespread device support (CUDA / CPU / MPS on Windows, Linux, and Mac). Inference is production-ready through an OpenAI-compatible endpoint, ONNX, WebUI, CLI, and ComfyUI nodes. The base 1.1 model is ekwek/Soprano-1.1-80M on HuggingFace; a fine-tuning toolkit (soprano-factory) was released January 13 2026 alongside the 1.1 weights — the 1.1 release reports 95% fewer hallucinations and a 63% preference rate over 1.0 (Soprano-80M). A live demo runs on ekwek/Soprano-TTS HF Space.
Release Date: December 22, 2025
| Feature | Value |
|---|---|
| Parameters | 80M (Soprano-1.1-80M) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English (US/UK family voices, per HF Space) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Sample Rate | 32,000 Hz |
| Inference Targets | OpenAI-compatible endpoint, ONNX, WebUI, CLI, ComfyUI |
| Performance Cpu | up to 20× real-time |
| Performance Gpu | up to 2000× real-time |
| Memory | <1 GB at inference |
| Text Length | infinite (automatic text splitting) |
| Devices | CUDA, CPU, MPS (Windows, Linux, Mac) |
| Training Toolkit | soprano-factory (https://github.com/ekwek1/soprano-factory) |
| History 1 1 | Soprano-1.1-80M released 2026-01-14 (95% fewer hallucinations; 63% preference over 1.0) |
| History 1 0 | Soprano-80M released 2025-12-22 |
Features: The defining trade-off of this release is extreme on-device
efficiency at sub-100M scale: an 80M-parameter backbone hits
<250 ms CPU latency and <15 ms GPU for lossless streaming
while keeping inference within <1 GB of memory — well under the
multi-billion-parameter budget that newer conversational TTS
systems require. The release pairs the model with soprano-factory
(open-source training/fine-tuning toolkit) so users can build their
own voices on top of the same backbone, and one installation can
drive OpenAI-compatible / ONNX / WebUI / CLI / ComfyUI inference.
The 1.1 update is a measured iteration: 95% fewer hallucinations
and a 63% preference over 1.0 at the same parameter budget, so the
measurable quality jump ships with no added inference cost.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: High-quality TTS synthesis system based on LLMs from ZhipuAI, supporting zero-shot voice cloning with Multi-Reward Reinforcement Learning.
Release Date: December 11, 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Echo is a 2.4B-parameter diffusion-based diffusion transformer (DiT) text-to-speech model. It conditions on target text and up to two minutes of speaker reference audio, generates Fish Speech S1-DAC latents, and decodes to 44.1 kHz audio. Output length is up to 30 seconds per segment. The model is fast at single-sample generation: on one A100, generating 30 seconds of audio from a 120-second prompt takes ~1.45 seconds (RTF < 0.05) — substantially faster than frontier autoregressive approaches at similar quality. The architecture is a deliberate pivot from the author's prior autoregressive-in-DAC-space model Parakeet, which struggled with semantic-consistency retries and weak voice cloning; Echo's diffusion approach trades off real-time interactivity for fast, high-fidelity zero-shot voice cloning in offline synthesis. Trained via the TPU Research Cloud (TRC). Demo (preview) hosted on jordand/echo-tts-preview HF Space; base model on jordand/echo-tts-base.
Release Date: December 4, 2025
| Feature | Value |
|---|---|
| Parameters | 2.4B (DiT) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English (per demo samples) |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Architecture | diffusion transformer (DiT) in Fish Speech S1-DAC latent space |
| Max Segment Duration | 30 s |
| Sample Rate | 44,100 Hz |
| Speaker Reference Max | 120 s |
| Performance A100 Rt | 30 s output in ~1.45 s (RTF < 0.05) |
| Audio Codec | Fish Speech S1-DAC |
| Prior Model | Parakeet (autoregressive in DAC space) |
| Training Infrastructure | TPU Research Cloud (TRC) |
| Inference Requirements | CUDA-capable GPU with at least 8 GB VRAM; Python 3.10+ |
| Sampler | euler CFG with independent guidances for text (3.0) and speaker (8.0); 40 steps; sequence_length 640 default |
| License Clarification | MIT (per GH repo license) |
Features: Echo is a deliberate next-step pivot from autoregressive-in-DAC-space TTS to a full diffusion approach. The author's prior model, Parakeet, generated DAC tokens autoregressively but suffered the classic AR weakness — semantic-consistency retries — and weak voice cloning. Echo keeps Fish Speech S1-DAC latents (so the audio representation is the same proven codec) but moves the generator upstream to a 2.4B DiT operating directly on those latents, conditioned on a long (up to 2-minute) speaker reference. The result: 30-second outputs in ~1.45 s on a single A100 (RTF < 0.05) with high-fidelity zero-shot cloning — fast enough that the "diffusion is too slow" objection no longer applies at the segment length that matters for offline content generation, while the AR class's retry-induced inconsistency is gone by construction.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: Real-time TTS model from Microsoft with streaming text input and ultra-low latency (~300ms).
Release Date: December 3, 2025
| Feature | Value |
|---|---|
| Parameters | 0.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Multilingual |
| Streaming | ✅ |
| License | ![MIT][license-mit] |
| Max Duration | ~10 minutes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Advanced TTS system based on LLMs for zero-shot multilingual speech synthesis from FunAudioLLM.
Release Date: December 2025
| Feature | Value |
|---|---|
| Parameters | 0.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 9 + 18+ Chinese dialects |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Liquid AI's first end-to-end audio foundation model with low latency and real-time conversation.
Release Date: November 28, 2025
| Feature | Value |
|---|---|
| Parameters | 1.5B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ✅ |
| License | ![LFM][license-lfm] |
Links: ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Marvis is a conversational real-time streaming TTS from Marvis-Labs. The architecture inherits Sesame's CSM-1B (Conversational Speech Model): a 250M-parameter multimodal backbone that processes interleaved text + audio tokens and a smaller 60M-parameter audio decoder that models the remaining 31 RVQ codebook levels to reconstruct high-quality speech from the backbone's representations. Audio tokens come from Kyutai's mimi codec (RVQ tokens). The dual-transformer split — semantic backbone + small decoder — yields sub-second latency, and the model is built for on-edge / on-device deployment (Apple Silicon / iPad / iPhone / Mac). Two operational choices distinguish Marvis:
Release Date: November 6, 2025
| Feature | Value |
|---|---|
| Parameters | 250M (multimodal backbone) + 60M (audio decoder) = 310M total |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English, French, German |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | dual-transformer CSM-1B (Conversational Speech Model) — multimodal backbone + audio decoder |
| Audio Codec | Kyutai mimi codec (RVQ tokens; backbone models codebook 0, decoder models codebook 1–31) |
| Quantized Size | ~500 MB (4-bit MLX) |
| Training Dataset | amphion/Emilia-Dataset |
| Library Name | transformers, mlx, mlx-audio |
| Inference Cli | mlx_audio.tts.generate --model Marvis-AI/marvis-tts-250m-v0.2 --stream --text "..." [--ref_audio ./x.wav] |
| Variants In Collection | 250m-v0.2, 250m-v0.2-MLX-{4bit,6bit,8bit}, 100m-v0.2 (+ MLX variants), 250m-v0.2-transformers |
| Emits Text Chunking | no (full-sequence contextual processing) |
Features: Two operational choices make Marvis stand out among conversational TTS. First, no regex chunking: most streaming TTS engines pre-split sentences by regex patterns before feeding them to the generator, which can disrupt flow / intonation; Marvis processes the entire text contextually, treating the text as a single interleaved multimodal sequence. Second, the dual-transformer CSM-1B design — a 250M backbone for codebook 0 (semantic) and a smaller 60M audio decoder for codebooks 1-31 (acoustic) — produces a quantized footprint of ~500 MB, enabling on-device Apple-Silicon inference (iPad / iPhone / Mac) with real-time streaming. The architecture makes a high-quality CSM-style TTS with zero-shot cloning actually deployable at the edge, while the official collection's 4 / 6 / 8-bit MLX variants let users trade footprint for fidelity on a per-device basis.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: AI-Enhanced Text-to-Speech System with Intelligent Optimization and self-learning capabilities.
Release Date: November 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Multi-Speaker | yes (1-4 speakers) |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: State-of-the-art speech model for expressive voice generation with natural language voice control.
Release Date: November 2025
| Feature | Value |
|---|---|
| Parameters | 3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | English (Multi-accent) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: 3B-parameter LLM-based RL audio model specialized in expressive and iterative audio editing.
Release Date: November 2025
| Feature | Value |
|---|---|
| Parameters | 3B (4B BF16) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Mandarin, English, Sichuanese, Cantonese, Japanese, Korean |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: KaniTTS is a 370M-parameter two-stage text-to-speech model from nineninesix-ai. The architecture pairs an LFM-2 backbone LLM (Liquid Foundation Model v2 — a non-transformer, structured state-space architecture) with a neural audio codec for output waveform synthesis. The LLM generates compressed audio-token representations and the codec renders them to 22 kHz waveforms, yielding low-latency generation: ~1 s to produce 15 s of audio on a single RTX 5080, with 2 GB GPU VRAM at inference, and MOS 4.3 / WER < 5% quality on the project's benchmarks. Languages covered: English, German, Chinese, Korean, Arabic, Spanish across multiple per-language voices (English, German, Chinese, Korean, Arabic, Spanish each ship a 400M checkpoint; Japanese ships a 370M "Expo-2025-Osaka" variant; a multilingual 370M checkpoint is also available). MLX variants exist for Apple-Silicon inference. The codec is the same author's nemo-nano-codec-22kHz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported to ~12.5 fps / 0.6 kbps). The model is part of the nineninesix-ai family alongside Gepard.
Release Date: September 30, 2025
| Feature | Value |
|---|---|
| Parameters | 370M (kani-tts-370m multilingual); 400M per-language (en / de / zh / ko / ar / es); 370M (expo2025-osaka-ja) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English, German, Chinese, Korean, Arabic, Spanish (multilingual 370M checkpoint); Japanese (Expo-2025-Osaka variant) |
| Streaming | ✅ |
| License | ![LFM][license-lfm] |
| Sample Rate | 22,000 Hz |
| Backbone Llm | LFM-2 (Liquid Foundation Model; non-transformer structured state-space architecture) |
| Audio Codec | nineninesix/nemo-nano-codec-22khz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported) |
| Performance Rt 5080 | ~1 s for 15 s audio on RTX 5080 |
| Memory | 2 GB GPU VRAM at inference |
| Quality Mos | 4.3 / 5 (naturalness) |
| Quality Wer | <5% (accuracy) |
| Training Dataset | ~80k hours (LibriTTS, Common Voice, Emilia) |
| Training Hardware | 8x H100 GPUs, 45 hours on Lambda AI |
| Per Language Models | kani-tts-400m-{en,zh,de,ar,es,ko} on HuggingFace |
| Pretrained Checkpoints | 0.2-pt (450M), 0.3-pt (400M) for custom posttraining / fine-tuning |
| Mlx Variants | kani-tts-370m-MLX (Apple Silicon) |
| Arxiv | 2505.20506 |
Features: KaniTTS's design choice worth flagging: it pairs a non-transformer
LFM-2 backbone (Liquid Foundation Model — structured state-space
rather than attention) with a neural audio codec for output at
the 370M scale. The choice lets the model hit a ~1 s / 15 s
audio generation rate on a 2 GB GPU VRAM budget — sub-1B
parameters, sub-entry-tier GPU requirement, but still multilingual
across six languages. The two-stage approach (LLM → codec) is
conventional; what's less conventional is the choice of a
state-space backbone over the usual transformer decoder at this
scale, hitting latency / VRAM numbers that open up sub-1B real-time
TTS on consumer-grade hardware. The same author ships soprano-factory-style
companion assets (pretrained v0.2-pt / v0.3-pt checkpoints, a
NeMo NanoCodec MLX port) to lower the bar for fine-tuning on custom
datasets.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: This is an unofficial, work-in-progress LoRA fine-tuning toolkit for the VibeVoice TTS / speech model (1.5B-base and 7B-base checkpoints). The base VibeVoice checkpoints are the same ones covered in this list's a separate entry: audio-conditioned diffusion TTS for spoken dialogue, streaming, etc. This toolkit takes pretrained VibeVoice weights + a paired (text, audio, optional reference-audio) dataset and trains a LoRA adapter against two losses simultaneously:
Release Date: September 16, 2025
| Feature | Value |
|---|---|
| Parameters | 1.5B (LoRA-adapted) / 7B (LoRA-adapted) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | inherits base VibeVoice coverage |
| Streaming | not directly (toolkit output is a LoRA adapter; the adapter inherits VibeVoice inference shape) |
| License | ![MIT][license-mit] |
| Loss Text | masked cross-entropy on text tokens |
| Loss Acoustic | diffusion MSE on acoustic latents |
| Hardware 1 5B | ≥16 GB VRAM |
| Hardware 7B | ≥48 GB VRAM |
| Transformers Version | 4.51.3 (known-good; other versions may break on Qwen2 architecture) |
| Tested Docker Image | runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04 |
| Audio Target Format | 24 kHz audio (paired dataset of target-audio + transcripts + optional reference-audio prompts) |
| Training Entrypoint | python -m src.finetune_vibevoice_lora --model_name_or_path aoi-ot/VibeVoice-Large --processor_name_or_path src/vibevoice/processor --dataset_name <your/dataset> --text_column_name text [--voice_column_name audio_ref] |
| Supports Hf Dataset Loader | yes |
| Output | LoRA adapter compatible with VibeVoice base |
Features: The dual-loss trick is the technical center of this toolkit. Naive LoRA fine-tuning of a unified TTS model often specializes the synthesis but silently damages the text LLM capability the base inherited from its Qwen-class backbone; the "masked CE on text tokens + diffusion MSE on acoustic latents" two-headed loss preserves both competencies at training time. Pair that with the careful pinning of Transformers 4.51.3 (other versions break on the Qwen2 architecture) and a documented minimum-VRAM budget per base size (16 GB for 1.5B, 48 GB for 7B), and you get a reproducible recipe for community fine-tuning of VibeVoice — something the official Microsoft VibeVoice release doesn't ship out-of-the-box.
Links: ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Tokenizer-free TTS system for context-aware speech generation and true-to-life voice cloning.
Release Date: September 16, 2025
| Feature | Value |
|---|---|
| Parameters | 640M-800M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Long-form streaming TTS system for multi-speaker dialogue generation with stable, natural speech.
Release Date: September 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | EN, ZH, JP, KO, FR, DE, RU |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Multi-Speaker | yes (4 speakers) |
| Max Duration | 3 minutes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: NVIDIA ADLR's fully open-source Large Audio Language Model with state-of-the-art audio understanding. Audio Flamingo Next (AF-Next) is the latest generation featuring stronger general audio understanding, longer context support, and timestamp-grounded reasoning.
Release Date: July 2025 (AF3), 2026 (AF-Next)
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ❌ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | Multi-lingual |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Context | Up to 30 minutes |
Features: Key Innovation (AF-Next): Staged curriculum training with GRPO-based RL post-training. Three specialized checkpoints: Instruct, Think (reasoning), and Captioner. Temporal Audio Chain-of-Thought grounding intermediate reasoning to timestamps.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Fast and high-quality zero-shot TTS models based on flow matching.
Release Date: June 16, 2025
| Feature | Value |
|---|---|
| Parameters | 123M |
| Languages | Chinese, English |
| License | ![Apache 2.0][license-apache-2.0] |
| Zero-Shot-Cloning | yes |
| Dialogue | yes |
Links: ![GitHub][link-github] ![Website][link-website] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: State-of-the-art open source TTS and voice cloning model that generates natural, realistic, and emotionally rich speech.
Release Date: May 31, 2025 (v1.5.1)
| Feature | Value |
|---|---|
| Parameters | 4B (S1), 0.5B (S1-mini) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 8 (EN, JP, KO, ZH, FR, DE, AR, ES) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Rtf | ~1:7 |
Links: ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Family of SOTA open-source TTS models by Resemble AI, covering a single-language English line plus a multilingual V3 release that brings broader language coverage, more consistent speaker similarity, reduced hallucinations, and more natural conversational speech across 23+ languages.
Release Date: April 24, 2025
| Feature | Value |
|---|---|
| Parameters | 500M (Llama backbone, 0.5B) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 23+ |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
Features: First open-source TTS model with explicit emotion exaggeration control, plus an alignment-informed inference pipeline and a watermarked decoder. Multilingual V3 narrows the quality gap to closed systems like ElevenLabs on cross-language voice cloning while staying under 1B parameters.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: SOTA open-source TTS built on Llama-3b backbone demonstrating emergent capabilities of LLMs for speech synthesis.
Release Date: April 2025
| Feature | Value |
|---|---|
| Parameters | 3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Multilingual |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Advanced zero-shot speech synthesis with Sparse Alignment Enhanced Latent Diffusion Transformer.
Release Date: March 22, 2025
| Feature | Value |
|---|---|
| Parameters | 0.45B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Efficient LLM-Based TTS Model with Single-Stream Decoupled Speech Tokens, built on Qwen2.5.
Release Date: March 2025
| Feature | Value |
|---|---|
| Parameters | 0.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Production-ready open-source framework for intelligent speech interaction with unified speech comprehension and generation.
Release Date: February 17, 2025
| Feature | Value |
|---|---|
| Parameters | 130B (Chat), 3B (TTS) |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English, Japanese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Kokoro is an open-weight Text-to-Speech model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.
Release Date: January 27, 2025 (v1.0)
| Feature | Value |
|---|---|
| Parameters | 82M |
| Architecture | StyleTTS 2, ISTFTNet |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 8 (54 voices) |
| Streaming | ✅ |
| Cost | <$0.06 per hour of audio |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: KokoClone is a fast, real-time compatible multilingual voice cloning system built on top of Kokoro-ONNX. It enables users to type text in multiple languages, provide a short 3-10 second reference audio clip, and instantly generate speech in that same voice.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | 82M (Base: Kokoro-ONNX) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | 7 (En, Hi, Fr, Ja, Zh, It, Pt, Es) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Lightweight ZipVoice-based TTS model for high quality voice cloning at speeds exceeding 150x realtime.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | - |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Rtf | 150x |
| Vram | 1GB |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Audio Language Model by Xiaomi functioning as a Few-Shot Learner with SOTA audio understanding.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | Multi-lingual |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: SOTA Multi-Speaker TTS model for generating realistic long-form podcasts with dialectal diversity.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Mandarin, English, Cantonese, Sichuanese, Henanese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Max Duration | 90+ minutes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Advanced on-device Vietnamese TTS model with instant voice cloning from 3-5 seconds of reference audio.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | 0.3B-0.6B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | Vietnamese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: 1.6B parameter TTS model by Nari Labs for generating ultra-realistic dialogue in one pass.
Release Date: June 27, 2024
| Feature | Value |
|---|---|
| Parameters | 1.6B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: MeloTTS is a high-quality multi-lingual text-to-speech library from MyShell.ai in collaboration with MIT, supporting English (American, British, Indian, Australian, and a default accent), Spanish, French, Chinese (with mixed Chinese–English capability), Japanese, and Korean. Built on VITS / VITS2 / Bert-VITS2 family work and packaged with both a Python API and a Web UI, it runs fast enough for CPU real-time inference.
Release Date: February 19, 2024
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English (American, British, Indian, Australian, Default), Spanish, French, Chinese, Japanese, Korean |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Base | VITS / VITS2 / Bert-VITS2 family |
| Mixed Chinese English | yes |
Features: A multi-accent multilingual TTS library that ships both a Python API and a Web UI on top of the VITS-style architecture, with explicit English-accent coverage (American, British, Indian, Australian, Default) and mixed Chinese–English output — designed for fast CPU real-time inference without requiring GPU servers.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Open-source audio foundation model by Moonshot AI for audio understanding, generation, and conversation.
Release Date: 2024
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | Multi-lingual |
| Streaming | ✅ |
| License | ![MIT][license-mit] ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: eSpeak NG is a compact open-source software text-to-speech synthesizer for Linux, Windows, Android, and other operating systems. It supports more than 100 languages and accents, is a fork of Jonathan Duddington's original eSpeak engine, and uses the formant synthesis method: the model produces speech by explicitly computing the acoustic resonances (formants) of each phoneme, not by concatenating human-speech recordings. The trade is well-known — the speech is clear and usable at high playback speeds, but not as natural or smooth as larger neural or concatenative synthesizers. The compensation is size: the program and its data, including many languages, total a few megabytes. Other synthesis methods supported: Klatt formant synthesis and MBROLA diphone back-end via the documented integration.
Release Date: December 8, 2015
| Feature | Value |
|---|---|
| Parameters | n/a (formant-synthesis engine; not a neural model) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | 100+ languages and accents (see docs/languages.md) |
| Streaming | ✅ |
| License | ![GPL 3.0][license-gpl-3.0] |
| Synthesis Method | formant synthesis (primary); Klatt formant synthesis (secondary); MBROLA diphone backend (optional) |
| Footprint | a few MB (program + data + many languages) |
| Audio Output | WAV file (CLI), direct playback, or shared-library API |
| Input Formats | text from file / stdin (CLI), SSML (partial), HTML (partial) |
| Packages | CLI (espeak-ng man page), shared library (libespeak-ng), SAPI5 Windows module |
| Supersedes | eSpeak (Jonathan Duddington's original engine) |
| Downstream Usage | G2P / phonemizer for neural TTS pipelines (e.g. sanoTTS bundles espeak-ng for its duration model) |
| Platforms | Linux, Windows, Android, Solaris, Mac OS X |
Features: eSpeak-NG is the canonical reference implementation of compact multi-language formant synthesis. Its 100+-language coverage in a few megabytes, plus SSML / SAPI5 / MBROLA / shared-library / CLI surfaces, are matched only by neural TTS systems that are orders of magnitude larger. The reason it belongs in a list whose other entries are neural TTS systems is its continued quiet role in the neural stack as a G2P / phonemizer front-end — the phoneme inventory and grapheme-to-phoneme rules that sanoTTS and similar sub-1B neural TTS engines bundle are often just a port of eSpeak-NG's language-data files. So even if the formant-synthesis audio output itself has been surpassed for naturalness, the phoneme infrastructure underneath many of the smaller neural TTS entries on this list still traces back to eSpeak-NG.
Links: ![GitHub][link-github]
· · · · · · · · · · · · · ·
Models that can generate audio from multiple input modalities (video, text, image, audio). These are unified frameworks for multimodal audio synthesis.
| Model | Text | Video | Audio | Max Duration | Sample Rate | License |
|---|---|---|---|---|---|---|
| MiDashengLM-Gen | ✅ | ❌ | ❌ | — | 16 kHz | ![Apache 2.0][license-apache-2.0] |
| ScenA | ✅ | ❌ | ✅ | — | — | ![Other][license-other] |
| Nemotron-Labs-Audex-2B | ✅ | ❌ | ✅ | — | — | ![NVIDIA NC][license-nvidia-noncommercial] |
| Nemotron-Labs-Audex-30B-A3B | ✅ | ❌ | ✅ | — | — | ![NVIDIA NC][license-nvidia-noncommercial] |
| MOSS-SoundEffect | ✅ | — | — | 30 s | 48 kHz | ![Apache 2.0][license-apache-2.0] |
| Omni2Sound (Omni2Audio) | ✅ | ✅ | ✅ | — | — | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| ControlFoley | ✅ | ✅ | ✅ | — | 44,100 Hz | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Woosh | ✅ | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
| Chroma-4B | ✅ | ❌ | ✅ | — | — | ![Apache 2.0][license-apache-2.0] |
| Uni-MoE (Audio) | ✅ | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
| AudioX / Audio-Omni | ✅ | ✅ | ✅ | — | — | ![Apache 2.0][license-apache-2.0] ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| HunyuanVideo-Foley | ✅ | ✅ | — | — | 48 kHz | ![Research Only][license-research-only] |
| PrismAudio | — | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
| ThinkSound | ✅ | — | ✅ | — | — | ![Apache 2.0][license-apache-2.0] |
| MMAudio | ✅ | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
Description: MiDashengLM-Gen (MiDasheng Language Model for Generation) is an end-to-end framework for unified audio-scene generation from Xiaomi. Built on a pre-trained LLM and the Dasheng audio tokenizer, it couples per-token conditional flow matching with autoregressive generation to produce coherent 16 kHz audio that simultaneously blends speech, music, sound effects and environmental acoustics from a structured text description. It supports 9 languages with emotion control and approaches dedicated TTS intelligibility on speech (Seed-TTS English WER drops from 12.15% to 2.79%) while retaining mixed-audio scene capability, and extends competitively to multilingual settings.
Release Date: August 12, 2026
| Feature | Value |
|---|---|
| Text | ✅ |
| Video | ❌ |
| Image | ❌ |
| Audio | ❌ |
| Sample Rate | 16 kHz |
| Languages | 9 |
| Emotion Control | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 1.7B (Qwen3-1.7B backbone) |
| Architecture | DashengTokenizer (768-dim @25Hz) + Qwen3-1.7B + flow-matching DiT (16 layers, hidden 2048) |
Features: LLM-conditioned high-dimensional (768-dim @25Hz) audio latents generated without quantization artifacts; audio-text alignment pre-training maps latents into the LLM token space before generation; a learned stop head enables variable-length truncation. First end-to-end trained model for general text-to-audio-scene generation.
Links: ![Demo][link-demo] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: ScenA generates multi-speaker audio scenes — dialogue and conversation with sound effects and ambience — from a text prompt, conditioned on one or more reference-audio clips that set the speakers' voices. Unlike prior multi-speaker dialogue systems it uses no per-turn tags, multi-stream transcripts, or speaker embeddings: a free-form natural-language prompt alone describes the scene. The text prompt determines which reference voice speaks where, allowing overlapping speech, spontaneous paralinguistic events, and scene-level ambient sound — all inherited from the in-the-wild text-to-audio pretraining distribution. The architecture is an audio-only, reference-conditioned flow-matching DiT built on the LTX-2 backbone (~4B parameters, 48 layers). Reference latents are concatenated into the token sequence and distinguished by lightweight identity-aware positional encodings. The training tackles a specifically identified "Reference Shortcut" failure mode — under standard noise schedules the model can identify the matching reference by noisy-target acoustic similarity, bypassing the text prompt — by using a high-noise-biased timestep distribution that forces reliance on the prompt for speaker assignment. Evaluator: CoVoMix2-Dialogue benchmark. Project page, code, paper, and HuggingFace checkpoint are linked below.
Release Date: July 7, 2026
| Feature | Value |
|---|---|
| Parameters | ~4B (DiT, 48 layers; built on LTX-2 architecture) |
| Text | ✅ |
| Video | ❌ |
| Audio | ✅ |
| Max Duration | not stated (scene-level generation) |
| Sample Rate | (not stated; inherits LTX-2 audio VAE) |
| Voice Cloning | ✅ |
| Multi Speaker | yes |
| Ambient Sound | yes (SFX, room acoustics, overlapping speech) |
| Architecture | flow-matching DiT (LTX-2 backbone, audio-only) |
| Speaker Assignment | natural language (no per-turn tags / identity encoders) |
| Training Fix | high-noise-biased timestep distribution (defeats Reference Shortcut) |
| Text Encoder | google/gemma-3-12b-it |
| Audio Vae | bundled (~365 MB; encodes+decodes so full LTX-2 not needed) |
| Checkpoint Size | ~8.2 GB (scena.safetensors) + ~365 MB (audio_vae.safetensors) |
| License | ![Other][license-other] |
| Training Data | in-the-wild text-to-audio pretrained, then reference-conditioned fine-tune |
| Evaluation | CoVoMix2-Dialogue (speaker-binding metrics) |
Features: The "Reference Shortcut" failure-mode identification is the technical center of the work: under standard diffusion noise schedules, a multi-speaker reference-conditioned model can match each reference to the noisy-target segment by acoustic similarity alone, bypassing the text prompt entirely. ScenA's high-noise-biased timestep distribution forces the model to rely on the prompt for speaker assignment at training time. Combined with the absence of any per-turn speaker structure (tags / transcripts / identity encoders) and the prompt's role as the only speaker-routing signal, this yields multi-speaker conversational scenes with overlapping speech, paralinguistic events, and ambient texture that previous structured-supervision multi-speaker systems filter out by design.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Nemotron-Labs-Audex-2B is NVIDIA's smaller sibling of the Audex unified audio-text LLM. Like the 30B-A3B flagship, the 2B is a single model family that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). It is built on the same audio-vocabulary-extended transformer stack as the 30B-A3B but at a densely-parameterized 2B scale (no MoE), so the compute and memory footprint are lowered to a budget tractable on more modest hardware. The 2B checkpoint is the project-tagged SFT variant in the Audex collection — instruction-tuned and ready for inference. Both sizes preserve text reasoning, alignment, knowledge, long-context, and agentic capabilities of the text backbone while adding discrete-token audio I/O.
Release Date: July 6, 2026
| Feature | Value |
|---|---|
| Parameters | 2B (dense; SFT fine-tune, instruct + reasoning-ready) |
| Text | ✅ |
| Video | ❌ |
| Audio | ✅ |
| Modalities | text + audio (input and output) |
| Max Duration | not stated |
| Sample Rate | not stated (decoder output) |
| Voice Cloning | ❌ |
| Audio Understanding | yes (audio QA, classification) |
| Asr | ✅ |
| Speech Translation | yes |
| Text To Speech | yes |
| Text To Audio | yes |
| Speech To Speech Generation | yes |
| Reasoning Mode | yes (thinking + instruct modes inherited from text backbone) |
| License | ![NVIDIA NC][license-nvidia-noncommercial] |
| Pipeline Tag | text-generation |
| Library Name | transformers |
| Derived From | same family as Nemotron-Labs-Audex-30B-A3B |
| Companion 30B | nvidia/Nemotron-Labs-Audex-30B-A3B (MoE: 30B total, 3B active) |
| Spaces | nvidia/Nemotron-Labs-Audex, WaveCut/Nemotron-Labs-Audex, hugging-apps/nemotron-labs-audex-2b |
| Createdat | 2026-07-06T16:21:07Z |
| Downloads | ~2.4k |
Features: The 2B sibling matters because it preserves the central thesis of the Audex paper — unified audio-text LLM intelligence without regressing on text intelligence — while dropping the parameter budget substantially. The 30B-A3B MoE hits a 1M-context, agentic flagship tier; the 2B dense version is the same audio-aware architecture extended down to a budget that doesn't require a high-end MoE serving stack. The pair lets users choose on deployment cost rather than on capability sub-selection: the 2B ships the same audio-to-audio + text-to-audio + audio-understanding
Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Nemotron-Labs-Audex-30B-A3B is NVIDIA's unified audio-text LLM — a single model that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). Built on Nemotron-Cascade-2-30B-A3B (text-only MoE: 30B parameters, 3B active), Audex extends the vocabulary with discrete audio tokens for speech / general-audio output and adds an audio encoder for speech / general-audio input. Runs in thinking and instruct (non-thinking) modes and supports up to a 1M-token context length — preserving text-reasoning, alignment, knowledge, long-context, and agentic capabilities of the backbone while gaining audio tasks.
Release Date: July 6, 2026
| Feature | Value |
|---|---|
| Parameters | 30B MoE (3B active) |
| Modalities | audio (input and output) |
| Audio Understanding | yes |
| Asr | ✅ |
| Speech Translation | yes |
| Text To Speech | yes |
| Text To Audio | yes |
| Speech To Speech Generation | yes |
| Voice Cloning | ❌ |
| License | ![NVIDIA NC][license-nvidia-noncommercial] |
| Languages | English |
| Modes | thinking, instruct (non-thinking) |
| Context Length | 1M tokens |
| Template | ChatML (with <think>…</think> for thinking mode) |
| Inference | vLLM 0.20.0 (recommended) or transformers >= 4.53.0 (mamba-ssm + causal-conv1d required) |
Features: First-class audio I/O for a 30B/3B-active text LLM: extended vocabulary with discrete audio tokens for outputting speech and general audio, plus an audio encoder for input — so the same backbone keeps its strong text reasoning (alignment, knowledge, long-context) and adds ASR + speech translation + TTS + audio generation + S2S without retraining. The MoE form (30B routes, 3B active) keeps inference tractable for a single pipeline that does both.
Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection]
· · · · · · · · · · · · · ·
Description: MOSS-SoundEffect is the dedicated text-to-sound model in the OpenMOSS / MOSI.AI MOSS-TTS family. It turns natural-language captions into high-fidelity non-speech audio (ambience, urban scenes, creatures, human actions, and short music-like clips).
Release Date: May 25, 2026
| Feature | Value |
|---|---|
| Type | Text-to-Sound / SFX generation |
| Conditioning | Text |
| Max Duration | 30 seconds |
| Sample Rate | 48 kHz |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | DiT + Flow Matching + DAC VAE + Qwen3 text encoder |
| Parameters | 1.3B (DiT variant 1.3B) |
| Languages | English, Chinese |
| Inference Defaults | 100 flow-match steps, cfg 4.0, sigma_shift 5.0 |
| Library | diffusers |
Features: Replaces the discrete-token autoregressive v1 (which bottlenecked on vocabulary) with a continuous-latent DiT + Flow Matching paired with a DAC VAE — yielding 30 s stable audio, bilingual English + Chinese prompts, and a clean CFG/sigma-shift inference schedule (cfg 4.0, shift 5.0) that works straight out of the box on the diffusers library.
Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Omni2Sound — also written Omni2Audio on the project page — is a unified VT2A / V2A / T2A framework and a CVPR 2026 Highlight. A single Diffusion Transformer (DiT) backbone with a decoupled two-branch conditioning design:
Release Date: April 20, 2026
| Feature | Value |
|---|---|
| Conditioning | Text / Video / Text+Video |
| Modalities | Video, Audio |
| Asr | ❌ |
| Voice Cloning | ❌ |
| Text | ✅ |
| Video | ✅ |
| Image | ❌ |
| Audio | ✅ |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Tasks | VT2A, V2A, T2A (single model) |
| Architecture | DiT + decoupled Semantic / Temporal branches + 3-stage progressive training |
| Pipeline Tag | text-to-audio |
Features: One single model that is SOTA on three distinct tasks (VT2A, V2A, T2A) without a separate model per mode — decoupled semantic and temporal conditioning let the same DiT backbone handle text-only, video-only, and text+video conditioning by cleanly omitting the missing modality rather than padding it, which is what most prior unified VA models had to do.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Benchmark][link-benchmark]
· · · · · · · · · · · · · ·
Description: ControlFoley (Xiaomi MiLM Plus) is a unified controllable video-to-audio (foley) generation model. It supports four conditioning combinations under one architecture:
Release Date: April 13, 2026
| Feature | Value |
|---|---|
| Conditioning | Text / Video / Text + Video / Video + Reference Audio |
| Modalities | Video (visual), Audio (foley) |
| Asr | ❌ |
| Voice Cloning | ❌ |
| Text | ✅ |
| Video | ✅ |
| Image | ❌ |
| Audio | ✅ |
| Sample Rate | 44,100 Hz |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Pipeline Tag | text-to-audio |
| Library | diffusers |
| Cross Modal Conflict | handled via modality-specific control (no explicit router) |
| Inference Skill | ClawHub ControlFoley Audio Generator |
| Upcoming | ComfyUI nodes (in preparation, expanding to V2A / TV2A / TC-V2A / AC-V2A / T2A) |
Features: Modality-specific cross-modal conflict resolution in a single generative stack: text governs semantics, reference audio governs timbre/acoustic style, and video governs temporal synchronization. Rather than routing to a single user-trusted modality, the model decouples control axes so an input disagreement (video shows a dog barking, text asks for a cat) is decomposed into a coherent output that respects each modality's responsibility. Trained with all-modality dropout for modality-robustness, ControlFoley is the first foley system that brings all four conditioning modes — T2A, V2A, TV2A, AC-V2A — under one model.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv] ![Skill][link-skill]
· · · · · · · · · · · · · ·
Description: Sony AI's sound effect foundation model for text-to-audio and video-to-audio generation. Includes Woosh-AE (audio encoder/decoder), Woosh-Flow/DFlow (T2A), and Woosh-VFlow/DVFlow (V2A) with distilled fast inference variants.
Release Date: 2026
| Feature | Value |
|---|---|
| Architecture | Flow-based generative models |
| Text | ✅ |
| Video | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Audio-Encoding | yes |
| Fast-Inference | yes (Distilled models) |
Features: Optimized for sound effects (not general audio) with both public and private model versions. Video-conditioned generation without requiring captions. Competitive with Stable Audio Open and TangoFlux.
Links: ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Chroma 1.0 (FlashLabs' Chroma-4B on HuggingFace) is the first open-source, real-time, end-to-end spoken dialogue model that achieves both sub-second end-to-end latency and high-fidelity personalized voice cloning. The pipeline is end-to-end — no separate ASR → LLM → TTS stitch — speech goes in, speech comes out. The architectural centerpiece is an interleaved text-audio token schedule (1 text : 2 audio) that supports streaming generation, so the model can begin emitting audio while the user is still talking (broken-off turns / barge-in handled). Experimental results from the project's paper:
Release Date: November 28, 2025
| Feature | Value |
|---|---|
| Parameters | 4B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Languages | English (per benchmark reporting) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | end-to-end spoken-dialogue LLM; interleaved text-audio token schedule (1:2); custom_code modules |
| Pipeline Tag | any-to-any (HF classification) |
| Audio Tokenization | chroma tokenizer (RVQ-style per project's tag) |
| Latency Rtf | 0.43 (speech out ~2.3× wall-clock) |
| Speaker Similarity | +10.96% relative improvement over human baseline |
| Inference Library | transformers (custom_code) |
| Correlations With Larger Class | matches full dialogue turn at streaming latency |
| Pretrained | yes (safetensors weights) |
| Hf Space Demos | hysts/Chroma-4B, Pnevka/Chroma-4B |
| History | paper arXiv 2601.11141 (2026-01) |
Features: Two bets together produce the dual property that no prior open-source spoken-dialogue model has hit simultaneously. First, an interleaved text-audio token schedule (1:2) — text tokens and audio tokens are interleaved at a fixed 1:2 ratio through the sequence, which gives the model a structured place to emit audio while still consuming user audio + text context, supporting sub-second end-to-end latency without a separate ASR / LLM / TTS pipeline. Second, personalized voice cloning baked into the spoke-dialogue model — the cloned voice is not bolted on top by a separate TTS stage (as is the default pattern), it's in-model at the audio-token-generation layer. The empirical payoff is a 10.96% relative speaker-similarity gain over the human baseline (i.e. the cloned voice is closer to the reference speaker than two of the same human speaker's recordings are to each other), while hitting RTF 0.43 — a floor that prior systems exceeded either in latency (no streaming) or in cloning fidelity (parrot the speaker poorly), rarely both.
Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: MoE-based omnimodal model with voice cloning, TTS, T2M (text-to-music), and V2M (video-to-music).
Release Date: October 16, 2025 (Uni-MoE-Audio)
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Text | ✅ |
| Video | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Dynamic-Routing | yes |
Links: ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Audio-Omni is the first end-to-end framework unifying understanding, generation, and editing across general sound, music, and speech domains. Presented at SIGGRAPH 2026. AudioX is a unified framework integrating text, video, image, and audio conditions.
Release Date: March 2025 (AudioX), 2026 (Audio-Omni)
| Feature | Value |
|---|---|
| Parameters | - |
| Text | ✅ |
| Video | ✅ |
| Audio | ✅ |
| License | ![Apache 2.0][license-apache-2.0] ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
Features: First unified framework covering all three audio domains. Combines frozen multimodal LLM (Qwen2.5-Omni) with trainable Diffusion Transformer for high-fidelity synthesis. Any-to-any audio processing.
Links: ![GitHub][link-github] ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Tencent's end-to-end video sound effect generation model for professional-grade AI Foley sound generation. Analyzes footage and creates immersive audio that matches the visual content perfectly.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Sample Rate | 48 kHz |
| Text | ✅ |
| Video | ✅ |
| License | ![Research Only][license-research-only] |
| High-Quality-Foley | yes |
| Context-Aware | yes |
Links: ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Video-to-Audio generation framework with Reinforcement Learning and specialized Chain-of-Thought (CoT) planning. Decomposes reasoning into four specialized modules (Semantic, Temporal, Aesthetic, Spatial CoT) for comprehensive video understanding. Built upon ThinkSound.
Release Date: 2025 (ICLR 2026)
| Feature | Value |
|---|---|
| Parameters | 518M |
| Video | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Cot-Planning | yes (4 modules) |
| Multi-Dimensional-Rl | yes |
| Fast-Grpo | yes (Hybrid ODE-SDE) |
| Inference-Time | 0.63 seconds |
Features: Performance Benchmarks:
| Metric | VGGSound | AudioCanvas |
|---|---|---|
| Semantic (CLAP) | 0.47 | 0.52 |
| Temporal (DeSync↓) | 0.41 | 0.36 |
| Aesthetic (MOS-Q) | 4.21±0.35 | 4.12±0.28 |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Unified Any2Audio generation framework with flow matching guided by Chain-of-Thought (CoT) reasoning. Supports generating or editing audio from video, text, audio, or their combinations. Accepted to NeurIPS 2025.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Text | ✅ |
| Audio | ✅ |
| License | ![Research Only][license-research-only] ![Apache 2.0][license-apache-2.0] |
| Cot-Driven-Reasoning | yes |
| Interactive-Object-Centric-Editing | yes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Multimodal joint training framework for high-quality synchronized audio generation from video and/or text inputs. State-of-the-art open source model for generating sounds for videos, images, and text prompts.
Release Date: December 2024 (CVPR 2025)
| Feature | Value |
|---|---|
| Parameters | - |
| Text | ✅ |
| Video | ✅ |
| Image | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Synchronized-Audio | yes |
| Multimodal-Joint-Training | yes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
| Model | Type | Bandwidth Extension | Inpainting | License |
|---|---|---|---|---|
| RE-USE | Universal Speech Enhancement | ✅ | ❌ | ![NVIDIA NC][license-nvidia-noncommercial] |
| NovaSR | Audio Super-Resolution | ✅ | ❌ | ![Apache 2.0][license-apache-2.0] |
| QuarkAudio-UniSE | Universal Speech Enhancement | ❌ | ❌ | ![Apache 2.0][license-apache-2.0] |
| PASE | Speech Enhancement | ❌ | ❌ | ![Apache 2.0][license-apache-2.0] |
| DTT-BSR | Music Source Restoration | ❌ | ❌ | ![MIT][license-mit] |
| NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges) | High-Resolution Audio Restoration | ✅ | ✅ | ![NVIDIA NC][license-nvidia-noncommercial] |
| ZipEnhancer | Acoustic Noise Suppression | ❌ | ❌ | ![Apache 2.0][license-apache-2.0] |
| AudioSR | Audio Super-Resolution | ✅ | ❌ | ![Apache 2.0][license-apache-2.0] |
Description: RE-USE (RE-…), NVIDIA's multilingual universal speech enhancement model, targets distortion–perception trade-off by training a single model that balances listening quality against fidelity to the underlying linguistic / speaker / emotional content. Designed to restore diverse degraded speech while leaving everything else (content, identity, prosody, accent, paralinguistic attributes) intact.
Release Date: March 17, 2026
| Feature | Value |
|---|---|
| Type | Universal Speech Enhancement |
| Bandwidth Extension | ✅ |
| Inpainting | ❌ |
| Sample Rate | 8 / 16 / 22.05 / 24 / 32 / 44.1 / 48 kHz (multi-rate input) |
| Architecture | Mamba-SSM backbone |
| Degradation Coverage | additive noise, reverberation, clipping, bandwidth limit, codec artifacts, packet loss, low-quality mics |
| Language Agnostic | yes |
| License | ![NVIDIA NC][license-nvidia-noncommercial] |
Features: A single Mamba-SSM model that handles seven different input sample rates (no resampling pre-step), covers a broad degradation menu in one checkpoint, stays language-agnostic without per-language training, and explicitly balances distortion reduction against fidelity to the input speech — addressing the universal-SE trade-off that earlier single-purpose enhancers couldn't.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: NovaSR is a tiny audio upsampler (~52 kB parameter count) that bandwidth-extends 16 kHz input up to 48 kHz. Public card on YatharthS/NovaSR advertises realtime factors around 3500× on A100, making it a candidate for real-time on-device super-resolution where model size dominates latency. Inference path is small enough to fit in CPU memory; the use case is speech-bandwidth extension without GPU.
Release Date: January 6, 2026
| Feature | Value |
|---|---|
| Type | Audio Super-Resolution (16 kHz → 48 kHz) |
| Bandwidth Extension | ✅ |
| Inpainting | ❌ |
| Channels | mono |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 52 kB |
| Streamable | yes (low VRAM / runs without GPU) |
| Realtime Factor | ~3500× (A100) |
Features: A 52 kB-parameter Upsampler that hits ~3500× realtime on GPU and runs on CPU — pushing bandwidth extension below the size / latency envelope where a typical neural upsampler is unacceptable (real-time on-device speech enhancement).
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: UniSE is a unified, prompt-free autoregressive speech-enhancement framework built on a decoder-only language model. A single model performs multiple speech-enhancement tasks — speech restoration (SR / denoising), target-speaker extraction (TSE), source separation (SS), and acoustic echo cancellation (AEC, in development) — without explicit task-specific instructions or prompt conditioning; the language model infers the task from the input context. Stack: WavLM as the feature extractor, BiCodec as the discrete codec, and a decoder-only LM as the middle autoregressive backbone. Outputs reconstructed waveform from predicted discrete token sequences.
Release Date: December 22, 2025
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ❌ |
| Streaming | ❌ |
| Languages | English (paper demo) |
| License | ![Apache 2.0][license-apache-2.0] |
| Tasks | Speech Restoration, Target Speaker Extraction, Source Separation, AEC (developing) |
| Architecture | WavLM (feature extractor) + BiCodec (discrete codec) + decoder-only AR-LM |
| Unified | yes (single model handles SE, SR, TSE, SS without explicit task prompts) |
| Prompt Free | yes (LM infers task from input context) |
| Dataset Signals | noise + reverb + packet-loss + clean (configurable per task) |
| Training | Speech-enhancement SFT, then multitask joint training |
Features: A single decoder-only LM that learns the speech-enhancement task distribution and infers which task to perform from the input context — eliminating the need for task-specific prompts, modules, or fine-tuning when switching between denoising, target-speaker extraction, and separation. Built as an autoregressive discrete-token predictor over a WavLM-extracted / BiCodec-quantised representation, it moves the speech-enhancement workflow from a zoo of specialist models into one generalist.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: PASE (Phonologically Anchored Speech Enhancer) is a generative speech-enhancement model from Cisco Collaboration AI that removes noise and reverberation while preserving linguistic content and speaker identity. It uses two fine-tuned WavLM-derived components:
Release Date: November 8, 2025
| Feature | Value |
|---|---|
| Type | Speech Enhancement |
| Bandwidth Extension | ❌ |
| Inpainting | ❌ |
| Sample Rate | 16 kHz mono |
| Architecture | Denoising WavLM (DRD from WavLM-Large) + Dual-Stream Vocoder (phonetic + acoustic) |
| Finetuned From | WavLM-Large |
| Training Data | DN5/DNS5 challenge clean + noise, LibriTTS, VCTK, OpenSLR26+28 RIRs |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Anchors enhancement to phonology instead of spectrum: by reconstructing from a phonetic stream and a separate acoustic stream (per DeWavLM's two representations), PASE keeps the words intact even when the spectrum is severely degraded — substantially lowering hallucinations while still regaining perceptual quality.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: DTT-BSR (DTTNet with BandSequence and RoPE) is a music-source-restoration challenge submission from team AC/DC (Wuhan University) to ICASSP 2026. It is built inside the official MSR-Kit GAN framework, where the baseline generator is replaced by a DTTNet-style time-frequency U-Net and augmented at the bottleneck with:
Release Date: October 16, 2025
| Feature | Value |
|---|---|
| Type | Music Source Restoration |
| Bandwidth Extension | ❌ |
| Inpainting | ❌ |
| Architecture | DTTNet TFC-TDF U-Net (complex STFT) + Improved Dual-Path BandSplitRNN block + RoPE-Transformer |
| Input | complex STFT (real + imag channels; n_fft=2048, hop=512) |
| Discriminator | Multi-Frequency Discriminator (baseline) |
| Framework | MSR-Kit GAN (reconstruction + adversarial + feature-matching losses) |
| License | ![MIT][license-mit] |
Features: Treats music-source restoration as a complex-STFT time-frequency U-Net enhancement at the bottleneck: keep the strong DTTNet dual-path TFC-TDF structure for local spectral patterns, then layer in BandSplitRNN-style sub-band recurrence + RoPE self-attention so the generator can model long-range, cross-band harmonic structure that ordinary GAN baselines miss — critical for restoring non-vocal stems cleanly.
Links: ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: A2SB is NVIDIA's audio-to-audio Schrödinger Bridge diffusion model for high-resolution (44.1 kHz) music restoration. It is the first long-audio restoration model that can restore hour-long inputs without boundary artifacts, and it's end-to-end — predicting waveform outputs directly withou
Truncated — view the full README on GitHub.
37 commits
List of open-source TTS, voice cloning, and music generation models
494
37 commits
updated Sep 10, 2026
A curated list of open-source Text-to-Speech (TTS) and voice cloning models. Models are sorted by release date (newest first).
| Model | Voice Cloning | ASR | Languages | Streaming | License |
|---|---|---|---|---|---|
| AuK | ✅ | ❌ | — | — | ![MIT][license-mit] |
| AuK-Flash | ✅ | ❌ | — | — | ![MIT][license-mit] |
| rumik-oss 1 | ❌ | ❌ | 22 Indic languages + English | — | ![Other][license-other] |
| Irodori-TTS-v4.1-Anime | — | ❌ | Japanese | — | ![MIT][license-mit] |
| ICE-012 Audio | ✅ | ❌ | 590 | ✅ | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| TontaubeV1 | ✅ | ❌ | 7 | ✅ | ![Other][license-other] |
| Breeze TTS 2 | ✅ | ❌ | 2 | ✅ | ![Other][license-other] |
| Sopro v2 Turbo | ✅ | ❌ | 4 | ✅ | ![Apache 2.0][license-apache-2.0] |
| CuteTTS | ✅ | ❌ | 5 | ✅ | ![Apache 2.0][license-apache-2.0] |
| Rynsan TTS | — | ❌ | Khasi, Garo, Pnar, English, Hindi | — | ![CC BY 4.0][license-cc-by-4.0] |
| Audio8 TTS Preview 0.1B | ✅ | ❌ | 8 | — | ![Other][license-other] |
| Kiseki-TTS | ❌ | ✅ | Japanese | — | ![MIT][license-mit] |
| FireRedTTS3 | ✅ | ❌ | 24 | — | ![Apache 2.0][license-apache-2.0] |
| Audio8-TTS-Preview-0.6b | ✅ | ❌ | Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish | ❌ | ![Apache 2.0][license-apache-2.0] |
| NeuTTS-2E | ❌ | ❌ | English | ✅ | ![Other][license-other] |
| Scylla's Band | ❌ | ❌ | en_us, en_gb, es, it | ✅ | ![Apache 2.0][license-apache-2.0] |
| sanoTTS | ❌ | ❌ | English, Nepali, Hindi, Vietnamese, Indonesian, Chinese | ❌ | ![Other][license-other] |
| FreyaTTS | ❌ | ❌ | Turkish | ❌ | ![Apache 2.0][license-apache-2.0] |
| Inflect-Nano-v2 | ❌ | ❌ | English | ❌ | ![Apache 2.0][license-apache-2.0] |
| Gepard | ✅ | ❌ | English, Spanish, Portuguese, Dutch | ✅ | ![Apache 2.0][license-apache-2.0] |
| Higgs Audio v3 TTS | ✅ | ❌ | 102 | ✅ | ![Research Only][license-research-only] |
| dots.tts | ✅ | ❌ | Multilingual | ❌ | ![Apache 2.0][license-apache-2.0] |
| Confucius4-TTS | ✅ | ❌ | 14 | ❌ | ![Apache 2.0][license-apache-2.0] |
| WavTTS | ✅ | ❌ | English, Chinese | ❌ | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| MOSS-TTS | ✅ | ❌ | 31 | ✅ | ![Apache 2.0][license-apache-2.0] |
| VoxFlash-TTS | ✅ | ❌ | Chinese, English | ❌ | ![Apache 2.0][license-apache-2.0] |
| Miso TTS | ✅ | ❌ | English | ❌ | ![MIT][license-mit] |
| Raon-OpenTTS-1B | ✅ | ❌ | English | — | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| OronTTS | ✅ | ❌ | Mongolian, Kazakh | ❌ | ![MIT][license-mit] |
| Supertonic 3 | ✅ | ❌ | 31 | ✅ | ![OpenRAIL-M][license-openrail-m] |
| Scenema Audio | ✅ | ❌ | English, German, French, Spanish, Italian, Portuguese, Japanese, Chinese, Korean, Russian, Arabic, Hindi, Swahili | ❌ | ![Other][license-other] |
| Dramabox | ✅ | ❌ | English | ❌ | ![Other][license-other] |
| Sarashina2.2-TTS | ✅ | ❌ | Japanese, English | ❌ | ![Research Only][license-research-only] |
| LongCat-AudioDiT | ✅ | ❌ | Chinese, English | ❌ | ![MIT][license-mit] |
| SILMA TTS | ✅ | ❌ | Arabic, English | — | ![Apache 2.0][license-apache-2.0] |
| Fish Audio S2 Pro | ✅ | ❌ | 80+ | ✅ | ![Research Only][license-research-only] |
| LongCat-Next | ✅ | ✅ | Chinese, English | ✅ | ![MIT][license-mit] |
| Voxtral-4B-TTS | ✅ | ❌ | 9 | ✅ | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Blue (Light Blue) TTS | ✅ | ❌ | Hebrew, English, Spanish, Italian, German | ❌ | ![MIT][license-mit] |
| KittenTTS | ✅ | ❌ | English, Multiple | ✅ | ![Apache 2.0][license-apache-2.0] |
| Ming-omni-tts | ✅ | ❌ | Chinese, English | ❌ | ![Apache 2.0][license-apache-2.0] |
| SoulX-Singer | ✅ | ❌ | Mandarin, English, Cantonese | ✅ | ![Apache 2.0][license-apache-2.0] |
| SoproTTS | ✅ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Qwen3-TTS | ✅ | ❌ | 10 | ✅ | ![Apache 2.0][license-apache-2.0] |
| TADA | ❌ | ❌ | English | ❌ | ![Other][license-other] |
| Irodori-TTS-500M-v2 | ✅ | ❌ | Japanese | ❌ | ![MIT][license-mit] |
| KugelAudio | ✅ | ❌ | 23 European languages | ✅ | ![MIT][license-mit] |
| LEMAS-TTS | ✅ | ❌ | 10 | ❌ | ![Apache 2.0][license-apache-2.0] |
| MioTTS-2.6B | ✅ | ❌ | English, Japanese | ✅ | ![LFM][license-lfm] |
| MOSS-TTS-Nano | ✅ | ❌ | 20 | ✅ | ![Apache 2.0][license-apache-2.0] |
| NeuTTS | ✅ | ❌ | English, Spanish, German, French | ✅ | ![Apache 2.0][license-apache-2.0] |
| OmniVoice | ✅ | ❌ | 600+ | ❌ | ![Apache 2.0][license-apache-2.0] |
| T5Gemma-TTS | ✅ | ❌ | English, Chinese, Japanese | ❌ | ![MIT][license-mit] |
| TinyTTS | ❌ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| VoxCPM2 | ✅ | ❌ | 30 | ✅ | ![Apache 2.0][license-apache-2.0] |
| Soprano | ❌ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| GLM-TTS | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Echo-TTS | ✅ | ❌ | English | ❌ | ![MIT][license-mit] |
| VibeVoice-Realtime | ✅ | ❌ | Multilingual | ✅ | ![MIT][license-mit] |
| Fun-CosyVoice 3.0 | ✅ | ❌ | 9 + 18+ Chinese dialects | ✅ | ![Apache 2.0][license-apache-2.0] |
| LFM2-Audio-1.5B | ✅ | ✅ | English | ✅ | ![LFM][license-lfm] |
| Marvis-TTS | ✅ | ❌ | English, French, German | ✅ | ![Apache 2.0][license-apache-2.0] |
| IndexTTS2 | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Maya1 | ✅ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Step-Audio-EditX | ✅ | ❌ | Mandarin, English, Sichuanese, Cantonese, Japanese, Korean | ✅ | ![Apache 2.0][license-apache-2.0] |
| KaniTTS | ❌ | ❌ | English, German, Chinese, Korean, Arabic, Spanish | ✅ | ![LFM][license-lfm] |
| VibeVoice-Finetuning | ❌ | ❌ | — | ❌ | ![MIT][license-mit] |
| VoxCPM | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| FireRedTTS2 | ✅ | ❌ | EN, ZH, JP, KO, FR, DE, RU | ✅ | ![Apache 2.0][license-apache-2.0] |
| Audio Flamingo 3 (AF3) / Audio Flamingo Next | ❌ | ✅ | Multi-lingual | ✅ | ![Apache 2.0][license-apache-2.0] |
| ZipVoice | — | — | Chinese, English | — | ![Apache 2.0][license-apache-2.0] |
| Fish Speech | ✅ | ❌ | 8 | ✅ | ![Apache 2.0][license-apache-2.0] |
| Chatterbox | ✅ | ❌ | 23+ | ❌ | ![MIT][license-mit] |
| Orpheus-TTS | ✅ | ❌ | Multilingual | ✅ | ![Apache 2.0][license-apache-2.0] |
| MegaTTS3 | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Spark-TTS | ✅ | ❌ | Chinese, English | ✅ | ![Apache 2.0][license-apache-2.0] |
| Step-Audio | ✅ | ✅ | Chinese, English, Japanese | ✅ | ![Apache 2.0][license-apache-2.0] |
| Kokoro-82M | ✅ | ❌ | 8 | ✅ | ![Apache 2.0][license-apache-2.0] |
| KokoClone | ✅ | ❌ | 7 | ✅ | ![Apache 2.0][license-apache-2.0] |
| LuxTTS | ✅ | ❌ | - | ✅ | ![Apache 2.0][license-apache-2.0] |
| MiMo-Audio | ✅ | ✅ | Multi-lingual | ✅ | ![Apache 2.0][license-apache-2.0] |
| SoulX-Podcast | ✅ | ❌ | Mandarin, English, Cantonese, Sichuanese, Henanese | ✅ | ![Apache 2.0][license-apache-2.0] |
| VieNeu-TTS | ✅ | ❌ | Vietnamese | ✅ | ![Apache 2.0][license-apache-2.0] |
| Dia | ✅ | ❌ | English | ✅ | ![Apache 2.0][license-apache-2.0] |
| MeloTTS | ❌ | ❌ | English, Spanish, French, Chinese, Japanese, Korean | ❌ | ![MIT][license-mit] |
| Kimi-Audio | ✅ | ✅ | Multi-lingual | ✅ | ![MIT][license-mit] ![Apache 2.0][license-apache-2.0] |
| eSpeak-NG | ❌ | ❌ | 100+ | ✅ | ![Other][license-other] |
Description: AuK is a 1.5B foundation model from Tencent for speech generation and editing, trained on millions of hours of diverse audio data. Through a single natural-language instruction interface it unifies an unusually broad task set: zero-shot TTS (speak text in the reference voice) and instruct TTS (voice from a description alone, no reference), content editing (rewrite what is said; even lyric editing that preserves melody and voice), acoustic editing (pitch by semitones, speed, volume), paralinguistic editing (emotion, timbre, de-accent, nonverbal sounds, whisper conversion), and enhancement & separation (denoise/dereverberate, speech separation, music/vocal separation, target-speaker extraction). Architecture: a diffusion transformer with layer-fusion weights, conditioned by a Qwen2.5-Omni-3B MLLM encoder and a separate VAE (loaded at runtime). Day-0 SGLang-Omni serving support, Gradio and ComfyUI integrations, and a task Cookbook are provided. Released under MIT.
Release Date: September 9, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| License | ![MIT][license-mit] |
| Parameters | 1.5B |
| Architecture | diffusion transformer + layer fusion, Qwen2.5-Omni-3B MLLM encoder, separate VAE |
| Variants | AuK (this, base) + AuK-Flash (distilled, 4-step inference) |
| Editing | content, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion |
| Enhancement Separation | speech enhancement, speech separation, music separation, target speaker extraction |
| Deployment | SGLang-Omni (day-0), Gradio, ComfyUI |
Features: Unifies generation and the full editing/enhancement/separation spectrum in one instruction-following model — most systems pick one lane (TTS, or editing, or separation); AuK does zero-shot + instruct TTS, lyric rewriting with melody preservation, emotion/timbre/de-accent/whisper paralinguistic edits, and source separation through the same natural-language interface. A diffusion transformer with layer fusion, distilled into a 4-step AuK-Flash variant for fast inference.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| ComfyUI-AuK | ComfyUI node | ComfyUI-AuK |
· · · · · · · · · · · · · ·
Description: AuK-Flash is the distilled variant of AuK, Tencent's 1.5B foundation model for speech generation and editing, optimized for fast 4-step inference. It exposes the same natural-language instruction interface as the base model: zero-shot TTS (reference voice) and instruct TTS (voice description, no reference), content and lyric editing, pitch/speed/volume acoustic edits, emotion/timbre/de-accent/nonverbal/whisper paralinguistic edits, plus speech enhancement and speech/music/target-speaker separation. Architecture matches the base: diffusion transformer with layer-fusion weights, Qwen2.5-Omni-3B MLLM encoder, and a separate runtime-loaded VAE. Released under MIT.
Release Date: September 9, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| License | ![MIT][license-mit] |
| Parameters | 1.5B |
| Architecture | diffusion transformer + layer fusion (distilled to 4 inference steps), Qwen2.5-Omni-3B MLLM encoder, separate VAE |
| Base Model | tencent/AuK |
| Editing | content, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion |
| Enhancement Separation | speech enhancement, speech separation, music separation, target speaker extraction |
| Deployment | SGLang-Omni, Gradio, ComfyUI |
Features: Distills the AuK foundation model's diffusion transformer down to 4 inference steps, making the full generate-and-edit capability set (including enhancement and separation) practical for interactive use — traded against the base model's maximum quality, with the two variants loadable side-by-side from the same codebase.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: rumik-oss 1 is a 3B multilingual text-to-speech model from rumik ai, trained on fewer than 70,000 hours of speech while performing competitively with existing TTS models. It covers 22 Indic languages in their native scripts and romanized forms plus English, supporting both single-language and code-switched synthesis. Delivery is conditioned via <description="..."> tags (tone, accent, pace), with inline vocalization control (<laugh>, <chuckle>, <sigh>). It extends CohereLabs/tiny-aya-fire with discrete speech tokens from the mimi codec: following the flattened codec-token formulation used in llama-mimi, text conditioning and audio generation share a single autoregressive sequence, predicting eight codebook tokens per frame before advancing, with the frozen mimi decoder reconstructing the 24 kHz waveform. The model ships with 4 fixed voices (Ira, Aisha, Siya, Zoya) that perform equally well across all 22 languages; there is no zero-shot voice cloning. Licensed under Cohere's CC-BY-NC-4.0 with acceptable-use addendum (research and non-commercial use only).
Release Date: September 6, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | 22 Indic languages + English (native scripts and romanized; code-switching supported) |
| License | ![Other][license-other] |
| Parameters | 3B (3,381,533,697 BF16) |
| Architecture | CohereLabs/tiny-aya-fire backbone + flattened mimi codec tokens (8 codebooks/frame, llama-mimi formulation) |
| Audio Codec | kyutai/mimi (frozen decoder), 24 kHz output |
| Pronunciation | ✅ |
| Highlights | description-conditioned delivery (<description> tags), inline vocalizations (<laugh>/<chuckle>/<sigh>) |
| Variants | rumik-oss-1 (post-trained), rumik-oss-1-base (speaker-conditioned pre-post-training) |
Features: Brings competitive multilingual TTS to 22 Indic languages with under 70k training hours, using a flattened mimi codec-token formulation (single autoregressive sequence for text conditioning + audio) on the tiny-aya-fire backbone. Code-switched synthesis, description-conditioned delivery, and inline vocalization tags are first-class capabilities, and its 4 voices perform equally well across all 22 languages — unusual, as most TTS voices are language-specific.
Links: ![HuggingFace][link-huggingface] ![Blog][link-blog] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Irodori-TTS-v4.1-Anime is a Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small using anime-style speech data. Because the base model's annotation pipeline is not publicly documented, the fine-tuning data was annotated independently — so caption conditioning and emoji controls may behave differently from the base model. The full-precision checkpoint (0.8B params, F32) ships at the repository root, with quantized variants (int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic) in subdirectories. It follows the base model's MIT License and ethical restrictions; inference uses the original Irodori-TTS repository.
Release Date: September 4, 2026
| Feature | Value |
|---|---|
| Voice Cloning | — |
| Asr | ❌ |
| Languages | Japanese |
| License | ![MIT][license-mit] |
| Parameters | ~0.8B (766,052,385 F32) |
| Architecture | Irodori-TTS (Aratako) fine-tune; caption-conditioned with emoji controls |
| Base Model | Aratako/Irodori-TTS-v4.1-Small |
| Variants | int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic |
Features: A community fine-tune that ports the Irodori-TTS line into the anime-voice domain using an independently built annotation pipeline (since the base model's is undocumented), and ships the result with five ready-made quantization variants (int4/int8/fp8) for efficient inference.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: ICE-012 Audio is a multilingual text-to-speech model from DarkPs (a FanuonAI organization) with streaming output and reference-based voice cloning. Its defining trait is breadth of language coverage — 590 language names/variants are accepted (name or 2–3-letter ID, with a language-agnostic fallback), including 13 Arabic dialects ("Lahgtna" variants) alongside the full ISO list. It introduces an active acoustic adapter — conditioning codec embeddings before the backbone and refining hidden states after it. Voice is controllable along six axes: gender (male/female), age (child → elderly), pitch (5 levels), accent (10 English accents), style (e.g. whisper), and speed (0.5–2.0×), plus an --auto-voice mode where the model picks a voice automatically. The checkpoint is ~714M parameters (F16) and runs via transformers with trust_remote_code=True. Released under CC BY-NC 4.0.
Release Date: August 29, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 590 names/variants (incl. 13 Arabic Lahgtna dialects; language-agnostic fallback) |
| Streaming | ✅ |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Parameters | ~714M (714,409,993 F16) |
| Architecture | causal LM with active acoustic adapter (conditions codec embeddings pre-backbone, refines hidden states post-backbone) |
| Voice Controls | gender, age, pitch, accent, style, speed; auto-voice mode |
Features: The active acoustic adapter wraps the backbone on both sides — conditioning codec embeddings before it and refining hidden states after — while a six-axis voice-control space (gender/age/pitch/accent/style/speed plus auto-voice) and 590-language coverage make it one of the broadest single-checkpoint TTS releases for dialect and minority-language synthesis.
Links: ![HuggingFace][link-huggingface] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: TontaubeV1 is a multilingual text-to-speech model from TontaubeAI (craitech) designed for expressive voice cloning, long-form generation, and low-latency streaming. Its release contains four causal codebook predictors: CB0 generates semantic audio and duration from text, while progressively smaller CB1–CB3 add acoustic detail. CB0 uses a Qwen3-1.7B-derived transformer trunk and CB1–CB3 progressively shallower Qwen3-0.6B-derived trunks, each with a two-layer audio-token head. The four output streams are decoded with DualCodec, and the inference path uses VibeVoice's acoustic encoder/decoder for continuous reconstruction and streaming. It ships bundled synthetic voices plus zero-shot cloning from up to 60 s of reference audio, with public speaking styles audiobook, conversational, and agentic. Released under the Tontaube Community Model License 1.0, which is explicitly not open-source.
Release Date: August 26, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 7 (English, German primary; Spanish, French, Italian, Dutch, Portuguese secondary) |
| Streaming | ✅ |
| License | ![Other][license-other] |
| Parameters | ~2.87B (2,873,962,498; CB0 1.83B + CB1 449M + CB2 327M + CB3 269M) |
| Architecture | 4-stage Qwen3-derived codebook cascade (CB0–CB3) + DualCodec + VibeVoice decode |
| Styles | audiobook, conversational, agentic |
Features: The four-stage codebook cascade (CB0 semantic+duration → CB1–CB3 progressive acoustic refinement) lets a single multilingual model deliver expressive, long-form, low-latency speech with strong zero-shot cloning. On the 1,088 English zero-shot Seed-TTS examples it posts 1.66% mean utterance-level WER (measured with Whisper large-v3 at semantic temperature 0.6), and the RTX-5090 streaming path reaches ~200 ms to first encoded audio.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Breeze TTS 2 is an open-weight text-to-speech model from BreezeBlue / RESONIA built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following supports reference-free voice design (create a voice from a text description) and reference-guided voice direction (clone a voice while steering tone, emotion, pace, delivery), alongside standard reference-audio voice cloning. Ultra-low-latency streaming reaches 0.32 RTF (≈3.1× real time with the warmed-up fast path) and under 40 ms time-to-first-audio on an NVIDIA H100, emitting 24 kHz PCM. Source code is Apache-2.0; model weights are governed by the BreezeBlue Research and Non-Commercial License (commercial use needs written authorization from RESONIA).
Release Date: August 25, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 2 (English, Chinese) |
| Streaming | ✅ |
| License | ![Other][license-other] |
| Parameters | 3B (3,466,363,713) |
| Architecture | seq2seq backbone + depth decoder + codec with CUDA-graph fast path (no named backbone disclosed) |
| Highlights | #1 open-weight on Artificial Analysis TTS leaderboard; vocal events inline (laugh/cough) |
Features: Pairs natural-language voice control with real-time streaming: a single model handles reference-free voice design (no reference audio needed) and voice direction (clone + steer prosody), and ships a CUDA-graph fast path that hits sub-40 ms TTFA at ~3.1× real time on H100 — open-weight quality that the authors claim exceeds frontier proprietary TTS.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| ComfyUI-Breeze-TTS-2 | ComfyUI node | ComfyUI-Breeze-TTS-2 |
· · · · · · · · · · · · · ·
Description: Sopro (Portuguese for "breath") is a lightweight voice-cloning text-to-speech family. This repo ships sopro-v2-turbo, a 120M-parameter open model that streams and runs comfortably on a laptop CPU or in the browser (ONNX runtime), reaching SOTA-level intelligibility against much larger systems. It supports zero-shot voice cloning from 5–20 s of reference audio, four languages (English, European Portuguese, French, German), and a streaming path with ~300 ms time-to-first-audio on a laptop CPU (0.24 RTF offline / 0.21 RTF streaming on an M3 CPU, 0.07 RTF on H100). Released under Apache-2.0.
Release Date: August 25, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 4 (English, European Portuguese, French, German) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 120M (121,574,193) |
| Deployment | in-browser ONNX runtime; int8 AR weights on CPU; causal vocoder |
| Architecture | autoregressive TTS + chunked-attention streaming path + causal vocoder (F5-TTS/CosyVoice/Vocos lineage acknowledged) |
Features: Packs SOTA-level intelligibility into a 120M footprint that runs in the browser or on a laptop CPU, with a chunked-attention + causal-vocoder streaming path (~300 ms TTFA) — making zero-shot multilingual voice cloning practical for on-device and edge deployment rather than GPU-only serving.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: CuteTTS is a lightweight (~230M-parameter) continuous autoregressive TTS model from OPPO that models continuous latents rather than discrete codec tokens, running efficiently on GPUs, CPUs, and Apple silicon. It delivers ultra-low latency — ~40 ms to the first audio chunk and ~9× real-time throughput on an RTX 4090 — with strong speech quality and zero-shot voice cloning (best-in-comparison 78.9 SIM on LibriSpeech test-clean). Multilingual support covers English, Chinese, French, German, and Spanish. A distilled variant (CuteTTS-distill) trades slight quality for further efficiency. Ships with a web demo, Python API, and CLI.
Release Date: August 24, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 5 (English, Chinese, French, German, Spanish) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | ~230M |
| Architecture | continuous autoregressive modeling of latents + speaker encoder + audio VAE (discrete-codec-free design) |
| Variants | CuteTTS, CuteTTS-distill |
Features: Autoregressively models continuous latent audio representations instead of discrete codec tokens, eliminating codebook-related artifacts and quantization loss at only ~230M parameters. Combined with a lightweight speaker encoder and audio VAE, this yields best-of-class speaker similarity among compared open models and ~40 ms first-chunk latency while remaining practical for CPU/Apple-silicon inference.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Rynsan TTS is a multilingual text-to-speech model that extends k2-fsa/OmniVoice to support Khasi, Garo, and Pnar — languages of Meghalaya, India that have historically had limited representation in modern speech technology. Rather than building a system from scratch, Rynsan retains the multilingual capabilities of the base model while adding speech data for these low-resource Khasic languages. Developed under the Tynrai AI initiative, its broader goal is accessible speech technology for the diverse languages and dialects of Meghalaya, supporting their preservation and use in voice-based applications. A live demo is available at ri.tynrai.in/demo. The repository is gated (manual access approval).
Release Date: August 22, 2026
| Feature | Value |
|---|---|
| Asr | ❌ |
| Languages | 5+ (English, Hindi + extension languages Khasi kha, Garo grt, Pnar pbv; base OmniVoice supports more) |
| License | ![CC BY 4.0][license-cc-by-4.0] |
| Parameters | ~0.61B |
| Architecture | OmniVoice (k2-fsa) multilingual TTS, extended fine-tune |
| Base Model | k2-fsa/OmniVoice |
| Developer | Toiar / Tynrai AI |
Features: Extends a modern multilingual TTS foundation to three substantially under-resourced Khasic languages — a rare production-oriented entry for indigenous-language speech tech, aimed at accessibility, education, and language preservation rather than benchmark leadership.
Links: ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Audio8 TTS Preview 0.1B is the smallest release in the Audio8 TTS family ("the smallest zero-shot TTS worth running"): a ~170M-parameter generative model plus a separate ~120M-parameter codec decoder, making the complete audio generation stack much smaller than most modern multilingual TTS systems. It supports speech generation and zero-shot voice cloning (reference audio + matching transcript). Primary languages are Chinese and English, with German, Spanish, French, Italian, Japanese, and Korean as experimental/multilingual-evaluation targets. Released under the custom Audio8 Community License v1.0: non-commercial use is free, and commercial use is free only for entities with annual revenue under US$2M.
Release Date: August 19, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 8 (Chinese + English primary; de/es/fr/it/ja/ko experimental) |
| License | ![Other][license-other] |
| Parameters | ~0.17B main model (+ ~120M codec decoder) |
| Architecture | Audio8 Falcon H1 DualAR — slow AR (semantic tokens) + fast AR (codec codebooks), 10 codebooks × 4096 entries |
| Audio Codec | bundled 44.1 kHz neural codec (~21.5 frames/s) |
| Context | up to 2,048 packed text/audio positions |
| Variants | 0.1B (this), 0.6B |
Features: Packs practical zero-shot cloning into a ~170M-parameter model using an Falcon-H1-derived DualAR design (slow AR predicts semantic tokens per frame; fast AR predicts the frame's 10 codebooks conditioned on the slow hidden state). On Seed-TTS it posts EN WER 1.662 at only ~0.17B — within reach of 4B+ systems — and ships with its own 44.1 kHz codec so no external codec checkpoint is needed.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Kiseki-TTS is a small, fast Japanese text-to-speech model from telecomadm1145 built on top of Qwen/Qwen3-TTS-Tokenizer-12Hz. It generates discrete neural audio codec tokens at 12.5 Hz (4–6× fewer autoregressive steps than 50–75 Hz codecs) and decodes them to waveform with the Qwen3 TTS codec. The acoustic decoder is a linear-time Mamba2 SSM rather than a self-attention stack, so generation cost is constant per frame — memory does not grow with utterance length and there is no KV cache to manage. Because TTS and ASR were trained jointly in a single multi-task run, the same checkpoint also performs ASR (Japanese speech → text, reading only codec layer 0). It is a single-domain voice (ASMR-style Japanese training data) with no speaker conditioning or voice cloning.
Release Date: August 15, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ✅ |
| Languages | Japanese only (ja) |
| License | ![MIT][license-mit] |
| Parameters | ~0.41B (0.33B backbone + 78M audio branch) |
| Architecture | Transformer encoder (12 layers, bidirectional self-attention) + cross-attention → Mamba2 SSM decoder (6 layers, no causal self-attention) |
| Audio Codec | Qwen3-TTS-Tokenizer-12Hz (12.5 Hz, 16 quantizer layers) |
| Base Model | Kiseki-1.1-0.3B (seq2seq translation model) |
| Training Data | telecomadm1145/asmr_archive_qwentts_encoded |
Features: The decoder deliberately omits causal self-attention — temporal context is carried entirely by the Mamba2 recurrent state while text conditioning enters through cross-attention whose K/V are computed once during prefill. This yields O(1) state per frame (a fixed SSM tensor plus a 3-frame conv window) instead of an O(T) KV cache, so long-form synthesis degrades gracefully past the ~41 s training ceiling instead of hitting a memory cliff. Combined with the 12.5 Hz codec and a shared multi-token-prediction head that resolves all 16 codebook layers in one trunk pass, the model is both compute- and memory-bandwidth-bound rather than quadratic in length.
Links: ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: FireRedTTS3 is a unified speech generation and editing system from the FireRed Team built on semantically enriched continuous speech representations. It ships in two variants: FireRedTTS3-Base (zero-shot voice cloning across 24 languages and 21 Chinese dialects) and FireRedTTS3-Instruct (natural-language voice design and combined semantic + acoustic speech editing in one model). Beyond cloning, it supports instruction-based voice design (no reference audio needed) and editing operations such as insertion / deletion / substitution (semantic) and speed / pitch / volume changes (acoustic).
Release Date: August 5, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 24 (plus 21 Chinese dialects) |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | Qwen3 backbone + patch-level diffusion autoregressive (DiTAR) + RedAE codec + CAM++ speaker encoder |
| Variants | Base (cloning), Instruct (cloning + voice design + editing) |
Features: Represents speech with semantically enriched continuous (non-quantized) representations, enabling a single system to do zero-shot cloning, text-driven voice design, and fine-grained semantic + acoustic editing. On Seed-TTS-eval it reaches an average WER/CER of 3.04% with 78.8% speaker similarity; MiniMax-MLS-Test average SIM 84.8%.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| FireRedTTS3-ComfyUI | ComfyUI node | FireRedTTS3-ComfyUI |
· · · · · · · · · · · · · ·
Description: Audio8 TTS Preview 0.6B is a 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning. It uses a DualAR architecture inspired by Fish Audio S2 Pro: a slow AR transformer predicts one semantic token per audio frame, and a fast AR transformer predicts the frame's codec codebooks conditioned on the slow hidden state and preceding codebooks. The bundled 44.1 kHz neural audio codec handles both reference-audio encoding and waveform decoding — no additional codec checkpoint is required. The model supports 11 recommended languages (Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish) with zero-shot voice cloning from a reference audio clip + matching transcript.
Release Date: July 28, 2026
| Feature | Value |
|---|---|
| Parameters | 601,159,424 (0.6B, excluding the codec) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish (11) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | DualAR (slow AR + fast AR), inspired by Fish Audio S2 Pro |
| Slow Ar | 24 layers, width 896, 14 attention heads, 2 KV heads |
| Fast Ar | 4 layers, width 896, 14 attention heads, 2 KV heads |
| Acoustic Tokens | 10 codebooks, 4,096 entries per codebook |
| Codec | 44.1 kHz, 2,048 samples per model frame (~21.5 frames/s), bundled (no external codec needed) |
| Context Length | up to 2,048 packed text/audio positions |
| Sample Rate | 44,100 Hz |
| Inference | transformers with trust_remote_code=True; CUDA-capable GPU recommended |
| Dependencies | torch>=2.5.0, torchaudio>=2.5.0, transformers>=4.57.0,<5, soundfile, safetensors |
| Preview Status | language coverage intentionally limited; broader multilingual + Chinese dialect support planned |
| Library Name | transformers (custom_code) |
| Pipeline Tag | text-to-speech |
| Createdat | 2026-07-28T07:53:00Z |
Features: The DualAR architecture is the technical centerpiece: rather than a single autoregressive decoder predicting all codebook levels sequentially (the standard codec-LLM TTS pattern), Audio8 splits the work into a slow AR that predicts one semantic token per audio frame and a fast AR that predicts the frame's remaining codec codebooks conditioned on the slow hidden state. This separation lets the semantic-level reasoning happen at the slow AR's 24-layer depth while the acoustic codebook prediction stays lightweight at 4 layers — reducing the total compute per frame without sacrificing semantic quality. The bundled 44.1 kHz codec (no external codec checkpoint needed) and the 10-codebook / 4,096-entry acoustic token design give the model self-contained high-fidelity output at a compact 0.6B scale, making it one of the smallest multilingual zero-shot-cloning TTS systems shipping at 44.1 kHz.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: NeuTTS-2E is a super-fast, highly realistic, on-device emotional text-to-speech model from Neuphonic. It is the next generation after NeuTTS Air / Nano (which continue to ship for multilingual + zero-shot-cloning contexts) — narrowed in scope to an English-only alpha focused on:
Release Date: July 21, 2026
| Feature | Value |
|---|---|
| Parameters | 0.2B (compact LM backbone + codec) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English (English-only alpha) |
| Streaming | ✅ |
| License | ![Other][license-other] |
| Backbone | compact LM backbone tuned for emotional TTS token generation |
| Codec | efficient codec (compact, paired with the LM) |
| Speakers | 4 fixed (emily, paul, sophie, steven) |
| Emotions | 6 + neutral (angry, disgusted, fearful, happy, sad, surprised, neutral) |
| Emotion Control Mode | single-argument selection (no composable multi-axis axes like Scylla's Band) |
| Input Format | text only — no phonemizer, no system dependencies |
| On Device | yes (laptop-class CPU real-time / better-than-real-time) |
| Distribution Formats | safetensors (torch), Q4 GGUF, Q8 GGUF |
| Formats In Collection | neuphonic/neutts-2e (safetensors), neuphonic/neutts-2e-q4-gguf (smallest footprint), neuphonic/neutts-2e-q8-gguf (mid-tier compression) |
| Gguf Features | imatrix, conversational, endpoints_compatible |
| Pipeline Tag | text-to-speech |
| Library Name | (HF tag does not declare transformers / safetensors stem beyond safetensors itself) |
| Downloads | 194 / 241 / 216 (torch / q4 / q8) |
| Intended Use | embedded voice agents, on-device assistants, toys, privacy-sensitive applications |
| Comparison With Air Nano | Air/Nano continue to ship for zero-shot cloning + multilingual contexts; 2E is the next-gen focused English emotional variant |
| Safety Note | model is alpha; legitimate project landing is neuphonic.com (not neutts.com) |
Features: The technical center of NeuTTS-2E is maximum speed per parameter
at on-device budgets — the 0.2B LM + codec pair delivers
real-time-or-better on laptop-class CPUs while exposing
discrete categorical emotion control (angry / disgusted /
fearful / happy / sad / surprised / neutral) plus a
fixed four-speaker cast for consistency in agent / toy /
accessibility voice personas. Two design choices distinguish it from
the surrounding TTS field:
First, the categorical emotion surface is single-axis and
discrete (one emotion per call), not the continuous multi-axis
composable vector surface used by models like Scylla's Band
([neurotica base + 6-axis continuous strengths]). The project's
positioning — production-grade agents + toys + accessibility —
benefits from a one-argument API where emotion="happy" is the
explicit operational state. The release locks emotional mode at
generation time, which simplifies downstream filtering / guardrails.
Second, the distribution-shape design (one model, three
deployment formats) is a deliberate on-device-first posture: the
safetensors torch build for max-quality GPU/server; Q8 GGUF for
mid-tier compression; Q4 GGUF for the small-footprint embedded
target. All three are direct llama.cpp-compatible drops of the
same model — no retraining-per-format — letting users pick size vs
quality at deployment time without changing the production API.
The combined CPU-first + GGUF-first design pattern is the
opposite of the cloud-first TTS systems in this list — and is what
makes 2E suitable for embedded voice agents, toys, and
privacy-sensitive applications where audio + text must remain
on-device.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Collection][link-collection]
· · · · · · · · · · · · · ·
Description: Scylla's Band is a multilingual, multi-voice, expressive TTS model from Spybyscript, designed specifically for local and self-hosted inference through ONNX Runtime (with an experimental LiteRT backend for explicit native / mobile use). The architecture is a continuous-latent TTS family:
Release Date: July 19, 2026
| Feature | Value |
|---|---|
| Parameters | not stated (architecture: 4-layer duration predictor (192 hidden) + 12-layer rectified-flow acoustic generator (512 hidden, AdaLN, QK norm)) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | en_us, en_gb, es, it (4 public text-input languages) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Sample Rate | 24,000 Hz |
| Managed Voices | 10 (ariadne, felix, gwen, ink, max, orpheus, rex, scylla, stone, tuesday) |
| Voice Default Locale | en_us for most; ink / orpheus / tuesday default to en_gb |
| Voice Style Dim | 128 (style features) + 32 (prosody features) |
| Affect Axes | 6 (calm, joy, anger, sadness, sarcasm, questioning — all continuous in [0, 1]) |
| Affect Overlay Axes | sarcasm, questioning (mixable with any core delivery) |
| Affect Cfg Scope | duration + acoustic-flow prediction (preserves voice / reference) |
| Encoders Default | ONNX Runtime (Python CLI / Python API / Android sample / libscyllasband native) |
| Encoders Experimental | LiteRT (experimental / explicit-selection) |
| Cli Quality Default | 8-step Heun sampling |
| Graph Budgets | 512 G2P text tokens / 512 phone frames / 640 latent frames |
| Latent Target Buckets | 256 / 384 / 512 / 640 (smallest-fit selection) |
| Vocoder | Scylla's Band acoustic adapter + frozen charactr/vocos-mel-24khz |
| Hop Lengths | 256 (waveform) / 512 (latent) |
| Text Frontend | phrase-level multilingual G2P (74-phone vocabulary) |
| Span Context | 3 segments over up to 768 phones with 512-dim context state |
| Prefix Context | up to 24 acoustic latent frames from preceding chunk |
| Long Form Features | boundary metadata + punctuation pause floors + prefix-latent carryover + span context |
| Group Speak Input | [voice], [voice:language], [voice:language:axis=value,...] annotations |
| Bundle Contract | 1.0.0 / scyllasband-duration-flow |
| Intended Use | single-voice speech synthesis (10 voices); en/es/it; long-form narration; multi-voice dialogue from tagged text; continuous affect control; ONNX desktop/server; ONNX + LiteRT native/mobile |
| Not Intended | arbitrary-speaker cloning / impersonation / fraud / deceptive speech |
| Distributions | training data, trainer checkpoints, and export tooling not distributed |
| Cli Commands | download, validate-bundle, list-voices, normalize-text, speak, group-speak, stream, plan |
| Library Name | onnxruntime (tags include onnx, tflite, litert, duration-flow) |
Features: Three design decisions distinguish Scylla's Band in the multilingual
TTS class. First, decoupling duration and acoustic flow as
separate rectified-flow stages — duration is a 192-hidden, 4-layer
predictor operating on a 512-phone window, acoustic latents a
512-hidden, 12-layer AdaLN / QK-norm generator at 24-dim. This split
lets affect-CFG act on both stages independently while retaining
voice / reference conditioning, supporting the 6-axis continuous
composability. Second, 6 affect axes (with sarcasm and
questioning as overlays mixed with any core delivery) instead of
mutually-exclusive discrete emotion classes — calm=0.5, joy=0.5 is
a valid input, and axes stay in [0, 1] so multi-axis states are
expressible without combinatorial blow-up. Third, the ONNX-first
runtime design with libscyllasband native + an experimental
LiteRT backend sits at a budget most neural TTS systems don't
target — the 8-step Heun default and 512/512/640 fixed graph budget
keep the model usable on CPU and mobile, and the inference-only
release surface (training data + checkpoints not distributed) is the
complement of the latency / mobile inference focus.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: sanoTTS is the smallest known neural text-to-speech family. The name sano (सानो) is Nepali for "small". Each voice weighs 294k to 2.3M parameters — smaller than the smallest voice in prior families (TinyTTS at 1.62M; Inflect Nano at 4.63M; Kokoro at 82M) and the family fits in under 4 MB per voice with zero runtime dependencies (the espeak-ng phonemizer is bundled). Voices run real-time on a ~$3 ESP32-S3 microcontroller (output through a GPIO into an LM386 and a speaker) and live in the browser via WebAssembly — no server, no upload, no NPU. The full neural stack is duration → acoustic → decoder, quantized to int8, with the espeak-ng phonemizer included. 11 voices across 6 languages ship: English, Nepali, Hindi, Vietnamese, Indonesian, and Chinese — including the 294k heart-nano voice (337 KB) and the mel-based heart / heart-nano pair that predicts a 100-band spectrogram rendered by a noise-fed ConvNeXt + iSTFT decoder at 24 kHz. The inference runtime was relicensed MIT (September 2026); the project as a whole remains GPLv3 via espeak-ng. The project page at ampixa.github.io/sanoTTS hosts a live browser synthesis demo for every voice.
Release Date: July 13, 2026
| Feature | Value |
|---|---|
| Parameters | 294k–2.3M per voice (smallest = the 294k "heart-nano" voice, 337 KB) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English, Nepali, Hindi, Vietnamese, Indonesian, Chinese (6 languages, 11 voices) |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Architecture | full neural stack — duration model → acoustic model → decoder |
| Quantization | int8 (W8/A12, corr 0.9995+; piperlite portable C99) |
| Runtime Microcontroller | ESP32-S3 (real-time RTF 0.41, GPIO → LM386 → speaker) |
| Runtime Browser | WebAssembly (no server, no upload, no NPU) |
| Runtime Footprint | under 4 MB per voice, zero dependencies |
| Voices | 11 (English: amy / kristin / hfc / amy-1p1m / amy-1p8m / robot / heart / heart-nano; one voice each for NE / VI / ID / ZH + shared lang voices) |
| Phonemizer | espeak-ng (bundled) |
| License Split | inference runtime MIT; project overall GPL-3.0 (copyleft from espeak-ng) |
| Library Name | sanotts |
| Training Method | distillation (per voice) |
Features: The hard constraint — be the smallest neural TTS family known,
real-time on a $3 microcontroller — drives the entire stack.
Conventional sub-100M TTS systems are too large for an ESP32's flash
and RAM. sanoTTS keeps the full duration → acoustic → decoder
neural pipeline (no espeak-NG-only fallback, no concatenative
hybrid), quantizes everything to int8, and bundles the phonemizer
so the whole voice ships in under 4 MB with zero runtime dependencies.
The newest heart / heart-nano voices switch to a mel-based recipe
(100-band spectrogram + noise-fed ConvNeXt + iSTFT decoder at 24 kHz),
bringing the smallest voice down to 294k parameters / 337 KB —
a per-voice footprint 100× smaller than Kokoro and 2× smaller than
TinyTTS while still leading SCOREQ / UTMOS in the sub-15M class —
and the demo synthesizes every voice live in the browser via
WASM, so the smallest-known neural TTS is also the only one that
runs unattended on a $3 chip and a $0 web page.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: FreyaTTS is a 183M-parameter Turkish text-to-speech model. It is tokenizer-free at the character level — 92 symbols in its Turkish vocabulary — so there is no phonemizer or G2P step in either training or inference. Speech is generated with a non-autoregressive conditional flow-matching DiT in the frozen AudioVAE2 latent space (25 Hz, 64-dim latents, 16 kHz encode / 48 kHz decode). Training runs from scratch on Turkish speech: a pretraining stage followed by SFT stage 1/2 for voice lock and short-utterance coverage. Output is 48 kHz mono. On the project's Freya-TR-Eval benchmark the model reports WER 8.0% / CER 3.0%, ranking 3rd of 7 among open sub-1B Turkish TTS systems — a deliberate single-target-speaker, no-cloning design choice for a focused foundation release. The evaluation dataset is freya-tr-eval.
Release Date: July 7, 2026
| Feature | Value |
|---|---|
| Parameters | 183.2M |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | Turkish (tr) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | conditional flow-matching diffusion transformer (DiT), non-autoregressive, 32-step Euler ODE, no CFG |
| Tokenizer | character-level (92 Turkish symbols; no phonemizer, no G2P) |
| Latent Space | frozen AudioVAE2 (Apache-2.0, openbmb/VoxCPM2), 64-dim at 25 Hz |
| Codec Io | 16 kHz encode / 48 kHz decode |
| Sample Rate | 48,000 Hz |
| Training | from scratch on Turkish speech; pretraining + SFT stage 1/2 (voice lock + short-utterance coverage) |
| Evaluation | Freya-TR-Eval — WER 8.0% / CER 3.0%, 3rd of 7 open sub-1B Turkish TTS |
| Library Name | freyatts |
Features: Two design choices are worth flagging. First, tokenizer-free character-level Turkish: by training directly on the 92-symbol Turkish alphabet with no phonemizer or G2P grapheme-to-phoneme step, the model removes a dependency that is fragile for agglutinative Turkish morphology and that often degrades quality when ported to low-resource Turkic relatives. Second, non-autoregressive conditional flow-matching in a frozen AudioVAE2 latent space: the 25 Hz / 64-dim bottleneck keeps the DiT small (183M) while inheriting a separately-trained audio codec's representation, letting a focused single-language-non-multilingual release ship at a fraction of the parameter budget of multilingual foundation TTS systems. The deliberate "no cloning, single target speaker" choice is a scope-lowering move that lets the foundation release put all its capacity into Turkish speech quality rather than spread it across zero-shot speaker adaptation.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Inflect-Nano-v2 is a complete local text-to-waveform speech synthesis model with 3,966,721 deployable parameters — under 4M total. It is a VITS-architecture fixed-voice English TTS designed for CPU or CUDA inference with deterministic seeds, long-text handling, and 24 kHz mono output. The full FP32 checkpoint is 15.97 MB, making it one of the smallest complete neural TTS systems that produces natural-sounding speech without a separate vocoder or phonemizer dependency. The model ships with a public adaptation toolkit for preparing data, auditing train/validation splits, adapting a fixed voice or language, resuming training, evaluating checkpoints, and exporting PyTorch or ONNX packages. A sibling Inflect-Micro-v2 (9.36M parameters) prioritizes quality below 10M; Nano prioritizes footprint below 4M. Both share one public API.
Release Date: June 25, 2026
| Feature | Value |
|---|---|
| Parameters | 3,966,721 (3.97M deployable) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | VITS (end-to-end text-to-waveform) |
| Sample Rate | 24,000 Hz |
| Footprint | 15.97 MB FP32 |
| Inference | CPU or CUDA; PyTorch + ONNX export |
| Determinism | deterministic seeds for reproducible generation |
| Long Text | automatic text splitting and handling |
| Input Format | text (no phonemizer or system dependencies) |
| Adaptation Toolkit | data prep, split auditing, voice/language adaptation, training resume, checkpoint eval, PyTorch/ONNX export |
| Sibling Model | Inflect-Micro-v2 (9.36M params, quality-prioritized below 10M) |
| Api | one public API across Micro and Nano sizes |
| Library | pytorch |
| Metrics | WER |
| Inference False On Hf | yes (no hosted HF inference endpoint; local-only) |
Features: Inflect-Nano-v2's defining constraint is completeness under 4M parameters: the entire text-to-waveform pipeline — no separate vocoder, no phonemizer, no system dependencies — fits in 3.97M deployable parameters and a 15.97 MB FP32 checkpoint. This is smaller than even sanoTTS's smallest voice (745k) when measured by complete-pipeline footprint, though sanoTTS ships per-voice weights rather than a single fixed-voice checkpoint. The VITS end-to-end architecture is the enabler: by folding the acoustic model and vocoder into a single jointly-trained network, Inflect avoids the multi-stage pipeline overhead that makes most neural TTS systems larger. The public adaptation toolkit extends the fixed-voice design into a customizable platform — users can prepare data, adapt a voice or language, resume training, and export PyTorch or ONNX packages — making the 4M-parameter footprint a starting point for domain-specific TTS rather than a dead-end fixed-voice release.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: GEnerative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue. Gepard is built for low-latency, high-throughput streaming conversation: the model starts speaking the moment text begins arriving, generating audio piece by piece instead of waiting for a full sentence. It is a single decoder-only autoregressive language model built on Qwen3.5 (14 layers, hidden 1024, 8 heads) with ≈556M total parameters (backbone + audio interface + voice-cloning compressor). Audio is produced through NVIDIA NeMo NanoCodec — Finite Scalar Quantization at 22.05 kHz, 21.5 frames/s, 1.89 kbps — with the full 32-channel FSQ frame sampled in one step. Reports ~25× real time on a single RTX 5090 with first-audio-chunk latency around 50 ms; a 96 GB Blackwell card serves up to 256 concurrent conversations. CFG refinement is baked into the weights so quality gain comes at no extra two-pass cost at inference, though the two-pass mode is still selectable as a quality dial.
Release Date: June 22, 2026
| Feature | Value |
|---|---|
| Parameters | ~556M (555,694,169; Qwen3.5 backbone + audio interface + voice-cloning compressor) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | English (US/UK), Spanish (es-MX), Portuguese (pt-BR), Dutch (NL) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Audio Codec | NVIDIA NeMo NanoCodec (FSQ, 22.05 kHz, 21.5 fps, 1.89 kbps; NVIDIA Open Model License) |
| Sample Rate | 22,050 Hz |
| Backbone | Qwen3.5 full-attention transformer (14 layers, hidden 1024, 8 heads; ~500M params) |
| Inference | vLLM |
| Throughput | 256 conversations on one 96 GB Blackwell (RTX Pro 6000) GPU |
| Benchmark | Seed-TTS-eval leader on perceived quality (NISQA-MOS 4.25, NOI 4.16, COL 4.16, DIS 4.51) trading some WER/SIM |
Features: A prosody-aware autoregressive single-pass frame generator: the whole 32-channel FSQ audio frame is sampled in one step (no depth transformer), and CFG quality refinement is baked into the weights rather than incurred at inference as a two-pass cost — so the publicly reported TTFA of ~50 ms and 25× real time on a single RTX 5090 represent the quality-on path, not a cheap-fast preview. Voice cloning is decoupled into a separate up-front compressor, which means cloning is "free" at run-time once the reference clip is encoded — a structural choice that supports serving hundreds of conversations per GPU. A stop-head weight update (2026-08-06) fixed premature stopping at sentence boundaries and lifted the effective duration ceiling; on Seed-TTS-eval Gepard leads the compared systems on perceived quality (NISQA-MOS 4.25) while trading some speaker similarity and WER for its streaming-first design.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo] ![Paper][link-paper] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Boson AI's flagship conversational TTS: an ~4B autoregressive decoder over interleaved text and audio tokens from the Higgs Tokenizer (8 codebooks at 25 fps / 24 kHz). Built for voice chat rather than narration, it covers 102 languages with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.
Release Date: June 4, 2026
| Feature | Value |
|---|---|
| Parameters | 4B (BF16, 36 layers, hidden=2560, GQA 32/8) |
| Architecture | Autoregressive decoder (Qwen3-style) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 102 (85 with WER/CER <5, 17 between 5-10) |
| Streaming | ✅ |
| Audio Output | 24 kHz |
| License | ![Research Only][license-research-only] |
Features: Interleaved text/audio token modelling with a delay-pattern multi-codebook embedding/head: a single autoregressive stack emits both modalities and supports inline <|category:value|> control tags (emotion/style/sfx/prosody) inserted at any point in the target text.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive TTS system from Rednote-HiLab. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE, with no discrete tokens anywhere in the pipeline. It achieves the best average performance on Seed-TTS-Eval (WER 0.94 / 1.30 / 6.60 on zh / en / zh-hard) and the highest speaker similarity on a 24-language MiniMax multilingual benchmark, with broad cross-lingual voice cloning.
Release Date: June 3, 2026
| Feature | Value |
|---|---|
| Parameters | 2B (semantic encoder + LLM + AR flow-matching acoustic head) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Multilingual (24+ languages; zh / en focus) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Sample Rate | 48 kHz |
| Tokenizer | 48 kHz AudioVAE (continuous, no discrete tokens) |
Features: A fully continuous autoregressive pipeline that keeps generation in waveform-latent space end-to-end (no discrete-code phase), pairing an LLM-side semantic encoder with an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE — yielding SOTA seed-TTS-Eval scores and the strongest speaker-similarity number (83.9 avg) on the 24-language MiniMax multilingual benchmark.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Confucius4-TTS is an LLM-based text-to-speech system from NetEase Youdao designed for multilingual and cross-lingual synthesis. It uses a speech encoder + LLM (Text2Semantic) + flow-matching Semantic2Acoustic architecture that allows zero-shot voice cloning without a required reference transcript and explicit cross-lingual voice transfer with unaccented output across languages. Covers Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese with code-switching and emotion transfer.
Release Date: June 2, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | 14 (zh, en, ja, ko, de, fr, es, id, vi, th, pt, it, ru, ms) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | speech encoder + LLM (T2S) + flow-matching head (S2A) |
Features: Cross-lingual voice transfer without accent drift: the same reference voice stays consistent when the speaker switches languages — backed by a speech encoder + LLM backbone pipeline (T2S) with a flow-matching acoustic decoder (S2A) and training that bundles 14 languages with code-switched, emotion-preserving decoding.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: WavTTS is an end-to-end zero-shot TTS framework that synthesizes speech directly in the raw waveform space — explicitly skipping the intermediate mel-spectrogram, VAE-latent, or codec-token representations that most modern TTS stacks use. It is built on a flow-matching diffusion transformer (DiT) with waveform patchification, multi-scale mel-spectrogram supervision, and an optimized noise schedule. Forked from F5-TTS at the codebase level but replaces the whole acoustic pipeline.
Release Date: May 28, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English, Chinese |
| Streaming | ❌ |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] ![MIT][license-mit] |
| Sample Rate | 16 kHz |
| Training Data | Emilia |
| Architecture | Flow-matching DiT, raw waveform patchification, multi-scale mel supervision |
| Training Steps | 1.2M |
Features: Skip every intermediate waveform representation (no mel, no VAE, no codec tokens): a flow-matching DiT produces raw-waveform patches directly, supervised at multiple mel scales and an optimized noise schedule — yielding high-quality zero-shot TTS at 16 kHz from a single end-to-end stack.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: MOSS-TTS is a production-grade Text-to-Speech foundation model developed by the OpenMOSS Team and MOSI.AI. The current public v1.5 release preserves the original 1.0 capabilities — zero-shot voice cloning, long-form speech generation, token-level duration control, Pinyin/IPA pronunciation supervision, multilingual synthesis, and code-switching — and extends multilingual continued training from 20 languages to 31 languages including Cantonese, Dutch, Finnish, Hindi, Macedonian, Malay, Romanian, Swahili, Tagalog, Thai, and Vietnamese. v1.5 improves speaker similarity, reduces cloning variance on long-reference / short-text scenarios, follows punctuation-driven prosody more reliably, and adds explicit inline pause markers (e.g., [pause 3.2s]).
Release Date: May 25, 2026
| Feature | Value |
|---|---|
| Parameters | 8B (Delay), 1.7B (Local) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 31 (extended from v1.0's 20) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Max Duration | 1 hour |
| Pause Control | yes (inline markers like [pause 3.2s]) |
| Lang Tag Control | yes (set language= in user message) |
Features: v1.5 widens MOSS-TTS from 20 → 31 languages with stronger per-language multilingual synthesis (control via a language tag in the user message), more stable cloning under long-reference / short-text conditions, punctuation-driven prosody that holds up across long sentences, and explicit inline pause tokens ([pause 3.2s]) for scripted narration control.
Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: VoxFlash-TTS is a zero-shot voice-cloning text-to-speech engine built around extreme latent compression. The VAE encodes 24 kHz waveforms into a 9 frames/s latent space — roughly 8× more compressed than EnCodec (75 fps) and 2.4× more than Stable Audio (21.5 fps). Generating 10 s of audio therefore requires the diffusion model to produce just 90 latent vectors rather than hundreds or thousands of tokens, with downstream quadratic savings in attention cost. A ConvNeXtV2-based phoneme encoder followed by a novel coarse-alignment algorithm (cheaper than cross-attention) maps text into the latent sequence; a modern diffusion head then iteratively refines speech latents that the lightweight VAE decoder renders back to waveforms. The architecture targets low-latency, low-resource deployment — consumer-grade GPUs and edge devices — with Chinese and English zero-shot cloning. The project card lists inference: false on HF (no hosted inference endpoint), but the project page at voxflash.github.io carries the abstract, demo examples, and ablations.
Release Date: May 22, 2026
| Feature | Value |
|---|---|
| Parameters | not stated (ConvNeXtV2 phoneme encoder + diffusion head + lightweight VAE decoder) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Chinese, English |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Audio Codec | VoxFlash VAE (9 Hz / 9 fps latent, 24 kHz input) |
| Compression Ratio | ~8× tighter than EnCodec (75 fps), ~2.4× tighter than Stable Audio (21.5 fps) |
| Phoneme Encoder | ConvNeXtV2 + coarse-alignment algorithm (no cross-attention) |
| Diffusion Head | modern multi-step iterative refinement |
| Decoder | lightweight VAE decoder |
| Sample Rate | 24,000 Hz |
| Inference | local CUDA ≥ 12.3.2; no HF hosted endpoint |
| Training Dataset | seed-tts-eval |
| Metrics | word_error_rate, speaker_similarity |
Features: The central technical move is compressing the audio latent space to 9 frames/s instead of the conventional 75 fps (EnCodec) or 21.5 fps (Stable Audio). This is not a quantization tweak — it is a temporal-decimation architectural choice that shrinks the sequence length the diffusion model has to traverse, and because attention cost scales quadratically with sequence length the end-to-end compute drops by orders of magnitude. Combined with a coarse-alignment phoneme-to-latent map that avoids cross-attention entirely (using a ConvNeXtV2 phoneme encoder instead), VoxFlash hits millisecond-level inference latency on consumer-grade and edge hardware for zero-shot Chinese + English cloning, where conventional latent-diffusion TTS systems are too slow for real-time edge deployment.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Miso TTS 8B is a text-to-speech model from Miso Labs built on the Sesame Conversational Speech Model (CSM) architecture. A large Llama-3.2-style backbone consumes text/audio-frame embeddings and predicts codebook 0 of the Mimi audio token stream, while a smaller 300M autoregressive audio decoder predicts codebooks 1–31 in codebook depth. The model is designed for high-quality conversational speech and voice continuation from a short prompt audio clip.
Release Date: May 21, 2026
| Feature | Value |
|---|---|
| Parameters | 8B (backbone llama-3.2-style) + 300M (audio decoder) = 8.3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Architecture | Sesame-style CSM (two transformer stack: backbone + audio decoder) |
| Audio Tokenizer | Mimi (32 codebooks, vocab 2051, max seq 2048) |
| Library | pytorch |
Features: A two-transformer Sesame-style CSM (Llama 8B backbone consumes text + audio frames and produces backbone codebook-0 prediction; a 300M audio decoder autoregresses over codebook depth via Mimi's 32-codebook stack) — letting the larger backbone spend capacity on linguistic / speaker conditioning while a leaner decoder handles fine-grained codebook-by-codebook generation.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Raon-OpenTTS is an open-data, open-weight zero-shot TTS system from KRAFTON that performs on par with state-of-the-art closed-data models. This is the 1B variant (1048M parameters). Both model weights and training data are public: Raon-OpenTTS-Core is 510.1K hours of English speech, quality-filtered from the 615K-hour public Raon-OpenTTS-Pool using combined DNSMOS, WER, and VAD rank-based filtering. It ranks 1st or 2nd in WER and SIM among recent zero-shot TTS models on Seed-TTS-Eval and CV3-Eval, and achieves the best average WER/SIM on Raon-OpenTTS-Eval across Clean, Noisy, Wild, and Expressive regimes. A smaller Raon-OpenTTS-0.3B variant is also available.
Release Date: May 21, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English only (trained on 11 English speech datasets) |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Parameters | 1048M |
| Architecture | DiT (Diffusion Transformer) based on F5-TTS with flow matching; dim=1408, depth=28, heads=24 |
| Audio Output | 80-ch mel-spectrogram at 16 kHz, HiFi-GAN vocoder (LibriTTS) |
| Training Data | Raon-OpenTTS-Core (510.1K hours), 520K updates on 48× B200 |
Features: Demonstrates that fully open data + open weights can match proprietary SOTA: on Seed-TTS-Eval it reaches 1.78 WER / 0.749 SIM (vs Qwen3-TTS 1.46/0.715 at 1.7B), and best overall robustness (WER 2.81 / SIM 0.695) across four acoustic regimes on its own Raon-OpenTTS-Eval benchmark. The pipeline pairs large-scale rank-based data curation (DNSMOS + WER + VAD filtering of a 615K-hour pool) with an efficient F5-TTS-derived DiT.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Dataset][link-dataset]
Additional Tools:
| Tool | Type | Link |
|---|---|---|
| ComfyUI-Raon-OpenTTS | ComfyUI node | ComfyUI-Raon-OpenTTS |
· · · · · · · · · · · · · ·
Description: OronTTS is a non-autoregressive text-to-speech model from btsee, an F5-TTS fork specialized for Mongolian (Khalkha Cyrillic) and Kazakh (Cyrillic). It uses Flow Matching + Diffusion Transformer + Vocos (dim 1024, depth 22, 16 heads, vocab 65, 24 kHz sample rate), trained on the btsee/mbspeech_mn corpus (3,846 Mongolian speech samples) and outputs zero-shot synthesis from a short reference audio + language tag.
Release Date: May 16, 2026
| Feature | Value |
|---|---|
| Parameters | (not stated) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | Mongolian (Khalkha Cyrillic), Kazakh (Cyrillic) |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Architecture | F5-TTS (OT-CFM + DiT + Vocos) |
| Dim | 1024 |
| Depth | 22 |
| Heads | 16 |
| Vocab Size | 65 |
| Sample Rate | 24000 Hz |
| Mel Bins | 100 |
| Training Data | btsee/mbspeech_mn (3,846 Mongolian speech samples) |
Features: F5-TTS re-purposed for low-resource Cyrillic languages (Mongolian + Kazakh) — non-autoregressive flow-matching DiT over a tight 65-word vocab. Trained on a small (~3.8k sample) Mongolian corpus; the architecture is small enough that Khalkha Cyrillic and Kazakh Cyrillic share the same checkpoint via the lang tag at inference time.
Links: ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Supertonic 3 is the third-generation open-weight release from Supertone. It is a lightweight, on-device text-to-speech system that runs with ONNX Runtime entirely on the user's machine (no network, no API call) and ships as a Python SDK (pip install supertonic). Compared with the Supertonic 2 base (5 languages, 66 M params), v3 expands to 31 languages and adds expression tags (<laugh>, <breath>, <sigh>), more stable reading on long utterances, and higher speaker similarity across the core language set.
Release Date: May 6, 2026
| Feature | Value |
|---|---|
| Parameters | (not stated on card; Supertonic 2 baseline 66 M — likely similar or smaller weight class) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | 31 (expanded from Supertonic 2's 5) |
| Streaming | ✅ |
| License | ![OpenRAIL-M][license-openrail-m] |
| On Device | yes (ONNX Runtime, no cloud call) |
| Expression Tags | yes (<laugh>, <breath>, <sigh>) |
Features: A more compact on-device multilingual TTS: ONNX-Runtime inference everywhere, 31 languages from a single small open-weight encoder, and discrete expression tags that the decoder interprets inline — without a separate speaker-emotion control path or a cloud-rendered audio round-trip.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi]
· · · · · · · · · · · · · ·
Description: Scenema Audio is a zero-shot expressive voice cloning and speech generation model from ScenemaAI. It is built on an audio diffusion transformer extracted from the audio branch of Lightricks' LTX 2.3 (a 22B audiovisual model) — keeping the in-the-wild acoustic quality the bigger model learned while specializing for speech output. Generation is prompt-driven: a <speak> tag carries a voice description, gender, optional scene (ambient audio around the voice), and language; an <action> tag shifts emotional state mid-generation. Action tags cover rage, grief, joy, fear, exhaustion; voice prompt can describe timbre/pitch/breathiness/rasp/resonance plus character archetypes ("Tony Soprano having a breakdown"). Supports zero-shot voice cloning from 10-20 seconds of reference audio with some emotional variability, automatic long-form narration by splitting text and maintaining voice continuity, and 13 multilingual built-ins.
Release Date: April 26, 2026
| Feature | Value |
|---|---|
| Parameters | (audio diffusion transformer of LTX 2.3, weights ~9.8 GB bf16 / ~4.9 GB INT8 + ~6.7 GB pipeline) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | 13 (en, de, fr, es, it, pt, ja, zh, ko, ru, ar, hi, sw) |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Parent Model | Lightricks LTX-2.3 (audio branch) |
| Prompt Format | <speak voice=… gender=… scene=… language=…> XML with <action> tag for shifting emotion |
| Long Form Narration | yes (auto-splits text while preserving voice continuity) |
| Quantized | yes (INT8 weights at ~4.9 GB, identical quality) |
Features: A standalone audio diffusion transformer extracted from a much bigger multimodal source: the model inherits how people actually sound in real scenes (angry, laughing, whispering, crying, exhausted, terrified) and exposes that capacity through a <speak> + <action> prompt interface — emotional state shifts within a single generation, instead of being a token-level or speaker-level conditioning problem.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Dramabox is Resemble AI's expressive TTS, distributed under the LTX-2 Community License. It is an IC-LoRA fine-tune of the LTX-2.3 3.3B audio-only branch (Diffusion Transformer + flow matching), conditioned on Gemma 3 12B text embeddings. Generation is prompt-driven: speaker identity, emotion, delivery, laughs, sighs, breaths, pauses, and transitions are all expressed inside a natural-language description, with an optional 10-second voice reference that clones the target timbre.
Release Date: April 17, 2026
| Feature | Value |
|---|---|
| Parameters | 3.3B (LTX-2.3 audio backbone, IC-LoRA fine-tune) + 12B Gemma 3 text encoder (conditioning only) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Base Model | Lightricks/LTX-2.3 (audio branch) |
| Architecture | DiT + flow matching, IC-LoRA fine-tune, Gemma 3 12B text embeddings |
| Inference Time | ~2.5 s / generation (warm server) |
Features: IC-LoRA fine-tune of LTX-2.3's audio branch leaves the heavy text-understanding work to Gemma 3 12B and lets the DiT do the expressive rendering — so what's normally multimodal-stage orchestration collapses into a single prompt-driven TTS where speaker identity, emotion, and delivery are encoded in the prompt itself, and the timbre comes from a 10-second voice reference when present.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Sarashina2.2-TTS is a Japanese-centric text-to-speech system from SB Intuitions built on a large language model. It supports both Japanese and English, delivers strong pronunciation accuracy on Japanese text through large-scale end-to-end training, and reproduces a speaker's voice, speaking style, and acoustic characteristics from a short reference clip (zero-shot). Training data is sourced exclusively from legitimately acquired, properly licensed speech archives per the Sarashina Model NonCommercial License Agreement v2.0 (released April 24, 2026).
Release Date: April 16, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | Japanese, English |
| Streaming | ❌ |
| License | ![Research Only][license-research-only] |
| Base Model | sbintuitions/sarashina2.2-0.5b-instruct-v0.1 |
| Cross Lingual | yes (Japanese ↔ English, code switching) |
Features: Japanese-optimized TTS fine-tuned on responsibly-licensed Japanese training corpora with explicit cross-lingual code-switching to English in a single utterance; reference audio carries speaking style and speaker identity together, so the same prompt yields narration, broadcast, conversation, or customer-service delivery without separate style conditioning.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: State-of-the-art diffusion-based TTS model operating directly in waveform latent space. Developed by Meituan's LongCat team, it requires only a Waveform VAE and Diffusion backbone, effectively mitigating compounding errors.
Release Date: March 30, 2026
| Feature | Value |
|---|---|
| Parameters | 1B / 3.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | Chinese, English |
| Streaming | ❌ |
| Audio Output | 24000 Hz |
| License | ![MIT][license-mit] |
Features: Adaptive Projection Guidance (APG) replaces traditional classifier-free guidance for elevated generation quality. Outperforms Seed-TTS on zero-shot voice cloning benchmarks.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: SILMA TTS v1 is a high-performance, 150M-parameter bilingual (Arabic/English) TTS model developed by SILMA AI. Built on the F5-TTS diffusion architecture, it was pretrained from scratch using tens of thousands of hours of high-quality public and proprietary data. It supports instant voice cloning with less than 8 seconds of reference audio (the reference transcript can also be left empty — it is transcribed on the fly), full support for Arabic Tashkeel (diacritics, auto-enriched via CATT when absent), NeMo-based text normalization, and an RTF around 0.12 on an RTX 4090. Released under a commercial-friendly license: code MIT, model weights Apache-2.0. The model is 100% compatible with F5-TTS v1.1.7 tooling for inference and fine-tuning.
Release Date: March 13, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | 2 (Arabic MSA/Fusha + English) |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 150M |
| Architecture | F5-TTS Diffusion Transformer with flow matching (pretrained from scratch, F5-TTS v1.1.7-compatible) |
| Pronunciation | ✅ |
| Cost | RTF ≈ 0.12 (RTX 4090) |
Features: Brings native-level Arabic synthesis to a 150M footprint: one of the smallest open F5-TTS-family models pretrained from scratch rather than fine-tuned, with first-class Arabic handling (Tashkeel-aware pronunciation via CATT enrichment, NeMo text normalization) alongside English, plus instant zero-shot cloning under fully permissive licensing (Apache-2.0 weights / MIT code).
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Fish Audio S2 Pro is a leading text-to-speech model with fine-grained inline control of prosody and emotion. It combines reinforcement learning alignment with a dual-autoregressive architecture for high-quality speech synthesis.
Release Date: March 10, 2026
| Feature | Value |
|---|---|
| Parameters | ~10 GB (BF16) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 80+ (Tier 1: En, Zh, Jp) |
| Streaming | ✅ |
| License | ![Research Only][license-research-only] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Native multimodal foundation model by Meituan LongCat Team processing text, vision, and audio under a single autoregressive objective. Industrial-strength model with strong speech synthesis and voice cloning.
Release Date: March 2026
| Feature | Value |
|---|---|
| Parameters | 3B (MoE A3B) |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| Audio Output | 24 kHz |
| License | ![MIT][license-mit] |
Features: Discrete Native Autoregression Paradigm (DiNA) unifying modalities in shared discrete token space. Combines visual understanding, generation, and audio processing in single model.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Frontier, open-weights text-to-speech model developed by Mistral AI. Designed to be fast, instantly adaptable, and produces lifelike speech with natural prosody and emotional range.
Release Date: March 2026
| Feature | Value |
|---|---|
| Parameters | 4B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | 9 (En, Fr, Es, De, It, Pt, Nl, Ar, Hi) |
| Streaming | ✅ |
| Audio Output | 24 kHz |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
Links: ![HuggingFace][link-huggingface] ![Demo][link-demo] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: BlueTTS (project page: lightbluetts.com) is a multilingual text-to-speech library. Built around slim ONNX graphs that run on ONNX Runtime with first-class CPU support and optional accelerators — OpenVINO (Intel), CUDA ORT (NVIDIA), TensorRT, and ONNX Runtime stock CPU. Targets five languages — Hebrew, English, Spanish, Italian, German — including inline mixed-language with XML-style tags in the text prompt. Inference is deliverable as a PyPI package (blue-onnx), a Rust crate, or directly from the pinned ONNX graphs on the HF Hub; the v2 release ships a slimmed opset-17 ONNX bundle (notmax123/blue-onnx-v2) that's intended for both FP32 production and the experimental INT8 weight-only fallback.
Release Date: February 27, 2026
| Feature | Value |
|---|---|
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | Hebrew, English, Spanish, Italian, German |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Runtime | ONNX Runtime (stock CPU; OpenVINO / CUDA / TensorRT optional) |
| Speed | "fastest open-source TTS" (per project description) |
| Graph Format | ONNX opset 17 (slim, full-precision; experimental weight-only INT8 fallback) |
| Distribution | PyPI + HuggingFace + Rust |
Features: A CPU-first multilingual TTS that ships both slimmed ONNX graphs and a Python package where the same code path runs on stock CPU ONNX Runtime by default — and optionally accelerates on OpenVINO / CUDA ORT / TensorRT — so deployment doesn't gate on GPU availability. Languages include Hebrew (with explicit G2P normalization) — a comparatively rare open-source TTS target — plus standard European languages, all from MIT-licensed weights distributed via both Hugging Face and PyPI.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![PyPI][link-pypi] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: KittenTTS is an open-source realistic text-to-speech model designed for lightweight deployment. It is a state-of-the-art TTS model under 25MB with just 15 million parameters, running without GPU on any device.
Release Date: February 24, 2026 (v0.8.1)
| Feature | Value |
|---|---|
| Parameters | 15M-80M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | English, Multiple |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Ming-omni-tts is a high-performance unified audio generation model in the Ming 2.0 series. It uses a custom 12.5 Hz continuous tokenizer and a Patch-by-Patch compression strategy that drives the LLM inference frame rate down to 3.1 Hz, enabling fine-grained control over speech rate, pitch, volume, emotion, and dialect (notably Cantonese at ~93 % accuracy). It supports 100+ premium built-in voices plus zero-shot voice design from natural-language prompts and is the first autoregressive model that jointly generates speech, ambient sound, and music in a single channel.
Release Date: February 11, 2026
| Feature | Value |
|---|---|
| Parameters | 16.8B (3B active, MoE; A3B) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | Chinese, English, Cantonese |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Patch-by-Patch compression drives the inference frame rate to 3.1 Hz, drastically cutting LLM-side latency for podcast-style audio while preserving naturalness. A custom 12.5 Hz continuous tokenizer plus a DiT head jointly produce speech, ambient sound, and music in a single output channel — an "in-the-scene" listening experience rather than TTS-on-top-of-a-track.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: SoulX-Singer is a high-fidelity, zero-shot singing voice synthesis model for generating realistic singing voices for unseen singers without fine-tuning.
Release Date: February 6, 2026
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Mandarin, English, Cantonese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: SoproTTS is a lightweight English text-to-speech model with zero-shot voice cloning. It uses dilated convolutions (WaveNet-style) and lightweight cross-attention layers instead of the common Transformer architecture.
Release Date: February 4, 2026 (v1.5)
| Feature | Value |
|---|---|
| Parameters | 135M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Rtf | 0.05 (CPU M3) |
| Training-Cost | ~$100 |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Qwen3-TTS is an open-source series of Text-to-Speech models developed by Alibaba Cloud. Supports stable, expressive, and streaming speech generation with free-form voice design.
Release Date: January 22, 2026
| Feature | Value |
|---|---|
| Parameters | 0.6B-1.7B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 10 (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: TADA is a unified speech-language model from Hume AI built around a Text-Acoustic Dual-Alignment tokenizer: for every text/subword token there is exactly one corresponding speech vector, so the audio stream stays 1:1 aligned with text. As a TTS model, each autoregressive step covers one text token and dynamically determines the duration and prosody for that token, breaking the fixed-frames-per-second constraint that drives most modern TTS backbones. As a speech-language model, it generates a text token and the speech for the preceding token in the same dual step.
Release Date: January 12, 2026
| Feature | Value |
|---|---|
| Parameters | 1B (Llama 3.2 1B base) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ❌ |
| License | ![Other][license-other] |
| Base Model | meta-llama/Llama-3.2-1B |
| Tokenization | 1:1 text–acoustic dual alignment (one speech vector per text token) |
| Dynamic Duration | yes (each autoregressive step covers one text token, duration is determined per-token) |
Features: A dual-alignment speech–text tokenizer that decouples autoregression from a fixed audio frame rate: each text token owns exactly one speech vector, and the model synthesizes the whole segment for that token in one step, regardless of how long the spoken form is — eliminating transcript hallucination and the latency overhead of constant-frame-rate codecs while staying as compact as Llama 3.2 1B.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi] ![Paper][link-paper] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: Japanese Text-to-Speech model based on Rectified Flow Diffusion Transformer. Features emoji-based style and sound effect control by embedding emojis in input text for expressive speech generation.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 500M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | Japanese |
| Streaming | ❌ |
| Audio Output | 48kHz waveform |
| License | ![MIT][license-mit] |
Features: Key Feature: Emoji annotation control - insert specific emojis into text to control speaking styles, emotions, and sound effects.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Open-source TTS for European languages with 7B parameters. Outperformed ElevenLabs in human preference testing.
Release Date: Early 2026
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 23 European languages |
| Streaming | ✅ |
| License | ![MIT][license-mit] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Part of the LEMAS (Large-scale Extensible Multilingual Audio Suite) project. Zero-shot multilingual TTS with 0.3B parameters supporting 10 languages with word-level precise editing capabilities.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 0.3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 10 (zh/en/de/fr/es/pt/it/ru/id/vi) |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Special-Feature | Word-level editing (LEMAS-Edit) |
Features: Built on 150,000+ hours of multilingual speech data with word-level timestamps. Includes LEMAS-Edit for precise word-level speech editing via masked token infilling.
Links: ![Website][link-website] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Lightweight, high-speed LLM-based TTS model for English and Japanese with minimal resource usage.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 2.6B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | English, Japanese |
| Streaming | ✅ |
| License | ![LFM][license-lfm] |
| Rtf | 0.135-0.145 |
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Ultra-lightweight open-source multilingual speech generation model with only 0.1B parameters. Designed for realtime speech generation that runs directly on CPU without GPU.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 0.1B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | 20 |
| Streaming | ✅ |
| Audio Output | 48 kHz Stereo |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Pure autoregressive architecture with MOSS-Audio-Tokenizer-Nano. Compresses audio to 12.5 Hz token stream using RVQ with 16 codebooks. Runs on 4-core CPU.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: NeuTTS is a collection of open-source on-device TTS models with instant voice cloning. Built off LLM backbones with GGUF format quantizations for efficient on-device deployment.
Release Date: Early 2026
| Feature | Value |
|---|---|
| Parameters | 360M (Air), 120M (Nano) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | English, Spanish, German, French |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| On-Device | yes (GGUF quantizations) |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Massive multilingual zero-shot TTS model scaling to 600+ languages. Uses diffusion language model-style discrete non-autoregressive architecture with single-stage text-to-acoustic mapping.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 600+ |
| Streaming | ❌ |
| License | ![Apache 2.0][license-apache-2.0] |
| Training-Data | 581k hours |
Features: Simplified single-stage architecture vs conventional two-stage pipelines. Full-codebook random masking strategy with LLM initialization for superior intelligibility. Noise-robust prompt processing.
Links: ![Website][link-website] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Multilingual TTS model with voice cloning and duration control, built on the T5Gemma encoder-decoder LLM architecture. Supports batch generation for multiple audio variations.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 2B-2B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | English, Chinese, Japanese |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Vram | 7.6-10.6 GB |
Features: PM-RoPE positional encoding with XCodec2 audio codec. Low-VRAM options with CPU offloading. Batch inference efficiency with single encoder pass.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: The smallest English TTS model with only 1.6 million parameters. End-to-end neural network achieving ~53x real-time synthesis speed on CPU via ONNX optimization.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | ~3.4 MB (ONNX FP16) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Ultra-compact architecture optimized for CPU-only deployment. Multi-platform support via Python and Node.js APIs. Works on laptops, edge devices, and embedded systems.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: OpenBMB's next-generation tokenizer-free diffusion autoregressive TTS model with 2 billion parameters. Supports 30 languages with automatic detection, voice design from text descriptions, and high-fidelity voice cloning.
Release Date: 2026
| Feature | Value |
|---|---|
| Parameters | 2B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 30 (+ 9 Chinese dialects) |
| Streaming | ✅ |
| Audio Output | 48 kHz |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Tokenizer-free design with LocEnc → TSLM → RALM → LocDiT pipeline. Built-in super-resolution via AudioVAE V2 for 48kHz output.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Soprano is an ultra-lightweight, on-device text-to-speech (TTS) model designed for expressive, high-fidelity speech synthesis at unprecedented speed. The 1.1 release ships an 80M-parameter backbone that achieves up to 20× real-time generation on CPU and 2000× real-time on GPU, with lossless streaming (<250 ms latency on CPU, <15 ms on GPU), <1 GB memory usage at inference, and infinite generation length (automatic text splitting). Output sample rate is 32 kHz, with widespread device support (CUDA / CPU / MPS on Windows, Linux, and Mac). Inference is production-ready through an OpenAI-compatible endpoint, ONNX, WebUI, CLI, and ComfyUI nodes. The base 1.1 model is ekwek/Soprano-1.1-80M on HuggingFace; a fine-tuning toolkit (soprano-factory) was released January 13 2026 alongside the 1.1 weights — the 1.1 release reports 95% fewer hallucinations and a 63% preference rate over 1.0 (Soprano-80M). A live demo runs on ekwek/Soprano-TTS HF Space.
Release Date: December 22, 2025
| Feature | Value |
|---|---|
| Parameters | 80M (Soprano-1.1-80M) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English (US/UK family voices, per HF Space) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Sample Rate | 32,000 Hz |
| Inference Targets | OpenAI-compatible endpoint, ONNX, WebUI, CLI, ComfyUI |
| Performance Cpu | up to 20× real-time |
| Performance Gpu | up to 2000× real-time |
| Memory | <1 GB at inference |
| Text Length | infinite (automatic text splitting) |
| Devices | CUDA, CPU, MPS (Windows, Linux, Mac) |
| Training Toolkit | soprano-factory (https://github.com/ekwek1/soprano-factory) |
| History 1 1 | Soprano-1.1-80M released 2026-01-14 (95% fewer hallucinations; 63% preference over 1.0) |
| History 1 0 | Soprano-80M released 2025-12-22 |
Features: The defining trade-off of this release is extreme on-device
efficiency at sub-100M scale: an 80M-parameter backbone hits
<250 ms CPU latency and <15 ms GPU for lossless streaming
while keeping inference within <1 GB of memory — well under the
multi-billion-parameter budget that newer conversational TTS
systems require. The release pairs the model with soprano-factory
(open-source training/fine-tuning toolkit) so users can build their
own voices on top of the same backbone, and one installation can
drive OpenAI-compatible / ONNX / WebUI / CLI / ComfyUI inference.
The 1.1 update is a measured iteration: 95% fewer hallucinations
and a 63% preference over 1.0 at the same parameter budget, so the
measurable quality jump ships with no added inference cost.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: High-quality TTS synthesis system based on LLMs from ZhipuAI, supporting zero-shot voice cloning with Multi-Reward Reinforcement Learning.
Release Date: December 11, 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Echo is a 2.4B-parameter diffusion-based diffusion transformer (DiT) text-to-speech model. It conditions on target text and up to two minutes of speaker reference audio, generates Fish Speech S1-DAC latents, and decodes to 44.1 kHz audio. Output length is up to 30 seconds per segment. The model is fast at single-sample generation: on one A100, generating 30 seconds of audio from a 120-second prompt takes ~1.45 seconds (RTF < 0.05) — substantially faster than frontier autoregressive approaches at similar quality. The architecture is a deliberate pivot from the author's prior autoregressive-in-DAC-space model Parakeet, which struggled with semantic-consistency retries and weak voice cloning; Echo's diffusion approach trades off real-time interactivity for fast, high-fidelity zero-shot voice cloning in offline synthesis. Trained via the TPU Research Cloud (TRC). Demo (preview) hosted on jordand/echo-tts-preview HF Space; base model on jordand/echo-tts-base.
Release Date: December 4, 2025
| Feature | Value |
|---|---|
| Parameters | 2.4B (DiT) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English (per demo samples) |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Architecture | diffusion transformer (DiT) in Fish Speech S1-DAC latent space |
| Max Segment Duration | 30 s |
| Sample Rate | 44,100 Hz |
| Speaker Reference Max | 120 s |
| Performance A100 Rt | 30 s output in ~1.45 s (RTF < 0.05) |
| Audio Codec | Fish Speech S1-DAC |
| Prior Model | Parakeet (autoregressive in DAC space) |
| Training Infrastructure | TPU Research Cloud (TRC) |
| Inference Requirements | CUDA-capable GPU with at least 8 GB VRAM; Python 3.10+ |
| Sampler | euler CFG with independent guidances for text (3.0) and speaker (8.0); 40 steps; sequence_length 640 default |
| License Clarification | MIT (per GH repo license) |
Features: Echo is a deliberate next-step pivot from autoregressive-in-DAC-space TTS to a full diffusion approach. The author's prior model, Parakeet, generated DAC tokens autoregressively but suffered the classic AR weakness — semantic-consistency retries — and weak voice cloning. Echo keeps Fish Speech S1-DAC latents (so the audio representation is the same proven codec) but moves the generator upstream to a 2.4B DiT operating directly on those latents, conditioned on a long (up to 2-minute) speaker reference. The result: 30-second outputs in ~1.45 s on a single A100 (RTF < 0.05) with high-fidelity zero-shot cloning — fast enough that the "diffusion is too slow" objection no longer applies at the segment length that matters for offline content generation, while the AR class's retry-induced inconsistency is gone by construction.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]
· · · · · · · · · · · · · ·
Description: Real-time TTS model from Microsoft with streaming text input and ultra-low latency (~300ms).
Release Date: December 3, 2025
| Feature | Value |
|---|---|
| Parameters | 0.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Multilingual |
| Streaming | ✅ |
| License | ![MIT][license-mit] |
| Max Duration | ~10 minutes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Advanced TTS system based on LLMs for zero-shot multilingual speech synthesis from FunAudioLLM.
Release Date: December 2025
| Feature | Value |
|---|---|
| Parameters | 0.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 9 + 18+ Chinese dialects |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Liquid AI's first end-to-end audio foundation model with low latency and real-time conversation.
Release Date: November 28, 2025
| Feature | Value |
|---|---|
| Parameters | 1.5B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ✅ |
| License | ![LFM][license-lfm] |
Links: ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Marvis is a conversational real-time streaming TTS from Marvis-Labs. The architecture inherits Sesame's CSM-1B (Conversational Speech Model): a 250M-parameter multimodal backbone that processes interleaved text + audio tokens and a smaller 60M-parameter audio decoder that models the remaining 31 RVQ codebook levels to reconstruct high-quality speech from the backbone's representations. Audio tokens come from Kyutai's mimi codec (RVQ tokens). The dual-transformer split — semantic backbone + small decoder — yields sub-second latency, and the model is built for on-edge / on-device deployment (Apple Silicon / iPad / iPhone / Mac). Two operational choices distinguish Marvis:
Release Date: November 6, 2025
| Feature | Value |
|---|---|
| Parameters | 250M (multimodal backbone) + 60M (audio decoder) = 310M total |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Languages | English, French, German |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | dual-transformer CSM-1B (Conversational Speech Model) — multimodal backbone + audio decoder |
| Audio Codec | Kyutai mimi codec (RVQ tokens; backbone models codebook 0, decoder models codebook 1–31) |
| Quantized Size | ~500 MB (4-bit MLX) |
| Training Dataset | amphion/Emilia-Dataset |
| Library Name | transformers, mlx, mlx-audio |
| Inference Cli | mlx_audio.tts.generate --model Marvis-AI/marvis-tts-250m-v0.2 --stream --text "..." [--ref_audio ./x.wav] |
| Variants In Collection | 250m-v0.2, 250m-v0.2-MLX-{4bit,6bit,8bit}, 100m-v0.2 (+ MLX variants), 250m-v0.2-transformers |
| Emits Text Chunking | no (full-sequence contextual processing) |
Features: Two operational choices make Marvis stand out among conversational TTS. First, no regex chunking: most streaming TTS engines pre-split sentences by regex patterns before feeding them to the generator, which can disrupt flow / intonation; Marvis processes the entire text contextually, treating the text as a single interleaved multimodal sequence. Second, the dual-transformer CSM-1B design — a 250M backbone for codebook 0 (semantic) and a smaller 60M audio decoder for codebooks 1-31 (acoustic) — produces a quantized footprint of ~500 MB, enabling on-device Apple-Silicon inference (iPad / iPhone / Mac) with real-time streaming. The architecture makes a high-quality CSM-style TTS with zero-shot cloning actually deployable at the edge, while the official collection's 4 / 6 / 8-bit MLX variants let users trade footprint for fidelity on a per-device basis.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: AI-Enhanced Text-to-Speech System with Intelligent Optimization and self-learning capabilities.
Release Date: November 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Multi-Speaker | yes (1-4 speakers) |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: State-of-the-art speech model for expressive voice generation with natural language voice control.
Release Date: November 2025
| Feature | Value |
|---|---|
| Parameters | 3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | English (Multi-accent) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: 3B-parameter LLM-based RL audio model specialized in expressive and iterative audio editing.
Release Date: November 2025
| Feature | Value |
|---|---|
| Parameters | 3B (4B BF16) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Mandarin, English, Sichuanese, Cantonese, Japanese, Korean |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: KaniTTS is a 370M-parameter two-stage text-to-speech model from nineninesix-ai. The architecture pairs an LFM-2 backbone LLM (Liquid Foundation Model v2 — a non-transformer, structured state-space architecture) with a neural audio codec for output waveform synthesis. The LLM generates compressed audio-token representations and the codec renders them to 22 kHz waveforms, yielding low-latency generation: ~1 s to produce 15 s of audio on a single RTX 5080, with 2 GB GPU VRAM at inference, and MOS 4.3 / WER < 5% quality on the project's benchmarks. Languages covered: English, German, Chinese, Korean, Arabic, Spanish across multiple per-language voices (English, German, Chinese, Korean, Arabic, Spanish each ship a 400M checkpoint; Japanese ships a 370M "Expo-2025-Osaka" variant; a multilingual 370M checkpoint is also available). MLX variants exist for Apple-Silicon inference. The codec is the same author's nemo-nano-codec-22kHz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported to ~12.5 fps / 0.6 kbps). The model is part of the nineninesix-ai family alongside Gepard.
Release Date: September 30, 2025
| Feature | Value |
|---|---|
| Parameters | 370M (kani-tts-370m multilingual); 400M per-language (en / de / zh / ko / ar / es); 370M (expo2025-osaka-ja) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English, German, Chinese, Korean, Arabic, Spanish (multilingual 370M checkpoint); Japanese (Expo-2025-Osaka variant) |
| Streaming | ✅ |
| License | ![LFM][license-lfm] |
| Sample Rate | 22,000 Hz |
| Backbone Llm | LFM-2 (Liquid Foundation Model; non-transformer structured state-space architecture) |
| Audio Codec | nineninesix/nemo-nano-codec-22khz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported) |
| Performance Rt 5080 | ~1 s for 15 s audio on RTX 5080 |
| Memory | 2 GB GPU VRAM at inference |
| Quality Mos | 4.3 / 5 (naturalness) |
| Quality Wer | <5% (accuracy) |
| Training Dataset | ~80k hours (LibriTTS, Common Voice, Emilia) |
| Training Hardware | 8x H100 GPUs, 45 hours on Lambda AI |
| Per Language Models | kani-tts-400m-{en,zh,de,ar,es,ko} on HuggingFace |
| Pretrained Checkpoints | 0.2-pt (450M), 0.3-pt (400M) for custom posttraining / fine-tuning |
| Mlx Variants | kani-tts-370m-MLX (Apple Silicon) |
| Arxiv | 2505.20506 |
Features: KaniTTS's design choice worth flagging: it pairs a non-transformer
LFM-2 backbone (Liquid Foundation Model — structured state-space
rather than attention) with a neural audio codec for output at
the 370M scale. The choice lets the model hit a ~1 s / 15 s
audio generation rate on a 2 GB GPU VRAM budget — sub-1B
parameters, sub-entry-tier GPU requirement, but still multilingual
across six languages. The two-stage approach (LLM → codec) is
conventional; what's less conventional is the choice of a
state-space backbone over the usual transformer decoder at this
scale, hitting latency / VRAM numbers that open up sub-1B real-time
TTS on consumer-grade hardware. The same author ships soprano-factory-style
companion assets (pretrained v0.2-pt / v0.3-pt checkpoints, a
NeMo NanoCodec MLX port) to lower the bar for fine-tuning on custom
datasets.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: This is an unofficial, work-in-progress LoRA fine-tuning toolkit for the VibeVoice TTS / speech model (1.5B-base and 7B-base checkpoints). The base VibeVoice checkpoints are the same ones covered in this list's a separate entry: audio-conditioned diffusion TTS for spoken dialogue, streaming, etc. This toolkit takes pretrained VibeVoice weights + a paired (text, audio, optional reference-audio) dataset and trains a LoRA adapter against two losses simultaneously:
Release Date: September 16, 2025
| Feature | Value |
|---|---|
| Parameters | 1.5B (LoRA-adapted) / 7B (LoRA-adapted) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | inherits base VibeVoice coverage |
| Streaming | not directly (toolkit output is a LoRA adapter; the adapter inherits VibeVoice inference shape) |
| License | ![MIT][license-mit] |
| Loss Text | masked cross-entropy on text tokens |
| Loss Acoustic | diffusion MSE on acoustic latents |
| Hardware 1 5B | ≥16 GB VRAM |
| Hardware 7B | ≥48 GB VRAM |
| Transformers Version | 4.51.3 (known-good; other versions may break on Qwen2 architecture) |
| Tested Docker Image | runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04 |
| Audio Target Format | 24 kHz audio (paired dataset of target-audio + transcripts + optional reference-audio prompts) |
| Training Entrypoint | python -m src.finetune_vibevoice_lora --model_name_or_path aoi-ot/VibeVoice-Large --processor_name_or_path src/vibevoice/processor --dataset_name <your/dataset> --text_column_name text [--voice_column_name audio_ref] |
| Supports Hf Dataset Loader | yes |
| Output | LoRA adapter compatible with VibeVoice base |
Features: The dual-loss trick is the technical center of this toolkit. Naive LoRA fine-tuning of a unified TTS model often specializes the synthesis but silently damages the text LLM capability the base inherited from its Qwen-class backbone; the "masked CE on text tokens + diffusion MSE on acoustic latents" two-headed loss preserves both competencies at training time. Pair that with the careful pinning of Transformers 4.51.3 (other versions break on the Qwen2 architecture) and a documented minimum-VRAM budget per base size (16 GB for 1.5B, 48 GB for 7B), and you get a reproducible recipe for community fine-tuning of VibeVoice — something the official Microsoft VibeVoice release doesn't ship out-of-the-box.
Links: ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Tokenizer-free TTS system for context-aware speech generation and true-to-life voice cloning.
Release Date: September 16, 2025
| Feature | Value |
|---|---|
| Parameters | 640M-800M |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Long-form streaming TTS system for multi-speaker dialogue generation with stable, natural speech.
Release Date: September 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | EN, ZH, JP, KO, FR, DE, RU |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Multi-Speaker | yes (4 speakers) |
| Max Duration | 3 minutes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: NVIDIA ADLR's fully open-source Large Audio Language Model with state-of-the-art audio understanding. Audio Flamingo Next (AF-Next) is the latest generation featuring stronger general audio understanding, longer context support, and timestamp-grounded reasoning.
Release Date: July 2025 (AF3), 2026 (AF-Next)
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ❌ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | Multi-lingual |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Context | Up to 30 minutes |
Features: Key Innovation (AF-Next): Staged curriculum training with GRPO-based RL post-training. Three specialized checkpoints: Instruct, Think (reasoning), and Captioner. Temporal Audio Chain-of-Thought grounding intermediate reasoning to timestamps.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Fast and high-quality zero-shot TTS models based on flow matching.
Release Date: June 16, 2025
| Feature | Value |
|---|---|
| Parameters | 123M |
| Languages | Chinese, English |
| License | ![Apache 2.0][license-apache-2.0] |
| Zero-Shot-Cloning | yes |
| Dialogue | yes |
Links: ![GitHub][link-github] ![Website][link-website] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: State-of-the-art open source TTS and voice cloning model that generates natural, realistic, and emotionally rich speech.
Release Date: May 31, 2025 (v1.5.1)
| Feature | Value |
|---|---|
| Parameters | 4B (S1), 0.5B (S1-mini) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 8 (EN, JP, KO, ZH, FR, DE, AR, ES) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Rtf | ~1:7 |
Links: ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Family of SOTA open-source TTS models by Resemble AI, covering a single-language English line plus a multilingual V3 release that brings broader language coverage, more consistent speaker similarity, reduced hallucinations, and more natural conversational speech across 23+ languages.
Release Date: April 24, 2025
| Feature | Value |
|---|---|
| Parameters | 500M (Llama backbone, 0.5B) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 23+ |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
Features: First open-source TTS model with explicit emotion exaggeration control, plus an alignment-informed inference pipeline and a watermarked decoder. Multilingual V3 narrows the quality gap to closed systems like ElevenLabs on cross-language voice cloning while staying under 1B parameters.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: SOTA open-source TTS built on Llama-3b backbone demonstrating emergent capabilities of LLMs for speech synthesis.
Release Date: April 2025
| Feature | Value |
|---|---|
| Parameters | 3B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Multilingual |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![Website][link-website]
· · · · · · · · · · · · · ·
Description: Advanced zero-shot speech synthesis with Sparse Alignment Enhanced Latent Diffusion Transformer.
Release Date: March 22, 2025
| Feature | Value |
|---|---|
| Parameters | 0.45B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Efficient LLM-Based TTS Model with Single-Stream Decoupled Speech Tokens, built on Qwen2.5.
Release Date: March 2025
| Feature | Value |
|---|---|
| Parameters | 0.5B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Production-ready open-source framework for intelligent speech interaction with unified speech comprehension and generation.
Release Date: February 17, 2025
| Feature | Value |
|---|---|
| Parameters | 130B (Chat), 3B (TTS) |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Chinese, English, Japanese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Kokoro is an open-weight Text-to-Speech model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.
Release Date: January 27, 2025 (v1.0)
| Feature | Value |
|---|---|
| Parameters | 82M |
| Architecture | StyleTTS 2, ISTFTNet |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | 8 (54 voices) |
| Streaming | ✅ |
| Cost | <$0.06 per hour of audio |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: KokoClone is a fast, real-time compatible multilingual voice cloning system built on top of Kokoro-ONNX. It enables users to type text in multiple languages, provide a short 3-10 second reference audio clip, and instantly generate speech in that same voice.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | 82M (Base: Kokoro-ONNX) |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ✅ |
| Languages | 7 (En, Hi, Fr, Ja, Zh, It, Pt, Es) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Lightweight ZipVoice-based TTS model for high quality voice cloning at speeds exceeding 150x realtime.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ❌ |
| Emotion Control | ❌ |
| Languages | - |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Rtf | 150x |
| Vram | 1GB |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Audio Language Model by Xiaomi functioning as a Few-Shot Learner with SOTA audio understanding.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | Multi-lingual |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: SOTA Multi-Speaker TTS model for generating realistic long-form podcasts with dialectal diversity.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | Mandarin, English, Cantonese, Sichuanese, Henanese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Max Duration | 90+ minutes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Advanced on-device Vietnamese TTS model with instant voice cloning from 3-5 seconds of reference audio.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | 0.3B-0.6B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ❌ |
| Languages | Vietnamese |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: 1.6B parameter TTS model by Nari Labs for generating ultra-realistic dialogue in one pass.
Release Date: June 27, 2024
| Feature | Value |
|---|---|
| Parameters | 1.6B |
| Voice Cloning | ✅ |
| Asr | ❌ |
| Pronunciation | ✅ |
| Emotion Control | ✅ |
| Languages | English |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: MeloTTS is a high-quality multi-lingual text-to-speech library from MyShell.ai in collaboration with MIT, supporting English (American, British, Indian, Australian, and a default accent), Spanish, French, Chinese (with mixed Chinese–English capability), Japanese, and Korean. Built on VITS / VITS2 / Bert-VITS2 family work and packaged with both a Python API and a Web UI, it runs fast enough for CPU real-time inference.
Release Date: February 19, 2024
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | English (American, British, Indian, Australian, Default), Spanish, French, Chinese, Japanese, Korean |
| Streaming | ❌ |
| License | ![MIT][license-mit] |
| Base | VITS / VITS2 / Bert-VITS2 family |
| Mixed Chinese English | yes |
Features: A multi-accent multilingual TTS library that ships both a Python API and a Web UI on top of the VITS-style architecture, with explicit English-accent coverage (American, British, Indian, Australian, Default) and mixed Chinese–English output — designed for fast CPU real-time inference without requiring GPU servers.
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: Open-source audio foundation model by Moonshot AI for audio understanding, generation, and conversation.
Release Date: 2024
| Feature | Value |
|---|---|
| Parameters | 7B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Emotion Control | ✅ |
| Languages | Multi-lingual |
| Streaming | ✅ |
| License | ![MIT][license-mit] ![Apache 2.0][license-apache-2.0] |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: eSpeak NG is a compact open-source software text-to-speech synthesizer for Linux, Windows, Android, and other operating systems. It supports more than 100 languages and accents, is a fork of Jonathan Duddington's original eSpeak engine, and uses the formant synthesis method: the model produces speech by explicitly computing the acoustic resonances (formants) of each phoneme, not by concatenating human-speech recordings. The trade is well-known — the speech is clear and usable at high playback speeds, but not as natural or smooth as larger neural or concatenative synthesizers. The compensation is size: the program and its data, including many languages, total a few megabytes. Other synthesis methods supported: Klatt formant synthesis and MBROLA diphone back-end via the documented integration.
Release Date: December 8, 2015
| Feature | Value |
|---|---|
| Parameters | n/a (formant-synthesis engine; not a neural model) |
| Voice Cloning | ❌ |
| Asr | ❌ |
| Languages | 100+ languages and accents (see docs/languages.md) |
| Streaming | ✅ |
| License | ![GPL 3.0][license-gpl-3.0] |
| Synthesis Method | formant synthesis (primary); Klatt formant synthesis (secondary); MBROLA diphone backend (optional) |
| Footprint | a few MB (program + data + many languages) |
| Audio Output | WAV file (CLI), direct playback, or shared-library API |
| Input Formats | text from file / stdin (CLI), SSML (partial), HTML (partial) |
| Packages | CLI (espeak-ng man page), shared library (libespeak-ng), SAPI5 Windows module |
| Supersedes | eSpeak (Jonathan Duddington's original engine) |
| Downstream Usage | G2P / phonemizer for neural TTS pipelines (e.g. sanoTTS bundles espeak-ng for its duration model) |
| Platforms | Linux, Windows, Android, Solaris, Mac OS X |
Features: eSpeak-NG is the canonical reference implementation of compact multi-language formant synthesis. Its 100+-language coverage in a few megabytes, plus SSML / SAPI5 / MBROLA / shared-library / CLI surfaces, are matched only by neural TTS systems that are orders of magnitude larger. The reason it belongs in a list whose other entries are neural TTS systems is its continued quiet role in the neural stack as a G2P / phonemizer front-end — the phoneme inventory and grapheme-to-phoneme rules that sanoTTS and similar sub-1B neural TTS engines bundle are often just a port of eSpeak-NG's language-data files. So even if the formant-synthesis audio output itself has been surpassed for naturalness, the phoneme infrastructure underneath many of the smaller neural TTS entries on this list still traces back to eSpeak-NG.
Links: ![GitHub][link-github]
· · · · · · · · · · · · · ·
Models that can generate audio from multiple input modalities (video, text, image, audio). These are unified frameworks for multimodal audio synthesis.
| Model | Text | Video | Audio | Max Duration | Sample Rate | License |
|---|---|---|---|---|---|---|
| MiDashengLM-Gen | ✅ | ❌ | ❌ | — | 16 kHz | ![Apache 2.0][license-apache-2.0] |
| ScenA | ✅ | ❌ | ✅ | — | — | ![Other][license-other] |
| Nemotron-Labs-Audex-2B | ✅ | ❌ | ✅ | — | — | ![NVIDIA NC][license-nvidia-noncommercial] |
| Nemotron-Labs-Audex-30B-A3B | ✅ | ❌ | ✅ | — | — | ![NVIDIA NC][license-nvidia-noncommercial] |
| MOSS-SoundEffect | ✅ | — | — | 30 s | 48 kHz | ![Apache 2.0][license-apache-2.0] |
| Omni2Sound (Omni2Audio) | ✅ | ✅ | ✅ | — | — | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| ControlFoley | ✅ | ✅ | ✅ | — | 44,100 Hz | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Woosh | ✅ | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
| Chroma-4B | ✅ | ❌ | ✅ | — | — | ![Apache 2.0][license-apache-2.0] |
| Uni-MoE (Audio) | ✅ | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
| AudioX / Audio-Omni | ✅ | ✅ | ✅ | — | — | ![Apache 2.0][license-apache-2.0] ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| HunyuanVideo-Foley | ✅ | ✅ | — | — | 48 kHz | ![Research Only][license-research-only] |
| PrismAudio | — | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
| ThinkSound | ✅ | — | ✅ | — | — | ![Apache 2.0][license-apache-2.0] |
| MMAudio | ✅ | ✅ | — | — | — | ![Apache 2.0][license-apache-2.0] |
Description: MiDashengLM-Gen (MiDasheng Language Model for Generation) is an end-to-end framework for unified audio-scene generation from Xiaomi. Built on a pre-trained LLM and the Dasheng audio tokenizer, it couples per-token conditional flow matching with autoregressive generation to produce coherent 16 kHz audio that simultaneously blends speech, music, sound effects and environmental acoustics from a structured text description. It supports 9 languages with emotion control and approaches dedicated TTS intelligibility on speech (Seed-TTS English WER drops from 12.15% to 2.79%) while retaining mixed-audio scene capability, and extends competitively to multilingual settings.
Release Date: August 12, 2026
| Feature | Value |
|---|---|
| Text | ✅ |
| Video | ❌ |
| Image | ❌ |
| Audio | ❌ |
| Sample Rate | 16 kHz |
| Languages | 9 |
| Emotion Control | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 1.7B (Qwen3-1.7B backbone) |
| Architecture | DashengTokenizer (768-dim @25Hz) + Qwen3-1.7B + flow-matching DiT (16 layers, hidden 2048) |
Features: LLM-conditioned high-dimensional (768-dim @25Hz) audio latents generated without quantization artifacts; audio-text alignment pre-training maps latents into the LLM token space before generation; a learned stop head enables variable-length truncation. First end-to-end trained model for general text-to-audio-scene generation.
Links: ![Demo][link-demo] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: ScenA generates multi-speaker audio scenes — dialogue and conversation with sound effects and ambience — from a text prompt, conditioned on one or more reference-audio clips that set the speakers' voices. Unlike prior multi-speaker dialogue systems it uses no per-turn tags, multi-stream transcripts, or speaker embeddings: a free-form natural-language prompt alone describes the scene. The text prompt determines which reference voice speaks where, allowing overlapping speech, spontaneous paralinguistic events, and scene-level ambient sound — all inherited from the in-the-wild text-to-audio pretraining distribution. The architecture is an audio-only, reference-conditioned flow-matching DiT built on the LTX-2 backbone (~4B parameters, 48 layers). Reference latents are concatenated into the token sequence and distinguished by lightweight identity-aware positional encodings. The training tackles a specifically identified "Reference Shortcut" failure mode — under standard noise schedules the model can identify the matching reference by noisy-target acoustic similarity, bypassing the text prompt — by using a high-noise-biased timestep distribution that forces reliance on the prompt for speaker assignment. Evaluator: CoVoMix2-Dialogue benchmark. Project page, code, paper, and HuggingFace checkpoint are linked below.
Release Date: July 7, 2026
| Feature | Value |
|---|---|
| Parameters | ~4B (DiT, 48 layers; built on LTX-2 architecture) |
| Text | ✅ |
| Video | ❌ |
| Audio | ✅ |
| Max Duration | not stated (scene-level generation) |
| Sample Rate | (not stated; inherits LTX-2 audio VAE) |
| Voice Cloning | ✅ |
| Multi Speaker | yes |
| Ambient Sound | yes (SFX, room acoustics, overlapping speech) |
| Architecture | flow-matching DiT (LTX-2 backbone, audio-only) |
| Speaker Assignment | natural language (no per-turn tags / identity encoders) |
| Training Fix | high-noise-biased timestep distribution (defeats Reference Shortcut) |
| Text Encoder | google/gemma-3-12b-it |
| Audio Vae | bundled (~365 MB; encodes+decodes so full LTX-2 not needed) |
| Checkpoint Size | ~8.2 GB (scena.safetensors) + ~365 MB (audio_vae.safetensors) |
| License | ![Other][license-other] |
| Training Data | in-the-wild text-to-audio pretrained, then reference-conditioned fine-tune |
| Evaluation | CoVoMix2-Dialogue (speaker-binding metrics) |
Features: The "Reference Shortcut" failure-mode identification is the technical center of the work: under standard diffusion noise schedules, a multi-speaker reference-conditioned model can match each reference to the noisy-target segment by acoustic similarity alone, bypassing the text prompt entirely. ScenA's high-noise-biased timestep distribution forces the model to rely on the prompt for speaker assignment at training time. Combined with the absence of any per-turn speaker structure (tags / transcripts / identity encoders) and the prompt's role as the only speaker-routing signal, this yields multi-speaker conversational scenes with overlapping speech, paralinguistic events, and ambient texture that previous structured-supervision multi-speaker systems filter out by design.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: Nemotron-Labs-Audex-2B is NVIDIA's smaller sibling of the Audex unified audio-text LLM. Like the 30B-A3B flagship, the 2B is a single model family that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). It is built on the same audio-vocabulary-extended transformer stack as the 30B-A3B but at a densely-parameterized 2B scale (no MoE), so the compute and memory footprint are lowered to a budget tractable on more modest hardware. The 2B checkpoint is the project-tagged SFT variant in the Audex collection — instruction-tuned and ready for inference. Both sizes preserve text reasoning, alignment, knowledge, long-context, and agentic capabilities of the text backbone while adding discrete-token audio I/O.
Release Date: July 6, 2026
| Feature | Value |
|---|---|
| Parameters | 2B (dense; SFT fine-tune, instruct + reasoning-ready) |
| Text | ✅ |
| Video | ❌ |
| Audio | ✅ |
| Modalities | text + audio (input and output) |
| Max Duration | not stated |
| Sample Rate | not stated (decoder output) |
| Voice Cloning | ❌ |
| Audio Understanding | yes (audio QA, classification) |
| Asr | ✅ |
| Speech Translation | yes |
| Text To Speech | yes |
| Text To Audio | yes |
| Speech To Speech Generation | yes |
| Reasoning Mode | yes (thinking + instruct modes inherited from text backbone) |
| License | ![NVIDIA NC][license-nvidia-noncommercial] |
| Pipeline Tag | text-generation |
| Library Name | transformers |
| Derived From | same family as Nemotron-Labs-Audex-30B-A3B |
| Companion 30B | nvidia/Nemotron-Labs-Audex-30B-A3B (MoE: 30B total, 3B active) |
| Spaces | nvidia/Nemotron-Labs-Audex, WaveCut/Nemotron-Labs-Audex, hugging-apps/nemotron-labs-audex-2b |
| Createdat | 2026-07-06T16:21:07Z |
| Downloads | ~2.4k |
Features: The 2B sibling matters because it preserves the central thesis of the Audex paper — unified audio-text LLM intelligence without regressing on text intelligence — while dropping the parameter budget substantially. The 30B-A3B MoE hits a 1M-context, agentic flagship tier; the 2B dense version is the same audio-aware architecture extended down to a budget that doesn't require a high-end MoE serving stack. The pair lets users choose on deployment cost rather than on capability sub-selection: the 2B ships the same audio-to-audio + text-to-audio + audio-understanding
Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Nemotron-Labs-Audex-30B-A3B is NVIDIA's unified audio-text LLM — a single model that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). Built on Nemotron-Cascade-2-30B-A3B (text-only MoE: 30B parameters, 3B active), Audex extends the vocabulary with discrete audio tokens for speech / general-audio output and adds an audio encoder for speech / general-audio input. Runs in thinking and instruct (non-thinking) modes and supports up to a 1M-token context length — preserving text-reasoning, alignment, knowledge, long-context, and agentic capabilities of the backbone while gaining audio tasks.
Release Date: July 6, 2026
| Feature | Value |
|---|---|
| Parameters | 30B MoE (3B active) |
| Modalities | audio (input and output) |
| Audio Understanding | yes |
| Asr | ✅ |
| Speech Translation | yes |
| Text To Speech | yes |
| Text To Audio | yes |
| Speech To Speech Generation | yes |
| Voice Cloning | ❌ |
| License | ![NVIDIA NC][license-nvidia-noncommercial] |
| Languages | English |
| Modes | thinking, instruct (non-thinking) |
| Context Length | 1M tokens |
| Template | ChatML (with <think>…</think> for thinking mode) |
| Inference | vLLM 0.20.0 (recommended) or transformers >= 4.53.0 (mamba-ssm + causal-conv1d required) |
Features: First-class audio I/O for a 30B/3B-active text LLM: extended vocabulary with discrete audio tokens for outputting speech and general audio, plus an audio encoder for input — so the same backbone keeps its strong text reasoning (alignment, knowledge, long-context) and adds ASR + speech translation + TTS + audio generation + S2S without retraining. The MoE form (30B routes, 3B active) keeps inference tractable for a single pipeline that does both.
Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection]
· · · · · · · · · · · · · ·
Description: MOSS-SoundEffect is the dedicated text-to-sound model in the OpenMOSS / MOSI.AI MOSS-TTS family. It turns natural-language captions into high-fidelity non-speech audio (ambience, urban scenes, creatures, human actions, and short music-like clips).
Release Date: May 25, 2026
| Feature | Value |
|---|---|
| Type | Text-to-Sound / SFX generation |
| Conditioning | Text |
| Max Duration | 30 seconds |
| Sample Rate | 48 kHz |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | DiT + Flow Matching + DAC VAE + Qwen3 text encoder |
| Parameters | 1.3B (DiT variant 1.3B) |
| Languages | English, Chinese |
| Inference Defaults | 100 flow-match steps, cfg 4.0, sigma_shift 5.0 |
| Library | diffusers |
Features: Replaces the discrete-token autoregressive v1 (which bottlenecked on vocabulary) with a continuous-latent DiT + Flow Matching paired with a DAC VAE — yielding 30 s stable audio, bilingual English + Chinese prompts, and a clean CFG/sigma-shift inference schedule (cfg 4.0, shift 5.0) that works straight out of the box on the diffusers library.
Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: Omni2Sound — also written Omni2Audio on the project page — is a unified VT2A / V2A / T2A framework and a CVPR 2026 Highlight. A single Diffusion Transformer (DiT) backbone with a decoupled two-branch conditioning design:
Release Date: April 20, 2026
| Feature | Value |
|---|---|
| Conditioning | Text / Video / Text+Video |
| Modalities | Video, Audio |
| Asr | ❌ |
| Voice Cloning | ❌ |
| Text | ✅ |
| Video | ✅ |
| Image | ❌ |
| Audio | ✅ |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Tasks | VT2A, V2A, T2A (single model) |
| Architecture | DiT + decoupled Semantic / Temporal branches + 3-stage progressive training |
| Pipeline Tag | text-to-audio |
Features: One single model that is SOTA on three distinct tasks (VT2A, V2A, T2A) without a separate model per mode — decoupled semantic and temporal conditioning let the same DiT backbone handle text-only, video-only, and text+video conditioning by cleanly omitting the missing modality rather than padding it, which is what most prior unified VA models had to do.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Benchmark][link-benchmark]
· · · · · · · · · · · · · ·
Description: ControlFoley (Xiaomi MiLM Plus) is a unified controllable video-to-audio (foley) generation model. It supports four conditioning combinations under one architecture:
Release Date: April 13, 2026
| Feature | Value |
|---|---|
| Conditioning | Text / Video / Text + Video / Video + Reference Audio |
| Modalities | Video (visual), Audio (foley) |
| Asr | ❌ |
| Voice Cloning | ❌ |
| Text | ✅ |
| Video | ✅ |
| Image | ❌ |
| Audio | ✅ |
| Sample Rate | 44,100 Hz |
| License | ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
| Pipeline Tag | text-to-audio |
| Library | diffusers |
| Cross Modal Conflict | handled via modality-specific control (no explicit router) |
| Inference Skill | ClawHub ControlFoley Audio Generator |
| Upcoming | ComfyUI nodes (in preparation, expanding to V2A / TV2A / TC-V2A / AC-V2A / T2A) |
Features: Modality-specific cross-modal conflict resolution in a single generative stack: text governs semantics, reference audio governs timbre/acoustic style, and video governs temporal synchronization. Rather than routing to a single user-trusted modality, the model decouples control axes so an input disagreement (video shows a dog barking, text asks for a cat) is decomposed into a coherent output that respects each modality's responsibility. Trained with all-modality dropout for modality-robustness, ControlFoley is the first foley system that brings all four conditioning modes — T2A, V2A, TV2A, AC-V2A — under one model.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv] ![Skill][link-skill]
· · · · · · · · · · · · · ·
Description: Sony AI's sound effect foundation model for text-to-audio and video-to-audio generation. Includes Woosh-AE (audio encoder/decoder), Woosh-Flow/DFlow (T2A), and Woosh-VFlow/DVFlow (V2A) with distilled fast inference variants.
Release Date: 2026
| Feature | Value |
|---|---|
| Architecture | Flow-based generative models |
| Text | ✅ |
| Video | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Audio-Encoding | yes |
| Fast-Inference | yes (Distilled models) |
Features: Optimized for sound effects (not general audio) with both public and private model versions. Video-conditioned generation without requiring captions. Competitive with Stable Audio Open and TangoFlux.
Links: ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Chroma 1.0 (FlashLabs' Chroma-4B on HuggingFace) is the first open-source, real-time, end-to-end spoken dialogue model that achieves both sub-second end-to-end latency and high-fidelity personalized voice cloning. The pipeline is end-to-end — no separate ASR → LLM → TTS stitch — speech goes in, speech comes out. The architectural centerpiece is an interleaved text-audio token schedule (1 text : 2 audio) that supports streaming generation, so the model can begin emitting audio while the user is still talking (broken-off turns / barge-in handled). Experimental results from the project's paper:
Release Date: November 28, 2025
| Feature | Value |
|---|---|
| Parameters | 4B |
| Voice Cloning | ✅ |
| Asr | ✅ |
| Languages | English (per benchmark reporting) |
| Streaming | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Architecture | end-to-end spoken-dialogue LLM; interleaved text-audio token schedule (1:2); custom_code modules |
| Pipeline Tag | any-to-any (HF classification) |
| Audio Tokenization | chroma tokenizer (RVQ-style per project's tag) |
| Latency Rtf | 0.43 (speech out ~2.3× wall-clock) |
| Speaker Similarity | +10.96% relative improvement over human baseline |
| Inference Library | transformers (custom_code) |
| Correlations With Larger Class | matches full dialogue turn at streaming latency |
| Pretrained | yes (safetensors weights) |
| Hf Space Demos | hysts/Chroma-4B, Pnevka/Chroma-4B |
| History | paper arXiv 2601.11141 (2026-01) |
Features: Two bets together produce the dual property that no prior open-source spoken-dialogue model has hit simultaneously. First, an interleaved text-audio token schedule (1:2) — text tokens and audio tokens are interleaved at a fixed 1:2 ratio through the sequence, which gives the model a structured place to emit audio while still consuming user audio + text context, supporting sub-second end-to-end latency without a separate ASR / LLM / TTS pipeline. Second, personalized voice cloning baked into the spoke-dialogue model — the cloned voice is not bolted on top by a separate TTS stage (as is the default pattern), it's in-model at the audio-token-generation layer. The empirical payoff is a 10.96% relative speaker-similarity gain over the human baseline (i.e. the cloned voice is closer to the reference speaker than two of the same human speaker's recordings are to each other), while hitting RTF 0.43 — a floor that prior systems exceeded either in latency (no streaming) or in cloning fidelity (parrot the speaker poorly), rarely both.
Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: MoE-based omnimodal model with voice cloning, TTS, T2M (text-to-music), and V2M (video-to-music).
Release Date: October 16, 2025 (Uni-MoE-Audio)
| Feature | Value |
|---|---|
| Parameters | - |
| Voice Cloning | ✅ |
| Text | ✅ |
| Video | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Dynamic-Routing | yes |
Links: ![GitHub][link-github] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Audio-Omni is the first end-to-end framework unifying understanding, generation, and editing across general sound, music, and speech domains. Presented at SIGGRAPH 2026. AudioX is a unified framework integrating text, video, image, and audio conditions.
Release Date: March 2025 (AudioX), 2026 (Audio-Omni)
| Feature | Value |
|---|---|
| Parameters | - |
| Text | ✅ |
| Video | ✅ |
| Audio | ✅ |
| License | ![Apache 2.0][license-apache-2.0] ![CC BY-NC 4.0][license-cc-by-nc-4.0] |
Features: First unified framework covering all three audio domains. Combines frozen multimodal LLM (Qwen2.5-Omni) with trainable Diffusion Transformer for high-fidelity synthesis. Any-to-any audio processing.
Links: ![GitHub][link-github] ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Tencent's end-to-end video sound effect generation model for professional-grade AI Foley sound generation. Analyzes footage and creates immersive audio that matches the visual content perfectly.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Sample Rate | 48 kHz |
| Text | ✅ |
| Video | ✅ |
| License | ![Research Only][license-research-only] |
| High-Quality-Foley | yes |
| Context-Aware | yes |
Links: ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Video-to-Audio generation framework with Reinforcement Learning and specialized Chain-of-Thought (CoT) planning. Decomposes reasoning into four specialized modules (Semantic, Temporal, Aesthetic, Spatial CoT) for comprehensive video understanding. Built upon ThinkSound.
Release Date: 2025 (ICLR 2026)
| Feature | Value |
|---|---|
| Parameters | 518M |
| Video | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Cot-Planning | yes (4 modules) |
| Multi-Dimensional-Rl | yes |
| Fast-Grpo | yes (Hybrid ODE-SDE) |
| Inference-Time | 0.63 seconds |
Features: Performance Benchmarks:
| Metric | VGGSound | AudioCanvas |
|---|---|---|
| Semantic (CLAP) | 0.47 | 0.52 |
| Temporal (DeSync↓) | 0.41 | 0.36 |
| Aesthetic (MOS-Q) | 4.21±0.35 | 4.12±0.28 |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
Description: Unified Any2Audio generation framework with flow matching guided by Chain-of-Thought (CoT) reasoning. Supports generating or editing audio from video, text, audio, or their combinations. Accepted to NeurIPS 2025.
Release Date: 2025
| Feature | Value |
|---|---|
| Parameters | - |
| Text | ✅ |
| Audio | ✅ |
| License | ![Research Only][license-research-only] ![Apache 2.0][license-apache-2.0] |
| Cot-Driven-Reasoning | yes |
| Interactive-Object-Centric-Editing | yes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]
· · · · · · · · · · · · · ·
Description: Multimodal joint training framework for high-quality synchronized audio generation from video and/or text inputs. State-of-the-art open source model for generating sounds for videos, images, and text prompts.
Release Date: December 2024 (CVPR 2025)
| Feature | Value |
|---|---|
| Parameters | - |
| Text | ✅ |
| Video | ✅ |
| Image | ✅ |
| License | ![Apache 2.0][license-apache-2.0] |
| Synchronized-Audio | yes |
| Multimodal-Joint-Training | yes |
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]
· · · · · · · · · · · · · ·
| Model | Type | Bandwidth Extension | Inpainting | License |
|---|---|---|---|---|
| RE-USE | Universal Speech Enhancement | ✅ | ❌ | ![NVIDIA NC][license-nvidia-noncommercial] |
| NovaSR | Audio Super-Resolution | ✅ | ❌ | ![Apache 2.0][license-apache-2.0] |
| QuarkAudio-UniSE | Universal Speech Enhancement | ❌ | ❌ | ![Apache 2.0][license-apache-2.0] |
| PASE | Speech Enhancement | ❌ | ❌ | ![Apache 2.0][license-apache-2.0] |
| DTT-BSR | Music Source Restoration | ❌ | ❌ | ![MIT][license-mit] |
| NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges) | High-Resolution Audio Restoration | ✅ | ✅ | ![NVIDIA NC][license-nvidia-noncommercial] |
| ZipEnhancer | Acoustic Noise Suppression | ❌ | ❌ | ![Apache 2.0][license-apache-2.0] |
| AudioSR | Audio Super-Resolution | ✅ | ❌ | ![Apache 2.0][license-apache-2.0] |
Description: RE-USE (RE-…), NVIDIA's multilingual universal speech enhancement model, targets distortion–perception trade-off by training a single model that balances listening quality against fidelity to the underlying linguistic / speaker / emotional content. Designed to restore diverse degraded speech while leaving everything else (content, identity, prosody, accent, paralinguistic attributes) intact.
Release Date: March 17, 2026
| Feature | Value |
|---|---|
| Type | Universal Speech Enhancement |
| Bandwidth Extension | ✅ |
| Inpainting | ❌ |
| Sample Rate | 8 / 16 / 22.05 / 24 / 32 / 44.1 / 48 kHz (multi-rate input) |
| Architecture | Mamba-SSM backbone |
| Degradation Coverage | additive noise, reverberation, clipping, bandwidth limit, codec artifacts, packet loss, low-quality mics |
| Language Agnostic | yes |
| License | ![NVIDIA NC][license-nvidia-noncommercial] |
Features: A single Mamba-SSM model that handles seven different input sample rates (no resampling pre-step), covers a broad degradation menu in one checkpoint, stays language-agnostic without per-language training, and explicitly balances distortion reduction against fidelity to the input speech — addressing the universal-SE trade-off that earlier single-purpose enhancers couldn't.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: NovaSR is a tiny audio upsampler (~52 kB parameter count) that bandwidth-extends 16 kHz input up to 48 kHz. Public card on YatharthS/NovaSR advertises realtime factors around 3500× on A100, making it a candidate for real-time on-device super-resolution where model size dominates latency. Inference path is small enough to fit in CPU memory; the use case is speech-bandwidth extension without GPU.
Release Date: January 6, 2026
| Feature | Value |
|---|---|
| Type | Audio Super-Resolution (16 kHz → 48 kHz) |
| Bandwidth Extension | ✅ |
| Inpainting | ❌ |
| Channels | mono |
| License | ![Apache 2.0][license-apache-2.0] |
| Parameters | 52 kB |
| Streamable | yes (low VRAM / runs without GPU) |
| Realtime Factor | ~3500× (A100) |
Features: A 52 kB-parameter Upsampler that hits ~3500× realtime on GPU and runs on CPU — pushing bandwidth extension below the size / latency envelope where a typical neural upsampler is unacceptable (real-time on-device speech enhancement).
Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]
· · · · · · · · · · · · · ·
Description: UniSE is a unified, prompt-free autoregressive speech-enhancement framework built on a decoder-only language model. A single model performs multiple speech-enhancement tasks — speech restoration (SR / denoising), target-speaker extraction (TSE), source separation (SS), and acoustic echo cancellation (AEC, in development) — without explicit task-specific instructions or prompt conditioning; the language model infers the task from the input context. Stack: WavLM as the feature extractor, BiCodec as the discrete codec, and a decoder-only LM as the middle autoregressive backbone. Outputs reconstructed waveform from predicted discrete token sequences.
Release Date: December 22, 2025
| Feature | Value |
|---|---|
| Voice Cloning | ❌ |
| Asr | ❌ |
| Streaming | ❌ |
| Languages | English (paper demo) |
| License | ![Apache 2.0][license-apache-2.0] |
| Tasks | Speech Restoration, Target Speaker Extraction, Source Separation, AEC (developing) |
| Architecture | WavLM (feature extractor) + BiCodec (discrete codec) + decoder-only AR-LM |
| Unified | yes (single model handles SE, SR, TSE, SS without explicit task prompts) |
| Prompt Free | yes (LM infers task from input context) |
| Dataset Signals | noise + reverb + packet-loss + clean (configurable per task) |
| Training | Speech-enhancement SFT, then multitask joint training |
Features: A single decoder-only LM that learns the speech-enhancement task distribution and infers which task to perform from the input context — eliminating the need for task-specific prompts, modules, or fine-tuning when switching between denoising, target-speaker extraction, and separation. Built as an autoregressive discrete-token predictor over a WavLM-extracted / BiCodec-quantised representation, it moves the speech-enhancement workflow from a zoo of specialist models into one generalist.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: PASE (Phonologically Anchored Speech Enhancer) is a generative speech-enhancement model from Cisco Collaboration AI that removes noise and reverberation while preserving linguistic content and speaker identity. It uses two fine-tuned WavLM-derived components:
Release Date: November 8, 2025
| Feature | Value |
|---|---|
| Type | Speech Enhancement |
| Bandwidth Extension | ❌ |
| Inpainting | ❌ |
| Sample Rate | 16 kHz mono |
| Architecture | Denoising WavLM (DRD from WavLM-Large) + Dual-Stream Vocoder (phonetic + acoustic) |
| Finetuned From | WavLM-Large |
| Training Data | DN5/DNS5 challenge clean + noise, LibriTTS, VCTK, OpenSLR26+28 RIRs |
| License | ![Apache 2.0][license-apache-2.0] |
Features: Anchors enhancement to phonology instead of spectrum: by reconstructing from a phonetic stream and a separate acoustic stream (per DeWavLM's two representations), PASE keeps the words intact even when the spectrum is severely degraded — substantially lowering hallucinations while still regaining perceptual quality.
Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]
· · · · · · · · · · · · · ·
Description: DTT-BSR (DTTNet with BandSequence and RoPE) is a music-source-restoration challenge submission from team AC/DC (Wuhan University) to ICASSP 2026. It is built inside the official MSR-Kit GAN framework, where the baseline generator is replaced by a DTTNet-style time-frequency U-Net and augmented at the bottleneck with:
Release Date: October 16, 2025
| Feature | Value |
|---|---|
| Type | Music Source Restoration |
| Bandwidth Extension | ❌ |
| Inpainting | ❌ |
| Architecture | DTTNet TFC-TDF U-Net (complex STFT) + Improved Dual-Path BandSplitRNN block + RoPE-Transformer |
| Input | complex STFT (real + imag channels; n_fft=2048, hop=512) |
| Discriminator | Multi-Frequency Discriminator (baseline) |
| Framework | MSR-Kit GAN (reconstruction + adversarial + feature-matching losses) |
| License | ![MIT][license-mit] |
Features: Treats music-source restoration as a complex-STFT time-frequency U-Net enhancement at the bottleneck: keep the strong DTTNet dual-path TFC-TDF structure for local spectral patterns, then layer in BandSplitRNN-style sub-band recurrence + RoPE self-attention so the generator can model long-range, cross-band harmonic structure that ordinary GAN baselines miss — critical for restoring non-vocal stems cleanly.
Links: ![GitHub][link-github]
· · · · · · · · · · · · · ·
Description: A2SB is NVIDIA's audio-to-audio Schrödinger Bridge diffusion model for high-resolution (44.1 kHz) music restoration. It is the first long-audio restoration model that can restore hour-long inputs without boundary artifacts, and it's end-to-end — predicting waveform outputs directly withou
Truncated — view the full README on GitHub.
37 commits