wildminder/awesome-ai-voice

List of open-source TTS, voice cloning, and music generation models

494

37 commits

updated Sep 10, 2026

See the code

README

Awesome TTS & Voice Generation Models

A curated list of open-source Text-to-Speech (TTS) and voice cloning models. Models are sorted by release date (newest first).

logo-tts2


Table of Contents


Text-to-Speech (TTS) Models

TTS Quick Comparison

ModelVoice CloningASRLanguagesStreamingLicense
AuK![MIT][license-mit]
AuK-Flash![MIT][license-mit]
rumik-oss 122 Indic languages + English![Other][license-other]
Irodori-TTS-v4.1-AnimeJapanese![MIT][license-mit]
ICE-012 Audio590![CC BY-NC 4.0][license-cc-by-nc-4.0]
TontaubeV17![Other][license-other]
Breeze TTS 22![Other][license-other]
Sopro v2 Turbo4![Apache 2.0][license-apache-2.0]
CuteTTS5![Apache 2.0][license-apache-2.0]
Rynsan TTSKhasi, Garo, Pnar, English, Hindi![CC BY 4.0][license-cc-by-4.0]
Audio8 TTS Preview 0.1B8![Other][license-other]
Kiseki-TTSJapanese![MIT][license-mit]
FireRedTTS324![Apache 2.0][license-apache-2.0]
Audio8-TTS-Preview-0.6bCantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish![Apache 2.0][license-apache-2.0]
NeuTTS-2EEnglish![Other][license-other]
Scylla's Banden_us, en_gb, es, it![Apache 2.0][license-apache-2.0]
sanoTTSEnglish, Nepali, Hindi, Vietnamese, Indonesian, Chinese![Other][license-other]
FreyaTTSTurkish![Apache 2.0][license-apache-2.0]
Inflect-Nano-v2English![Apache 2.0][license-apache-2.0]
GepardEnglish, Spanish, Portuguese, Dutch![Apache 2.0][license-apache-2.0]
Higgs Audio v3 TTS102![Research Only][license-research-only]
dots.ttsMultilingual![Apache 2.0][license-apache-2.0]
Confucius4-TTS14![Apache 2.0][license-apache-2.0]
WavTTSEnglish, Chinese![CC BY-NC 4.0][license-cc-by-nc-4.0]
MOSS-TTS31![Apache 2.0][license-apache-2.0]
VoxFlash-TTSChinese, English![Apache 2.0][license-apache-2.0]
Miso TTSEnglish![MIT][license-mit]
Raon-OpenTTS-1BEnglish![CC BY-NC 4.0][license-cc-by-nc-4.0]
OronTTSMongolian, Kazakh![MIT][license-mit]
Supertonic 331![OpenRAIL-M][license-openrail-m]
Scenema AudioEnglish, German, French, Spanish, Italian, Portuguese, Japanese, Chinese, Korean, Russian, Arabic, Hindi, Swahili![Other][license-other]
DramaboxEnglish![Other][license-other]
Sarashina2.2-TTSJapanese, English![Research Only][license-research-only]
LongCat-AudioDiTChinese, English![MIT][license-mit]
SILMA TTSArabic, English![Apache 2.0][license-apache-2.0]
Fish Audio S2 Pro80+![Research Only][license-research-only]
LongCat-NextChinese, English![MIT][license-mit]
Voxtral-4B-TTS9![CC BY-NC 4.0][license-cc-by-nc-4.0]
Blue (Light Blue) TTSHebrew, English, Spanish, Italian, German![MIT][license-mit]
KittenTTSEnglish, Multiple![Apache 2.0][license-apache-2.0]
Ming-omni-ttsChinese, English![Apache 2.0][license-apache-2.0]
SoulX-SingerMandarin, English, Cantonese![Apache 2.0][license-apache-2.0]
SoproTTSEnglish![Apache 2.0][license-apache-2.0]
Qwen3-TTS10![Apache 2.0][license-apache-2.0]
TADAEnglish![Other][license-other]
Irodori-TTS-500M-v2Japanese![MIT][license-mit]
KugelAudio23 European languages![MIT][license-mit]
LEMAS-TTS10![Apache 2.0][license-apache-2.0]
MioTTS-2.6BEnglish, Japanese![LFM][license-lfm]
MOSS-TTS-Nano20![Apache 2.0][license-apache-2.0]
NeuTTSEnglish, Spanish, German, French![Apache 2.0][license-apache-2.0]
OmniVoice600+![Apache 2.0][license-apache-2.0]
T5Gemma-TTSEnglish, Chinese, Japanese![MIT][license-mit]
TinyTTSEnglish![Apache 2.0][license-apache-2.0]
VoxCPM230![Apache 2.0][license-apache-2.0]
SopranoEnglish![Apache 2.0][license-apache-2.0]
GLM-TTSChinese, English![Apache 2.0][license-apache-2.0]
Echo-TTSEnglish![MIT][license-mit]
VibeVoice-RealtimeMultilingual![MIT][license-mit]
Fun-CosyVoice 3.09 + 18+ Chinese dialects![Apache 2.0][license-apache-2.0]
LFM2-Audio-1.5BEnglish![LFM][license-lfm]
Marvis-TTSEnglish, French, German![Apache 2.0][license-apache-2.0]
IndexTTS2Chinese, English![Apache 2.0][license-apache-2.0]
Maya1English![Apache 2.0][license-apache-2.0]
Step-Audio-EditXMandarin, English, Sichuanese, Cantonese, Japanese, Korean![Apache 2.0][license-apache-2.0]
KaniTTSEnglish, German, Chinese, Korean, Arabic, Spanish![LFM][license-lfm]
VibeVoice-Finetuning![MIT][license-mit]
VoxCPMChinese, English![Apache 2.0][license-apache-2.0]
FireRedTTS2EN, ZH, JP, KO, FR, DE, RU![Apache 2.0][license-apache-2.0]
Audio Flamingo 3 (AF3) / Audio Flamingo NextMulti-lingual![Apache 2.0][license-apache-2.0]
ZipVoiceChinese, English![Apache 2.0][license-apache-2.0]
Fish Speech8![Apache 2.0][license-apache-2.0]
Chatterbox23+![MIT][license-mit]
Orpheus-TTSMultilingual![Apache 2.0][license-apache-2.0]
MegaTTS3Chinese, English![Apache 2.0][license-apache-2.0]
Spark-TTSChinese, English![Apache 2.0][license-apache-2.0]
Step-AudioChinese, English, Japanese![Apache 2.0][license-apache-2.0]
Kokoro-82M8![Apache 2.0][license-apache-2.0]
KokoClone7![Apache 2.0][license-apache-2.0]
LuxTTS-![Apache 2.0][license-apache-2.0]
MiMo-AudioMulti-lingual![Apache 2.0][license-apache-2.0]
SoulX-PodcastMandarin, English, Cantonese, Sichuanese, Henanese![Apache 2.0][license-apache-2.0]
VieNeu-TTSVietnamese![Apache 2.0][license-apache-2.0]
DiaEnglish![Apache 2.0][license-apache-2.0]
MeloTTSEnglish, Spanish, French, Chinese, Japanese, Korean![MIT][license-mit]
Kimi-AudioMulti-lingual![MIT][license-mit]
![Apache 2.0][license-apache-2.0]
eSpeak-NG100+![Other][license-other]
AuK

AuK

Description: AuK is a 1.5B foundation model from Tencent for speech generation and editing, trained on millions of hours of diverse audio data. Through a single natural-language instruction interface it unifies an unusually broad task set: zero-shot TTS (speak text in the reference voice) and instruct TTS (voice from a description alone, no reference), content editing (rewrite what is said; even lyric editing that preserves melody and voice), acoustic editing (pitch by semitones, speed, volume), paralinguistic editing (emotion, timbre, de-accent, nonverbal sounds, whisper conversion), and enhancement & separation (denoise/dereverberate, speech separation, music/vocal separation, target-speaker extraction). Architecture: a diffusion transformer with layer-fusion weights, conditioned by a Qwen2.5-Omni-3B MLLM encoder and a separate VAE (loaded at runtime). Day-0 SGLang-Omni serving support, Gradio and ComfyUI integrations, and a task Cookbook are provided. Released under MIT.

Release Date: September 9, 2026

FeatureValue
Voice Cloning
Asr
License![MIT][license-mit]
Parameters1.5B
Architecturediffusion transformer + layer fusion, Qwen2.5-Omni-3B MLLM encoder, separate VAE
VariantsAuK (this, base) + AuK-Flash (distilled, 4-step inference)
Editingcontent, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion
Enhancement Separationspeech enhancement, speech separation, music separation, target speaker extraction
DeploymentSGLang-Omni (day-0), Gradio, ComfyUI

Features: Unifies generation and the full editing/enhancement/separation spectrum in one instruction-following model — most systems pick one lane (TTS, or editing, or separation); AuK does zero-shot + instruct TTS, lyric rewriting with melody preservation, emotion/timbre/de-accent/whisper paralinguistic edits, and source separation through the same natural-language interface. A diffusion transformer with layer fusion, distilled into a 4-step AuK-Flash variant for fast inference.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]

Additional Tools:

ToolTypeLink
ComfyUI-AuKComfyUI nodeComfyUI-AuK

· · · · · · · · · · · · · ·

AuK-Flash

AuK-Flash

Description: AuK-Flash is the distilled variant of AuK, Tencent's 1.5B foundation model for speech generation and editing, optimized for fast 4-step inference. It exposes the same natural-language instruction interface as the base model: zero-shot TTS (reference voice) and instruct TTS (voice description, no reference), content and lyric editing, pitch/speed/volume acoustic edits, emotion/timbre/de-accent/nonverbal/whisper paralinguistic edits, plus speech enhancement and speech/music/target-speaker separation. Architecture matches the base: diffusion transformer with layer-fusion weights, Qwen2.5-Omni-3B MLLM encoder, and a separate runtime-loaded VAE. Released under MIT.

Release Date: September 9, 2026

FeatureValue
Voice Cloning
Asr
License![MIT][license-mit]
Parameters1.5B
Architecturediffusion transformer + layer fusion (distilled to 4 inference steps), Qwen2.5-Omni-3B MLLM encoder, separate VAE
Base Modeltencent/AuK
Editingcontent, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion
Enhancement Separationspeech enhancement, speech separation, music separation, target speaker extraction
DeploymentSGLang-Omni, Gradio, ComfyUI

Features: Distills the AuK foundation model's diffusion transformer down to 4 inference steps, making the full generate-and-edit capability set (including enhancement and separation) practical for interactive use — traded against the base model's maximum quality, with the two variants loadable side-by-side from the same codebase.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]

· · · · · · · · · · · · · ·

rumik-oss 1

rumik-oss 1

Description: rumik-oss 1 is a 3B multilingual text-to-speech model from rumik ai, trained on fewer than 70,000 hours of speech while performing competitively with existing TTS models. It covers 22 Indic languages in their native scripts and romanized forms plus English, supporting both single-language and code-switched synthesis. Delivery is conditioned via <description="..."> tags (tone, accent, pace), with inline vocalization control (<laugh>, <chuckle>, <sigh>). It extends CohereLabs/tiny-aya-fire with discrete speech tokens from the mimi codec: following the flattened codec-token formulation used in llama-mimi, text conditioning and audio generation share a single autoregressive sequence, predicting eight codebook tokens per frame before advancing, with the frozen mimi decoder reconstructing the 24 kHz waveform. The model ships with 4 fixed voices (Ira, Aisha, Siya, Zoya) that perform equally well across all 22 languages; there is no zero-shot voice cloning. Licensed under Cohere's CC-BY-NC-4.0 with acceptable-use addendum (research and non-commercial use only).

Release Date: September 6, 2026

FeatureValue
Voice Cloning
Asr
Languages22 Indic languages + English (native scripts and romanized; code-switching supported)
License![Other][license-other]
Parameters3B (3,381,533,697 BF16)
ArchitectureCohereLabs/tiny-aya-fire backbone + flattened mimi codec tokens (8 codebooks/frame, llama-mimi formulation)
Audio Codeckyutai/mimi (frozen decoder), 24 kHz output
Pronunciation
Highlightsdescription-conditioned delivery (<description> tags), inline vocalizations (<laugh>/<chuckle>/<sigh>)
Variantsrumik-oss-1 (post-trained), rumik-oss-1-base (speaker-conditioned pre-post-training)

Features: Brings competitive multilingual TTS to 22 Indic languages with under 70k training hours, using a flattened mimi codec-token formulation (single autoregressive sequence for text conditioning + audio) on the tiny-aya-fire backbone. Code-switched synthesis, description-conditioned delivery, and inline vocalization tags are first-class capabilities, and its 4 voices perform equally well across all 22 languages — unusual, as most TTS voices are language-specific.

Links: ![HuggingFace][link-huggingface] ![Blog][link-blog] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Irodori-TTS-v4.1-Anime

Irodori-TTS-v4.1-Anime

Description: Irodori-TTS-v4.1-Anime is a Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small using anime-style speech data. Because the base model's annotation pipeline is not publicly documented, the fine-tuning data was annotated independently — so caption conditioning and emoji controls may behave differently from the base model. The full-precision checkpoint (0.8B params, F32) ships at the repository root, with quantized variants (int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic) in subdirectories. It follows the base model's MIT License and ethical restrictions; inference uses the original Irodori-TTS repository.

Release Date: September 4, 2026

FeatureValue
Voice Cloning
Asr
LanguagesJapanese
License![MIT][license-mit]
Parameters~0.8B (766,052,385 F32)
ArchitectureIrodori-TTS (Aratako) fine-tune; caption-conditioned with emoji controls
Base ModelAratako/Irodori-TTS-v4.1-Small
Variantsint8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic

Features: A community fine-tune that ports the Irodori-TTS line into the anime-voice domain using an independently built annotation pipeline (since the base model's is undocumented), and ships the result with five ready-made quantization variants (int4/int8/fp8) for efficient inference.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

ICE-012 Audio

ICE-012 Audio

Description: ICE-012 Audio is a multilingual text-to-speech model from DarkPs (a FanuonAI organization) with streaming output and reference-based voice cloning. Its defining trait is breadth of language coverage — 590 language names/variants are accepted (name or 2–3-letter ID, with a language-agnostic fallback), including 13 Arabic dialects ("Lahgtna" variants) alongside the full ISO list. It introduces an active acoustic adapter — conditioning codec embeddings before the backbone and refining hidden states after it. Voice is controllable along six axes: gender (male/female), age (child → elderly), pitch (5 levels), accent (10 English accents), style (e.g. whisper), and speed (0.5–2.0×), plus an --auto-voice mode where the model picks a voice automatically. The checkpoint is ~714M parameters (F16) and runs via transformers with trust_remote_code=True. Released under CC BY-NC 4.0.

Release Date: August 29, 2026

FeatureValue
Voice Cloning
Asr
Languages590 names/variants (incl. 13 Arabic Lahgtna dialects; language-agnostic fallback)
Streaming
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
Parameters~714M (714,409,993 F16)
Architecturecausal LM with active acoustic adapter (conditions codec embeddings pre-backbone, refines hidden states post-backbone)
Voice Controlsgender, age, pitch, accent, style, speed; auto-voice mode

Features: The active acoustic adapter wraps the backbone on both sides — conditioning codec embeddings before it and refining hidden states after — while a six-axis voice-control space (gender/age/pitch/accent/style/speed plus auto-voice) and 590-language coverage make it one of the broadest single-checkpoint TTS releases for dialect and minority-language synthesis.

Links: ![HuggingFace][link-huggingface] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

TontaubeV1

TontaubeV1

Description: TontaubeV1 is a multilingual text-to-speech model from TontaubeAI (craitech) designed for expressive voice cloning, long-form generation, and low-latency streaming. Its release contains four causal codebook predictors: CB0 generates semantic audio and duration from text, while progressively smaller CB1–CB3 add acoustic detail. CB0 uses a Qwen3-1.7B-derived transformer trunk and CB1–CB3 progressively shallower Qwen3-0.6B-derived trunks, each with a two-layer audio-token head. The four output streams are decoded with DualCodec, and the inference path uses VibeVoice's acoustic encoder/decoder for continuous reconstruction and streaming. It ships bundled synthetic voices plus zero-shot cloning from up to 60 s of reference audio, with public speaking styles audiobook, conversational, and agentic. Released under the Tontaube Community Model License 1.0, which is explicitly not open-source.

Release Date: August 26, 2026

FeatureValue
Voice Cloning
Asr
Languages7 (English, German primary; Spanish, French, Italian, Dutch, Portuguese secondary)
Streaming
License![Other][license-other]
Parameters~2.87B (2,873,962,498; CB0 1.83B + CB1 449M + CB2 327M + CB3 269M)
Architecture4-stage Qwen3-derived codebook cascade (CB0–CB3) + DualCodec + VibeVoice decode
Stylesaudiobook, conversational, agentic

Features: The four-stage codebook cascade (CB0 semantic+duration → CB1–CB3 progressive acoustic refinement) lets a single multilingual model deliver expressive, long-form, low-latency speech with strong zero-shot cloning. On the 1,088 English zero-shot Seed-TTS examples it posts 1.66% mean utterance-level WER (measured with Whisper large-v3 at semantic temperature 0.6), and the RTX-5090 streaming path reaches ~200 ms to first encoded audio.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Breeze TTS 2

Breeze TTS 2

Description: Breeze TTS 2 is an open-weight text-to-speech model from BreezeBlue / RESONIA built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following supports reference-free voice design (create a voice from a text description) and reference-guided voice direction (clone a voice while steering tone, emotion, pace, delivery), alongside standard reference-audio voice cloning. Ultra-low-latency streaming reaches 0.32 RTF (≈3.1× real time with the warmed-up fast path) and under 40 ms time-to-first-audio on an NVIDIA H100, emitting 24 kHz PCM. Source code is Apache-2.0; model weights are governed by the BreezeBlue Research and Non-Commercial License (commercial use needs written authorization from RESONIA).

Release Date: August 25, 2026

FeatureValue
Voice Cloning
Asr
Languages2 (English, Chinese)
Streaming
License![Other][license-other]
Parameters3B (3,466,363,713)
Architectureseq2seq backbone + depth decoder + codec with CUDA-graph fast path (no named backbone disclosed)
Highlights#1 open-weight on Artificial Analysis TTS leaderboard; vocal events inline (laugh/cough)

Features: Pairs natural-language voice control with real-time streaming: a single model handles reference-free voice design (no reference audio needed) and voice direction (clone + steer prosody), and ships a CUDA-graph fast path that hits sub-40 ms TTFA at ~3.1× real time on H100 — open-weight quality that the authors claim exceeds frontier proprietary TTS.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]

Additional Tools:

ToolTypeLink
ComfyUI-Breeze-TTS-2ComfyUI nodeComfyUI-Breeze-TTS-2

· · · · · · · · · · · · · ·

Sopro v2 Turbo

Sopro v2 Turbo

Description: Sopro (Portuguese for "breath") is a lightweight voice-cloning text-to-speech family. This repo ships sopro-v2-turbo, a 120M-parameter open model that streams and runs comfortably on a laptop CPU or in the browser (ONNX runtime), reaching SOTA-level intelligibility against much larger systems. It supports zero-shot voice cloning from 5–20 s of reference audio, four languages (English, European Portuguese, French, German), and a streaming path with ~300 ms time-to-first-audio on a laptop CPU (0.24 RTF offline / 0.21 RTF streaming on an M3 CPU, 0.07 RTF on H100). Released under Apache-2.0.

Release Date: August 25, 2026

FeatureValue
Voice Cloning
Asr
Languages4 (English, European Portuguese, French, German)
Streaming
License![Apache 2.0][license-apache-2.0]
Parameters120M (121,574,193)
Deploymentin-browser ONNX runtime; int8 AR weights on CPU; causal vocoder
Architectureautoregressive TTS + chunked-attention streaming path + causal vocoder (F5-TTS/CosyVoice/Vocos lineage acknowledged)

Features: Packs SOTA-level intelligibility into a 120M footprint that runs in the browser or on a laptop CPU, with a chunked-attention + causal-vocoder streaming path (~300 ms TTFA) — making zero-shot multilingual voice cloning practical for on-device and edge deployment rather than GPU-only serving.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]

· · · · · · · · · · · · · ·

CuteTTS

CuteTTS

Description: CuteTTS is a lightweight (~230M-parameter) continuous autoregressive TTS model from OPPO that models continuous latents rather than discrete codec tokens, running efficiently on GPUs, CPUs, and Apple silicon. It delivers ultra-low latency — ~40 ms to the first audio chunk and ~9× real-time throughput on an RTX 4090 — with strong speech quality and zero-shot voice cloning (best-in-comparison 78.9 SIM on LibriSpeech test-clean). Multilingual support covers English, Chinese, French, German, and Spanish. A distilled variant (CuteTTS-distill) trades slight quality for further efficiency. Ships with a web demo, Python API, and CLI.

Release Date: August 24, 2026

FeatureValue
Voice Cloning
Asr
Languages5 (English, Chinese, French, German, Spanish)
Streaming
License![Apache 2.0][license-apache-2.0]
Parameters~230M
Architecturecontinuous autoregressive modeling of latents + speaker encoder + audio VAE (discrete-codec-free design)
VariantsCuteTTS, CuteTTS-distill

Features: Autoregressively models continuous latent audio representations instead of discrete codec tokens, eliminating codebook-related artifacts and quantization loss at only ~230M parameters. Combined with a lightweight speaker encoder and audio VAE, this yields best-of-class speaker similarity among compared open models and ~40 ms first-chunk latency while remaining practical for CPU/Apple-silicon inference.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Rynsan TTS

Rynsan TTS

Description: Rynsan TTS is a multilingual text-to-speech model that extends k2-fsa/OmniVoice to support Khasi, Garo, and Pnar — languages of Meghalaya, India that have historically had limited representation in modern speech technology. Rather than building a system from scratch, Rynsan retains the multilingual capabilities of the base model while adding speech data for these low-resource Khasic languages. Developed under the Tynrai AI initiative, its broader goal is accessible speech technology for the diverse languages and dialects of Meghalaya, supporting their preservation and use in voice-based applications. A live demo is available at ri.tynrai.in/demo. The repository is gated (manual access approval).

Release Date: August 22, 2026

FeatureValue
Asr
Languages5+ (English, Hindi + extension languages Khasi kha, Garo grt, Pnar pbv; base OmniVoice supports more)
License![CC BY 4.0][license-cc-by-4.0]
Parameters~0.61B
ArchitectureOmniVoice (k2-fsa) multilingual TTS, extended fine-tune
Base Modelk2-fsa/OmniVoice
DeveloperToiar / Tynrai AI

Features: Extends a modern multilingual TTS foundation to three substantially under-resourced Khasic languages — a rare production-oriented entry for indigenous-language speech tech, aimed at accessibility, education, and language preservation rather than benchmark leadership.

Links: ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Audio8 TTS Preview 0.1B

Audio8 TTS Preview 0.1B

Description: Audio8 TTS Preview 0.1B is the smallest release in the Audio8 TTS family ("the smallest zero-shot TTS worth running"): a ~170M-parameter generative model plus a separate ~120M-parameter codec decoder, making the complete audio generation stack much smaller than most modern multilingual TTS systems. It supports speech generation and zero-shot voice cloning (reference audio + matching transcript). Primary languages are Chinese and English, with German, Spanish, French, Italian, Japanese, and Korean as experimental/multilingual-evaluation targets. Released under the custom Audio8 Community License v1.0: non-commercial use is free, and commercial use is free only for entities with annual revenue under US$2M.

Release Date: August 19, 2026

FeatureValue
Voice Cloning
Asr
Languages8 (Chinese + English primary; de/es/fr/it/ja/ko experimental)
License![Other][license-other]
Parameters~0.17B main model (+ ~120M codec decoder)
ArchitectureAudio8 Falcon H1 DualAR — slow AR (semantic tokens) + fast AR (codec codebooks), 10 codebooks × 4096 entries
Audio Codecbundled 44.1 kHz neural codec (~21.5 frames/s)
Contextup to 2,048 packed text/audio positions
Variants0.1B (this), 0.6B

Features: Packs practical zero-shot cloning into a ~170M-parameter model using an Falcon-H1-derived DualAR design (slow AR predicts semantic tokens per frame; fast AR predicts the frame's 10 codebooks conditioned on the slow hidden state). On Seed-TTS it posts EN WER 1.662 at only ~0.17B — within reach of 4B+ systems — and ships with its own 44.1 kHz codec so no external codec checkpoint is needed.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

Kiseki-TTS

Kiseki-TTS

Description: Kiseki-TTS is a small, fast Japanese text-to-speech model from telecomadm1145 built on top of Qwen/Qwen3-TTS-Tokenizer-12Hz. It generates discrete neural audio codec tokens at 12.5 Hz (4–6× fewer autoregressive steps than 50–75 Hz codecs) and decodes them to waveform with the Qwen3 TTS codec. The acoustic decoder is a linear-time Mamba2 SSM rather than a self-attention stack, so generation cost is constant per frame — memory does not grow with utterance length and there is no KV cache to manage. Because TTS and ASR were trained jointly in a single multi-task run, the same checkpoint also performs ASR (Japanese speech → text, reading only codec layer 0). It is a single-domain voice (ASMR-style Japanese training data) with no speaker conditioning or voice cloning.

Release Date: August 15, 2026

FeatureValue
Voice Cloning
Asr
LanguagesJapanese only (ja)
License![MIT][license-mit]
Parameters~0.41B (0.33B backbone + 78M audio branch)
ArchitectureTransformer encoder (12 layers, bidirectional self-attention) + cross-attention → Mamba2 SSM decoder (6 layers, no causal self-attention)
Audio CodecQwen3-TTS-Tokenizer-12Hz (12.5 Hz, 16 quantizer layers)
Base ModelKiseki-1.1-0.3B (seq2seq translation model)
Training Datatelecomadm1145/asmr_archive_qwentts_encoded

Features: The decoder deliberately omits causal self-attention — temporal context is carried entirely by the Mamba2 recurrent state while text conditioning enters through cross-attention whose K/V are computed once during prefill. This yields O(1) state per frame (a fixed SSM tensor plus a 3-frame conv window) instead of an O(T) KV cache, so long-form synthesis degrades gracefully past the ~41 s training ceiling instead of hitting a memory cliff. Combined with the 12.5 Hz codec and a shared multi-token-prediction head that resolves all 16 codebook layers in one trunk pass, the model is both compute- and memory-bandwidth-bound rather than quadratic in length.

Links: ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

FireRedTTS3

FireRedTTS3

Description: FireRedTTS3 is a unified speech generation and editing system from the FireRed Team built on semantically enriched continuous speech representations. It ships in two variants: FireRedTTS3-Base (zero-shot voice cloning across 24 languages and 21 Chinese dialects) and FireRedTTS3-Instruct (natural-language voice design and combined semantic + acoustic speech editing in one model). Beyond cloning, it supports instruction-based voice design (no reference audio needed) and editing operations such as insertion / deletion / substitution (semantic) and speed / pitch / volume changes (acoustic).

Release Date: August 5, 2026

FeatureValue
Voice Cloning
Asr
Languages24 (plus 21 Chinese dialects)
License![Apache 2.0][license-apache-2.0]
ArchitectureQwen3 backbone + patch-level diffusion autoregressive (DiTAR) + RedAE codec + CAM++ speaker encoder
VariantsBase (cloning), Instruct (cloning + voice design + editing)

Features: Represents speech with semantically enriched continuous (non-quantized) representations, enabling a single system to do zero-shot cloning, text-driven voice design, and fine-grained semantic + acoustic editing. On Seed-TTS-eval it reaches an average WER/CER of 3.04% with 78.8% speaker similarity; MiniMax-MLS-Test average SIM 84.8%.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

Additional Tools:

ToolTypeLink
FireRedTTS3-ComfyUIComfyUI nodeFireRedTTS3-ComfyUI

· · · · · · · · · · · · · ·

Audio8-TTS-Preview-0.6b

Audio8-TTS-Preview-0.6b

Description: Audio8 TTS Preview 0.6B is a 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning. It uses a DualAR architecture inspired by Fish Audio S2 Pro: a slow AR transformer predicts one semantic token per audio frame, and a fast AR transformer predicts the frame's codec codebooks conditioned on the slow hidden state and preceding codebooks. The bundled 44.1 kHz neural audio codec handles both reference-audio encoding and waveform decoding — no additional codec checkpoint is required. The model supports 11 recommended languages (Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish) with zero-shot voice cloning from a reference audio clip + matching transcript.

Release Date: July 28, 2026

FeatureValue
Parameters601,159,424 (0.6B, excluding the codec)
Voice Cloning
Asr
LanguagesCantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish (11)
Streaming
License![Apache 2.0][license-apache-2.0]
ArchitectureDualAR (slow AR + fast AR), inspired by Fish Audio S2 Pro
Slow Ar24 layers, width 896, 14 attention heads, 2 KV heads
Fast Ar4 layers, width 896, 14 attention heads, 2 KV heads
Acoustic Tokens10 codebooks, 4,096 entries per codebook
Codec44.1 kHz, 2,048 samples per model frame (~21.5 frames/s), bundled (no external codec needed)
Context Lengthup to 2,048 packed text/audio positions
Sample Rate44,100 Hz
Inferencetransformers with trust_remote_code=True; CUDA-capable GPU recommended
Dependenciestorch>=2.5.0, torchaudio>=2.5.0, transformers>=4.57.0,<5, soundfile, safetensors
Preview Statuslanguage coverage intentionally limited; broader multilingual + Chinese dialect support planned
Library Nametransformers (custom_code)
Pipeline Tagtext-to-speech
Createdat2026-07-28T07:53:00Z

Features: The DualAR architecture is the technical centerpiece: rather than a single autoregressive decoder predicting all codebook levels sequentially (the standard codec-LLM TTS pattern), Audio8 splits the work into a slow AR that predicts one semantic token per audio frame and a fast AR that predicts the frame's remaining codec codebooks conditioned on the slow hidden state. This separation lets the semantic-level reasoning happen at the slow AR's 24-layer depth while the acoustic codebook prediction stays lightweight at 4 layers — reducing the total compute per frame without sacrificing semantic quality. The bundled 44.1 kHz codec (no external codec checkpoint needed) and the 10-codebook / 4,096-entry acoustic token design give the model self-contained high-fidelity output at a compact 0.6B scale, making it one of the smallest multilingual zero-shot-cloning TTS systems shipping at 44.1 kHz.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

NeuTTS-2E

NeuTTS-2E

Description: NeuTTS-2E is a super-fast, highly realistic, on-device emotional text-to-speech model from Neuphonic. It is the next generation after NeuTTS Air / Nano (which continue to ship for multilingual + zero-shot-cloning contexts) — narrowed in scope to an English-only alpha focused on:

Release Date: July 21, 2026

FeatureValue
Parameters0.2B (compact LM backbone + codec)
Voice Cloning
Asr
LanguagesEnglish (English-only alpha)
Streaming
License![Other][license-other]
Backbonecompact LM backbone tuned for emotional TTS token generation
Codecefficient codec (compact, paired with the LM)
Speakers4 fixed (emily, paul, sophie, steven)
Emotions6 + neutral (angry, disgusted, fearful, happy, sad, surprised, neutral)
Emotion Control Modesingle-argument selection (no composable multi-axis axes like Scylla's Band)
Input Formattext only — no phonemizer, no system dependencies
On Deviceyes (laptop-class CPU real-time / better-than-real-time)
Distribution Formatssafetensors (torch), Q4 GGUF, Q8 GGUF
Formats In Collectionneuphonic/neutts-2e (safetensors), neuphonic/neutts-2e-q4-gguf (smallest footprint), neuphonic/neutts-2e-q8-gguf (mid-tier compression)
Gguf Featuresimatrix, conversational, endpoints_compatible
Pipeline Tagtext-to-speech
Library Name(HF tag does not declare transformers / safetensors stem beyond safetensors itself)
Downloads194 / 241 / 216 (torch / q4 / q8)
Intended Useembedded voice agents, on-device assistants, toys, privacy-sensitive applications
Comparison With Air NanoAir/Nano continue to ship for zero-shot cloning + multilingual contexts; 2E is the next-gen focused English emotional variant
Safety Notemodel is alpha; legitimate project landing is neuphonic.com (not neutts.com)

Features: The technical center of NeuTTS-2E is maximum speed per parameter at on-device budgets — the 0.2B LM + codec pair delivers real-time-or-better on laptop-class CPUs while exposing discrete categorical emotion control (angry / disgusted / fearful / happy / sad / surprised / neutral) plus a fixed four-speaker cast for consistency in agent / toy / accessibility voice personas. Two design choices distinguish it from the surrounding TTS field:

First, the categorical emotion surface is single-axis and discrete (one emotion per call), not the continuous multi-axis composable vector surface used by models like Scylla's Band ([neurotica base + 6-axis continuous strengths]). The project's positioning — production-grade agents + toys + accessibility — benefits from a one-argument API where emotion="happy" is the explicit operational state. The release locks emotional mode at generation time, which simplifies downstream filtering / guardrails.

Second, the distribution-shape design (one model, three deployment formats) is a deliberate on-device-first posture: the safetensors torch build for max-quality GPU/server; Q8 GGUF for mid-tier compression; Q4 GGUF for the small-footprint embedded target. All three are direct llama.cpp-compatible drops of the same model — no retraining-per-format — letting users pick size vs quality at deployment time without changing the production API. The combined CPU-first + GGUF-first design pattern is the opposite of the cloud-first TTS systems in this list — and is what makes 2E suitable for embedded voice agents, toys, and privacy-sensitive applications where audio + text must remain on-device.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Collection][link-collection]

· · · · · · · · · · · · · ·

Scylla's Band

Scylla's Band

Description: Scylla's Band is a multilingual, multi-voice, expressive TTS model from Spybyscript, designed specifically for local and self-hosted inference through ONNX Runtime (with an experimental LiteRT backend for explicit native / mobile use). The architecture is a continuous-latent TTS family:

Release Date: July 19, 2026

FeatureValue
Parametersnot stated (architecture: 4-layer duration predictor (192 hidden) + 12-layer rectified-flow acoustic generator (512 hidden, AdaLN, QK norm))
Voice Cloning
Asr
Languagesen_us, en_gb, es, it (4 public text-input languages)
Streaming
License![Apache 2.0][license-apache-2.0]
Sample Rate24,000 Hz
Managed Voices10 (ariadne, felix, gwen, ink, max, orpheus, rex, scylla, stone, tuesday)
Voice Default Localeen_us for most; ink / orpheus / tuesday default to en_gb
Voice Style Dim128 (style features) + 32 (prosody features)
Affect Axes6 (calm, joy, anger, sadness, sarcasm, questioning — all continuous in [0, 1])
Affect Overlay Axessarcasm, questioning (mixable with any core delivery)
Affect Cfg Scopeduration + acoustic-flow prediction (preserves voice / reference)
Encoders DefaultONNX Runtime (Python CLI / Python API / Android sample / libscyllasband native)
Encoders ExperimentalLiteRT (experimental / explicit-selection)
Cli Quality Default8-step Heun sampling
Graph Budgets512 G2P text tokens / 512 phone frames / 640 latent frames
Latent Target Buckets256 / 384 / 512 / 640 (smallest-fit selection)
VocoderScylla's Band acoustic adapter + frozen charactr/vocos-mel-24khz
Hop Lengths256 (waveform) / 512 (latent)
Text Frontendphrase-level multilingual G2P (74-phone vocabulary)
Span Context3 segments over up to 768 phones with 512-dim context state
Prefix Contextup to 24 acoustic latent frames from preceding chunk
Long Form Featuresboundary metadata + punctuation pause floors + prefix-latent carryover + span context
Group Speak Input[voice], [voice:language], [voice:language:axis=value,...] annotations
Bundle Contract1.0.0 / scyllasband-duration-flow
Intended Usesingle-voice speech synthesis (10 voices); en/es/it; long-form narration; multi-voice dialogue from tagged text; continuous affect control; ONNX desktop/server; ONNX + LiteRT native/mobile
Not Intendedarbitrary-speaker cloning / impersonation / fraud / deceptive speech
Distributionstraining data, trainer checkpoints, and export tooling not distributed
Cli Commandsdownload, validate-bundle, list-voices, normalize-text, speak, group-speak, stream, plan
Library Nameonnxruntime (tags include onnx, tflite, litert, duration-flow)

Features: Three design decisions distinguish Scylla's Band in the multilingual TTS class. First, decoupling duration and acoustic flow as separate rectified-flow stages — duration is a 192-hidden, 4-layer predictor operating on a 512-phone window, acoustic latents a 512-hidden, 12-layer AdaLN / QK-norm generator at 24-dim. This split lets affect-CFG act on both stages independently while retaining voice / reference conditioning, supporting the 6-axis continuous composability. Second, 6 affect axes (with sarcasm and questioning as overlays mixed with any core delivery) instead of mutually-exclusive discrete emotion classes — calm=0.5, joy=0.5 is a valid input, and axes stay in [0, 1] so multi-axis states are expressible without combinatorial blow-up. Third, the ONNX-first runtime design with libscyllasband native + an experimental LiteRT backend sits at a budget most neural TTS systems don't target — the 8-step Heun default and 512/512/640 fixed graph budget keep the model usable on CPU and mobile, and the inference-only release surface (training data + checkpoints not distributed) is the complement of the latency / mobile inference focus.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

sanoTTS

sanoTTS

Description: sanoTTS is the smallest known neural text-to-speech family. The name sano (सानो) is Nepali for "small". Each voice weighs 294k to 2.3M parameters — smaller than the smallest voice in prior families (TinyTTS at 1.62M; Inflect Nano at 4.63M; Kokoro at 82M) and the family fits in under 4 MB per voice with zero runtime dependencies (the espeak-ng phonemizer is bundled). Voices run real-time on a ~$3 ESP32-S3 microcontroller (output through a GPIO into an LM386 and a speaker) and live in the browser via WebAssembly — no server, no upload, no NPU. The full neural stack is duration → acoustic → decoder, quantized to int8, with the espeak-ng phonemizer included. 11 voices across 6 languages ship: English, Nepali, Hindi, Vietnamese, Indonesian, and Chinese — including the 294k heart-nano voice (337 KB) and the mel-based heart / heart-nano pair that predicts a 100-band spectrogram rendered by a noise-fed ConvNeXt + iSTFT decoder at 24 kHz. The inference runtime was relicensed MIT (September 2026); the project as a whole remains GPLv3 via espeak-ng. The project page at ampixa.github.io/sanoTTS hosts a live browser synthesis demo for every voice.

Release Date: July 13, 2026

FeatureValue
Parameters294k–2.3M per voice (smallest = the 294k "heart-nano" voice, 337 KB)
Voice Cloning
Asr
LanguagesEnglish, Nepali, Hindi, Vietnamese, Indonesian, Chinese (6 languages, 11 voices)
Streaming
License![Other][license-other]
Architecturefull neural stack — duration model → acoustic model → decoder
Quantizationint8 (W8/A12, corr 0.9995+; piperlite portable C99)
Runtime MicrocontrollerESP32-S3 (real-time RTF 0.41, GPIO → LM386 → speaker)
Runtime BrowserWebAssembly (no server, no upload, no NPU)
Runtime Footprintunder 4 MB per voice, zero dependencies
Voices11 (English: amy / kristin / hfc / amy-1p1m / amy-1p8m / robot / heart / heart-nano; one voice each for NE / VI / ID / ZH + shared lang voices)
Phonemizerespeak-ng (bundled)
License Splitinference runtime MIT; project overall GPL-3.0 (copyleft from espeak-ng)
Library Namesanotts
Training Methoddistillation (per voice)

Features: The hard constraint — be the smallest neural TTS family known, real-time on a $3 microcontroller — drives the entire stack. Conventional sub-100M TTS systems are too large for an ESP32's flash and RAM. sanoTTS keeps the full duration → acoustic → decoder neural pipeline (no espeak-NG-only fallback, no concatenative hybrid), quantizes everything to int8, and bundles the phonemizer so the whole voice ships in under 4 MB with zero runtime dependencies. The newest heart / heart-nano voices switch to a mel-based recipe (100-band spectrogram + noise-fed ConvNeXt + iSTFT decoder at 24 kHz), bringing the smallest voice down to 294k parameters / 337 KB — a per-voice footprint 100× smaller than Kokoro and 2× smaller than TinyTTS while still leading SCOREQ / UTMOS in the sub-15M class — and the demo synthesizes every voice live in the browser via WASM, so the smallest-known neural TTS is also the only one that runs unattended on a $3 chip and a $0 web page.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

FreyaTTS

FreyaTTS

Description: FreyaTTS is a 183M-parameter Turkish text-to-speech model. It is tokenizer-free at the character level — 92 symbols in its Turkish vocabulary — so there is no phonemizer or G2P step in either training or inference. Speech is generated with a non-autoregressive conditional flow-matching DiT in the frozen AudioVAE2 latent space (25 Hz, 64-dim latents, 16 kHz encode / 48 kHz decode). Training runs from scratch on Turkish speech: a pretraining stage followed by SFT stage 1/2 for voice lock and short-utterance coverage. Output is 48 kHz mono. On the project's Freya-TR-Eval benchmark the model reports WER 8.0% / CER 3.0%, ranking 3rd of 7 among open sub-1B Turkish TTS systems — a deliberate single-target-speaker, no-cloning design choice for a focused foundation release. The evaluation dataset is freya-tr-eval.

Release Date: July 7, 2026

FeatureValue
Parameters183.2M
Voice Cloning
Asr
LanguagesTurkish (tr)
Streaming
License![Apache 2.0][license-apache-2.0]
Architectureconditional flow-matching diffusion transformer (DiT), non-autoregressive, 32-step Euler ODE, no CFG
Tokenizercharacter-level (92 Turkish symbols; no phonemizer, no G2P)
Latent Spacefrozen AudioVAE2 (Apache-2.0, openbmb/VoxCPM2), 64-dim at 25 Hz
Codec Io16 kHz encode / 48 kHz decode
Sample Rate48,000 Hz
Trainingfrom scratch on Turkish speech; pretraining + SFT stage 1/2 (voice lock + short-utterance coverage)
EvaluationFreya-TR-Eval — WER 8.0% / CER 3.0%, 3rd of 7 open sub-1B Turkish TTS
Library Namefreyatts

Features: Two design choices are worth flagging. First, tokenizer-free character-level Turkish: by training directly on the 92-symbol Turkish alphabet with no phonemizer or G2P grapheme-to-phoneme step, the model removes a dependency that is fragile for agglutinative Turkish morphology and that often degrades quality when ported to low-resource Turkic relatives. Second, non-autoregressive conditional flow-matching in a frozen AudioVAE2 latent space: the 25 Hz / 64-dim bottleneck keeps the DiT small (183M) while inheriting a separately-trained audio codec's representation, letting a focused single-language-non-multilingual release ship at a fraction of the parameter budget of multilingual foundation TTS systems. The deliberate "no cloning, single target speaker" choice is a scope-lowering move that lets the foundation release put all its capacity into Turkish speech quality rather than spread it across zero-shot speaker adaptation.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Inflect-Nano-v2

Inflect-Nano-v2

Description: Inflect-Nano-v2 is a complete local text-to-waveform speech synthesis model with 3,966,721 deployable parameters — under 4M total. It is a VITS-architecture fixed-voice English TTS designed for CPU or CUDA inference with deterministic seeds, long-text handling, and 24 kHz mono output. The full FP32 checkpoint is 15.97 MB, making it one of the smallest complete neural TTS systems that produces natural-sounding speech without a separate vocoder or phonemizer dependency. The model ships with a public adaptation toolkit for preparing data, auditing train/validation splits, adapting a fixed voice or language, resuming training, evaluating checkpoints, and exporting PyTorch or ONNX packages. A sibling Inflect-Micro-v2 (9.36M parameters) prioritizes quality below 10M; Nano prioritizes footprint below 4M. Both share one public API.

Release Date: June 25, 2026

FeatureValue
Parameters3,966,721 (3.97M deployable)
Voice Cloning
Asr
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]
ArchitectureVITS (end-to-end text-to-waveform)
Sample Rate24,000 Hz
Footprint15.97 MB FP32
InferenceCPU or CUDA; PyTorch + ONNX export
Determinismdeterministic seeds for reproducible generation
Long Textautomatic text splitting and handling
Input Formattext (no phonemizer or system dependencies)
Adaptation Toolkitdata prep, split auditing, voice/language adaptation, training resume, checkpoint eval, PyTorch/ONNX export
Sibling ModelInflect-Micro-v2 (9.36M params, quality-prioritized below 10M)
Apione public API across Micro and Nano sizes
Librarypytorch
MetricsWER
Inference False On Hfyes (no hosted HF inference endpoint; local-only)

Features: Inflect-Nano-v2's defining constraint is completeness under 4M parameters: the entire text-to-waveform pipeline — no separate vocoder, no phonemizer, no system dependencies — fits in 3.97M deployable parameters and a 15.97 MB FP32 checkpoint. This is smaller than even sanoTTS's smallest voice (745k) when measured by complete-pipeline footprint, though sanoTTS ships per-voice weights rather than a single fixed-voice checkpoint. The VITS end-to-end architecture is the enabler: by folding the acoustic model and vocoder into a single jointly-trained network, Inflect avoids the multi-stage pipeline overhead that makes most neural TTS systems larger. The public adaptation toolkit extends the fixed-voice design into a customizable platform — users can prepare data, adapt a voice or language, resume training, and export PyTorch or ONNX packages — making the 4M-parameter footprint a starting point for domain-specific TTS rather than a dead-end fixed-voice release.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Gepard

Gepard

Description: GEnerative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue. Gepard is built for low-latency, high-throughput streaming conversation: the model starts speaking the moment text begins arriving, generating audio piece by piece instead of waiting for a full sentence. It is a single decoder-only autoregressive language model built on Qwen3.5 (14 layers, hidden 1024, 8 heads) with ≈556M total parameters (backbone + audio interface + voice-cloning compressor). Audio is produced through NVIDIA NeMo NanoCodec — Finite Scalar Quantization at 22.05 kHz, 21.5 frames/s, 1.89 kbps — with the full 32-channel FSQ frame sampled in one step. Reports ~25× real time on a single RTX 5090 with first-audio-chunk latency around 50 ms; a 96 GB Blackwell card serves up to 256 concurrent conversations. CFG refinement is baked into the weights so quality gain comes at no extra two-pass cost at inference, though the two-pass mode is still selectable as a quality dial.

Release Date: June 22, 2026

FeatureValue
Parameters~556M (555,694,169; Qwen3.5 backbone + audio interface + voice-cloning compressor)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish (US/UK), Spanish (es-MX), Portuguese (pt-BR), Dutch (NL)
Streaming
License![Apache 2.0][license-apache-2.0]
Audio CodecNVIDIA NeMo NanoCodec (FSQ, 22.05 kHz, 21.5 fps, 1.89 kbps; NVIDIA Open Model License)
Sample Rate22,050 Hz
BackboneQwen3.5 full-attention transformer (14 layers, hidden 1024, 8 heads; ~500M params)
InferencevLLM
Throughput256 conversations on one 96 GB Blackwell (RTX Pro 6000) GPU
BenchmarkSeed-TTS-eval leader on perceived quality (NISQA-MOS 4.25, NOI 4.16, COL 4.16, DIS 4.51) trading some WER/SIM

Features: A prosody-aware autoregressive single-pass frame generator: the whole 32-channel FSQ audio frame is sampled in one step (no depth transformer), and CFG quality refinement is baked into the weights rather than incurred at inference as a two-pass cost — so the publicly reported TTFA of ~50 ms and 25× real time on a single RTX 5090 represent the quality-on path, not a cheap-fast preview. Voice cloning is decoupled into a separate up-front compressor, which means cloning is "free" at run-time once the reference clip is encoded — a structural choice that supports serving hundreds of conversations per GPU. A stop-head weight update (2026-08-06) fixed premature stopping at sentence boundaries and lifted the effective duration ceiling; on Seed-TTS-eval Gepard leads the compared systems on perceived quality (NISQA-MOS 4.25) while trading some speaker similarity and WER for its streaming-first design.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo] ![Paper][link-paper] ![Website][link-website]

· · · · · · · · · · · · · ·

Higgs Audio v3 TTS

Higgs Audio v3 TTS

Description: Boson AI's flagship conversational TTS: an ~4B autoregressive decoder over interleaved text and audio tokens from the Higgs Tokenizer (8 codebooks at 25 fps / 24 kHz). Built for voice chat rather than narration, it covers 102 languages with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

Release Date: June 4, 2026

FeatureValue
Parameters4B (BF16, 36 layers, hidden=2560, GQA 32/8)
ArchitectureAutoregressive decoder (Qwen3-style)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages102 (85 with WER/CER <5, 17 between 5-10)
Streaming
Audio Output24 kHz
License![Research Only][license-research-only]

Features: Interleaved text/audio token modelling with a delay-pattern multi-codebook embedding/head: a single autoregressive stack emits both modalities and supports inline <|category:value|> control tags (emotion/style/sfx/prosody) inserted at any point in the target text.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]

· · · · · · · · · · · · · ·

dots.tts

dots.tts

Description: dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive TTS system from Rednote-HiLab. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE, with no discrete tokens anywhere in the pipeline. It achieves the best average performance on Seed-TTS-Eval (WER 0.94 / 1.30 / 6.60 on zh / en / zh-hard) and the highest speaker similarity on a 24-language MiniMax multilingual benchmark, with broad cross-lingual voice cloning.

Release Date: June 3, 2026

FeatureValue
Parameters2B (semantic encoder + LLM + AR flow-matching acoustic head)
Voice Cloning
Asr
LanguagesMultilingual (24+ languages; zh / en focus)
Streaming
License![Apache 2.0][license-apache-2.0]
Sample Rate48 kHz
Tokenizer48 kHz AudioVAE (continuous, no discrete tokens)

Features: A fully continuous autoregressive pipeline that keeps generation in waveform-latent space end-to-end (no discrete-code phase), pairing an LLM-side semantic encoder with an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE — yielding SOTA seed-TTS-Eval scores and the strongest speaker-similarity number (83.9 avg) on the 24-language MiniMax multilingual benchmark.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Confucius4-TTS

Confucius4-TTS

Description: Confucius4-TTS is an LLM-based text-to-speech system from NetEase Youdao designed for multilingual and cross-lingual synthesis. It uses a speech encoder + LLM (Text2Semantic) + flow-matching Semantic2Acoustic architecture that allows zero-shot voice cloning without a required reference transcript and explicit cross-lingual voice transfer with unaccented output across languages. Covers Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese with code-switching and emotion transfer.

Release Date: June 2, 2026

FeatureValue
Voice Cloning
Asr
Emotion Control
Languages14 (zh, en, ja, ko, de, fr, es, id, vi, th, pt, it, ru, ms)
Streaming
License![Apache 2.0][license-apache-2.0]
Architecturespeech encoder + LLM (T2S) + flow-matching head (S2A)

Features: Cross-lingual voice transfer without accent drift: the same reference voice stays consistent when the speaker switches languages — backed by a speech encoder + LLM backbone pipeline (T2S) with a flow-matching acoustic decoder (S2A) and training that bundles 14 languages with code-switched, emotion-preserving decoding.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

WavTTS

WavTTS

Description: WavTTS is an end-to-end zero-shot TTS framework that synthesizes speech directly in the raw waveform space — explicitly skipping the intermediate mel-spectrogram, VAE-latent, or codec-token representations that most modern TTS stacks use. It is built on a flow-matching diffusion transformer (DiT) with waveform patchification, multi-scale mel-spectrogram supervision, and an optimized noise schedule. Forked from F5-TTS at the codebase level but replaces the whole acoustic pipeline.

Release Date: May 28, 2026

FeatureValue
Voice Cloning
Asr
LanguagesEnglish, Chinese
Streaming
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
![MIT][license-mit]
Sample Rate16 kHz
Training DataEmilia
ArchitectureFlow-matching DiT, raw waveform patchification, multi-scale mel supervision
Training Steps1.2M

Features: Skip every intermediate waveform representation (no mel, no VAE, no codec tokens): a flow-matching DiT produces raw-waveform patches directly, supervised at multiple mel scales and an optimized noise schedule — yielding high-quality zero-shot TTS at 16 kHz from a single end-to-end stack.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

MOSS-TTS

MOSS-TTS

Description: MOSS-TTS is a production-grade Text-to-Speech foundation model developed by the OpenMOSS Team and MOSI.AI. The current public v1.5 release preserves the original 1.0 capabilities — zero-shot voice cloning, long-form speech generation, token-level duration control, Pinyin/IPA pronunciation supervision, multilingual synthesis, and code-switching — and extends multilingual continued training from 20 languages to 31 languages including Cantonese, Dutch, Finnish, Hindi, Macedonian, Malay, Romanian, Swahili, Tagalog, Thai, and Vietnamese. v1.5 improves speaker similarity, reduces cloning variance on long-reference / short-text scenarios, follows punctuation-driven prosody more reliably, and adds explicit inline pause markers (e.g., [pause 3.2s]).

Release Date: May 25, 2026

FeatureValue
Parameters8B (Delay), 1.7B (Local)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages31 (extended from v1.0's 20)
Streaming
License![Apache 2.0][license-apache-2.0]
Max Duration1 hour
Pause Controlyes (inline markers like [pause 3.2s])
Lang Tag Controlyes (set language= in user message)

Features: v1.5 widens MOSS-TTS from 20 → 31 languages with stronger per-language multilingual synthesis (control via a language tag in the user message), more stable cloning under long-reference / short-text conditions, punctuation-driven prosody that holds up across long sentences, and explicit inline pause tokens ([pause 3.2s]) for scripted narration control.

Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Demo][link-demo]

· · · · · · · · · · · · · ·

VoxFlash-TTS

VoxFlash-TTS

Description: VoxFlash-TTS is a zero-shot voice-cloning text-to-speech engine built around extreme latent compression. The VAE encodes 24 kHz waveforms into a 9 frames/s latent space — roughly 8× more compressed than EnCodec (75 fps) and 2.4× more than Stable Audio (21.5 fps). Generating 10 s of audio therefore requires the diffusion model to produce just 90 latent vectors rather than hundreds or thousands of tokens, with downstream quadratic savings in attention cost. A ConvNeXtV2-based phoneme encoder followed by a novel coarse-alignment algorithm (cheaper than cross-attention) maps text into the latent sequence; a modern diffusion head then iteratively refines speech latents that the lightweight VAE decoder renders back to waveforms. The architecture targets low-latency, low-resource deployment — consumer-grade GPUs and edge devices — with Chinese and English zero-shot cloning. The project card lists inference: false on HF (no hosted inference endpoint), but the project page at voxflash.github.io carries the abstract, demo examples, and ablations.

Release Date: May 22, 2026

FeatureValue
Parametersnot stated (ConvNeXtV2 phoneme encoder + diffusion head + lightweight VAE decoder)
Voice Cloning
Asr
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]
Audio CodecVoxFlash VAE (9 Hz / 9 fps latent, 24 kHz input)
Compression Ratio~8× tighter than EnCodec (75 fps), ~2.4× tighter than Stable Audio (21.5 fps)
Phoneme EncoderConvNeXtV2 + coarse-alignment algorithm (no cross-attention)
Diffusion Headmodern multi-step iterative refinement
Decoderlightweight VAE decoder
Sample Rate24,000 Hz
Inferencelocal CUDA ≥ 12.3.2; no HF hosted endpoint
Training Datasetseed-tts-eval
Metricsword_error_rate, speaker_similarity

Features: The central technical move is compressing the audio latent space to 9 frames/s instead of the conventional 75 fps (EnCodec) or 21.5 fps (Stable Audio). This is not a quantization tweak — it is a temporal-decimation architectural choice that shrinks the sequence length the diffusion model has to traverse, and because attention cost scales quadratically with sequence length the end-to-end compute drops by orders of magnitude. Combined with a coarse-alignment phoneme-to-latent map that avoids cross-attention entirely (using a ConvNeXtV2 phoneme encoder instead), VoxFlash hits millisecond-level inference latency on consumer-grade and edge hardware for zero-shot Chinese + English cloning, where conventional latent-diffusion TTS systems are too slow for real-time edge deployment.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Miso TTS

Miso TTS

Description: Miso TTS 8B is a text-to-speech model from Miso Labs built on the Sesame Conversational Speech Model (CSM) architecture. A large Llama-3.2-style backbone consumes text/audio-frame embeddings and predicts codebook 0 of the Mimi audio token stream, while a smaller 300M autoregressive audio decoder predicts codebooks 1–31 in codebook depth. The model is designed for high-quality conversational speech and voice continuation from a short prompt audio clip.

Release Date: May 21, 2026

FeatureValue
Parameters8B (backbone llama-3.2-style) + 300M (audio decoder) = 8.3B
Voice Cloning
Asr
LanguagesEnglish
Streaming
License![MIT][license-mit]
ArchitectureSesame-style CSM (two transformer stack: backbone + audio decoder)
Audio TokenizerMimi (32 codebooks, vocab 2051, max seq 2048)
Librarypytorch

Features: A two-transformer Sesame-style CSM (Llama 8B backbone consumes text + audio frames and produces backbone codebook-0 prediction; a 300M audio decoder autoregresses over codebook depth via Mimi's 32-codebook stack) — letting the larger backbone spend capacity on linguistic / speaker conditioning while a leaner decoder handles fine-grained codebook-by-codebook generation.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Raon-OpenTTS-1B

Raon-OpenTTS-1B

Description: Raon-OpenTTS is an open-data, open-weight zero-shot TTS system from KRAFTON that performs on par with state-of-the-art closed-data models. This is the 1B variant (1048M parameters). Both model weights and training data are public: Raon-OpenTTS-Core is 510.1K hours of English speech, quality-filtered from the 615K-hour public Raon-OpenTTS-Pool using combined DNSMOS, WER, and VAD rank-based filtering. It ranks 1st or 2nd in WER and SIM among recent zero-shot TTS models on Seed-TTS-Eval and CV3-Eval, and achieves the best average WER/SIM on Raon-OpenTTS-Eval across Clean, Noisy, Wild, and Expressive regimes. A smaller Raon-OpenTTS-0.3B variant is also available.

Release Date: May 21, 2026

FeatureValue
Voice Cloning
Asr
LanguagesEnglish only (trained on 11 English speech datasets)
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
Parameters1048M
ArchitectureDiT (Diffusion Transformer) based on F5-TTS with flow matching; dim=1408, depth=28, heads=24
Audio Output80-ch mel-spectrogram at 16 kHz, HiFi-GAN vocoder (LibriTTS)
Training DataRaon-OpenTTS-Core (510.1K hours), 520K updates on 48× B200

Features: Demonstrates that fully open data + open weights can match proprietary SOTA: on Seed-TTS-Eval it reaches 1.78 WER / 0.749 SIM (vs Qwen3-TTS 1.46/0.715 at 1.7B), and best overall robustness (WER 2.81 / SIM 0.695) across four acoustic regimes on its own Raon-OpenTTS-Eval benchmark. The pipeline pairs large-scale rank-based data curation (DNSMOS + WER + VAD filtering of a 615K-hour pool) with an efficient F5-TTS-derived DiT.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Dataset][link-dataset]

Additional Tools:

ToolTypeLink
ComfyUI-Raon-OpenTTSComfyUI nodeComfyUI-Raon-OpenTTS

· · · · · · · · · · · · · ·

OronTTS

OronTTS

Description: OronTTS is a non-autoregressive text-to-speech model from btsee, an F5-TTS fork specialized for Mongolian (Khalkha Cyrillic) and Kazakh (Cyrillic). It uses Flow Matching + Diffusion Transformer + Vocos (dim 1024, depth 22, 16 heads, vocab 65, 24 kHz sample rate), trained on the btsee/mbspeech_mn corpus (3,846 Mongolian speech samples) and outputs zero-shot synthesis from a short reference audio + language tag.

Release Date: May 16, 2026

FeatureValue
Parameters(not stated)
Voice Cloning
Asr
LanguagesMongolian (Khalkha Cyrillic), Kazakh (Cyrillic)
Streaming
License![MIT][license-mit]
ArchitectureF5-TTS (OT-CFM + DiT + Vocos)
Dim1024
Depth22
Heads16
Vocab Size65
Sample Rate24000 Hz
Mel Bins100
Training Databtsee/mbspeech_mn (3,846 Mongolian speech samples)

Features: F5-TTS re-purposed for low-resource Cyrillic languages (Mongolian + Kazakh) — non-autoregressive flow-matching DiT over a tight 65-word vocab. Trained on a small (~3.8k sample) Mongolian corpus; the architecture is small enough that Khalkha Cyrillic and Kazakh Cyrillic share the same checkpoint via the lang tag at inference time.

Links: ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Supertonic 3

Supertonic 3

Description: Supertonic 3 is the third-generation open-weight release from Supertone. It is a lightweight, on-device text-to-speech system that runs with ONNX Runtime entirely on the user's machine (no network, no API call) and ships as a Python SDK (pip install supertonic). Compared with the Supertonic 2 base (5 languages, 66 M params), v3 expands to 31 languages and adds expression tags (<laugh>, <breath>, <sigh>), more stable reading on long utterances, and higher speaker similarity across the core language set.

Release Date: May 6, 2026

FeatureValue
Parameters(not stated on card; Supertonic 2 baseline 66 M — likely similar or smaller weight class)
Voice Cloning
Asr
Emotion Control
Languages31 (expanded from Supertonic 2's 5)
Streaming
License![OpenRAIL-M][license-openrail-m]
On Deviceyes (ONNX Runtime, no cloud call)
Expression Tagsyes (<laugh>, <breath>, <sigh>)

Features: A more compact on-device multilingual TTS: ONNX-Runtime inference everywhere, 31 languages from a single small open-weight encoder, and discrete expression tags that the decoder interprets inline — without a separate speaker-emotion control path or a cloud-rendered audio round-trip.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi]

· · · · · · · · · · · · · ·

Scenema Audio

Scenema Audio

Description: Scenema Audio is a zero-shot expressive voice cloning and speech generation model from ScenemaAI. It is built on an audio diffusion transformer extracted from the audio branch of Lightricks' LTX 2.3 (a 22B audiovisual model) — keeping the in-the-wild acoustic quality the bigger model learned while specializing for speech output. Generation is prompt-driven: a <speak> tag carries a voice description, gender, optional scene (ambient audio around the voice), and language; an <action> tag shifts emotional state mid-generation. Action tags cover rage, grief, joy, fear, exhaustion; voice prompt can describe timbre/pitch/breathiness/rasp/resonance plus character archetypes ("Tony Soprano having a breakdown"). Supports zero-shot voice cloning from 10-20 seconds of reference audio with some emotional variability, automatic long-form narration by splitting text and maintaining voice continuity, and 13 multilingual built-ins.

Release Date: April 26, 2026

FeatureValue
Parameters(audio diffusion transformer of LTX 2.3, weights ~9.8 GB bf16 / ~4.9 GB INT8 + ~6.7 GB pipeline)
Voice Cloning
Asr
Emotion Control
Languages13 (en, de, fr, es, it, pt, ja, zh, ko, ru, ar, hi, sw)
Streaming
License![Other][license-other]
Parent ModelLightricks LTX-2.3 (audio branch)
Prompt Format<speak voice=… gender=… scene=… language=…> XML with <action> tag for shifting emotion
Long Form Narrationyes (auto-splits text while preserving voice continuity)
Quantizedyes (INT8 weights at ~4.9 GB, identical quality)

Features: A standalone audio diffusion transformer extracted from a much bigger multimodal source: the model inherits how people actually sound in real scenes (angry, laughing, whispering, crying, exhausted, terrified) and exposes that capacity through a <speak> + <action> prompt interface — emotional state shifts within a single generation, instead of being a token-level or speaker-level conditioning problem.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Dramabox

Dramabox

Description: Dramabox is Resemble AI's expressive TTS, distributed under the LTX-2 Community License. It is an IC-LoRA fine-tune of the LTX-2.3 3.3B audio-only branch (Diffusion Transformer + flow matching), conditioned on Gemma 3 12B text embeddings. Generation is prompt-driven: speaker identity, emotion, delivery, laughs, sighs, breaths, pauses, and transitions are all expressed inside a natural-language description, with an optional 10-second voice reference that clones the target timbre.

Release Date: April 17, 2026

FeatureValue
Parameters3.3B (LTX-2.3 audio backbone, IC-LoRA fine-tune) + 12B Gemma 3 text encoder (conditioning only)
Voice Cloning
Asr
Emotion Control
LanguagesEnglish
Streaming
License![Other][license-other]
Base ModelLightricks/LTX-2.3 (audio branch)
ArchitectureDiT + flow matching, IC-LoRA fine-tune, Gemma 3 12B text embeddings
Inference Time~2.5 s / generation (warm server)

Features: IC-LoRA fine-tune of LTX-2.3's audio branch leaves the heavy text-understanding work to Gemma 3 12B and lets the DiT do the expressive rendering — so what's normally multimodal-stage orchestration collapses into a single prompt-driven TTS where speaker identity, emotion, and delivery are encoded in the prompt itself, and the timbre comes from a 10-second voice reference when present.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website]

· · · · · · · · · · · · · ·

Sarashina2.2-TTS

Sarashina2.2-TTS

Description: Sarashina2.2-TTS is a Japanese-centric text-to-speech system from SB Intuitions built on a large language model. It supports both Japanese and English, delivers strong pronunciation accuracy on Japanese text through large-scale end-to-end training, and reproduces a speaker's voice, speaking style, and acoustic characteristics from a short reference clip (zero-shot). Training data is sourced exclusively from legitimately acquired, properly licensed speech archives per the Sarashina Model NonCommercial License Agreement v2.0 (released April 24, 2026).

Release Date: April 16, 2026

FeatureValue
Voice Cloning
Asr
Emotion Control
LanguagesJapanese, English
Streaming
License![Research Only][license-research-only]
Base Modelsbintuitions/sarashina2.2-0.5b-instruct-v0.1
Cross Lingualyes (Japanese ↔ English, code switching)

Features: Japanese-optimized TTS fine-tuned on responsibly-licensed Japanese training corpora with explicit cross-lingual code-switching to English in a single utterance; reference audio carries speaking style and speaker identity together, so the same prompt yields narration, broadcast, conversation, or customer-service delivery without separate style conditioning.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]

· · · · · · · · · · · · · ·

LongCat-AudioDiT

LongCat-AudioDiT

Description: State-of-the-art diffusion-based TTS model operating directly in waveform latent space. Developed by Meituan's LongCat team, it requires only a Waveform VAE and Diffusion backbone, effectively mitigating compounding errors.

Release Date: March 30, 2026

FeatureValue
Parameters1B / 3.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
Audio Output24000 Hz
License![MIT][license-mit]

Features: Adaptive Projection Guidance (APG) replaces traditional classifier-free guidance for elevated generation quality. Outperforms Seed-TTS on zero-shot voice cloning benchmarks.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

SILMA TTS

SILMA TTS

Description: SILMA TTS v1 is a high-performance, 150M-parameter bilingual (Arabic/English) TTS model developed by SILMA AI. Built on the F5-TTS diffusion architecture, it was pretrained from scratch using tens of thousands of hours of high-quality public and proprietary data. It supports instant voice cloning with less than 8 seconds of reference audio (the reference transcript can also be left empty — it is transcribed on the fly), full support for Arabic Tashkeel (diacritics, auto-enriched via CATT when absent), NeMo-based text normalization, and an RTF around 0.12 on an RTX 4090. Released under a commercial-friendly license: code MIT, model weights Apache-2.0. The model is 100% compatible with F5-TTS v1.1.7 tooling for inference and fine-tuning.

Release Date: March 13, 2026

FeatureValue
Voice Cloning
Asr
Languages2 (Arabic MSA/Fusha + English)
License![Apache 2.0][license-apache-2.0]
Parameters150M
ArchitectureF5-TTS Diffusion Transformer with flow matching (pretrained from scratch, F5-TTS v1.1.7-compatible)
Pronunciation
CostRTF ≈ 0.12 (RTX 4090)

Features: Brings native-level Arabic synthesis to a 150M footprint: one of the smallest open F5-TTS-family models pretrained from scratch rather than fine-tuned, with first-class Arabic handling (Tashkeel-aware pronunciation via CATT enrichment, NeMo text normalization) alongside English, plus instant zero-shot cloning under fully permissive licensing (Apache-2.0 weights / MIT code).

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Fish Audio S2 Pro

Fish Audio S2 Pro

Description: Fish Audio S2 Pro is a leading text-to-speech model with fine-grained inline control of prosody and emotion. It combines reinforcement learning alignment with a dual-autoregressive architecture for high-quality speech synthesis.

Release Date: March 10, 2026

FeatureValue
Parameters~10 GB (BF16)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages80+ (Tier 1: En, Zh, Jp)
Streaming
License![Research Only][license-research-only]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

LongCat-Next

LongCat-Next

Description: Native multimodal foundation model by Meituan LongCat Team processing text, vision, and audio under a single autoregressive objective. Industrial-strength model with strong speech synthesis and voice cloning.

Release Date: March 2026

FeatureValue
Parameters3B (MoE A3B)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
Audio Output24 kHz
License![MIT][license-mit]

Features: Discrete Native Autoregression Paradigm (DiNA) unifying modalities in shared discrete token space. Combines visual understanding, generation, and audio processing in single model.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Voxtral-4B-TTS

Voxtral-4B-TTS

Description: Frontier, open-weights text-to-speech model developed by Mistral AI. Designed to be fast, instantly adaptable, and produces lifelike speech with natural prosody and emotional range.

Release Date: March 2026

FeatureValue
Parameters4B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages9 (En, Fr, Es, De, It, Pt, Nl, Ar, Hi)
Streaming
Audio Output24 kHz
License![CC BY-NC 4.0][license-cc-by-nc-4.0]

Links: ![HuggingFace][link-huggingface] ![Demo][link-demo] ![Blog][link-blog]

· · · · · · · · · · · · · ·

Blue (Light Blue) TTS

Blue (Light Blue) TTS

Description: BlueTTS (project page: lightbluetts.com) is a multilingual text-to-speech library. Built around slim ONNX graphs that run on ONNX Runtime with first-class CPU support and optional accelerators — OpenVINO (Intel), CUDA ORT (NVIDIA), TensorRT, and ONNX Runtime stock CPU. Targets five languages — Hebrew, English, Spanish, Italian, German — including inline mixed-language with XML-style tags in the text prompt. Inference is deliverable as a PyPI package (blue-onnx), a Rust crate, or directly from the pinned ONNX graphs on the HF Hub; the v2 release ships a slimmed opset-17 ONNX bundle (notmax123/blue-onnx-v2) that's intended for both FP32 production and the experimental INT8 weight-only fallback.

Release Date: February 27, 2026

FeatureValue
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesHebrew, English, Spanish, Italian, German
Streaming
License![MIT][license-mit]
RuntimeONNX Runtime (stock CPU; OpenVINO / CUDA / TensorRT optional)
Speed"fastest open-source TTS" (per project description)
Graph FormatONNX opset 17 (slim, full-precision; experimental weight-only INT8 fallback)
DistributionPyPI + HuggingFace + Rust

Features: A CPU-first multilingual TTS that ships both slimmed ONNX graphs and a Python package where the same code path runs on stock CPU ONNX Runtime by default — and optionally accelerates on OpenVINO / CUDA ORT / TensorRT — so deployment doesn't gate on GPU availability. Languages include Hebrew (with explicit G2P normalization) — a comparatively rare open-source TTS target — plus standard European languages, all from MIT-licensed weights distributed via both Hugging Face and PyPI.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![PyPI][link-pypi] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

KittenTTS

KittenTTS

Description: KittenTTS is an open-source realistic text-to-speech model designed for lightweight deployment. It is a state-of-the-art TTS model under 25MB with just 15 million parameters, running without GPU on any device.

Release Date: February 24, 2026 (v0.8.1)

FeatureValue
Parameters15M-80M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Multiple
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Ming-omni-tts

Ming-omni-tts

Description: Ming-omni-tts is a high-performance unified audio generation model in the Ming 2.0 series. It uses a custom 12.5 Hz continuous tokenizer and a Patch-by-Patch compression strategy that drives the LLM inference frame rate down to 3.1 Hz, enabling fine-grained control over speech rate, pitch, volume, emotion, and dialect (notably Cantonese at ~93 % accuracy). It supports 100+ premium built-in voices plus zero-shot voice design from natural-language prompts and is the first autoregressive model that jointly generates speech, ambient sound, and music in a single channel.

Release Date: February 11, 2026

FeatureValue
Parameters16.8B (3B active, MoE; A3B)
Voice Cloning
Asr
Emotion Control
LanguagesChinese, English, Cantonese
Streaming
License![Apache 2.0][license-apache-2.0]

Features: Patch-by-Patch compression drives the inference frame rate to 3.1 Hz, drastically cutting LLM-side latency for podcast-style audio while preserving naturalness. A custom 12.5 Hz continuous tokenizer plus a DiT head jointly produce speech, ambient sound, and music in a single output channel — an "in-the-scene" listening experience rather than TTS-on-top-of-a-track.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

SoulX-Singer

SoulX-Singer

Description: SoulX-Singer is a high-fidelity, zero-shot singing voice synthesis model for generating realistic singing voices for unseen singers without fine-tuning.

Release Date: February 6, 2026

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMandarin, English, Cantonese
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

SoproTTS

SoproTTS

Description: SoproTTS is a lightweight English text-to-speech model with zero-shot voice cloning. It uses dilated convolutions (WaveNet-style) and lightweight cross-attention layers instead of the common Transformer architecture.

Release Date: February 4, 2026 (v1.5)

FeatureValue
Parameters135M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]
Rtf0.05 (CPU M3)
Training-Cost~$100

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Qwen3-TTS

Qwen3-TTS

Description: Qwen3-TTS is an open-source series of Text-to-Speech models developed by Alibaba Cloud. Supports stable, expressive, and streaming speech generation with free-form voice design.

Release Date: January 22, 2026

FeatureValue
Parameters0.6B-1.7B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages10 (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian)
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

TADA

TADA

Description: TADA is a unified speech-language model from Hume AI built around a Text-Acoustic Dual-Alignment tokenizer: for every text/subword token there is exactly one corresponding speech vector, so the audio stream stays 1:1 aligned with text. As a TTS model, each autoregressive step covers one text token and dynamically determines the duration and prosody for that token, breaking the fixed-frames-per-second constraint that drives most modern TTS backbones. As a speech-language model, it generates a text token and the speech for the preceding token in the same dual step.

Release Date: January 12, 2026

FeatureValue
Parameters1B (Llama 3.2 1B base)
Voice Cloning
Asr
Emotion Control
LanguagesEnglish
Streaming
License![Other][license-other]
Base Modelmeta-llama/Llama-3.2-1B
Tokenization1:1 text–acoustic dual alignment (one speech vector per text token)
Dynamic Durationyes (each autoregressive step covers one text token, duration is determined per-token)

Features: A dual-alignment speech–text tokenizer that decouples autoregression from a fixed audio frame rate: each text token owns exactly one speech vector, and the model synthesizes the whole segment for that token in one step, regardless of how long the spoken form is — eliminating transcript hallucination and the latency overhead of constant-frame-rate codecs while staying as compact as Llama 3.2 1B.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi] ![Paper][link-paper] ![Blog][link-blog]

· · · · · · · · · · · · · ·

Irodori-TTS-500M-v2

Irodori-TTS-500M-v2

Description: Japanese Text-to-Speech model based on Rectified Flow Diffusion Transformer. Features emoji-based style and sound effect control by embedding emojis in input text for expressive speech generation.

Release Date: 2026

FeatureValue
Parameters500M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesJapanese
Streaming
Audio Output48kHz waveform
License![MIT][license-mit]

Features: Key Feature: Emoji annotation control - insert specific emojis into text to control speaking styles, emotions, and sound effects.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

KugelAudio

KugelAudio

Description: Open-source TTS for European languages with 7B parameters. Outperformed ElevenLabs in human preference testing.

Release Date: Early 2026

FeatureValue
Parameters7B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages23 European languages
Streaming
License![MIT][license-mit]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

LEMAS-TTS

LEMAS-TTS

Description: Part of the LEMAS (Large-scale Extensible Multilingual Audio Suite) project. Zero-shot multilingual TTS with 0.3B parameters supporting 10 languages with word-level precise editing capabilities.

Release Date: 2026

FeatureValue
Parameters0.3B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages10 (zh/en/de/fr/es/pt/it/ru/id/vi)
Streaming
License![Apache 2.0][license-apache-2.0]
Special-FeatureWord-level editing (LEMAS-Edit)

Features: Built on 150,000+ hours of multilingual speech data with word-level timestamps. Includes LEMAS-Edit for precise word-level speech editing via masked token infilling.

Links: ![Website][link-website] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

MioTTS-2.6B

MioTTS-2.6B

Description: Lightweight, high-speed LLM-based TTS model for English and Japanese with minimal resource usage.

Release Date: 2026

FeatureValue
Parameters2.6B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Japanese
Streaming
License![LFM][license-lfm]
Rtf0.135-0.145

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

MOSS-TTS-Nano

MOSS-TTS-Nano

Description: Ultra-lightweight open-source multilingual speech generation model with only 0.1B parameters. Designed for realtime speech generation that runs directly on CPU without GPU.

Release Date: 2026

FeatureValue
Parameters0.1B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages20
Streaming
Audio Output48 kHz Stereo
License![Apache 2.0][license-apache-2.0]

Features: Pure autoregressive architecture with MOSS-Audio-Tokenizer-Nano. Compresses audio to 12.5 Hz token stream using RVQ with 16 codebooks. Runs on 4-core CPU.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

NeuTTS

NeuTTS

Description: NeuTTS is a collection of open-source on-device TTS models with instant voice cloning. Built off LLM backbones with GGUF format quantizations for efficient on-device deployment.

Release Date: Early 2026

FeatureValue
Parameters360M (Air), 120M (Nano)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Spanish, German, French
Streaming
License![Apache 2.0][license-apache-2.0]
On-Deviceyes (GGUF quantizations)

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

OmniVoice

OmniVoice

Description: Massive multilingual zero-shot TTS model scaling to 600+ languages. Uses diffusion language model-style discrete non-autoregressive architecture with single-stage text-to-acoustic mapping.

Release Date: 2026

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages600+
Streaming
License![Apache 2.0][license-apache-2.0]
Training-Data581k hours

Features: Simplified single-stage architecture vs conventional two-stage pipelines. Full-codebook random masking strategy with LLM initialization for superior intelligibility. Noise-robust prompt processing.

Links: ![Website][link-website] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

T5Gemma-TTS

T5Gemma-TTS

Description: Multilingual TTS model with voice cloning and duration control, built on the T5Gemma encoder-decoder LLM architecture. Supports batch generation for multiple audio variations.

Release Date: 2026

FeatureValue
Parameters2B-2B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Chinese, Japanese
Streaming
License![MIT][license-mit]
Vram7.6-10.6 GB

Features: PM-RoPE positional encoding with XCodec2 audio codec. Low-VRAM options with CPU offloading. Batch inference efficiency with single encoder pass.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

TinyTTS

TinyTTS

Description: The smallest English TTS model with only 1.6 million parameters. End-to-end neural network achieving ~53x real-time synthesis speed on CPU via ONNX optimization.

Release Date: 2026

FeatureValue
Parameters~3.4 MB (ONNX FP16)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]

Features: Ultra-compact architecture optimized for CPU-only deployment. Multi-platform support via Python and Node.js APIs. Works on laptops, edge devices, and embedded systems.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

VoxCPM2

VoxCPM2

Description: OpenBMB's next-generation tokenizer-free diffusion autoregressive TTS model with 2 billion parameters. Supports 30 languages with automatic detection, voice design from text descriptions, and high-fidelity voice cloning.

Release Date: 2026

FeatureValue
Parameters2B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages30 (+ 9 Chinese dialects)
Streaming
Audio Output48 kHz
License![Apache 2.0][license-apache-2.0]

Features: Tokenizer-free design with LocEnc → TSLM → RALM → LocDiT pipeline. Built-in super-resolution via AudioVAE V2 for 48kHz output.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Soprano

Soprano

Description: Soprano is an ultra-lightweight, on-device text-to-speech (TTS) model designed for expressive, high-fidelity speech synthesis at unprecedented speed. The 1.1 release ships an 80M-parameter backbone that achieves up to 20× real-time generation on CPU and 2000× real-time on GPU, with lossless streaming (<250 ms latency on CPU, <15 ms on GPU), <1 GB memory usage at inference, and infinite generation length (automatic text splitting). Output sample rate is 32 kHz, with widespread device support (CUDA / CPU / MPS on Windows, Linux, and Mac). Inference is production-ready through an OpenAI-compatible endpoint, ONNX, WebUI, CLI, and ComfyUI nodes. The base 1.1 model is ekwek/Soprano-1.1-80M on HuggingFace; a fine-tuning toolkit (soprano-factory) was released January 13 2026 alongside the 1.1 weights — the 1.1 release reports 95% fewer hallucinations and a 63% preference rate over 1.0 (Soprano-80M). A live demo runs on ekwek/Soprano-TTS HF Space.

Release Date: December 22, 2025

FeatureValue
Parameters80M (Soprano-1.1-80M)
Voice Cloning
Asr
LanguagesEnglish (US/UK family voices, per HF Space)
Streaming
License![Apache 2.0][license-apache-2.0]
Sample Rate32,000 Hz
Inference TargetsOpenAI-compatible endpoint, ONNX, WebUI, CLI, ComfyUI
Performance Cpuup to 20× real-time
Performance Gpuup to 2000× real-time
Memory<1 GB at inference
Text Lengthinfinite (automatic text splitting)
DevicesCUDA, CPU, MPS (Windows, Linux, Mac)
Training Toolkitsoprano-factory (https://github.com/ekwek1/soprano-factory)
History 1 1Soprano-1.1-80M released 2026-01-14 (95% fewer hallucinations; 63% preference over 1.0)
History 1 0Soprano-80M released 2025-12-22

Features: The defining trade-off of this release is extreme on-device efficiency at sub-100M scale: an 80M-parameter backbone hits <250 ms CPU latency and <15 ms GPU for lossless streaming while keeping inference within <1 GB of memory — well under the multi-billion-parameter budget that newer conversational TTS systems require. The release pairs the model with soprano-factory (open-source training/fine-tuning toolkit) so users can build their own voices on top of the same backbone, and one installation can drive OpenAI-compatible / ONNX / WebUI / CLI / ComfyUI inference. The 1.1 update is a measured iteration: 95% fewer hallucinations and a 63% preference over 1.0 at the same parameter budget, so the measurable quality jump ships with no added inference cost.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

GLM-TTS

GLM-TTS

Description: High-quality TTS synthesis system based on LLMs from ZhipuAI, supporting zero-shot voice cloning with Multi-Reward Reinforcement Learning.

Release Date: December 11, 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Echo-TTS

Echo-TTS

Description: Echo is a 2.4B-parameter diffusion-based diffusion transformer (DiT) text-to-speech model. It conditions on target text and up to two minutes of speaker reference audio, generates Fish Speech S1-DAC latents, and decodes to 44.1 kHz audio. Output length is up to 30 seconds per segment. The model is fast at single-sample generation: on one A100, generating 30 seconds of audio from a 120-second prompt takes ~1.45 seconds (RTF < 0.05) — substantially faster than frontier autoregressive approaches at similar quality. The architecture is a deliberate pivot from the author's prior autoregressive-in-DAC-space model Parakeet, which struggled with semantic-consistency retries and weak voice cloning; Echo's diffusion approach trades off real-time interactivity for fast, high-fidelity zero-shot voice cloning in offline synthesis. Trained via the TPU Research Cloud (TRC). Demo (preview) hosted on jordand/echo-tts-preview HF Space; base model on jordand/echo-tts-base.

Release Date: December 4, 2025

FeatureValue
Parameters2.4B (DiT)
Voice Cloning
Asr
LanguagesEnglish (per demo samples)
Streaming
License![MIT][license-mit]
Architecturediffusion transformer (DiT) in Fish Speech S1-DAC latent space
Max Segment Duration30 s
Sample Rate44,100 Hz
Speaker Reference Max120 s
Performance A100 Rt30 s output in ~1.45 s (RTF < 0.05)
Audio CodecFish Speech S1-DAC
Prior ModelParakeet (autoregressive in DAC space)
Training InfrastructureTPU Research Cloud (TRC)
Inference RequirementsCUDA-capable GPU with at least 8 GB VRAM; Python 3.10+
Samplereuler CFG with independent guidances for text (3.0) and speaker (8.0); 40 steps; sequence_length 640 default
License ClarificationMIT (per GH repo license)

Features: Echo is a deliberate next-step pivot from autoregressive-in-DAC-space TTS to a full diffusion approach. The author's prior model, Parakeet, generated DAC tokens autoregressively but suffered the classic AR weakness — semantic-consistency retries — and weak voice cloning. Echo keeps Fish Speech S1-DAC latents (so the audio representation is the same proven codec) but moves the generator upstream to a 2.4B DiT operating directly on those latents, conditioned on a long (up to 2-minute) speaker reference. The result: 30-second outputs in ~1.45 s on a single A100 (RTF < 0.05) with high-fidelity zero-shot cloning — fast enough that the "diffusion is too slow" objection no longer applies at the segment length that matters for offline content generation, while the AR class's retry-induced inconsistency is gone by construction.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]

· · · · · · · · · · · · · ·

VibeVoice-Realtime

VibeVoice-Realtime

Description: Real-time TTS model from Microsoft with streaming text input and ultra-low latency (~300ms).

Release Date: December 3, 2025

FeatureValue
Parameters0.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMultilingual
Streaming
License![MIT][license-mit]
Max Duration~10 minutes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Fun-CosyVoice 3.0

Fun-CosyVoice 3.0

Description: Advanced TTS system based on LLMs for zero-shot multilingual speech synthesis from FunAudioLLM.

Release Date: December 2025

FeatureValue
Parameters0.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages9 + 18+ Chinese dialects
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

LFM2-Audio-1.5B

LFM2-Audio-1.5B

Description: Liquid AI's first end-to-end audio foundation model with low latency and real-time conversation.

Release Date: November 28, 2025

FeatureValue
Parameters1.5B
Voice Cloning
Asr
Emotion Control
LanguagesEnglish
Streaming
License![LFM][license-lfm]

Links: ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

Marvis-TTS

Marvis-TTS

Description: Marvis is a conversational real-time streaming TTS from Marvis-Labs. The architecture inherits Sesame's CSM-1B (Conversational Speech Model): a 250M-parameter multimodal backbone that processes interleaved text + audio tokens and a smaller 60M-parameter audio decoder that models the remaining 31 RVQ codebook levels to reconstruct high-quality speech from the backbone's representations. Audio tokens come from Kyutai's mimi codec (RVQ tokens). The dual-transformer split — semantic backbone + small decoder — yields sub-second latency, and the model is built for on-edge / on-device deployment (Apple Silicon / iPad / iPhone / Mac). Two operational choices distinguish Marvis:

Release Date: November 6, 2025

FeatureValue
Parameters250M (multimodal backbone) + 60M (audio decoder) = 310M total
Voice Cloning
Asr
LanguagesEnglish, French, German
Streaming
License![Apache 2.0][license-apache-2.0]
Architecturedual-transformer CSM-1B (Conversational Speech Model) — multimodal backbone + audio decoder
Audio CodecKyutai mimi codec (RVQ tokens; backbone models codebook 0, decoder models codebook 1–31)
Quantized Size~500 MB (4-bit MLX)
Training Datasetamphion/Emilia-Dataset
Library Nametransformers, mlx, mlx-audio
Inference Climlx_audio.tts.generate --model Marvis-AI/marvis-tts-250m-v0.2 --stream --text "..." [--ref_audio ./x.wav]
Variants In Collection250m-v0.2, 250m-v0.2-MLX-{4bit,6bit,8bit}, 100m-v0.2 (+ MLX variants), 250m-v0.2-transformers
Emits Text Chunkingno (full-sequence contextual processing)

Features: Two operational choices make Marvis stand out among conversational TTS. First, no regex chunking: most streaming TTS engines pre-split sentences by regex patterns before feeding them to the generator, which can disrupt flow / intonation; Marvis processes the entire text contextually, treating the text as a single interleaved multimodal sequence. Second, the dual-transformer CSM-1B design — a 250M backbone for codebook 0 (semantic) and a smaller 60M audio decoder for codebooks 1-31 (acoustic) — produces a quantized footprint of ~500 MB, enabling on-device Apple-Silicon inference (iPad / iPhone / Mac) with real-time streaming. The architecture makes a high-quality CSM-style TTS with zero-shot cloning actually deployable at the edge, while the official collection's 4 / 6 / 8-bit MLX variants let users trade footprint for fidelity on a per-device basis.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

IndexTTS2

IndexTTS2

Description: AI-Enhanced Text-to-Speech System with Intelligent Optimization and self-learning capabilities.

Release Date: November 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]
Multi-Speakeryes (1-4 speakers)

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Maya1

Maya1

Description: State-of-the-art speech model for expressive voice generation with natural language voice control.

Release Date: November 2025

FeatureValue
Parameters3B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish (Multi-accent)
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

Step-Audio-EditX

Step-Audio-EditX

Description: 3B-parameter LLM-based RL audio model specialized in expressive and iterative audio editing.

Release Date: November 2025

FeatureValue
Parameters3B (4B BF16)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMandarin, English, Sichuanese, Cantonese, Japanese, Korean
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

KaniTTS

KaniTTS

Description: KaniTTS is a 370M-parameter two-stage text-to-speech model from nineninesix-ai. The architecture pairs an LFM-2 backbone LLM (Liquid Foundation Model v2 — a non-transformer, structured state-space architecture) with a neural audio codec for output waveform synthesis. The LLM generates compressed audio-token representations and the codec renders them to 22 kHz waveforms, yielding low-latency generation: ~1 s to produce 15 s of audio on a single RTX 5080, with 2 GB GPU VRAM at inference, and MOS 4.3 / WER < 5% quality on the project's benchmarks. Languages covered: English, German, Chinese, Korean, Arabic, Spanish across multiple per-language voices (English, German, Chinese, Korean, Arabic, Spanish each ship a 400M checkpoint; Japanese ships a 370M "Expo-2025-Osaka" variant; a multilingual 370M checkpoint is also available). MLX variants exist for Apple-Silicon inference. The codec is the same author's nemo-nano-codec-22kHz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported to ~12.5 fps / 0.6 kbps). The model is part of the nineninesix-ai family alongside Gepard.

Release Date: September 30, 2025

FeatureValue
Parameters370M (kani-tts-370m multilingual); 400M per-language (en / de / zh / ko / ar / es); 370M (expo2025-osaka-ja)
Voice Cloning
Asr
LanguagesEnglish, German, Chinese, Korean, Arabic, Spanish (multilingual 370M checkpoint); Japanese (Expo-2025-Osaka variant)
Streaming
License![LFM][license-lfm]
Sample Rate22,000 Hz
Backbone LlmLFM-2 (Liquid Foundation Model; non-transformer structured state-space architecture)
Audio Codecnineninesix/nemo-nano-codec-22khz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported)
Performance Rt 5080~1 s for 15 s audio on RTX 5080
Memory2 GB GPU VRAM at inference
Quality Mos4.3 / 5 (naturalness)
Quality Wer<5% (accuracy)
Training Dataset~80k hours (LibriTTS, Common Voice, Emilia)
Training Hardware8x H100 GPUs, 45 hours on Lambda AI
Per Language Modelskani-tts-400m-{en,zh,de,ar,es,ko} on HuggingFace
Pretrained Checkpoints0.2-pt (450M), 0.3-pt (400M) for custom posttraining / fine-tuning
Mlx Variantskani-tts-370m-MLX (Apple Silicon)
Arxiv2505.20506

Features: KaniTTS's design choice worth flagging: it pairs a non-transformer LFM-2 backbone (Liquid Foundation Model — structured state-space rather than attention) with a neural audio codec for output at the 370M scale. The choice lets the model hit a ~1 s / 15 s audio generation rate on a 2 GB GPU VRAM budget — sub-1B parameters, sub-entry-tier GPU requirement, but still multilingual across six languages. The two-stage approach (LLM → codec) is conventional; what's less conventional is the choice of a state-space backbone over the usual transformer decoder at this scale, hitting latency / VRAM numbers that open up sub-1B real-time TTS on consumer-grade hardware. The same author ships soprano-factory-style companion assets (pretrained v0.2-pt / v0.3-pt checkpoints, a NeMo NanoCodec MLX port) to lower the bar for fine-tuning on custom datasets.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

VibeVoice-Finetuning

VibeVoice-Finetuning

Description: This is an unofficial, work-in-progress LoRA fine-tuning toolkit for the VibeVoice TTS / speech model (1.5B-base and 7B-base checkpoints). The base VibeVoice checkpoints are the same ones covered in this list's a separate entry: audio-conditioned diffusion TTS for spoken dialogue, streaming, etc. This toolkit takes pretrained VibeVoice weights + a paired (text, audio, optional reference-audio) dataset and trains a LoRA adapter against two losses simultaneously:

Release Date: September 16, 2025

FeatureValue
Parameters1.5B (LoRA-adapted) / 7B (LoRA-adapted)
Voice Cloning
Asr
Languagesinherits base VibeVoice coverage
Streamingnot directly (toolkit output is a LoRA adapter; the adapter inherits VibeVoice inference shape)
License![MIT][license-mit]
Loss Textmasked cross-entropy on text tokens
Loss Acousticdiffusion MSE on acoustic latents
Hardware 1 5B≥16 GB VRAM
Hardware 7B≥48 GB VRAM
Transformers Version4.51.3 (known-good; other versions may break on Qwen2 architecture)
Tested Docker Imagerunpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04
Audio Target Format24 kHz audio (paired dataset of target-audio + transcripts + optional reference-audio prompts)
Training Entrypointpython -m src.finetune_vibevoice_lora --model_name_or_path aoi-ot/VibeVoice-Large --processor_name_or_path src/vibevoice/processor --dataset_name <your/dataset> --text_column_name text [--voice_column_name audio_ref]
Supports Hf Dataset Loaderyes
OutputLoRA adapter compatible with VibeVoice base

Features: The dual-loss trick is the technical center of this toolkit. Naive LoRA fine-tuning of a unified TTS model often specializes the synthesis but silently damages the text LLM capability the base inherited from its Qwen-class backbone; the "masked CE on text tokens + diffusion MSE on acoustic latents" two-headed loss preserves both competencies at training time. Pair that with the careful pinning of Transformers 4.51.3 (other versions break on the Qwen2 architecture) and a documented minimum-VRAM budget per base size (16 GB for 1.5B, 48 GB for 7B), and you get a reproducible recipe for community fine-tuning of VibeVoice — something the official Microsoft VibeVoice release doesn't ship out-of-the-box.

Links: ![GitHub][link-github]

· · · · · · · · · · · · · ·

VoxCPM

VoxCPM

Description: Tokenizer-free TTS system for context-aware speech generation and true-to-life voice cloning.

Release Date: September 16, 2025

FeatureValue
Parameters640M-800M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

FireRedTTS2

FireRedTTS2

Description: Long-form streaming TTS system for multi-speaker dialogue generation with stable, natural speech.

Release Date: September 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEN, ZH, JP, KO, FR, DE, RU
Streaming
License![Apache 2.0][license-apache-2.0]
Multi-Speakeryes (4 speakers)
Max Duration3 minutes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Audio Flamingo 3 (AF3) / Audio Flamingo Next

Audio Flamingo 3 (AF3) / Audio Flamingo Next

Description: NVIDIA ADLR's fully open-source Large Audio Language Model with state-of-the-art audio understanding. Audio Flamingo Next (AF-Next) is the latest generation featuring stronger general audio understanding, longer context support, and timestamp-grounded reasoning.

Release Date: July 2025 (AF3), 2026 (AF-Next)

FeatureValue
Parameters7B
Voice Cloning
Asr
Emotion Control
LanguagesMulti-lingual
Streaming
License![Apache 2.0][license-apache-2.0]
ContextUp to 30 minutes

Features: Key Innovation (AF-Next): Staged curriculum training with GRPO-based RL post-training. Three specialized checkpoints: Instruct, Think (reasoning), and Captioner. Temporal Audio Chain-of-Thought grounding intermediate reasoning to timestamps.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

ZipVoice

ZipVoice

Description: Fast and high-quality zero-shot TTS models based on flow matching.

Release Date: June 16, 2025

FeatureValue
Parameters123M
LanguagesChinese, English
License![Apache 2.0][license-apache-2.0]
Zero-Shot-Cloningyes
Dialogueyes

Links: ![GitHub][link-github] ![Website][link-website] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Fish Speech

Fish Speech

Description: State-of-the-art open source TTS and voice cloning model that generates natural, realistic, and emotionally rich speech.

Release Date: May 31, 2025 (v1.5.1)

FeatureValue
Parameters4B (S1), 0.5B (S1-mini)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages8 (EN, JP, KO, ZH, FR, DE, AR, ES)
Streaming
License![Apache 2.0][license-apache-2.0]
Rtf~1:7

Links: ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Chatterbox

Chatterbox

Description: Family of SOTA open-source TTS models by Resemble AI, covering a single-language English line plus a multilingual V3 release that brings broader language coverage, more consistent speaker similarity, reduced hallucinations, and more natural conversational speech across 23+ languages.

Release Date: April 24, 2025

FeatureValue
Parameters500M (Llama backbone, 0.5B)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages23+
Streaming
License![MIT][license-mit]

Features: First open-source TTS model with explicit emotion exaggeration control, plus an alignment-informed inference pipeline and a watermarked decoder. Multilingual V3 narrows the quality gap to closed systems like ElevenLabs on cross-language voice cloning while staying under 1B parameters.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Orpheus-TTS

Orpheus-TTS

Description: SOTA open-source TTS built on Llama-3b backbone demonstrating emergent capabilities of LLMs for speech synthesis.

Release Date: April 2025

FeatureValue
Parameters3B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMultilingual
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

MegaTTS3

MegaTTS3

Description: Advanced zero-shot speech synthesis with Sparse Alignment Enhanced Latent Diffusion Transformer.

Release Date: March 22, 2025

FeatureValue
Parameters0.45B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Spark-TTS

Spark-TTS

Description: Efficient LLM-Based TTS Model with Single-Stream Decoupled Speech Tokens, built on Qwen2.5.

Release Date: March 2025

FeatureValue
Parameters0.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Step-Audio

Step-Audio

Description: Production-ready open-source framework for intelligent speech interaction with unified speech comprehension and generation.

Release Date: February 17, 2025

FeatureValue
Parameters130B (Chat), 3B (TTS)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English, Japanese
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Kokoro-82M

Kokoro-82M

Description: Kokoro is an open-weight Text-to-Speech model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.

Release Date: January 27, 2025 (v1.0)

FeatureValue
Parameters82M
ArchitectureStyleTTS 2, ISTFTNet
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages8 (54 voices)
Streaming
Cost<$0.06 per hour of audio
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

KokoClone

KokoClone

Description: KokoClone is a fast, real-time compatible multilingual voice cloning system built on top of Kokoro-ONNX. It enables users to type text in multiple languages, provide a short 3-10 second reference audio clip, and instantly generate speech in that same voice.

Release Date: 2025

FeatureValue
Parameters82M (Base: Kokoro-ONNX)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages7 (En, Hi, Fr, Ja, Zh, It, Pt, Es)
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

LuxTTS

LuxTTS

Description: Lightweight ZipVoice-based TTS model for high quality voice cloning at speeds exceeding 150x realtime.

Release Date: 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages-
Streaming
License![Apache 2.0][license-apache-2.0]
Rtf150x
Vram1GB

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

MiMo-Audio

MiMo-Audio

Description: Audio Language Model by Xiaomi functioning as a Few-Shot Learner with SOTA audio understanding.

Release Date: 2025

FeatureValue
Parameters7B
Voice Cloning
Asr
Emotion Control
LanguagesMulti-lingual
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

SoulX-Podcast

SoulX-Podcast

Description: SOTA Multi-Speaker TTS model for generating realistic long-form podcasts with dialectal diversity.

Release Date: 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMandarin, English, Cantonese, Sichuanese, Henanese
Streaming
License![Apache 2.0][license-apache-2.0]
Max Duration90+ minutes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

VieNeu-TTS

VieNeu-TTS

Description: Advanced on-device Vietnamese TTS model with instant voice cloning from 3-5 seconds of reference audio.

Release Date: 2025

FeatureValue
Parameters0.3B-0.6B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesVietnamese
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

Dia

Dia

Description: 1.6B parameter TTS model by Nari Labs for generating ultra-realistic dialogue in one pass.

Release Date: June 27, 2024

FeatureValue
Parameters1.6B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

MeloTTS

MeloTTS

Description: MeloTTS is a high-quality multi-lingual text-to-speech library from MyShell.ai in collaboration with MIT, supporting English (American, British, Indian, Australian, and a default accent), Spanish, French, Chinese (with mixed Chinese–English capability), Japanese, and Korean. Built on VITS / VITS2 / Bert-VITS2 family work and packaged with both a Python API and a Web UI, it runs fast enough for CPU real-time inference.

Release Date: February 19, 2024

FeatureValue
Voice Cloning
Asr
LanguagesEnglish (American, British, Indian, Australian, Default), Spanish, French, Chinese, Japanese, Korean
Streaming
License![MIT][license-mit]
BaseVITS / VITS2 / Bert-VITS2 family
Mixed Chinese Englishyes

Features: A multi-accent multilingual TTS library that ships both a Python API and a Web UI on top of the VITS-style architecture, with explicit English-accent coverage (American, British, Indian, Australian, Default) and mixed Chinese–English output — designed for fast CPU real-time inference without requiring GPU servers.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Kimi-Audio

Kimi-Audio

Description: Open-source audio foundation model by Moonshot AI for audio understanding, generation, and conversation.

Release Date: 2024

FeatureValue
Parameters7B
Voice Cloning
Asr
Emotion Control
LanguagesMulti-lingual
Streaming
License![MIT][license-mit]
![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

eSpeak-NG

eSpeak-NG

Description: eSpeak NG is a compact open-source software text-to-speech synthesizer for Linux, Windows, Android, and other operating systems. It supports more than 100 languages and accents, is a fork of Jonathan Duddington's original eSpeak engine, and uses the formant synthesis method: the model produces speech by explicitly computing the acoustic resonances (formants) of each phoneme, not by concatenating human-speech recordings. The trade is well-known — the speech is clear and usable at high playback speeds, but not as natural or smooth as larger neural or concatenative synthesizers. The compensation is size: the program and its data, including many languages, total a few megabytes. Other synthesis methods supported: Klatt formant synthesis and MBROLA diphone back-end via the documented integration.

Release Date: December 8, 2015

FeatureValue
Parametersn/a (formant-synthesis engine; not a neural model)
Voice Cloning
Asr
Languages100+ languages and accents (see docs/languages.md)
Streaming
License![GPL 3.0][license-gpl-3.0]
Synthesis Methodformant synthesis (primary); Klatt formant synthesis (secondary); MBROLA diphone backend (optional)
Footprinta few MB (program + data + many languages)
Audio OutputWAV file (CLI), direct playback, or shared-library API
Input Formatstext from file / stdin (CLI), SSML (partial), HTML (partial)
PackagesCLI (espeak-ng man page), shared library (libespeak-ng), SAPI5 Windows module
SupersedeseSpeak (Jonathan Duddington's original engine)
Downstream UsageG2P / phonemizer for neural TTS pipelines (e.g. sanoTTS bundles espeak-ng for its duration model)
PlatformsLinux, Windows, Android, Solaris, Mac OS X

Features: eSpeak-NG is the canonical reference implementation of compact multi-language formant synthesis. Its 100+-language coverage in a few megabytes, plus SSML / SAPI5 / MBROLA / shared-library / CLI surfaces, are matched only by neural TTS systems that are orders of magnitude larger. The reason it belongs in a list whose other entries are neural TTS systems is its continued quiet role in the neural stack as a G2P / phonemizer front-end — the phoneme inventory and grapheme-to-phoneme rules that sanoTTS and similar sub-1B neural TTS engines bundle are often just a port of eSpeak-NG's language-data files. So even if the formant-synthesis audio output itself has been surpassed for naturalness, the phoneme infrastructure underneath many of the smaller neural TTS entries on this list still traces back to eSpeak-NG.

Links: ![GitHub][link-github]

· · · · · · · · · · · · · ·


Anything to Audio

Models that can generate audio from multiple input modalities (video, text, image, audio). These are unified frameworks for multimodal audio synthesis.

Anything to Audio Quick Comparison

ModelTextVideoAudioMax DurationSample RateLicense
MiDashengLM-Gen16 kHz![Apache 2.0][license-apache-2.0]
ScenA![Other][license-other]
Nemotron-Labs-Audex-2B![NVIDIA NC][license-nvidia-noncommercial]
Nemotron-Labs-Audex-30B-A3B![NVIDIA NC][license-nvidia-noncommercial]
MOSS-SoundEffect30 s48 kHz![Apache 2.0][license-apache-2.0]
Omni2Sound (Omni2Audio)![CC BY-NC 4.0][license-cc-by-nc-4.0]
ControlFoley44,100 Hz![CC BY-NC 4.0][license-cc-by-nc-4.0]
Woosh![Apache 2.0][license-apache-2.0]
Chroma-4B![Apache 2.0][license-apache-2.0]
Uni-MoE (Audio)![Apache 2.0][license-apache-2.0]
AudioX / Audio-Omni![Apache 2.0][license-apache-2.0]
![CC BY-NC 4.0][license-cc-by-nc-4.0]
HunyuanVideo-Foley48 kHz![Research Only][license-research-only]
PrismAudio![Apache 2.0][license-apache-2.0]
ThinkSound![Apache 2.0][license-apache-2.0]
MMAudio![Apache 2.0][license-apache-2.0]
MiDashengLM-Gen

MiDashengLM-Gen

Description: MiDashengLM-Gen (MiDasheng Language Model for Generation) is an end-to-end framework for unified audio-scene generation from Xiaomi. Built on a pre-trained LLM and the Dasheng audio tokenizer, it couples per-token conditional flow matching with autoregressive generation to produce coherent 16 kHz audio that simultaneously blends speech, music, sound effects and environmental acoustics from a structured text description. It supports 9 languages with emotion control and approaches dedicated TTS intelligibility on speech (Seed-TTS English WER drops from 12.15% to 2.79%) while retaining mixed-audio scene capability, and extends competitively to multilingual settings.

Release Date: August 12, 2026

FeatureValue
Text
Video
Image
Audio
Sample Rate16 kHz
Languages9
Emotion Control
License![Apache 2.0][license-apache-2.0]
Parameters1.7B (Qwen3-1.7B backbone)
ArchitectureDashengTokenizer (768-dim @25Hz) + Qwen3-1.7B + flow-matching DiT (16 layers, hidden 2048)

Features: LLM-conditioned high-dimensional (768-dim @25Hz) audio latents generated without quantization artifacts; audio-text alignment pre-training maps latents into the LLM token space before generation; a learned stop head enables variable-length truncation. First end-to-end trained model for general text-to-audio-scene generation.

Links: ![Demo][link-demo] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

ScenA

ScenA

Description: ScenA generates multi-speaker audio scenes — dialogue and conversation with sound effects and ambience — from a text prompt, conditioned on one or more reference-audio clips that set the speakers' voices. Unlike prior multi-speaker dialogue systems it uses no per-turn tags, multi-stream transcripts, or speaker embeddings: a free-form natural-language prompt alone describes the scene. The text prompt determines which reference voice speaks where, allowing overlapping speech, spontaneous paralinguistic events, and scene-level ambient sound — all inherited from the in-the-wild text-to-audio pretraining distribution. The architecture is an audio-only, reference-conditioned flow-matching DiT built on the LTX-2 backbone (~4B parameters, 48 layers). Reference latents are concatenated into the token sequence and distinguished by lightweight identity-aware positional encodings. The training tackles a specifically identified "Reference Shortcut" failure mode — under standard noise schedules the model can identify the matching reference by noisy-target acoustic similarity, bypassing the text prompt — by using a high-noise-biased timestep distribution that forces reliance on the prompt for speaker assignment. Evaluator: CoVoMix2-Dialogue benchmark. Project page, code, paper, and HuggingFace checkpoint are linked below.

Release Date: July 7, 2026

FeatureValue
Parameters~4B (DiT, 48 layers; built on LTX-2 architecture)
Text
Video
Audio
Max Durationnot stated (scene-level generation)
Sample Rate(not stated; inherits LTX-2 audio VAE)
Voice Cloning
Multi Speakeryes
Ambient Soundyes (SFX, room acoustics, overlapping speech)
Architectureflow-matching DiT (LTX-2 backbone, audio-only)
Speaker Assignmentnatural language (no per-turn tags / identity encoders)
Training Fixhigh-noise-biased timestep distribution (defeats Reference Shortcut)
Text Encodergoogle/gemma-3-12b-it
Audio Vaebundled (~365 MB; encodes+decodes so full LTX-2 not needed)
Checkpoint Size~8.2 GB (scena.safetensors) + ~365 MB (audio_vae.safetensors)
License![Other][license-other]
Training Datain-the-wild text-to-audio pretrained, then reference-conditioned fine-tune
EvaluationCoVoMix2-Dialogue (speaker-binding metrics)

Features: The "Reference Shortcut" failure-mode identification is the technical center of the work: under standard diffusion noise schedules, a multi-speaker reference-conditioned model can match each reference to the noisy-target segment by acoustic similarity alone, bypassing the text prompt entirely. ScenA's high-noise-biased timestep distribution forces the model to rely on the prompt for speaker assignment at training time. Combined with the absence of any per-turn speaker structure (tags / transcripts / identity encoders) and the prompt's role as the only speaker-routing signal, this yields multi-speaker conversational scenes with overlapping speech, paralinguistic events, and ambient texture that previous structured-supervision multi-speaker systems filter out by design.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Nemotron-Labs-Audex-2B

Nemotron-Labs-Audex-2B

Description: Nemotron-Labs-Audex-2B is NVIDIA's smaller sibling of the Audex unified audio-text LLM. Like the 30B-A3B flagship, the 2B is a single model family that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). It is built on the same audio-vocabulary-extended transformer stack as the 30B-A3B but at a densely-parameterized 2B scale (no MoE), so the compute and memory footprint are lowered to a budget tractable on more modest hardware. The 2B checkpoint is the project-tagged SFT variant in the Audex collection — instruction-tuned and ready for inference. Both sizes preserve text reasoning, alignment, knowledge, long-context, and agentic capabilities of the text backbone while adding discrete-token audio I/O.

Release Date: July 6, 2026

FeatureValue
Parameters2B (dense; SFT fine-tune, instruct + reasoning-ready)
Text
Video
Audio
Modalitiestext + audio (input and output)
Max Durationnot stated
Sample Ratenot stated (decoder output)
Voice Cloning
Audio Understandingyes (audio QA, classification)
Asr
Speech Translationyes
Text To Speechyes
Text To Audioyes
Speech To Speech Generationyes
Reasoning Modeyes (thinking + instruct modes inherited from text backbone)
License![NVIDIA NC][license-nvidia-noncommercial]
Pipeline Tagtext-generation
Library Nametransformers
Derived Fromsame family as Nemotron-Labs-Audex-30B-A3B
Companion 30Bnvidia/Nemotron-Labs-Audex-30B-A3B (MoE: 30B total, 3B active)
Spacesnvidia/Nemotron-Labs-Audex, WaveCut/Nemotron-Labs-Audex, hugging-apps/nemotron-labs-audex-2b
Createdat2026-07-06T16:21:07Z
Downloads~2.4k

Features: The 2B sibling matters because it preserves the central thesis of the Audex paper — unified audio-text LLM intelligence without regressing on text intelligence — while dropping the parameter budget substantially. The 30B-A3B MoE hits a 1M-context, agentic flagship tier; the 2B dense version is the same audio-aware architecture extended down to a budget that doesn't require a high-end MoE serving stack. The pair lets users choose on deployment cost rather than on capability sub-selection: the 2B ships the same audio-to-audio + text-to-audio + audio-understanding

  • ASR + speech-translation coverage as the MoE flagship, just at the cost of longer-context / reasoning depth that the MoE was specifically tuned for. Both share the discrete-audio-token vocabulary extension of the text backbone so they can be reasoned about interchangeably.

Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Nemotron-Labs-Audex-30B-A3B

Nemotron-Labs-Audex-30B-A3B

Description: Nemotron-Labs-Audex-30B-A3B is NVIDIA's unified audio-text LLM — a single model that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). Built on Nemotron-Cascade-2-30B-A3B (text-only MoE: 30B parameters, 3B active), Audex extends the vocabulary with discrete audio tokens for speech / general-audio output and adds an audio encoder for speech / general-audio input. Runs in thinking and instruct (non-thinking) modes and supports up to a 1M-token context length — preserving text-reasoning, alignment, knowledge, long-context, and agentic capabilities of the backbone while gaining audio tasks.

Release Date: July 6, 2026

FeatureValue
Parameters30B MoE (3B active)
Modalitiesaudio (input and output)
Audio Understandingyes
Asr
Speech Translationyes
Text To Speechyes
Text To Audioyes
Speech To Speech Generationyes
Voice Cloning
License![NVIDIA NC][license-nvidia-noncommercial]
LanguagesEnglish
Modesthinking, instruct (non-thinking)
Context Length1M tokens
TemplateChatML (with <think>…</think> for thinking mode)
InferencevLLM 0.20.0 (recommended) or transformers >= 4.53.0 (mamba-ssm + causal-conv1d required)

Features: First-class audio I/O for a 30B/3B-active text LLM: extended vocabulary with discrete audio tokens for outputting speech and general audio, plus an audio encoder for input — so the same backbone keeps its strong text reasoning (alignment, knowledge, long-context) and adds ASR + speech translation + TTS + audio generation + S2S without retraining. The MoE form (30B routes, 3B active) keeps inference tractable for a single pipeline that does both.

Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection]

· · · · · · · · · · · · · ·

MOSS-SoundEffect

MOSS-SoundEffect

Description: MOSS-SoundEffect is the dedicated text-to-sound model in the OpenMOSS / MOSI.AI MOSS-TTS family. It turns natural-language captions into high-fidelity non-speech audio (ambience, urban scenes, creatures, human actions, and short music-like clips).

Release Date: May 25, 2026

FeatureValue
TypeText-to-Sound / SFX generation
ConditioningText
Max Duration30 seconds
Sample Rate48 kHz
License![Apache 2.0][license-apache-2.0]
ArchitectureDiT + Flow Matching + DAC VAE + Qwen3 text encoder
Parameters1.3B (DiT variant 1.3B)
LanguagesEnglish, Chinese
Inference Defaults100 flow-match steps, cfg 4.0, sigma_shift 5.0
Librarydiffusers

Features: Replaces the discrete-token autoregressive v1 (which bottlenecked on vocabulary) with a continuous-latent DiT + Flow Matching paired with a DAC VAE — yielding 30 s stable audio, bilingual English + Chinese prompts, and a clean CFG/sigma-shift inference schedule (cfg 4.0, shift 5.0) that works straight out of the box on the diffusers library.

Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

Omni2Sound (Omni2Audio)

Omni2Sound (Omni2Audio)

Description: Omni2Sound — also written Omni2Audio on the project page — is a unified VT2A / V2A / T2A framework and a CVPR 2026 Highlight. A single Diffusion Transformer (DiT) backbone with a decoupled two-branch conditioning design:

Release Date: April 20, 2026

FeatureValue
ConditioningText / Video / Text+Video
ModalitiesVideo, Audio
Asr
Voice Cloning
Text
Video
Image
Audio
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
TasksVT2A, V2A, T2A (single model)
ArchitectureDiT + decoupled Semantic / Temporal branches + 3-stage progressive training
Pipeline Tagtext-to-audio

Features: One single model that is SOTA on three distinct tasks (VT2A, V2A, T2A) without a separate model per mode — decoupled semantic and temporal conditioning let the same DiT backbone handle text-only, video-only, and text+video conditioning by cleanly omitting the missing modality rather than padding it, which is what most prior unified VA models had to do.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Benchmark][link-benchmark]

· · · · · · · · · · · · · ·

ControlFoley

ControlFoley

Description: ControlFoley (Xiaomi MiLM Plus) is a unified controllable video-to-audio (foley) generation model. It supports four conditioning combinations under one architecture:

Release Date: April 13, 2026

FeatureValue
ConditioningText / Video / Text + Video / Video + Reference Audio
ModalitiesVideo (visual), Audio (foley)
Asr
Voice Cloning
Text
Video
Image
Audio
Sample Rate44,100 Hz
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
Pipeline Tagtext-to-audio
Librarydiffusers
Cross Modal Conflicthandled via modality-specific control (no explicit router)
Inference SkillClawHub ControlFoley Audio Generator
UpcomingComfyUI nodes (in preparation, expanding to V2A / TV2A / TC-V2A / AC-V2A / T2A)

Features: Modality-specific cross-modal conflict resolution in a single generative stack: text governs semantics, reference audio governs timbre/acoustic style, and video governs temporal synchronization. Rather than routing to a single user-trusted modality, the model decouples control axes so an input disagreement (video shows a dog barking, text asks for a cat) is decomposed into a coherent output that respects each modality's responsibility. Trained with all-modality dropout for modality-robustness, ControlFoley is the first foley system that brings all four conditioning modes — T2A, V2A, TV2A, AC-V2A — under one model.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv] ![Skill][link-skill]

· · · · · · · · · · · · · ·

Woosh

Woosh

Description: Sony AI's sound effect foundation model for text-to-audio and video-to-audio generation. Includes Woosh-AE (audio encoder/decoder), Woosh-Flow/DFlow (T2A), and Woosh-VFlow/DVFlow (V2A) with distilled fast inference variants.

Release Date: 2026

FeatureValue
ArchitectureFlow-based generative models
Text
Video
License![Apache 2.0][license-apache-2.0]
Audio-Encodingyes
Fast-Inferenceyes (Distilled models)

Features: Optimized for sound effects (not general audio) with both public and private model versions. Video-conditioned generation without requiring captions. Competitive with Stable Audio Open and TangoFlux.

Links: ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Chroma-4B

Chroma-4B

Description: Chroma 1.0 (FlashLabs' Chroma-4B on HuggingFace) is the first open-source, real-time, end-to-end spoken dialogue model that achieves both sub-second end-to-end latency and high-fidelity personalized voice cloning. The pipeline is end-to-end — no separate ASR → LLM → TTS stitch — speech goes in, speech comes out. The architectural centerpiece is an interleaved text-audio token schedule (1 text : 2 audio) that supports streaming generation, so the model can begin emitting audio while the user is still talking (broken-off turns / barge-in handled). Experimental results from the project's paper:

Release Date: November 28, 2025

FeatureValue
Parameters4B
Voice Cloning
Asr
LanguagesEnglish (per benchmark reporting)
Streaming
License![Apache 2.0][license-apache-2.0]
Architectureend-to-end spoken-dialogue LLM; interleaved text-audio token schedule (1:2); custom_code modules
Pipeline Tagany-to-any (HF classification)
Audio Tokenizationchroma tokenizer (RVQ-style per project's tag)
Latency Rtf0.43 (speech out ~2.3× wall-clock)
Speaker Similarity+10.96% relative improvement over human baseline
Inference Librarytransformers (custom_code)
Correlations With Larger Classmatches full dialogue turn at streaming latency
Pretrainedyes (safetensors weights)
Hf Space Demoshysts/Chroma-4B, Pnevka/Chroma-4B
Historypaper arXiv 2601.11141 (2026-01)

Features: Two bets together produce the dual property that no prior open-source spoken-dialogue model has hit simultaneously. First, an interleaved text-audio token schedule (1:2) — text tokens and audio tokens are interleaved at a fixed 1:2 ratio through the sequence, which gives the model a structured place to emit audio while still consuming user audio + text context, supporting sub-second end-to-end latency without a separate ASR / LLM / TTS pipeline. Second, personalized voice cloning baked into the spoke-dialogue model — the cloned voice is not bolted on top by a separate TTS stage (as is the default pattern), it's in-model at the audio-token-generation layer. The empirical payoff is a 10.96% relative speaker-similarity gain over the human baseline (i.e. the cloned voice is closer to the reference speaker than two of the same human speaker's recordings are to each other), while hitting RTF 0.43 — a floor that prior systems exceeded either in latency (no streaming) or in cloning fidelity (parrot the speaker poorly), rarely both.

Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Uni-MoE (Audio)

Uni-MoE (Audio)

Description: MoE-based omnimodal model with voice cloning, TTS, T2M (text-to-music), and V2M (video-to-music).

Release Date: October 16, 2025 (Uni-MoE-Audio)

FeatureValue
Parameters-
Voice Cloning
Text
Video
License![Apache 2.0][license-apache-2.0]
Dynamic-Routingyes

Links: ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

AudioX / Audio-Omni

AudioX / Audio-Omni

Description: Audio-Omni is the first end-to-end framework unifying understanding, generation, and editing across general sound, music, and speech domains. Presented at SIGGRAPH 2026. AudioX is a unified framework integrating text, video, image, and audio conditions.

Release Date: March 2025 (AudioX), 2026 (Audio-Omni)

FeatureValue
Parameters-
Text
Video
Audio
License![Apache 2.0][license-apache-2.0]
![CC BY-NC 4.0][license-cc-by-nc-4.0]

Features: First unified framework covering all three audio domains. Combines frozen multimodal LLM (Qwen2.5-Omni) with trainable Diffusion Transformer for high-fidelity synthesis. Any-to-any audio processing.

Links: ![GitHub][link-github] ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

HunyuanVideo-Foley

HunyuanVideo-Foley

Description: Tencent's end-to-end video sound effect generation model for professional-grade AI Foley sound generation. Analyzes footage and creates immersive audio that matches the visual content perfectly.

Release Date: 2025

FeatureValue
Parameters-
Sample Rate48 kHz
Text
Video
License![Research Only][license-research-only]
High-Quality-Foleyyes
Context-Awareyes

Links: ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

PrismAudio

PrismAudio

Description: Video-to-Audio generation framework with Reinforcement Learning and specialized Chain-of-Thought (CoT) planning. Decomposes reasoning into four specialized modules (Semantic, Temporal, Aesthetic, Spatial CoT) for comprehensive video understanding. Built upon ThinkSound.

Release Date: 2025 (ICLR 2026)

FeatureValue
Parameters518M
Video
License![Apache 2.0][license-apache-2.0]
Cot-Planningyes (4 modules)
Multi-Dimensional-Rlyes
Fast-Grpoyes (Hybrid ODE-SDE)
Inference-Time0.63 seconds

Features: Performance Benchmarks:

MetricVGGSoundAudioCanvas
Semantic (CLAP)0.470.52
Temporal (DeSync↓)0.410.36
Aesthetic (MOS-Q)4.21±0.354.12±0.28

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

ThinkSound

ThinkSound

Description: Unified Any2Audio generation framework with flow matching guided by Chain-of-Thought (CoT) reasoning. Supports generating or editing audio from video, text, audio, or their combinations. Accepted to NeurIPS 2025.

Release Date: 2025

FeatureValue
Parameters-
Text
Audio
License![Research Only][license-research-only]
![Apache 2.0][license-apache-2.0]
Cot-Driven-Reasoningyes
Interactive-Object-Centric-Editingyes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

MMAudio

MMAudio

Description: Multimodal joint training framework for high-quality synchronized audio generation from video and/or text inputs. State-of-the-art open source model for generating sounds for videos, images, and text prompts.

Release Date: December 2024 (CVPR 2025)

FeatureValue
Parameters-
Text
Video
Image
License![Apache 2.0][license-apache-2.0]
Synchronized-Audioyes
Multimodal-Joint-Trainingyes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·


Audio Restoration & Enhancement

Audio Restoration & Enhancement Quick Comparison

ModelTypeBandwidth ExtensionInpaintingLicense
RE-USEUniversal Speech Enhancement![NVIDIA NC][license-nvidia-noncommercial]
NovaSRAudio Super-Resolution![Apache 2.0][license-apache-2.0]
QuarkAudio-UniSEUniversal Speech Enhancement![Apache 2.0][license-apache-2.0]
PASESpeech Enhancement![Apache 2.0][license-apache-2.0]
DTT-BSRMusic Source Restoration![MIT][license-mit]
NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges)High-Resolution Audio Restoration![NVIDIA NC][license-nvidia-noncommercial]
ZipEnhancerAcoustic Noise Suppression![Apache 2.0][license-apache-2.0]
AudioSRAudio Super-Resolution![Apache 2.0][license-apache-2.0]
RE-USE

RE-USE

Description: RE-USE (RE-…), NVIDIA's multilingual universal speech enhancement model, targets distortion–perception trade-off by training a single model that balances listening quality against fidelity to the underlying linguistic / speaker / emotional content. Designed to restore diverse degraded speech while leaving everything else (content, identity, prosody, accent, paralinguistic attributes) intact.

Release Date: March 17, 2026

FeatureValue
TypeUniversal Speech Enhancement
Bandwidth Extension
Inpainting
Sample Rate8 / 16 / 22.05 / 24 / 32 / 44.1 / 48 kHz (multi-rate input)
ArchitectureMamba-SSM backbone
Degradation Coverageadditive noise, reverberation, clipping, bandwidth limit, codec artifacts, packet loss, low-quality mics
Language Agnosticyes
License![NVIDIA NC][license-nvidia-noncommercial]

Features: A single Mamba-SSM model that handles seven different input sample rates (no resampling pre-step), covers a broad degradation menu in one checkpoint, stays language-agnostic without per-language training, and explicitly balances distortion reduction against fidelity to the input speech — addressing the universal-SE trade-off that earlier single-purpose enhancers couldn't.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

NovaSR

NovaSR

Description: NovaSR is a tiny audio upsampler (~52 kB parameter count) that bandwidth-extends 16 kHz input up to 48 kHz. Public card on YatharthS/NovaSR advertises realtime factors around 3500× on A100, making it a candidate for real-time on-device super-resolution where model size dominates latency. Inference path is small enough to fit in CPU memory; the use case is speech-bandwidth extension without GPU.

Release Date: January 6, 2026

FeatureValue
TypeAudio Super-Resolution (16 kHz → 48 kHz)
Bandwidth Extension
Inpainting
Channelsmono
License![Apache 2.0][license-apache-2.0]
Parameters52 kB
Streamableyes (low VRAM / runs without GPU)
Realtime Factor~3500× (A100)

Features: A 52 kB-parameter Upsampler that hits ~3500× realtime on GPU and runs on CPU — pushing bandwidth extension below the size / latency envelope where a typical neural upsampler is unacceptable (real-time on-device speech enhancement).

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

QuarkAudio-UniSE

QuarkAudio-UniSE

Description: UniSE is a unified, prompt-free autoregressive speech-enhancement framework built on a decoder-only language model. A single model performs multiple speech-enhancement tasks — speech restoration (SR / denoising), target-speaker extraction (TSE), source separation (SS), and acoustic echo cancellation (AEC, in development) — without explicit task-specific instructions or prompt conditioning; the language model infers the task from the input context. Stack: WavLM as the feature extractor, BiCodec as the discrete codec, and a decoder-only LM as the middle autoregressive backbone. Outputs reconstructed waveform from predicted discrete token sequences.

Release Date: December 22, 2025

FeatureValue
Voice Cloning
Asr
Streaming
LanguagesEnglish (paper demo)
License![Apache 2.0][license-apache-2.0]
TasksSpeech Restoration, Target Speaker Extraction, Source Separation, AEC (developing)
ArchitectureWavLM (feature extractor) + BiCodec (discrete codec) + decoder-only AR-LM
Unifiedyes (single model handles SE, SR, TSE, SS without explicit task prompts)
Prompt Freeyes (LM infers task from input context)
Dataset Signalsnoise + reverb + packet-loss + clean (configurable per task)
TrainingSpeech-enhancement SFT, then multitask joint training

Features: A single decoder-only LM that learns the speech-enhancement task distribution and infers which task to perform from the input context — eliminating the need for task-specific prompts, modules, or fine-tuning when switching between denoising, target-speaker extraction, and separation. Built as an autoregressive discrete-token predictor over a WavLM-extracted / BiCodec-quantised representation, it moves the speech-enhancement workflow from a zoo of specialist models into one generalist.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]

· · · · · · · · · · · · · ·

PASE

PASE

Description: PASE (Phonologically Anchored Speech Enhancer) is a generative speech-enhancement model from Cisco Collaboration AI that removes noise and reverberation while preserving linguistic content and speaker identity. It uses two fine-tuned WavLM-derived components:

Release Date: November 8, 2025

FeatureValue
TypeSpeech Enhancement
Bandwidth Extension
Inpainting
Sample Rate16 kHz mono
ArchitectureDenoising WavLM (DRD from WavLM-Large) + Dual-Stream Vocoder (phonetic + acoustic)
Finetuned FromWavLM-Large
Training DataDN5/DNS5 challenge clean + noise, LibriTTS, VCTK, OpenSLR26+28 RIRs
License![Apache 2.0][license-apache-2.0]

Features: Anchors enhancement to phonology instead of spectrum: by reconstructing from a phonetic stream and a separate acoustic stream (per DeWavLM's two representations), PASE keeps the words intact even when the spectrum is severely degraded — substantially lowering hallucinations while still regaining perceptual quality.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

DTT-BSR

DTT-BSR

Description: DTT-BSR (DTTNet with BandSequence and RoPE) is a music-source-restoration challenge submission from team AC/DC (Wuhan University) to ICASSP 2026. It is built inside the official MSR-Kit GAN framework, where the baseline generator is replaced by a DTTNet-style time-frequency U-Net and augmented at the bottleneck with:

Release Date: October 16, 2025

FeatureValue
TypeMusic Source Restoration
Bandwidth Extension
Inpainting
ArchitectureDTTNet TFC-TDF U-Net (complex STFT) + Improved Dual-Path BandSplitRNN block + RoPE-Transformer
Inputcomplex STFT (real + imag channels; n_fft=2048, hop=512)
DiscriminatorMulti-Frequency Discriminator (baseline)
FrameworkMSR-Kit GAN (reconstruction + adversarial + feature-matching losses)
License![MIT][license-mit]

Features: Treats music-source restoration as a complex-STFT time-frequency U-Net enhancement at the bottleneck: keep the strong DTTNet dual-path TFC-TDF structure for local spectral patterns, then layer in BandSplitRNN-style sub-band recurrence + RoPE self-attention so the generator can model long-range, cross-band harmonic structure that ordinary GAN baselines miss — critical for restoring non-vocal stems cleanly.

Links: ![GitHub][link-github]

· · · · · · · · · · · · · ·

NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges)

NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges)

Description: A2SB is NVIDIA's audio-to-audio Schrödinger Bridge diffusion model for high-resolution (44.1 kHz) music restoration. It is the first long-audio restoration model that can restore hour-long inputs without boundary artifacts, and it's end-to-end — predicting waveform outputs directly withou

Truncated — view the full README on GitHub.

ai-music
ai-voice
asr
music-generation
tts
voice-cloning

Contributors

wildminder

37 commits

wildminder/awesome-ai-voice

List of open-source TTS, voice cloning, and music generation models

494

37 commits

updated Sep 10, 2026

See the code

README

Awesome TTS & Voice Generation Models

A curated list of open-source Text-to-Speech (TTS) and voice cloning models. Models are sorted by release date (newest first).

logo-tts2


Table of Contents


Text-to-Speech (TTS) Models

TTS Quick Comparison

ModelVoice CloningASRLanguagesStreamingLicense
AuK![MIT][license-mit]
AuK-Flash![MIT][license-mit]
rumik-oss 122 Indic languages + English![Other][license-other]
Irodori-TTS-v4.1-AnimeJapanese![MIT][license-mit]
ICE-012 Audio590![CC BY-NC 4.0][license-cc-by-nc-4.0]
TontaubeV17![Other][license-other]
Breeze TTS 22![Other][license-other]
Sopro v2 Turbo4![Apache 2.0][license-apache-2.0]
CuteTTS5![Apache 2.0][license-apache-2.0]
Rynsan TTSKhasi, Garo, Pnar, English, Hindi![CC BY 4.0][license-cc-by-4.0]
Audio8 TTS Preview 0.1B8![Other][license-other]
Kiseki-TTSJapanese![MIT][license-mit]
FireRedTTS324![Apache 2.0][license-apache-2.0]
Audio8-TTS-Preview-0.6bCantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish![Apache 2.0][license-apache-2.0]
NeuTTS-2EEnglish![Other][license-other]
Scylla's Banden_us, en_gb, es, it![Apache 2.0][license-apache-2.0]
sanoTTSEnglish, Nepali, Hindi, Vietnamese, Indonesian, Chinese![Other][license-other]
FreyaTTSTurkish![Apache 2.0][license-apache-2.0]
Inflect-Nano-v2English![Apache 2.0][license-apache-2.0]
GepardEnglish, Spanish, Portuguese, Dutch![Apache 2.0][license-apache-2.0]
Higgs Audio v3 TTS102![Research Only][license-research-only]
dots.ttsMultilingual![Apache 2.0][license-apache-2.0]
Confucius4-TTS14![Apache 2.0][license-apache-2.0]
WavTTSEnglish, Chinese![CC BY-NC 4.0][license-cc-by-nc-4.0]
MOSS-TTS31![Apache 2.0][license-apache-2.0]
VoxFlash-TTSChinese, English![Apache 2.0][license-apache-2.0]
Miso TTSEnglish![MIT][license-mit]
Raon-OpenTTS-1BEnglish![CC BY-NC 4.0][license-cc-by-nc-4.0]
OronTTSMongolian, Kazakh![MIT][license-mit]
Supertonic 331![OpenRAIL-M][license-openrail-m]
Scenema AudioEnglish, German, French, Spanish, Italian, Portuguese, Japanese, Chinese, Korean, Russian, Arabic, Hindi, Swahili![Other][license-other]
DramaboxEnglish![Other][license-other]
Sarashina2.2-TTSJapanese, English![Research Only][license-research-only]
LongCat-AudioDiTChinese, English![MIT][license-mit]
SILMA TTSArabic, English![Apache 2.0][license-apache-2.0]
Fish Audio S2 Pro80+![Research Only][license-research-only]
LongCat-NextChinese, English![MIT][license-mit]
Voxtral-4B-TTS9![CC BY-NC 4.0][license-cc-by-nc-4.0]
Blue (Light Blue) TTSHebrew, English, Spanish, Italian, German![MIT][license-mit]
KittenTTSEnglish, Multiple![Apache 2.0][license-apache-2.0]
Ming-omni-ttsChinese, English![Apache 2.0][license-apache-2.0]
SoulX-SingerMandarin, English, Cantonese![Apache 2.0][license-apache-2.0]
SoproTTSEnglish![Apache 2.0][license-apache-2.0]
Qwen3-TTS10![Apache 2.0][license-apache-2.0]
TADAEnglish![Other][license-other]
Irodori-TTS-500M-v2Japanese![MIT][license-mit]
KugelAudio23 European languages![MIT][license-mit]
LEMAS-TTS10![Apache 2.0][license-apache-2.0]
MioTTS-2.6BEnglish, Japanese![LFM][license-lfm]
MOSS-TTS-Nano20![Apache 2.0][license-apache-2.0]
NeuTTSEnglish, Spanish, German, French![Apache 2.0][license-apache-2.0]
OmniVoice600+![Apache 2.0][license-apache-2.0]
T5Gemma-TTSEnglish, Chinese, Japanese![MIT][license-mit]
TinyTTSEnglish![Apache 2.0][license-apache-2.0]
VoxCPM230![Apache 2.0][license-apache-2.0]
SopranoEnglish![Apache 2.0][license-apache-2.0]
GLM-TTSChinese, English![Apache 2.0][license-apache-2.0]
Echo-TTSEnglish![MIT][license-mit]
VibeVoice-RealtimeMultilingual![MIT][license-mit]
Fun-CosyVoice 3.09 + 18+ Chinese dialects![Apache 2.0][license-apache-2.0]
LFM2-Audio-1.5BEnglish![LFM][license-lfm]
Marvis-TTSEnglish, French, German![Apache 2.0][license-apache-2.0]
IndexTTS2Chinese, English![Apache 2.0][license-apache-2.0]
Maya1English![Apache 2.0][license-apache-2.0]
Step-Audio-EditXMandarin, English, Sichuanese, Cantonese, Japanese, Korean![Apache 2.0][license-apache-2.0]
KaniTTSEnglish, German, Chinese, Korean, Arabic, Spanish![LFM][license-lfm]
VibeVoice-Finetuning![MIT][license-mit]
VoxCPMChinese, English![Apache 2.0][license-apache-2.0]
FireRedTTS2EN, ZH, JP, KO, FR, DE, RU![Apache 2.0][license-apache-2.0]
Audio Flamingo 3 (AF3) / Audio Flamingo NextMulti-lingual![Apache 2.0][license-apache-2.0]
ZipVoiceChinese, English![Apache 2.0][license-apache-2.0]
Fish Speech8![Apache 2.0][license-apache-2.0]
Chatterbox23+![MIT][license-mit]
Orpheus-TTSMultilingual![Apache 2.0][license-apache-2.0]
MegaTTS3Chinese, English![Apache 2.0][license-apache-2.0]
Spark-TTSChinese, English![Apache 2.0][license-apache-2.0]
Step-AudioChinese, English, Japanese![Apache 2.0][license-apache-2.0]
Kokoro-82M8![Apache 2.0][license-apache-2.0]
KokoClone7![Apache 2.0][license-apache-2.0]
LuxTTS-![Apache 2.0][license-apache-2.0]
MiMo-AudioMulti-lingual![Apache 2.0][license-apache-2.0]
SoulX-PodcastMandarin, English, Cantonese, Sichuanese, Henanese![Apache 2.0][license-apache-2.0]
VieNeu-TTSVietnamese![Apache 2.0][license-apache-2.0]
DiaEnglish![Apache 2.0][license-apache-2.0]
MeloTTSEnglish, Spanish, French, Chinese, Japanese, Korean![MIT][license-mit]
Kimi-AudioMulti-lingual![MIT][license-mit]
![Apache 2.0][license-apache-2.0]
eSpeak-NG100+![Other][license-other]
AuK

AuK

Description: AuK is a 1.5B foundation model from Tencent for speech generation and editing, trained on millions of hours of diverse audio data. Through a single natural-language instruction interface it unifies an unusually broad task set: zero-shot TTS (speak text in the reference voice) and instruct TTS (voice from a description alone, no reference), content editing (rewrite what is said; even lyric editing that preserves melody and voice), acoustic editing (pitch by semitones, speed, volume), paralinguistic editing (emotion, timbre, de-accent, nonverbal sounds, whisper conversion), and enhancement & separation (denoise/dereverberate, speech separation, music/vocal separation, target-speaker extraction). Architecture: a diffusion transformer with layer-fusion weights, conditioned by a Qwen2.5-Omni-3B MLLM encoder and a separate VAE (loaded at runtime). Day-0 SGLang-Omni serving support, Gradio and ComfyUI integrations, and a task Cookbook are provided. Released under MIT.

Release Date: September 9, 2026

FeatureValue
Voice Cloning
Asr
License![MIT][license-mit]
Parameters1.5B
Architecturediffusion transformer + layer fusion, Qwen2.5-Omni-3B MLLM encoder, separate VAE
VariantsAuK (this, base) + AuK-Flash (distilled, 4-step inference)
Editingcontent, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion
Enhancement Separationspeech enhancement, speech separation, music separation, target speaker extraction
DeploymentSGLang-Omni (day-0), Gradio, ComfyUI

Features: Unifies generation and the full editing/enhancement/separation spectrum in one instruction-following model — most systems pick one lane (TTS, or editing, or separation); AuK does zero-shot + instruct TTS, lyric rewriting with melody preservation, emotion/timbre/de-accent/whisper paralinguistic edits, and source separation through the same natural-language interface. A diffusion transformer with layer fusion, distilled into a 4-step AuK-Flash variant for fast inference.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]

Additional Tools:

ToolTypeLink
ComfyUI-AuKComfyUI nodeComfyUI-AuK

· · · · · · · · · · · · · ·

AuK-Flash

AuK-Flash

Description: AuK-Flash is the distilled variant of AuK, Tencent's 1.5B foundation model for speech generation and editing, optimized for fast 4-step inference. It exposes the same natural-language instruction interface as the base model: zero-shot TTS (reference voice) and instruct TTS (voice description, no reference), content and lyric editing, pitch/speed/volume acoustic edits, emotion/timbre/de-accent/nonverbal/whisper paralinguistic edits, plus speech enhancement and speech/music/target-speaker separation. Architecture matches the base: diffusion transformer with layer-fusion weights, Qwen2.5-Omni-3B MLLM encoder, and a separate runtime-loaded VAE. Released under MIT.

Release Date: September 9, 2026

FeatureValue
Voice Cloning
Asr
License![MIT][license-mit]
Parameters1.5B
Architecturediffusion transformer + layer fusion (distilled to 4 inference steps), Qwen2.5-Omni-3B MLLM encoder, separate VAE
Base Modeltencent/AuK
Editingcontent, lyric, pitch, speed, volume, emotion, timbre, de-accent, nonverbal, whisper conversion
Enhancement Separationspeech enhancement, speech separation, music separation, target speaker extraction
DeploymentSGLang-Omni, Gradio, ComfyUI

Features: Distills the AuK foundation model's diffusion transformer down to 4 inference steps, making the full generate-and-edit capability set (including enhancement and separation) practical for interactive use — traded against the base model's maximum quality, with the two variants loadable side-by-side from the same codebase.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo]

· · · · · · · · · · · · · ·

rumik-oss 1

rumik-oss 1

Description: rumik-oss 1 is a 3B multilingual text-to-speech model from rumik ai, trained on fewer than 70,000 hours of speech while performing competitively with existing TTS models. It covers 22 Indic languages in their native scripts and romanized forms plus English, supporting both single-language and code-switched synthesis. Delivery is conditioned via <description="..."> tags (tone, accent, pace), with inline vocalization control (<laugh>, <chuckle>, <sigh>). It extends CohereLabs/tiny-aya-fire with discrete speech tokens from the mimi codec: following the flattened codec-token formulation used in llama-mimi, text conditioning and audio generation share a single autoregressive sequence, predicting eight codebook tokens per frame before advancing, with the frozen mimi decoder reconstructing the 24 kHz waveform. The model ships with 4 fixed voices (Ira, Aisha, Siya, Zoya) that perform equally well across all 22 languages; there is no zero-shot voice cloning. Licensed under Cohere's CC-BY-NC-4.0 with acceptable-use addendum (research and non-commercial use only).

Release Date: September 6, 2026

FeatureValue
Voice Cloning
Asr
Languages22 Indic languages + English (native scripts and romanized; code-switching supported)
License![Other][license-other]
Parameters3B (3,381,533,697 BF16)
ArchitectureCohereLabs/tiny-aya-fire backbone + flattened mimi codec tokens (8 codebooks/frame, llama-mimi formulation)
Audio Codeckyutai/mimi (frozen decoder), 24 kHz output
Pronunciation
Highlightsdescription-conditioned delivery (<description> tags), inline vocalizations (<laugh>/<chuckle>/<sigh>)
Variantsrumik-oss-1 (post-trained), rumik-oss-1-base (speaker-conditioned pre-post-training)

Features: Brings competitive multilingual TTS to 22 Indic languages with under 70k training hours, using a flattened mimi codec-token formulation (single autoregressive sequence for text conditioning + audio) on the tiny-aya-fire backbone. Code-switched synthesis, description-conditioned delivery, and inline vocalization tags are first-class capabilities, and its 4 voices perform equally well across all 22 languages — unusual, as most TTS voices are language-specific.

Links: ![HuggingFace][link-huggingface] ![Blog][link-blog] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Irodori-TTS-v4.1-Anime

Irodori-TTS-v4.1-Anime

Description: Irodori-TTS-v4.1-Anime is a Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small using anime-style speech data. Because the base model's annotation pipeline is not publicly documented, the fine-tuning data was annotated independently — so caption conditioning and emoji controls may behave differently from the base model. The full-precision checkpoint (0.8B params, F32) ships at the repository root, with quantized variants (int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic) in subdirectories. It follows the base model's MIT License and ethical restrictions; inference uses the original Irodori-TTS repository.

Release Date: September 4, 2026

FeatureValue
Voice Cloning
Asr
LanguagesJapanese
License![MIT][license-mit]
Parameters~0.8B (766,052,385 F32)
ArchitectureIrodori-TTS (Aratako) fine-tune; caption-conditioned with emoji controls
Base ModelAratako/Irodori-TTS-v4.1-Small
Variantsint8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic

Features: A community fine-tune that ports the Irodori-TTS line into the anime-voice domain using an independently built annotation pipeline (since the base model's is undocumented), and ships the result with five ready-made quantization variants (int4/int8/fp8) for efficient inference.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

ICE-012 Audio

ICE-012 Audio

Description: ICE-012 Audio is a multilingual text-to-speech model from DarkPs (a FanuonAI organization) with streaming output and reference-based voice cloning. Its defining trait is breadth of language coverage — 590 language names/variants are accepted (name or 2–3-letter ID, with a language-agnostic fallback), including 13 Arabic dialects ("Lahgtna" variants) alongside the full ISO list. It introduces an active acoustic adapter — conditioning codec embeddings before the backbone and refining hidden states after it. Voice is controllable along six axes: gender (male/female), age (child → elderly), pitch (5 levels), accent (10 English accents), style (e.g. whisper), and speed (0.5–2.0×), plus an --auto-voice mode where the model picks a voice automatically. The checkpoint is ~714M parameters (F16) and runs via transformers with trust_remote_code=True. Released under CC BY-NC 4.0.

Release Date: August 29, 2026

FeatureValue
Voice Cloning
Asr
Languages590 names/variants (incl. 13 Arabic Lahgtna dialects; language-agnostic fallback)
Streaming
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
Parameters~714M (714,409,993 F16)
Architecturecausal LM with active acoustic adapter (conditions codec embeddings pre-backbone, refines hidden states post-backbone)
Voice Controlsgender, age, pitch, accent, style, speed; auto-voice mode

Features: The active acoustic adapter wraps the backbone on both sides — conditioning codec embeddings before it and refining hidden states after — while a six-axis voice-control space (gender/age/pitch/accent/style/speed plus auto-voice) and 590-language coverage make it one of the broadest single-checkpoint TTS releases for dialect and minority-language synthesis.

Links: ![HuggingFace][link-huggingface] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

TontaubeV1

TontaubeV1

Description: TontaubeV1 is a multilingual text-to-speech model from TontaubeAI (craitech) designed for expressive voice cloning, long-form generation, and low-latency streaming. Its release contains four causal codebook predictors: CB0 generates semantic audio and duration from text, while progressively smaller CB1–CB3 add acoustic detail. CB0 uses a Qwen3-1.7B-derived transformer trunk and CB1–CB3 progressively shallower Qwen3-0.6B-derived trunks, each with a two-layer audio-token head. The four output streams are decoded with DualCodec, and the inference path uses VibeVoice's acoustic encoder/decoder for continuous reconstruction and streaming. It ships bundled synthetic voices plus zero-shot cloning from up to 60 s of reference audio, with public speaking styles audiobook, conversational, and agentic. Released under the Tontaube Community Model License 1.0, which is explicitly not open-source.

Release Date: August 26, 2026

FeatureValue
Voice Cloning
Asr
Languages7 (English, German primary; Spanish, French, Italian, Dutch, Portuguese secondary)
Streaming
License![Other][license-other]
Parameters~2.87B (2,873,962,498; CB0 1.83B + CB1 449M + CB2 327M + CB3 269M)
Architecture4-stage Qwen3-derived codebook cascade (CB0–CB3) + DualCodec + VibeVoice decode
Stylesaudiobook, conversational, agentic

Features: The four-stage codebook cascade (CB0 semantic+duration → CB1–CB3 progressive acoustic refinement) lets a single multilingual model deliver expressive, long-form, low-latency speech with strong zero-shot cloning. On the 1,088 English zero-shot Seed-TTS examples it posts 1.66% mean utterance-level WER (measured with Whisper large-v3 at semantic temperature 0.6), and the RTX-5090 streaming path reaches ~200 ms to first encoded audio.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Breeze TTS 2

Breeze TTS 2

Description: Breeze TTS 2 is an open-weight text-to-speech model from BreezeBlue / RESONIA built for real-time interaction. It ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard while outperforming frontier proprietary systems. Its open-ended natural-language instruction-following supports reference-free voice design (create a voice from a text description) and reference-guided voice direction (clone a voice while steering tone, emotion, pace, delivery), alongside standard reference-audio voice cloning. Ultra-low-latency streaming reaches 0.32 RTF (≈3.1× real time with the warmed-up fast path) and under 40 ms time-to-first-audio on an NVIDIA H100, emitting 24 kHz PCM. Source code is Apache-2.0; model weights are governed by the BreezeBlue Research and Non-Commercial License (commercial use needs written authorization from RESONIA).

Release Date: August 25, 2026

FeatureValue
Voice Cloning
Asr
Languages2 (English, Chinese)
Streaming
License![Other][license-other]
Parameters3B (3,466,363,713)
Architectureseq2seq backbone + depth decoder + codec with CUDA-graph fast path (no named backbone disclosed)
Highlights#1 open-weight on Artificial Analysis TTS leaderboard; vocal events inline (laugh/cough)

Features: Pairs natural-language voice control with real-time streaming: a single model handles reference-free voice design (no reference audio needed) and voice direction (clone + steer prosody), and ships a CUDA-graph fast path that hits sub-40 ms TTFA at ~3.1× real time on H100 — open-weight quality that the authors claim exceeds frontier proprietary TTS.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]

Additional Tools:

ToolTypeLink
ComfyUI-Breeze-TTS-2ComfyUI nodeComfyUI-Breeze-TTS-2

· · · · · · · · · · · · · ·

Sopro v2 Turbo

Sopro v2 Turbo

Description: Sopro (Portuguese for "breath") is a lightweight voice-cloning text-to-speech family. This repo ships sopro-v2-turbo, a 120M-parameter open model that streams and runs comfortably on a laptop CPU or in the browser (ONNX runtime), reaching SOTA-level intelligibility against much larger systems. It supports zero-shot voice cloning from 5–20 s of reference audio, four languages (English, European Portuguese, French, German), and a streaming path with ~300 ms time-to-first-audio on a laptop CPU (0.24 RTF offline / 0.21 RTF streaming on an M3 CPU, 0.07 RTF on H100). Released under Apache-2.0.

Release Date: August 25, 2026

FeatureValue
Voice Cloning
Asr
Languages4 (English, European Portuguese, French, German)
Streaming
License![Apache 2.0][license-apache-2.0]
Parameters120M (121,574,193)
Deploymentin-browser ONNX runtime; int8 AR weights on CPU; causal vocoder
Architectureautoregressive TTS + chunked-attention streaming path + causal vocoder (F5-TTS/CosyVoice/Vocos lineage acknowledged)

Features: Packs SOTA-level intelligibility into a 120M footprint that runs in the browser or on a laptop CPU, with a chunked-attention + causal-vocoder streaming path (~300 ms TTFA) — making zero-shot multilingual voice cloning practical for on-device and edge deployment rather than GPU-only serving.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]

· · · · · · · · · · · · · ·

CuteTTS

CuteTTS

Description: CuteTTS is a lightweight (~230M-parameter) continuous autoregressive TTS model from OPPO that models continuous latents rather than discrete codec tokens, running efficiently on GPUs, CPUs, and Apple silicon. It delivers ultra-low latency — ~40 ms to the first audio chunk and ~9× real-time throughput on an RTX 4090 — with strong speech quality and zero-shot voice cloning (best-in-comparison 78.9 SIM on LibriSpeech test-clean). Multilingual support covers English, Chinese, French, German, and Spanish. A distilled variant (CuteTTS-distill) trades slight quality for further efficiency. Ships with a web demo, Python API, and CLI.

Release Date: August 24, 2026

FeatureValue
Voice Cloning
Asr
Languages5 (English, Chinese, French, German, Spanish)
Streaming
License![Apache 2.0][license-apache-2.0]
Parameters~230M
Architecturecontinuous autoregressive modeling of latents + speaker encoder + audio VAE (discrete-codec-free design)
VariantsCuteTTS, CuteTTS-distill

Features: Autoregressively models continuous latent audio representations instead of discrete codec tokens, eliminating codebook-related artifacts and quantization loss at only ~230M parameters. Combined with a lightweight speaker encoder and audio VAE, this yields best-of-class speaker similarity among compared open models and ~40 ms first-chunk latency while remaining practical for CPU/Apple-silicon inference.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Rynsan TTS

Rynsan TTS

Description: Rynsan TTS is a multilingual text-to-speech model that extends k2-fsa/OmniVoice to support Khasi, Garo, and Pnar — languages of Meghalaya, India that have historically had limited representation in modern speech technology. Rather than building a system from scratch, Rynsan retains the multilingual capabilities of the base model while adding speech data for these low-resource Khasic languages. Developed under the Tynrai AI initiative, its broader goal is accessible speech technology for the diverse languages and dialects of Meghalaya, supporting their preservation and use in voice-based applications. A live demo is available at ri.tynrai.in/demo. The repository is gated (manual access approval).

Release Date: August 22, 2026

FeatureValue
Asr
Languages5+ (English, Hindi + extension languages Khasi kha, Garo grt, Pnar pbv; base OmniVoice supports more)
License![CC BY 4.0][license-cc-by-4.0]
Parameters~0.61B
ArchitectureOmniVoice (k2-fsa) multilingual TTS, extended fine-tune
Base Modelk2-fsa/OmniVoice
DeveloperToiar / Tynrai AI

Features: Extends a modern multilingual TTS foundation to three substantially under-resourced Khasic languages — a rare production-oriented entry for indigenous-language speech tech, aimed at accessibility, education, and language preservation rather than benchmark leadership.

Links: ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Audio8 TTS Preview 0.1B

Audio8 TTS Preview 0.1B

Description: Audio8 TTS Preview 0.1B is the smallest release in the Audio8 TTS family ("the smallest zero-shot TTS worth running"): a ~170M-parameter generative model plus a separate ~120M-parameter codec decoder, making the complete audio generation stack much smaller than most modern multilingual TTS systems. It supports speech generation and zero-shot voice cloning (reference audio + matching transcript). Primary languages are Chinese and English, with German, Spanish, French, Italian, Japanese, and Korean as experimental/multilingual-evaluation targets. Released under the custom Audio8 Community License v1.0: non-commercial use is free, and commercial use is free only for entities with annual revenue under US$2M.

Release Date: August 19, 2026

FeatureValue
Voice Cloning
Asr
Languages8 (Chinese + English primary; de/es/fr/it/ja/ko experimental)
License![Other][license-other]
Parameters~0.17B main model (+ ~120M codec decoder)
ArchitectureAudio8 Falcon H1 DualAR — slow AR (semantic tokens) + fast AR (codec codebooks), 10 codebooks × 4096 entries
Audio Codecbundled 44.1 kHz neural codec (~21.5 frames/s)
Contextup to 2,048 packed text/audio positions
Variants0.1B (this), 0.6B

Features: Packs practical zero-shot cloning into a ~170M-parameter model using an Falcon-H1-derived DualAR design (slow AR predicts semantic tokens per frame; fast AR predicts the frame's 10 codebooks conditioned on the slow hidden state). On Seed-TTS it posts EN WER 1.662 at only ~0.17B — within reach of 4B+ systems — and ships with its own 44.1 kHz codec so no external codec checkpoint is needed.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

Kiseki-TTS

Kiseki-TTS

Description: Kiseki-TTS is a small, fast Japanese text-to-speech model from telecomadm1145 built on top of Qwen/Qwen3-TTS-Tokenizer-12Hz. It generates discrete neural audio codec tokens at 12.5 Hz (4–6× fewer autoregressive steps than 50–75 Hz codecs) and decodes them to waveform with the Qwen3 TTS codec. The acoustic decoder is a linear-time Mamba2 SSM rather than a self-attention stack, so generation cost is constant per frame — memory does not grow with utterance length and there is no KV cache to manage. Because TTS and ASR were trained jointly in a single multi-task run, the same checkpoint also performs ASR (Japanese speech → text, reading only codec layer 0). It is a single-domain voice (ASMR-style Japanese training data) with no speaker conditioning or voice cloning.

Release Date: August 15, 2026

FeatureValue
Voice Cloning
Asr
LanguagesJapanese only (ja)
License![MIT][license-mit]
Parameters~0.41B (0.33B backbone + 78M audio branch)
ArchitectureTransformer encoder (12 layers, bidirectional self-attention) + cross-attention → Mamba2 SSM decoder (6 layers, no causal self-attention)
Audio CodecQwen3-TTS-Tokenizer-12Hz (12.5 Hz, 16 quantizer layers)
Base ModelKiseki-1.1-0.3B (seq2seq translation model)
Training Datatelecomadm1145/asmr_archive_qwentts_encoded

Features: The decoder deliberately omits causal self-attention — temporal context is carried entirely by the Mamba2 recurrent state while text conditioning enters through cross-attention whose K/V are computed once during prefill. This yields O(1) state per frame (a fixed SSM tensor plus a 3-frame conv window) instead of an O(T) KV cache, so long-form synthesis degrades gracefully past the ~41 s training ceiling instead of hitting a memory cliff. Combined with the 12.5 Hz codec and a shared multi-token-prediction head that resolves all 16 codebook layers in one trunk pass, the model is both compute- and memory-bandwidth-bound rather than quadratic in length.

Links: ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

FireRedTTS3

FireRedTTS3

Description: FireRedTTS3 is a unified speech generation and editing system from the FireRed Team built on semantically enriched continuous speech representations. It ships in two variants: FireRedTTS3-Base (zero-shot voice cloning across 24 languages and 21 Chinese dialects) and FireRedTTS3-Instruct (natural-language voice design and combined semantic + acoustic speech editing in one model). Beyond cloning, it supports instruction-based voice design (no reference audio needed) and editing operations such as insertion / deletion / substitution (semantic) and speed / pitch / volume changes (acoustic).

Release Date: August 5, 2026

FeatureValue
Voice Cloning
Asr
Languages24 (plus 21 Chinese dialects)
License![Apache 2.0][license-apache-2.0]
ArchitectureQwen3 backbone + patch-level diffusion autoregressive (DiTAR) + RedAE codec + CAM++ speaker encoder
VariantsBase (cloning), Instruct (cloning + voice design + editing)

Features: Represents speech with semantically enriched continuous (non-quantized) representations, enabling a single system to do zero-shot cloning, text-driven voice design, and fine-grained semantic + acoustic editing. On Seed-TTS-eval it reaches an average WER/CER of 3.04% with 78.8% speaker similarity; MiniMax-MLS-Test average SIM 84.8%.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

Additional Tools:

ToolTypeLink
FireRedTTS3-ComfyUIComfyUI nodeFireRedTTS3-ComfyUI

· · · · · · · · · · · · · ·

Audio8-TTS-Preview-0.6b

Audio8-TTS-Preview-0.6b

Description: Audio8 TTS Preview 0.6B is a 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning. It uses a DualAR architecture inspired by Fish Audio S2 Pro: a slow AR transformer predicts one semantic token per audio frame, and a fast AR transformer predicts the frame's codec codebooks conditioned on the slow hidden state and preceding codebooks. The bundled 44.1 kHz neural audio codec handles both reference-audio encoding and waveform decoding — no additional codec checkpoint is required. The model supports 11 recommended languages (Cantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish) with zero-shot voice cloning from a reference audio clip + matching transcript.

Release Date: July 28, 2026

FeatureValue
Parameters601,159,424 (0.6B, excluding the codec)
Voice Cloning
Asr
LanguagesCantonese, Chinese, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Spanish (11)
Streaming
License![Apache 2.0][license-apache-2.0]
ArchitectureDualAR (slow AR + fast AR), inspired by Fish Audio S2 Pro
Slow Ar24 layers, width 896, 14 attention heads, 2 KV heads
Fast Ar4 layers, width 896, 14 attention heads, 2 KV heads
Acoustic Tokens10 codebooks, 4,096 entries per codebook
Codec44.1 kHz, 2,048 samples per model frame (~21.5 frames/s), bundled (no external codec needed)
Context Lengthup to 2,048 packed text/audio positions
Sample Rate44,100 Hz
Inferencetransformers with trust_remote_code=True; CUDA-capable GPU recommended
Dependenciestorch>=2.5.0, torchaudio>=2.5.0, transformers>=4.57.0,<5, soundfile, safetensors
Preview Statuslanguage coverage intentionally limited; broader multilingual + Chinese dialect support planned
Library Nametransformers (custom_code)
Pipeline Tagtext-to-speech
Createdat2026-07-28T07:53:00Z

Features: The DualAR architecture is the technical centerpiece: rather than a single autoregressive decoder predicting all codebook levels sequentially (the standard codec-LLM TTS pattern), Audio8 splits the work into a slow AR that predicts one semantic token per audio frame and a fast AR that predicts the frame's remaining codec codebooks conditioned on the slow hidden state. This separation lets the semantic-level reasoning happen at the slow AR's 24-layer depth while the acoustic codebook prediction stays lightweight at 4 layers — reducing the total compute per frame without sacrificing semantic quality. The bundled 44.1 kHz codec (no external codec checkpoint needed) and the 10-codebook / 4,096-entry acoustic token design give the model self-contained high-fidelity output at a compact 0.6B scale, making it one of the smallest multilingual zero-shot-cloning TTS systems shipping at 44.1 kHz.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

NeuTTS-2E

NeuTTS-2E

Description: NeuTTS-2E is a super-fast, highly realistic, on-device emotional text-to-speech model from Neuphonic. It is the next generation after NeuTTS Air / Nano (which continue to ship for multilingual + zero-shot-cloning contexts) — narrowed in scope to an English-only alpha focused on:

Release Date: July 21, 2026

FeatureValue
Parameters0.2B (compact LM backbone + codec)
Voice Cloning
Asr
LanguagesEnglish (English-only alpha)
Streaming
License![Other][license-other]
Backbonecompact LM backbone tuned for emotional TTS token generation
Codecefficient codec (compact, paired with the LM)
Speakers4 fixed (emily, paul, sophie, steven)
Emotions6 + neutral (angry, disgusted, fearful, happy, sad, surprised, neutral)
Emotion Control Modesingle-argument selection (no composable multi-axis axes like Scylla's Band)
Input Formattext only — no phonemizer, no system dependencies
On Deviceyes (laptop-class CPU real-time / better-than-real-time)
Distribution Formatssafetensors (torch), Q4 GGUF, Q8 GGUF
Formats In Collectionneuphonic/neutts-2e (safetensors), neuphonic/neutts-2e-q4-gguf (smallest footprint), neuphonic/neutts-2e-q8-gguf (mid-tier compression)
Gguf Featuresimatrix, conversational, endpoints_compatible
Pipeline Tagtext-to-speech
Library Name(HF tag does not declare transformers / safetensors stem beyond safetensors itself)
Downloads194 / 241 / 216 (torch / q4 / q8)
Intended Useembedded voice agents, on-device assistants, toys, privacy-sensitive applications
Comparison With Air NanoAir/Nano continue to ship for zero-shot cloning + multilingual contexts; 2E is the next-gen focused English emotional variant
Safety Notemodel is alpha; legitimate project landing is neuphonic.com (not neutts.com)

Features: The technical center of NeuTTS-2E is maximum speed per parameter at on-device budgets — the 0.2B LM + codec pair delivers real-time-or-better on laptop-class CPUs while exposing discrete categorical emotion control (angry / disgusted / fearful / happy / sad / surprised / neutral) plus a fixed four-speaker cast for consistency in agent / toy / accessibility voice personas. Two design choices distinguish it from the surrounding TTS field:

First, the categorical emotion surface is single-axis and discrete (one emotion per call), not the continuous multi-axis composable vector surface used by models like Scylla's Band ([neurotica base + 6-axis continuous strengths]). The project's positioning — production-grade agents + toys + accessibility — benefits from a one-argument API where emotion="happy" is the explicit operational state. The release locks emotional mode at generation time, which simplifies downstream filtering / guardrails.

Second, the distribution-shape design (one model, three deployment formats) is a deliberate on-device-first posture: the safetensors torch build for max-quality GPU/server; Q8 GGUF for mid-tier compression; Q4 GGUF for the small-footprint embedded target. All three are direct llama.cpp-compatible drops of the same model — no retraining-per-format — letting users pick size vs quality at deployment time without changing the production API. The combined CPU-first + GGUF-first design pattern is the opposite of the cloud-first TTS systems in this list — and is what makes 2E suitable for embedded voice agents, toys, and privacy-sensitive applications where audio + text must remain on-device.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Collection][link-collection]

· · · · · · · · · · · · · ·

Scylla's Band

Scylla's Band

Description: Scylla's Band is a multilingual, multi-voice, expressive TTS model from Spybyscript, designed specifically for local and self-hosted inference through ONNX Runtime (with an experimental LiteRT backend for explicit native / mobile use). The architecture is a continuous-latent TTS family:

Release Date: July 19, 2026

FeatureValue
Parametersnot stated (architecture: 4-layer duration predictor (192 hidden) + 12-layer rectified-flow acoustic generator (512 hidden, AdaLN, QK norm))
Voice Cloning
Asr
Languagesen_us, en_gb, es, it (4 public text-input languages)
Streaming
License![Apache 2.0][license-apache-2.0]
Sample Rate24,000 Hz
Managed Voices10 (ariadne, felix, gwen, ink, max, orpheus, rex, scylla, stone, tuesday)
Voice Default Localeen_us for most; ink / orpheus / tuesday default to en_gb
Voice Style Dim128 (style features) + 32 (prosody features)
Affect Axes6 (calm, joy, anger, sadness, sarcasm, questioning — all continuous in [0, 1])
Affect Overlay Axessarcasm, questioning (mixable with any core delivery)
Affect Cfg Scopeduration + acoustic-flow prediction (preserves voice / reference)
Encoders DefaultONNX Runtime (Python CLI / Python API / Android sample / libscyllasband native)
Encoders ExperimentalLiteRT (experimental / explicit-selection)
Cli Quality Default8-step Heun sampling
Graph Budgets512 G2P text tokens / 512 phone frames / 640 latent frames
Latent Target Buckets256 / 384 / 512 / 640 (smallest-fit selection)
VocoderScylla's Band acoustic adapter + frozen charactr/vocos-mel-24khz
Hop Lengths256 (waveform) / 512 (latent)
Text Frontendphrase-level multilingual G2P (74-phone vocabulary)
Span Context3 segments over up to 768 phones with 512-dim context state
Prefix Contextup to 24 acoustic latent frames from preceding chunk
Long Form Featuresboundary metadata + punctuation pause floors + prefix-latent carryover + span context
Group Speak Input[voice], [voice:language], [voice:language:axis=value,...] annotations
Bundle Contract1.0.0 / scyllasband-duration-flow
Intended Usesingle-voice speech synthesis (10 voices); en/es/it; long-form narration; multi-voice dialogue from tagged text; continuous affect control; ONNX desktop/server; ONNX + LiteRT native/mobile
Not Intendedarbitrary-speaker cloning / impersonation / fraud / deceptive speech
Distributionstraining data, trainer checkpoints, and export tooling not distributed
Cli Commandsdownload, validate-bundle, list-voices, normalize-text, speak, group-speak, stream, plan
Library Nameonnxruntime (tags include onnx, tflite, litert, duration-flow)

Features: Three design decisions distinguish Scylla's Band in the multilingual TTS class. First, decoupling duration and acoustic flow as separate rectified-flow stages — duration is a 192-hidden, 4-layer predictor operating on a 512-phone window, acoustic latents a 512-hidden, 12-layer AdaLN / QK-norm generator at 24-dim. This split lets affect-CFG act on both stages independently while retaining voice / reference conditioning, supporting the 6-axis continuous composability. Second, 6 affect axes (with sarcasm and questioning as overlays mixed with any core delivery) instead of mutually-exclusive discrete emotion classes — calm=0.5, joy=0.5 is a valid input, and axes stay in [0, 1] so multi-axis states are expressible without combinatorial blow-up. Third, the ONNX-first runtime design with libscyllasband native + an experimental LiteRT backend sits at a budget most neural TTS systems don't target — the 8-step Heun default and 512/512/640 fixed graph budget keep the model usable on CPU and mobile, and the inference-only release surface (training data + checkpoints not distributed) is the complement of the latency / mobile inference focus.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

sanoTTS

sanoTTS

Description: sanoTTS is the smallest known neural text-to-speech family. The name sano (सानो) is Nepali for "small". Each voice weighs 294k to 2.3M parameters — smaller than the smallest voice in prior families (TinyTTS at 1.62M; Inflect Nano at 4.63M; Kokoro at 82M) and the family fits in under 4 MB per voice with zero runtime dependencies (the espeak-ng phonemizer is bundled). Voices run real-time on a ~$3 ESP32-S3 microcontroller (output through a GPIO into an LM386 and a speaker) and live in the browser via WebAssembly — no server, no upload, no NPU. The full neural stack is duration → acoustic → decoder, quantized to int8, with the espeak-ng phonemizer included. 11 voices across 6 languages ship: English, Nepali, Hindi, Vietnamese, Indonesian, and Chinese — including the 294k heart-nano voice (337 KB) and the mel-based heart / heart-nano pair that predicts a 100-band spectrogram rendered by a noise-fed ConvNeXt + iSTFT decoder at 24 kHz. The inference runtime was relicensed MIT (September 2026); the project as a whole remains GPLv3 via espeak-ng. The project page at ampixa.github.io/sanoTTS hosts a live browser synthesis demo for every voice.

Release Date: July 13, 2026

FeatureValue
Parameters294k–2.3M per voice (smallest = the 294k "heart-nano" voice, 337 KB)
Voice Cloning
Asr
LanguagesEnglish, Nepali, Hindi, Vietnamese, Indonesian, Chinese (6 languages, 11 voices)
Streaming
License![Other][license-other]
Architecturefull neural stack — duration model → acoustic model → decoder
Quantizationint8 (W8/A12, corr 0.9995+; piperlite portable C99)
Runtime MicrocontrollerESP32-S3 (real-time RTF 0.41, GPIO → LM386 → speaker)
Runtime BrowserWebAssembly (no server, no upload, no NPU)
Runtime Footprintunder 4 MB per voice, zero dependencies
Voices11 (English: amy / kristin / hfc / amy-1p1m / amy-1p8m / robot / heart / heart-nano; one voice each for NE / VI / ID / ZH + shared lang voices)
Phonemizerespeak-ng (bundled)
License Splitinference runtime MIT; project overall GPL-3.0 (copyleft from espeak-ng)
Library Namesanotts
Training Methoddistillation (per voice)

Features: The hard constraint — be the smallest neural TTS family known, real-time on a $3 microcontroller — drives the entire stack. Conventional sub-100M TTS systems are too large for an ESP32's flash and RAM. sanoTTS keeps the full duration → acoustic → decoder neural pipeline (no espeak-NG-only fallback, no concatenative hybrid), quantizes everything to int8, and bundles the phonemizer so the whole voice ships in under 4 MB with zero runtime dependencies. The newest heart / heart-nano voices switch to a mel-based recipe (100-band spectrogram + noise-fed ConvNeXt + iSTFT decoder at 24 kHz), bringing the smallest voice down to 294k parameters / 337 KB — a per-voice footprint 100× smaller than Kokoro and 2× smaller than TinyTTS while still leading SCOREQ / UTMOS in the sub-15M class — and the demo synthesizes every voice live in the browser via WASM, so the smallest-known neural TTS is also the only one that runs unattended on a $3 chip and a $0 web page.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

FreyaTTS

FreyaTTS

Description: FreyaTTS is a 183M-parameter Turkish text-to-speech model. It is tokenizer-free at the character level — 92 symbols in its Turkish vocabulary — so there is no phonemizer or G2P step in either training or inference. Speech is generated with a non-autoregressive conditional flow-matching DiT in the frozen AudioVAE2 latent space (25 Hz, 64-dim latents, 16 kHz encode / 48 kHz decode). Training runs from scratch on Turkish speech: a pretraining stage followed by SFT stage 1/2 for voice lock and short-utterance coverage. Output is 48 kHz mono. On the project's Freya-TR-Eval benchmark the model reports WER 8.0% / CER 3.0%, ranking 3rd of 7 among open sub-1B Turkish TTS systems — a deliberate single-target-speaker, no-cloning design choice for a focused foundation release. The evaluation dataset is freya-tr-eval.

Release Date: July 7, 2026

FeatureValue
Parameters183.2M
Voice Cloning
Asr
LanguagesTurkish (tr)
Streaming
License![Apache 2.0][license-apache-2.0]
Architectureconditional flow-matching diffusion transformer (DiT), non-autoregressive, 32-step Euler ODE, no CFG
Tokenizercharacter-level (92 Turkish symbols; no phonemizer, no G2P)
Latent Spacefrozen AudioVAE2 (Apache-2.0, openbmb/VoxCPM2), 64-dim at 25 Hz
Codec Io16 kHz encode / 48 kHz decode
Sample Rate48,000 Hz
Trainingfrom scratch on Turkish speech; pretraining + SFT stage 1/2 (voice lock + short-utterance coverage)
EvaluationFreya-TR-Eval — WER 8.0% / CER 3.0%, 3rd of 7 open sub-1B Turkish TTS
Library Namefreyatts

Features: Two design choices are worth flagging. First, tokenizer-free character-level Turkish: by training directly on the 92-symbol Turkish alphabet with no phonemizer or G2P grapheme-to-phoneme step, the model removes a dependency that is fragile for agglutinative Turkish morphology and that often degrades quality when ported to low-resource Turkic relatives. Second, non-autoregressive conditional flow-matching in a frozen AudioVAE2 latent space: the 25 Hz / 64-dim bottleneck keeps the DiT small (183M) while inheriting a separately-trained audio codec's representation, letting a focused single-language-non-multilingual release ship at a fraction of the parameter budget of multilingual foundation TTS systems. The deliberate "no cloning, single target speaker" choice is a scope-lowering move that lets the foundation release put all its capacity into Turkish speech quality rather than spread it across zero-shot speaker adaptation.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Inflect-Nano-v2

Inflect-Nano-v2

Description: Inflect-Nano-v2 is a complete local text-to-waveform speech synthesis model with 3,966,721 deployable parameters — under 4M total. It is a VITS-architecture fixed-voice English TTS designed for CPU or CUDA inference with deterministic seeds, long-text handling, and 24 kHz mono output. The full FP32 checkpoint is 15.97 MB, making it one of the smallest complete neural TTS systems that produces natural-sounding speech without a separate vocoder or phonemizer dependency. The model ships with a public adaptation toolkit for preparing data, auditing train/validation splits, adapting a fixed voice or language, resuming training, evaluating checkpoints, and exporting PyTorch or ONNX packages. A sibling Inflect-Micro-v2 (9.36M parameters) prioritizes quality below 10M; Nano prioritizes footprint below 4M. Both share one public API.

Release Date: June 25, 2026

FeatureValue
Parameters3,966,721 (3.97M deployable)
Voice Cloning
Asr
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]
ArchitectureVITS (end-to-end text-to-waveform)
Sample Rate24,000 Hz
Footprint15.97 MB FP32
InferenceCPU or CUDA; PyTorch + ONNX export
Determinismdeterministic seeds for reproducible generation
Long Textautomatic text splitting and handling
Input Formattext (no phonemizer or system dependencies)
Adaptation Toolkitdata prep, split auditing, voice/language adaptation, training resume, checkpoint eval, PyTorch/ONNX export
Sibling ModelInflect-Micro-v2 (9.36M params, quality-prioritized below 10M)
Apione public API across Micro and Nano sizes
Librarypytorch
MetricsWER
Inference False On Hfyes (no hosted HF inference endpoint; local-only)

Features: Inflect-Nano-v2's defining constraint is completeness under 4M parameters: the entire text-to-waveform pipeline — no separate vocoder, no phonemizer, no system dependencies — fits in 3.97M deployable parameters and a 15.97 MB FP32 checkpoint. This is smaller than even sanoTTS's smallest voice (745k) when measured by complete-pipeline footprint, though sanoTTS ships per-voice weights rather than a single fixed-voice checkpoint. The VITS end-to-end architecture is the enabler: by folding the acoustic model and vocoder into a single jointly-trained network, Inflect avoids the multi-stage pipeline overhead that makes most neural TTS systems larger. The public adaptation toolkit extends the fixed-voice design into a customizable platform — users can prepare data, adapt a voice or language, resume training, and export PyTorch or ONNX packages — making the 4M-parameter footprint a starting point for domain-specific TTS rather than a dead-end fixed-voice release.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Gepard

Gepard

Description: GEnerative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue. Gepard is built for low-latency, high-throughput streaming conversation: the model starts speaking the moment text begins arriving, generating audio piece by piece instead of waiting for a full sentence. It is a single decoder-only autoregressive language model built on Qwen3.5 (14 layers, hidden 1024, 8 heads) with ≈556M total parameters (backbone + audio interface + voice-cloning compressor). Audio is produced through NVIDIA NeMo NanoCodec — Finite Scalar Quantization at 22.05 kHz, 21.5 frames/s, 1.89 kbps — with the full 32-channel FSQ frame sampled in one step. Reports ~25× real time on a single RTX 5090 with first-audio-chunk latency around 50 ms; a 96 GB Blackwell card serves up to 256 concurrent conversations. CFG refinement is baked into the weights so quality gain comes at no extra two-pass cost at inference, though the two-pass mode is still selectable as a quality dial.

Release Date: June 22, 2026

FeatureValue
Parameters~556M (555,694,169; Qwen3.5 backbone + audio interface + voice-cloning compressor)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish (US/UK), Spanish (es-MX), Portuguese (pt-BR), Dutch (NL)
Streaming
License![Apache 2.0][license-apache-2.0]
Audio CodecNVIDIA NeMo NanoCodec (FSQ, 22.05 kHz, 21.5 fps, 1.89 kbps; NVIDIA Open Model License)
Sample Rate22,050 Hz
BackboneQwen3.5 full-attention transformer (14 layers, hidden 1024, 8 heads; ~500M params)
InferencevLLM
Throughput256 conversations on one 96 GB Blackwell (RTX Pro 6000) GPU
BenchmarkSeed-TTS-eval leader on perceived quality (NISQA-MOS 4.25, NOI 4.16, COL 4.16, DIS 4.51) trading some WER/SIM

Features: A prosody-aware autoregressive single-pass frame generator: the whole 32-channel FSQ audio frame is sampled in one step (no depth transformer), and CFG quality refinement is baked into the weights rather than incurred at inference as a two-pass cost — so the publicly reported TTFA of ~50 ms and 25× real time on a single RTX 5090 represent the quality-on path, not a cheap-fast preview. Voice cloning is decoupled into a separate up-front compressor, which means cloning is "free" at run-time once the reference clip is encoded — a structural choice that supports serving hundreds of conversations per GPU. A stop-head weight update (2026-08-06) fixed premature stopping at sentence boundaries and lifted the effective duration ceiling; on Seed-TTS-eval Gepard leads the compared systems on perceived quality (NISQA-MOS 4.25) while trading some speaker similarity and WER for its streaming-first design.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Demo][link-demo] ![Paper][link-paper] ![Website][link-website]

· · · · · · · · · · · · · ·

Higgs Audio v3 TTS

Higgs Audio v3 TTS

Description: Boson AI's flagship conversational TTS: an ~4B autoregressive decoder over interleaved text and audio tokens from the Higgs Tokenizer (8 codebooks at 25 fps / 24 kHz). Built for voice chat rather than narration, it covers 102 languages with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

Release Date: June 4, 2026

FeatureValue
Parameters4B (BF16, 36 layers, hidden=2560, GQA 32/8)
ArchitectureAutoregressive decoder (Qwen3-style)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages102 (85 with WER/CER <5, 17 between 5-10)
Streaming
Audio Output24 kHz
License![Research Only][license-research-only]

Features: Interleaved text/audio token modelling with a delay-pattern multi-codebook embedding/head: a single autoregressive stack emits both modalities and supports inline <|category:value|> control tags (emotion/style/sfx/prosody) inserted at any point in the target text.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Blog][link-blog] ![Demo][link-demo]

· · · · · · · · · · · · · ·

dots.tts

dots.tts

Description: dots.tts is a 2B-parameter fully continuous, end-to-end autoregressive TTS system from Rednote-HiLab. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE, with no discrete tokens anywhere in the pipeline. It achieves the best average performance on Seed-TTS-Eval (WER 0.94 / 1.30 / 6.60 on zh / en / zh-hard) and the highest speaker similarity on a 24-language MiniMax multilingual benchmark, with broad cross-lingual voice cloning.

Release Date: June 3, 2026

FeatureValue
Parameters2B (semantic encoder + LLM + AR flow-matching acoustic head)
Voice Cloning
Asr
LanguagesMultilingual (24+ languages; zh / en focus)
Streaming
License![Apache 2.0][license-apache-2.0]
Sample Rate48 kHz
Tokenizer48 kHz AudioVAE (continuous, no discrete tokens)

Features: A fully continuous autoregressive pipeline that keeps generation in waveform-latent space end-to-end (no discrete-code phase), pairing an LLM-side semantic encoder with an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE — yielding SOTA seed-TTS-Eval scores and the strongest speaker-similarity number (83.9 avg) on the 24-language MiniMax multilingual benchmark.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Confucius4-TTS

Confucius4-TTS

Description: Confucius4-TTS is an LLM-based text-to-speech system from NetEase Youdao designed for multilingual and cross-lingual synthesis. It uses a speech encoder + LLM (Text2Semantic) + flow-matching Semantic2Acoustic architecture that allows zero-shot voice cloning without a required reference transcript and explicit cross-lingual voice transfer with unaccented output across languages. Covers Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese with code-switching and emotion transfer.

Release Date: June 2, 2026

FeatureValue
Voice Cloning
Asr
Emotion Control
Languages14 (zh, en, ja, ko, de, fr, es, id, vi, th, pt, it, ru, ms)
Streaming
License![Apache 2.0][license-apache-2.0]
Architecturespeech encoder + LLM (T2S) + flow-matching head (S2A)

Features: Cross-lingual voice transfer without accent drift: the same reference voice stays consistent when the speaker switches languages — backed by a speech encoder + LLM backbone pipeline (T2S) with a flow-matching acoustic decoder (S2A) and training that bundles 14 languages with code-switched, emotion-preserving decoding.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

WavTTS

WavTTS

Description: WavTTS is an end-to-end zero-shot TTS framework that synthesizes speech directly in the raw waveform space — explicitly skipping the intermediate mel-spectrogram, VAE-latent, or codec-token representations that most modern TTS stacks use. It is built on a flow-matching diffusion transformer (DiT) with waveform patchification, multi-scale mel-spectrogram supervision, and an optimized noise schedule. Forked from F5-TTS at the codebase level but replaces the whole acoustic pipeline.

Release Date: May 28, 2026

FeatureValue
Voice Cloning
Asr
LanguagesEnglish, Chinese
Streaming
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
![MIT][license-mit]
Sample Rate16 kHz
Training DataEmilia
ArchitectureFlow-matching DiT, raw waveform patchification, multi-scale mel supervision
Training Steps1.2M

Features: Skip every intermediate waveform representation (no mel, no VAE, no codec tokens): a flow-matching DiT produces raw-waveform patches directly, supervised at multiple mel scales and an optimized noise schedule — yielding high-quality zero-shot TTS at 16 kHz from a single end-to-end stack.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

MOSS-TTS

MOSS-TTS

Description: MOSS-TTS is a production-grade Text-to-Speech foundation model developed by the OpenMOSS Team and MOSI.AI. The current public v1.5 release preserves the original 1.0 capabilities — zero-shot voice cloning, long-form speech generation, token-level duration control, Pinyin/IPA pronunciation supervision, multilingual synthesis, and code-switching — and extends multilingual continued training from 20 languages to 31 languages including Cantonese, Dutch, Finnish, Hindi, Macedonian, Malay, Romanian, Swahili, Tagalog, Thai, and Vietnamese. v1.5 improves speaker similarity, reduces cloning variance on long-reference / short-text scenarios, follows punctuation-driven prosody more reliably, and adds explicit inline pause markers (e.g., [pause 3.2s]).

Release Date: May 25, 2026

FeatureValue
Parameters8B (Delay), 1.7B (Local)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages31 (extended from v1.0's 20)
Streaming
License![Apache 2.0][license-apache-2.0]
Max Duration1 hour
Pause Controlyes (inline markers like [pause 3.2s])
Lang Tag Controlyes (set language= in user message)

Features: v1.5 widens MOSS-TTS from 20 → 31 languages with stronger per-language multilingual synthesis (control via a language tag in the user message), more stable cloning under long-reference / short-text conditions, punctuation-driven prosody that holds up across long sentences, and explicit inline pause tokens ([pause 3.2s]) for scripted narration control.

Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Demo][link-demo]

· · · · · · · · · · · · · ·

VoxFlash-TTS

VoxFlash-TTS

Description: VoxFlash-TTS is a zero-shot voice-cloning text-to-speech engine built around extreme latent compression. The VAE encodes 24 kHz waveforms into a 9 frames/s latent space — roughly 8× more compressed than EnCodec (75 fps) and 2.4× more than Stable Audio (21.5 fps). Generating 10 s of audio therefore requires the diffusion model to produce just 90 latent vectors rather than hundreds or thousands of tokens, with downstream quadratic savings in attention cost. A ConvNeXtV2-based phoneme encoder followed by a novel coarse-alignment algorithm (cheaper than cross-attention) maps text into the latent sequence; a modern diffusion head then iteratively refines speech latents that the lightweight VAE decoder renders back to waveforms. The architecture targets low-latency, low-resource deployment — consumer-grade GPUs and edge devices — with Chinese and English zero-shot cloning. The project card lists inference: false on HF (no hosted inference endpoint), but the project page at voxflash.github.io carries the abstract, demo examples, and ablations.

Release Date: May 22, 2026

FeatureValue
Parametersnot stated (ConvNeXtV2 phoneme encoder + diffusion head + lightweight VAE decoder)
Voice Cloning
Asr
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]
Audio CodecVoxFlash VAE (9 Hz / 9 fps latent, 24 kHz input)
Compression Ratio~8× tighter than EnCodec (75 fps), ~2.4× tighter than Stable Audio (21.5 fps)
Phoneme EncoderConvNeXtV2 + coarse-alignment algorithm (no cross-attention)
Diffusion Headmodern multi-step iterative refinement
Decoderlightweight VAE decoder
Sample Rate24,000 Hz
Inferencelocal CUDA ≥ 12.3.2; no HF hosted endpoint
Training Datasetseed-tts-eval
Metricsword_error_rate, speaker_similarity

Features: The central technical move is compressing the audio latent space to 9 frames/s instead of the conventional 75 fps (EnCodec) or 21.5 fps (Stable Audio). This is not a quantization tweak — it is a temporal-decimation architectural choice that shrinks the sequence length the diffusion model has to traverse, and because attention cost scales quadratically with sequence length the end-to-end compute drops by orders of magnitude. Combined with a coarse-alignment phoneme-to-latent map that avoids cross-attention entirely (using a ConvNeXtV2 phoneme encoder instead), VoxFlash hits millisecond-level inference latency on consumer-grade and edge hardware for zero-shot Chinese + English cloning, where conventional latent-diffusion TTS systems are too slow for real-time edge deployment.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Miso TTS

Miso TTS

Description: Miso TTS 8B is a text-to-speech model from Miso Labs built on the Sesame Conversational Speech Model (CSM) architecture. A large Llama-3.2-style backbone consumes text/audio-frame embeddings and predicts codebook 0 of the Mimi audio token stream, while a smaller 300M autoregressive audio decoder predicts codebooks 1–31 in codebook depth. The model is designed for high-quality conversational speech and voice continuation from a short prompt audio clip.

Release Date: May 21, 2026

FeatureValue
Parameters8B (backbone llama-3.2-style) + 300M (audio decoder) = 8.3B
Voice Cloning
Asr
LanguagesEnglish
Streaming
License![MIT][license-mit]
ArchitectureSesame-style CSM (two transformer stack: backbone + audio decoder)
Audio TokenizerMimi (32 codebooks, vocab 2051, max seq 2048)
Librarypytorch

Features: A two-transformer Sesame-style CSM (Llama 8B backbone consumes text + audio frames and produces backbone codebook-0 prediction; a 300M audio decoder autoregresses over codebook depth via Mimi's 32-codebook stack) — letting the larger backbone spend capacity on linguistic / speaker conditioning while a leaner decoder handles fine-grained codebook-by-codebook generation.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Raon-OpenTTS-1B

Raon-OpenTTS-1B

Description: Raon-OpenTTS is an open-data, open-weight zero-shot TTS system from KRAFTON that performs on par with state-of-the-art closed-data models. This is the 1B variant (1048M parameters). Both model weights and training data are public: Raon-OpenTTS-Core is 510.1K hours of English speech, quality-filtered from the 615K-hour public Raon-OpenTTS-Pool using combined DNSMOS, WER, and VAD rank-based filtering. It ranks 1st or 2nd in WER and SIM among recent zero-shot TTS models on Seed-TTS-Eval and CV3-Eval, and achieves the best average WER/SIM on Raon-OpenTTS-Eval across Clean, Noisy, Wild, and Expressive regimes. A smaller Raon-OpenTTS-0.3B variant is also available.

Release Date: May 21, 2026

FeatureValue
Voice Cloning
Asr
LanguagesEnglish only (trained on 11 English speech datasets)
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
Parameters1048M
ArchitectureDiT (Diffusion Transformer) based on F5-TTS with flow matching; dim=1408, depth=28, heads=24
Audio Output80-ch mel-spectrogram at 16 kHz, HiFi-GAN vocoder (LibriTTS)
Training DataRaon-OpenTTS-Core (510.1K hours), 520K updates on 48× B200

Features: Demonstrates that fully open data + open weights can match proprietary SOTA: on Seed-TTS-Eval it reaches 1.78 WER / 0.749 SIM (vs Qwen3-TTS 1.46/0.715 at 1.7B), and best overall robustness (WER 2.81 / SIM 0.695) across four acoustic regimes on its own Raon-OpenTTS-Eval benchmark. The pipeline pairs large-scale rank-based data curation (DNSMOS + WER + VAD filtering of a 615K-hour pool) with an efficient F5-TTS-derived DiT.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv] ![Dataset][link-dataset]

Additional Tools:

ToolTypeLink
ComfyUI-Raon-OpenTTSComfyUI nodeComfyUI-Raon-OpenTTS

· · · · · · · · · · · · · ·

OronTTS

OronTTS

Description: OronTTS is a non-autoregressive text-to-speech model from btsee, an F5-TTS fork specialized for Mongolian (Khalkha Cyrillic) and Kazakh (Cyrillic). It uses Flow Matching + Diffusion Transformer + Vocos (dim 1024, depth 22, 16 heads, vocab 65, 24 kHz sample rate), trained on the btsee/mbspeech_mn corpus (3,846 Mongolian speech samples) and outputs zero-shot synthesis from a short reference audio + language tag.

Release Date: May 16, 2026

FeatureValue
Parameters(not stated)
Voice Cloning
Asr
LanguagesMongolian (Khalkha Cyrillic), Kazakh (Cyrillic)
Streaming
License![MIT][license-mit]
ArchitectureF5-TTS (OT-CFM + DiT + Vocos)
Dim1024
Depth22
Heads16
Vocab Size65
Sample Rate24000 Hz
Mel Bins100
Training Databtsee/mbspeech_mn (3,846 Mongolian speech samples)

Features: F5-TTS re-purposed for low-resource Cyrillic languages (Mongolian + Kazakh) — non-autoregressive flow-matching DiT over a tight 65-word vocab. Trained on a small (~3.8k sample) Mongolian corpus; the architecture is small enough that Khalkha Cyrillic and Kazakh Cyrillic share the same checkpoint via the lang tag at inference time.

Links: ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Supertonic 3

Supertonic 3

Description: Supertonic 3 is the third-generation open-weight release from Supertone. It is a lightweight, on-device text-to-speech system that runs with ONNX Runtime entirely on the user's machine (no network, no API call) and ships as a Python SDK (pip install supertonic). Compared with the Supertonic 2 base (5 languages, 66 M params), v3 expands to 31 languages and adds expression tags (<laugh>, <breath>, <sigh>), more stable reading on long utterances, and higher speaker similarity across the core language set.

Release Date: May 6, 2026

FeatureValue
Parameters(not stated on card; Supertonic 2 baseline 66 M — likely similar or smaller weight class)
Voice Cloning
Asr
Emotion Control
Languages31 (expanded from Supertonic 2's 5)
Streaming
License![OpenRAIL-M][license-openrail-m]
On Deviceyes (ONNX Runtime, no cloud call)
Expression Tagsyes (<laugh>, <breath>, <sigh>)

Features: A more compact on-device multilingual TTS: ONNX-Runtime inference everywhere, 31 languages from a single small open-weight encoder, and discrete expression tags that the decoder interprets inline — without a separate speaker-emotion control path or a cloud-rendered audio round-trip.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi]

· · · · · · · · · · · · · ·

Scenema Audio

Scenema Audio

Description: Scenema Audio is a zero-shot expressive voice cloning and speech generation model from ScenemaAI. It is built on an audio diffusion transformer extracted from the audio branch of Lightricks' LTX 2.3 (a 22B audiovisual model) — keeping the in-the-wild acoustic quality the bigger model learned while specializing for speech output. Generation is prompt-driven: a <speak> tag carries a voice description, gender, optional scene (ambient audio around the voice), and language; an <action> tag shifts emotional state mid-generation. Action tags cover rage, grief, joy, fear, exhaustion; voice prompt can describe timbre/pitch/breathiness/rasp/resonance plus character archetypes ("Tony Soprano having a breakdown"). Supports zero-shot voice cloning from 10-20 seconds of reference audio with some emotional variability, automatic long-form narration by splitting text and maintaining voice continuity, and 13 multilingual built-ins.

Release Date: April 26, 2026

FeatureValue
Parameters(audio diffusion transformer of LTX 2.3, weights ~9.8 GB bf16 / ~4.9 GB INT8 + ~6.7 GB pipeline)
Voice Cloning
Asr
Emotion Control
Languages13 (en, de, fr, es, it, pt, ja, zh, ko, ru, ar, hi, sw)
Streaming
License![Other][license-other]
Parent ModelLightricks LTX-2.3 (audio branch)
Prompt Format<speak voice=… gender=… scene=… language=…> XML with <action> tag for shifting emotion
Long Form Narrationyes (auto-splits text while preserving voice continuity)
Quantizedyes (INT8 weights at ~4.9 GB, identical quality)

Features: A standalone audio diffusion transformer extracted from a much bigger multimodal source: the model inherits how people actually sound in real scenes (angry, laughing, whispering, crying, exhausted, terrified) and exposes that capacity through a <speak> + <action> prompt interface — emotional state shifts within a single generation, instead of being a token-level or speaker-level conditioning problem.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Dramabox

Dramabox

Description: Dramabox is Resemble AI's expressive TTS, distributed under the LTX-2 Community License. It is an IC-LoRA fine-tune of the LTX-2.3 3.3B audio-only branch (Diffusion Transformer + flow matching), conditioned on Gemma 3 12B text embeddings. Generation is prompt-driven: speaker identity, emotion, delivery, laughs, sighs, breaths, pauses, and transitions are all expressed inside a natural-language description, with an optional 10-second voice reference that clones the target timbre.

Release Date: April 17, 2026

FeatureValue
Parameters3.3B (LTX-2.3 audio backbone, IC-LoRA fine-tune) + 12B Gemma 3 text encoder (conditioning only)
Voice Cloning
Asr
Emotion Control
LanguagesEnglish
Streaming
License![Other][license-other]
Base ModelLightricks/LTX-2.3 (audio branch)
ArchitectureDiT + flow matching, IC-LoRA fine-tune, Gemma 3 12B text embeddings
Inference Time~2.5 s / generation (warm server)

Features: IC-LoRA fine-tune of LTX-2.3's audio branch leaves the heavy text-understanding work to Gemma 3 12B and lets the DiT do the expressive rendering — so what's normally multimodal-stage orchestration collapses into a single prompt-driven TTS where speaker identity, emotion, and delivery are encoded in the prompt itself, and the timbre comes from a 10-second voice reference when present.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website]

· · · · · · · · · · · · · ·

Sarashina2.2-TTS

Sarashina2.2-TTS

Description: Sarashina2.2-TTS is a Japanese-centric text-to-speech system from SB Intuitions built on a large language model. It supports both Japanese and English, delivers strong pronunciation accuracy on Japanese text through large-scale end-to-end training, and reproduces a speaker's voice, speaking style, and acoustic characteristics from a short reference clip (zero-shot). Training data is sourced exclusively from legitimately acquired, properly licensed speech archives per the Sarashina Model NonCommercial License Agreement v2.0 (released April 24, 2026).

Release Date: April 16, 2026

FeatureValue
Voice Cloning
Asr
Emotion Control
LanguagesJapanese, English
Streaming
License![Research Only][license-research-only]
Base Modelsbintuitions/sarashina2.2-0.5b-instruct-v0.1
Cross Lingualyes (Japanese ↔ English, code switching)

Features: Japanese-optimized TTS fine-tuned on responsibly-licensed Japanese training corpora with explicit cross-lingual code-switching to English in a single utterance; reference audio carries speaking style and speaker identity together, so the same prompt yields narration, broadcast, conversation, or customer-service delivery without separate style conditioning.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]

· · · · · · · · · · · · · ·

LongCat-AudioDiT

LongCat-AudioDiT

Description: State-of-the-art diffusion-based TTS model operating directly in waveform latent space. Developed by Meituan's LongCat team, it requires only a Waveform VAE and Diffusion backbone, effectively mitigating compounding errors.

Release Date: March 30, 2026

FeatureValue
Parameters1B / 3.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
Audio Output24000 Hz
License![MIT][license-mit]

Features: Adaptive Projection Guidance (APG) replaces traditional classifier-free guidance for elevated generation quality. Outperforms Seed-TTS on zero-shot voice cloning benchmarks.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

SILMA TTS

SILMA TTS

Description: SILMA TTS v1 is a high-performance, 150M-parameter bilingual (Arabic/English) TTS model developed by SILMA AI. Built on the F5-TTS diffusion architecture, it was pretrained from scratch using tens of thousands of hours of high-quality public and proprietary data. It supports instant voice cloning with less than 8 seconds of reference audio (the reference transcript can also be left empty — it is transcribed on the fly), full support for Arabic Tashkeel (diacritics, auto-enriched via CATT when absent), NeMo-based text normalization, and an RTF around 0.12 on an RTX 4090. Released under a commercial-friendly license: code MIT, model weights Apache-2.0. The model is 100% compatible with F5-TTS v1.1.7 tooling for inference and fine-tuning.

Release Date: March 13, 2026

FeatureValue
Voice Cloning
Asr
Languages2 (Arabic MSA/Fusha + English)
License![Apache 2.0][license-apache-2.0]
Parameters150M
ArchitectureF5-TTS Diffusion Transformer with flow matching (pretrained from scratch, F5-TTS v1.1.7-compatible)
Pronunciation
CostRTF ≈ 0.12 (RTX 4090)

Features: Brings native-level Arabic synthesis to a 150M footprint: one of the smallest open F5-TTS-family models pretrained from scratch rather than fine-tuned, with first-class Arabic handling (Tashkeel-aware pronunciation via CATT enrichment, NeMo text normalization) alongside English, plus instant zero-shot cloning under fully permissive licensing (Apache-2.0 weights / MIT code).

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Fish Audio S2 Pro

Fish Audio S2 Pro

Description: Fish Audio S2 Pro is a leading text-to-speech model with fine-grained inline control of prosody and emotion. It combines reinforcement learning alignment with a dual-autoregressive architecture for high-quality speech synthesis.

Release Date: March 10, 2026

FeatureValue
Parameters~10 GB (BF16)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages80+ (Tier 1: En, Zh, Jp)
Streaming
License![Research Only][license-research-only]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

LongCat-Next

LongCat-Next

Description: Native multimodal foundation model by Meituan LongCat Team processing text, vision, and audio under a single autoregressive objective. Industrial-strength model with strong speech synthesis and voice cloning.

Release Date: March 2026

FeatureValue
Parameters3B (MoE A3B)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
Audio Output24 kHz
License![MIT][license-mit]

Features: Discrete Native Autoregression Paradigm (DiNA) unifying modalities in shared discrete token space. Combines visual understanding, generation, and audio processing in single model.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Voxtral-4B-TTS

Voxtral-4B-TTS

Description: Frontier, open-weights text-to-speech model developed by Mistral AI. Designed to be fast, instantly adaptable, and produces lifelike speech with natural prosody and emotional range.

Release Date: March 2026

FeatureValue
Parameters4B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages9 (En, Fr, Es, De, It, Pt, Nl, Ar, Hi)
Streaming
Audio Output24 kHz
License![CC BY-NC 4.0][license-cc-by-nc-4.0]

Links: ![HuggingFace][link-huggingface] ![Demo][link-demo] ![Blog][link-blog]

· · · · · · · · · · · · · ·

Blue (Light Blue) TTS

Blue (Light Blue) TTS

Description: BlueTTS (project page: lightbluetts.com) is a multilingual text-to-speech library. Built around slim ONNX graphs that run on ONNX Runtime with first-class CPU support and optional accelerators — OpenVINO (Intel), CUDA ORT (NVIDIA), TensorRT, and ONNX Runtime stock CPU. Targets five languages — Hebrew, English, Spanish, Italian, German — including inline mixed-language with XML-style tags in the text prompt. Inference is deliverable as a PyPI package (blue-onnx), a Rust crate, or directly from the pinned ONNX graphs on the HF Hub; the v2 release ships a slimmed opset-17 ONNX bundle (notmax123/blue-onnx-v2) that's intended for both FP32 production and the experimental INT8 weight-only fallback.

Release Date: February 27, 2026

FeatureValue
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesHebrew, English, Spanish, Italian, German
Streaming
License![MIT][license-mit]
RuntimeONNX Runtime (stock CPU; OpenVINO / CUDA / TensorRT optional)
Speed"fastest open-source TTS" (per project description)
Graph FormatONNX opset 17 (slim, full-precision; experimental weight-only INT8 fallback)
DistributionPyPI + HuggingFace + Rust

Features: A CPU-first multilingual TTS that ships both slimmed ONNX graphs and a Python package where the same code path runs on stock CPU ONNX Runtime by default — and optionally accelerates on OpenVINO / CUDA ORT / TensorRT — so deployment doesn't gate on GPU availability. Languages include Hebrew (with explicit G2P normalization) — a comparatively rare open-source TTS target — plus standard European languages, all from MIT-licensed weights distributed via both Hugging Face and PyPI.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![PyPI][link-pypi] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

KittenTTS

KittenTTS

Description: KittenTTS is an open-source realistic text-to-speech model designed for lightweight deployment. It is a state-of-the-art TTS model under 25MB with just 15 million parameters, running without GPU on any device.

Release Date: February 24, 2026 (v0.8.1)

FeatureValue
Parameters15M-80M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Multiple
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Ming-omni-tts

Ming-omni-tts

Description: Ming-omni-tts is a high-performance unified audio generation model in the Ming 2.0 series. It uses a custom 12.5 Hz continuous tokenizer and a Patch-by-Patch compression strategy that drives the LLM inference frame rate down to 3.1 Hz, enabling fine-grained control over speech rate, pitch, volume, emotion, and dialect (notably Cantonese at ~93 % accuracy). It supports 100+ premium built-in voices plus zero-shot voice design from natural-language prompts and is the first autoregressive model that jointly generates speech, ambient sound, and music in a single channel.

Release Date: February 11, 2026

FeatureValue
Parameters16.8B (3B active, MoE; A3B)
Voice Cloning
Asr
Emotion Control
LanguagesChinese, English, Cantonese
Streaming
License![Apache 2.0][license-apache-2.0]

Features: Patch-by-Patch compression drives the inference frame rate to 3.1 Hz, drastically cutting LLM-side latency for podcast-style audio while preserving naturalness. A custom 12.5 Hz continuous tokenizer plus a DiT head jointly produce speech, ambient sound, and music in a single output channel — an "in-the-scene" listening experience rather than TTS-on-top-of-a-track.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

SoulX-Singer

SoulX-Singer

Description: SoulX-Singer is a high-fidelity, zero-shot singing voice synthesis model for generating realistic singing voices for unseen singers without fine-tuning.

Release Date: February 6, 2026

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMandarin, English, Cantonese
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

SoproTTS

SoproTTS

Description: SoproTTS is a lightweight English text-to-speech model with zero-shot voice cloning. It uses dilated convolutions (WaveNet-style) and lightweight cross-attention layers instead of the common Transformer architecture.

Release Date: February 4, 2026 (v1.5)

FeatureValue
Parameters135M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]
Rtf0.05 (CPU M3)
Training-Cost~$100

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Qwen3-TTS

Qwen3-TTS

Description: Qwen3-TTS is an open-source series of Text-to-Speech models developed by Alibaba Cloud. Supports stable, expressive, and streaming speech generation with free-form voice design.

Release Date: January 22, 2026

FeatureValue
Parameters0.6B-1.7B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages10 (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian)
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

TADA

TADA

Description: TADA is a unified speech-language model from Hume AI built around a Text-Acoustic Dual-Alignment tokenizer: for every text/subword token there is exactly one corresponding speech vector, so the audio stream stays 1:1 aligned with text. As a TTS model, each autoregressive step covers one text token and dynamically determines the duration and prosody for that token, breaking the fixed-frames-per-second constraint that drives most modern TTS backbones. As a speech-language model, it generates a text token and the speech for the preceding token in the same dual step.

Release Date: January 12, 2026

FeatureValue
Parameters1B (Llama 3.2 1B base)
Voice Cloning
Asr
Emotion Control
LanguagesEnglish
Streaming
License![Other][license-other]
Base Modelmeta-llama/Llama-3.2-1B
Tokenization1:1 text–acoustic dual alignment (one speech vector per text token)
Dynamic Durationyes (each autoregressive step covers one text token, duration is determined per-token)

Features: A dual-alignment speech–text tokenizer that decouples autoregression from a fixed audio frame rate: each text token owns exactly one speech vector, and the model synthesizes the whole segment for that token in one step, regardless of how long the spoken form is — eliminating transcript hallucination and the latency overhead of constant-frame-rate codecs while staying as compact as Llama 3.2 1B.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![PyPI][link-pypi] ![Paper][link-paper] ![Blog][link-blog]

· · · · · · · · · · · · · ·

Irodori-TTS-500M-v2

Irodori-TTS-500M-v2

Description: Japanese Text-to-Speech model based on Rectified Flow Diffusion Transformer. Features emoji-based style and sound effect control by embedding emojis in input text for expressive speech generation.

Release Date: 2026

FeatureValue
Parameters500M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesJapanese
Streaming
Audio Output48kHz waveform
License![MIT][license-mit]

Features: Key Feature: Emoji annotation control - insert specific emojis into text to control speaking styles, emotions, and sound effects.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

KugelAudio

KugelAudio

Description: Open-source TTS for European languages with 7B parameters. Outperformed ElevenLabs in human preference testing.

Release Date: Early 2026

FeatureValue
Parameters7B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages23 European languages
Streaming
License![MIT][license-mit]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

LEMAS-TTS

LEMAS-TTS

Description: Part of the LEMAS (Large-scale Extensible Multilingual Audio Suite) project. Zero-shot multilingual TTS with 0.3B parameters supporting 10 languages with word-level precise editing capabilities.

Release Date: 2026

FeatureValue
Parameters0.3B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages10 (zh/en/de/fr/es/pt/it/ru/id/vi)
Streaming
License![Apache 2.0][license-apache-2.0]
Special-FeatureWord-level editing (LEMAS-Edit)

Features: Built on 150,000+ hours of multilingual speech data with word-level timestamps. Includes LEMAS-Edit for precise word-level speech editing via masked token infilling.

Links: ![Website][link-website] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

MioTTS-2.6B

MioTTS-2.6B

Description: Lightweight, high-speed LLM-based TTS model for English and Japanese with minimal resource usage.

Release Date: 2026

FeatureValue
Parameters2.6B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Japanese
Streaming
License![LFM][license-lfm]
Rtf0.135-0.145

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

MOSS-TTS-Nano

MOSS-TTS-Nano

Description: Ultra-lightweight open-source multilingual speech generation model with only 0.1B parameters. Designed for realtime speech generation that runs directly on CPU without GPU.

Release Date: 2026

FeatureValue
Parameters0.1B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages20
Streaming
Audio Output48 kHz Stereo
License![Apache 2.0][license-apache-2.0]

Features: Pure autoregressive architecture with MOSS-Audio-Tokenizer-Nano. Compresses audio to 12.5 Hz token stream using RVQ with 16 codebooks. Runs on 4-core CPU.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

NeuTTS

NeuTTS

Description: NeuTTS is a collection of open-source on-device TTS models with instant voice cloning. Built off LLM backbones with GGUF format quantizations for efficient on-device deployment.

Release Date: Early 2026

FeatureValue
Parameters360M (Air), 120M (Nano)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Spanish, German, French
Streaming
License![Apache 2.0][license-apache-2.0]
On-Deviceyes (GGUF quantizations)

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

OmniVoice

OmniVoice

Description: Massive multilingual zero-shot TTS model scaling to 600+ languages. Uses diffusion language model-style discrete non-autoregressive architecture with single-stage text-to-acoustic mapping.

Release Date: 2026

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages600+
Streaming
License![Apache 2.0][license-apache-2.0]
Training-Data581k hours

Features: Simplified single-stage architecture vs conventional two-stage pipelines. Full-codebook random masking strategy with LLM initialization for superior intelligibility. Noise-robust prompt processing.

Links: ![Website][link-website] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

T5Gemma-TTS

T5Gemma-TTS

Description: Multilingual TTS model with voice cloning and duration control, built on the T5Gemma encoder-decoder LLM architecture. Supports batch generation for multiple audio variations.

Release Date: 2026

FeatureValue
Parameters2B-2B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish, Chinese, Japanese
Streaming
License![MIT][license-mit]
Vram7.6-10.6 GB

Features: PM-RoPE positional encoding with XCodec2 audio codec. Low-VRAM options with CPU offloading. Batch inference efficiency with single encoder pass.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

TinyTTS

TinyTTS

Description: The smallest English TTS model with only 1.6 million parameters. End-to-end neural network achieving ~53x real-time synthesis speed on CPU via ONNX optimization.

Release Date: 2026

FeatureValue
Parameters~3.4 MB (ONNX FP16)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]

Features: Ultra-compact architecture optimized for CPU-only deployment. Multi-platform support via Python and Node.js APIs. Works on laptops, edge devices, and embedded systems.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

VoxCPM2

VoxCPM2

Description: OpenBMB's next-generation tokenizer-free diffusion autoregressive TTS model with 2 billion parameters. Supports 30 languages with automatic detection, voice design from text descriptions, and high-fidelity voice cloning.

Release Date: 2026

FeatureValue
Parameters2B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages30 (+ 9 Chinese dialects)
Streaming
Audio Output48 kHz
License![Apache 2.0][license-apache-2.0]

Features: Tokenizer-free design with LocEnc → TSLM → RALM → LocDiT pipeline. Built-in super-resolution via AudioVAE V2 for 48kHz output.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Soprano

Soprano

Description: Soprano is an ultra-lightweight, on-device text-to-speech (TTS) model designed for expressive, high-fidelity speech synthesis at unprecedented speed. The 1.1 release ships an 80M-parameter backbone that achieves up to 20× real-time generation on CPU and 2000× real-time on GPU, with lossless streaming (<250 ms latency on CPU, <15 ms on GPU), <1 GB memory usage at inference, and infinite generation length (automatic text splitting). Output sample rate is 32 kHz, with widespread device support (CUDA / CPU / MPS on Windows, Linux, and Mac). Inference is production-ready through an OpenAI-compatible endpoint, ONNX, WebUI, CLI, and ComfyUI nodes. The base 1.1 model is ekwek/Soprano-1.1-80M on HuggingFace; a fine-tuning toolkit (soprano-factory) was released January 13 2026 alongside the 1.1 weights — the 1.1 release reports 95% fewer hallucinations and a 63% preference rate over 1.0 (Soprano-80M). A live demo runs on ekwek/Soprano-TTS HF Space.

Release Date: December 22, 2025

FeatureValue
Parameters80M (Soprano-1.1-80M)
Voice Cloning
Asr
LanguagesEnglish (US/UK family voices, per HF Space)
Streaming
License![Apache 2.0][license-apache-2.0]
Sample Rate32,000 Hz
Inference TargetsOpenAI-compatible endpoint, ONNX, WebUI, CLI, ComfyUI
Performance Cpuup to 20× real-time
Performance Gpuup to 2000× real-time
Memory<1 GB at inference
Text Lengthinfinite (automatic text splitting)
DevicesCUDA, CPU, MPS (Windows, Linux, Mac)
Training Toolkitsoprano-factory (https://github.com/ekwek1/soprano-factory)
History 1 1Soprano-1.1-80M released 2026-01-14 (95% fewer hallucinations; 63% preference over 1.0)
History 1 0Soprano-80M released 2025-12-22

Features: The defining trade-off of this release is extreme on-device efficiency at sub-100M scale: an 80M-parameter backbone hits <250 ms CPU latency and <15 ms GPU for lossless streaming while keeping inference within <1 GB of memory — well under the multi-billion-parameter budget that newer conversational TTS systems require. The release pairs the model with soprano-factory (open-source training/fine-tuning toolkit) so users can build their own voices on top of the same backbone, and one installation can drive OpenAI-compatible / ONNX / WebUI / CLI / ComfyUI inference. The 1.1 update is a measured iteration: 95% fewer hallucinations and a 63% preference over 1.0 at the same parameter budget, so the measurable quality jump ships with no added inference cost.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo]

· · · · · · · · · · · · · ·

GLM-TTS

GLM-TTS

Description: High-quality TTS synthesis system based on LLMs from ZhipuAI, supporting zero-shot voice cloning with Multi-Reward Reinforcement Learning.

Release Date: December 11, 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Echo-TTS

Echo-TTS

Description: Echo is a 2.4B-parameter diffusion-based diffusion transformer (DiT) text-to-speech model. It conditions on target text and up to two minutes of speaker reference audio, generates Fish Speech S1-DAC latents, and decodes to 44.1 kHz audio. Output length is up to 30 seconds per segment. The model is fast at single-sample generation: on one A100, generating 30 seconds of audio from a 120-second prompt takes ~1.45 seconds (RTF < 0.05) — substantially faster than frontier autoregressive approaches at similar quality. The architecture is a deliberate pivot from the author's prior autoregressive-in-DAC-space model Parakeet, which struggled with semantic-consistency retries and weak voice cloning; Echo's diffusion approach trades off real-time interactivity for fast, high-fidelity zero-shot voice cloning in offline synthesis. Trained via the TPU Research Cloud (TRC). Demo (preview) hosted on jordand/echo-tts-preview HF Space; base model on jordand/echo-tts-base.

Release Date: December 4, 2025

FeatureValue
Parameters2.4B (DiT)
Voice Cloning
Asr
LanguagesEnglish (per demo samples)
Streaming
License![MIT][license-mit]
Architecturediffusion transformer (DiT) in Fish Speech S1-DAC latent space
Max Segment Duration30 s
Sample Rate44,100 Hz
Speaker Reference Max120 s
Performance A100 Rt30 s output in ~1.45 s (RTF < 0.05)
Audio CodecFish Speech S1-DAC
Prior ModelParakeet (autoregressive in DAC space)
Training InfrastructureTPU Research Cloud (TRC)
Inference RequirementsCUDA-capable GPU with at least 8 GB VRAM; Python 3.10+
Samplereuler CFG with independent guidances for text (3.0) and speaker (8.0); 40 steps; sequence_length 640 default
License ClarificationMIT (per GH repo license)

Features: Echo is a deliberate next-step pivot from autoregressive-in-DAC-space TTS to a full diffusion approach. The author's prior model, Parakeet, generated DAC tokens autoregressively but suffered the classic AR weakness — semantic-consistency retries — and weak voice cloning. Echo keeps Fish Speech S1-DAC latents (so the audio representation is the same proven codec) but moves the generator upstream to a 2.4B DiT operating directly on those latents, conditioned on a long (up to 2-minute) speaker reference. The result: 30-second outputs in ~1.45 s on a single A100 (RTF < 0.05) with high-fidelity zero-shot cloning — fast enough that the "diffusion is too slow" objection no longer applies at the segment length that matters for offline content generation, while the AR class's retry-induced inconsistency is gone by construction.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Blog][link-blog]

· · · · · · · · · · · · · ·

VibeVoice-Realtime

VibeVoice-Realtime

Description: Real-time TTS model from Microsoft with streaming text input and ultra-low latency (~300ms).

Release Date: December 3, 2025

FeatureValue
Parameters0.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMultilingual
Streaming
License![MIT][license-mit]
Max Duration~10 minutes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Fun-CosyVoice 3.0

Fun-CosyVoice 3.0

Description: Advanced TTS system based on LLMs for zero-shot multilingual speech synthesis from FunAudioLLM.

Release Date: December 2025

FeatureValue
Parameters0.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages9 + 18+ Chinese dialects
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

LFM2-Audio-1.5B

LFM2-Audio-1.5B

Description: Liquid AI's first end-to-end audio foundation model with low latency and real-time conversation.

Release Date: November 28, 2025

FeatureValue
Parameters1.5B
Voice Cloning
Asr
Emotion Control
LanguagesEnglish
Streaming
License![LFM][license-lfm]

Links: ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

Marvis-TTS

Marvis-TTS

Description: Marvis is a conversational real-time streaming TTS from Marvis-Labs. The architecture inherits Sesame's CSM-1B (Conversational Speech Model): a 250M-parameter multimodal backbone that processes interleaved text + audio tokens and a smaller 60M-parameter audio decoder that models the remaining 31 RVQ codebook levels to reconstruct high-quality speech from the backbone's representations. Audio tokens come from Kyutai's mimi codec (RVQ tokens). The dual-transformer split — semantic backbone + small decoder — yields sub-second latency, and the model is built for on-edge / on-device deployment (Apple Silicon / iPad / iPhone / Mac). Two operational choices distinguish Marvis:

Release Date: November 6, 2025

FeatureValue
Parameters250M (multimodal backbone) + 60M (audio decoder) = 310M total
Voice Cloning
Asr
LanguagesEnglish, French, German
Streaming
License![Apache 2.0][license-apache-2.0]
Architecturedual-transformer CSM-1B (Conversational Speech Model) — multimodal backbone + audio decoder
Audio CodecKyutai mimi codec (RVQ tokens; backbone models codebook 0, decoder models codebook 1–31)
Quantized Size~500 MB (4-bit MLX)
Training Datasetamphion/Emilia-Dataset
Library Nametransformers, mlx, mlx-audio
Inference Climlx_audio.tts.generate --model Marvis-AI/marvis-tts-250m-v0.2 --stream --text "..." [--ref_audio ./x.wav]
Variants In Collection250m-v0.2, 250m-v0.2-MLX-{4bit,6bit,8bit}, 100m-v0.2 (+ MLX variants), 250m-v0.2-transformers
Emits Text Chunkingno (full-sequence contextual processing)

Features: Two operational choices make Marvis stand out among conversational TTS. First, no regex chunking: most streaming TTS engines pre-split sentences by regex patterns before feeding them to the generator, which can disrupt flow / intonation; Marvis processes the entire text contextually, treating the text as a single interleaved multimodal sequence. Second, the dual-transformer CSM-1B design — a 250M backbone for codebook 0 (semantic) and a smaller 60M audio decoder for codebooks 1-31 (acoustic) — produces a quantized footprint of ~500 MB, enabling on-device Apple-Silicon inference (iPad / iPhone / Mac) with real-time streaming. The architecture makes a high-quality CSM-style TTS with zero-shot cloning actually deployable at the edge, while the official collection's 4 / 6 / 8-bit MLX variants let users trade footprint for fidelity on a per-device basis.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

IndexTTS2

IndexTTS2

Description: AI-Enhanced Text-to-Speech System with Intelligent Optimization and self-learning capabilities.

Release Date: November 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]
Multi-Speakeryes (1-4 speakers)

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Maya1

Maya1

Description: State-of-the-art speech model for expressive voice generation with natural language voice control.

Release Date: November 2025

FeatureValue
Parameters3B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish (Multi-accent)
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

Step-Audio-EditX

Step-Audio-EditX

Description: 3B-parameter LLM-based RL audio model specialized in expressive and iterative audio editing.

Release Date: November 2025

FeatureValue
Parameters3B (4B BF16)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMandarin, English, Sichuanese, Cantonese, Japanese, Korean
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

KaniTTS

KaniTTS

Description: KaniTTS is a 370M-parameter two-stage text-to-speech model from nineninesix-ai. The architecture pairs an LFM-2 backbone LLM (Liquid Foundation Model v2 — a non-transformer, structured state-space architecture) with a neural audio codec for output waveform synthesis. The LLM generates compressed audio-token representations and the codec renders them to 22 kHz waveforms, yielding low-latency generation: ~1 s to produce 15 s of audio on a single RTX 5080, with 2 GB GPU VRAM at inference, and MOS 4.3 / WER < 5% quality on the project's benchmarks. Languages covered: English, German, Chinese, Korean, Arabic, Spanish across multiple per-language voices (English, German, Chinese, Korean, Arabic, Spanish each ship a 400M checkpoint; Japanese ships a 370M "Expo-2025-Osaka" variant; a multilingual 370M checkpoint is also available). MLX variants exist for Apple-Silicon inference. The codec is the same author's nemo-nano-codec-22kHz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported to ~12.5 fps / 0.6 kbps). The model is part of the nineninesix-ai family alongside Gepard.

Release Date: September 30, 2025

FeatureValue
Parameters370M (kani-tts-370m multilingual); 400M per-language (en / de / zh / ko / ar / es); 370M (expo2025-osaka-ja)
Voice Cloning
Asr
LanguagesEnglish, German, Chinese, Korean, Arabic, Spanish (multilingual 370M checkpoint); Japanese (Expo-2025-Osaka variant)
Streaming
License![LFM][license-lfm]
Sample Rate22,000 Hz
Backbone LlmLFM-2 (Liquid Foundation Model; non-transformer structured state-space architecture)
Audio Codecnineninesix/nemo-nano-codec-22khz-0.6kbps-12.5fps-MLX (NVIDIA NeMo NanoCodec, MLX-ported)
Performance Rt 5080~1 s for 15 s audio on RTX 5080
Memory2 GB GPU VRAM at inference
Quality Mos4.3 / 5 (naturalness)
Quality Wer<5% (accuracy)
Training Dataset~80k hours (LibriTTS, Common Voice, Emilia)
Training Hardware8x H100 GPUs, 45 hours on Lambda AI
Per Language Modelskani-tts-400m-{en,zh,de,ar,es,ko} on HuggingFace
Pretrained Checkpoints0.2-pt (450M), 0.3-pt (400M) for custom posttraining / fine-tuning
Mlx Variantskani-tts-370m-MLX (Apple Silicon)
Arxiv2505.20506

Features: KaniTTS's design choice worth flagging: it pairs a non-transformer LFM-2 backbone (Liquid Foundation Model — structured state-space rather than attention) with a neural audio codec for output at the 370M scale. The choice lets the model hit a ~1 s / 15 s audio generation rate on a 2 GB GPU VRAM budget — sub-1B parameters, sub-entry-tier GPU requirement, but still multilingual across six languages. The two-stage approach (LLM → codec) is conventional; what's less conventional is the choice of a state-space backbone over the usual transformer decoder at this scale, hitting latency / VRAM numbers that open up sub-1B real-time TTS on consumer-grade hardware. The same author ships soprano-factory-style companion assets (pretrained v0.2-pt / v0.3-pt checkpoints, a NeMo NanoCodec MLX port) to lower the bar for fine-tuning on custom datasets.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

VibeVoice-Finetuning

VibeVoice-Finetuning

Description: This is an unofficial, work-in-progress LoRA fine-tuning toolkit for the VibeVoice TTS / speech model (1.5B-base and 7B-base checkpoints). The base VibeVoice checkpoints are the same ones covered in this list's a separate entry: audio-conditioned diffusion TTS for spoken dialogue, streaming, etc. This toolkit takes pretrained VibeVoice weights + a paired (text, audio, optional reference-audio) dataset and trains a LoRA adapter against two losses simultaneously:

Release Date: September 16, 2025

FeatureValue
Parameters1.5B (LoRA-adapted) / 7B (LoRA-adapted)
Voice Cloning
Asr
Languagesinherits base VibeVoice coverage
Streamingnot directly (toolkit output is a LoRA adapter; the adapter inherits VibeVoice inference shape)
License![MIT][license-mit]
Loss Textmasked cross-entropy on text tokens
Loss Acousticdiffusion MSE on acoustic latents
Hardware 1 5B≥16 GB VRAM
Hardware 7B≥48 GB VRAM
Transformers Version4.51.3 (known-good; other versions may break on Qwen2 architecture)
Tested Docker Imagerunpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04
Audio Target Format24 kHz audio (paired dataset of target-audio + transcripts + optional reference-audio prompts)
Training Entrypointpython -m src.finetune_vibevoice_lora --model_name_or_path aoi-ot/VibeVoice-Large --processor_name_or_path src/vibevoice/processor --dataset_name <your/dataset> --text_column_name text [--voice_column_name audio_ref]
Supports Hf Dataset Loaderyes
OutputLoRA adapter compatible with VibeVoice base

Features: The dual-loss trick is the technical center of this toolkit. Naive LoRA fine-tuning of a unified TTS model often specializes the synthesis but silently damages the text LLM capability the base inherited from its Qwen-class backbone; the "masked CE on text tokens + diffusion MSE on acoustic latents" two-headed loss preserves both competencies at training time. Pair that with the careful pinning of Transformers 4.51.3 (other versions break on the Qwen2 architecture) and a documented minimum-VRAM budget per base size (16 GB for 1.5B, 48 GB for 7B), and you get a reproducible recipe for community fine-tuning of VibeVoice — something the official Microsoft VibeVoice release doesn't ship out-of-the-box.

Links: ![GitHub][link-github]

· · · · · · · · · · · · · ·

VoxCPM

VoxCPM

Description: Tokenizer-free TTS system for context-aware speech generation and true-to-life voice cloning.

Release Date: September 16, 2025

FeatureValue
Parameters640M-800M
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

FireRedTTS2

FireRedTTS2

Description: Long-form streaming TTS system for multi-speaker dialogue generation with stable, natural speech.

Release Date: September 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEN, ZH, JP, KO, FR, DE, RU
Streaming
License![Apache 2.0][license-apache-2.0]
Multi-Speakeryes (4 speakers)
Max Duration3 minutes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Audio Flamingo 3 (AF3) / Audio Flamingo Next

Audio Flamingo 3 (AF3) / Audio Flamingo Next

Description: NVIDIA ADLR's fully open-source Large Audio Language Model with state-of-the-art audio understanding. Audio Flamingo Next (AF-Next) is the latest generation featuring stronger general audio understanding, longer context support, and timestamp-grounded reasoning.

Release Date: July 2025 (AF3), 2026 (AF-Next)

FeatureValue
Parameters7B
Voice Cloning
Asr
Emotion Control
LanguagesMulti-lingual
Streaming
License![Apache 2.0][license-apache-2.0]
ContextUp to 30 minutes

Features: Key Innovation (AF-Next): Staged curriculum training with GRPO-based RL post-training. Three specialized checkpoints: Instruct, Think (reasoning), and Captioner. Temporal Audio Chain-of-Thought grounding intermediate reasoning to timestamps.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Website][link-website]

· · · · · · · · · · · · · ·

ZipVoice

ZipVoice

Description: Fast and high-quality zero-shot TTS models based on flow matching.

Release Date: June 16, 2025

FeatureValue
Parameters123M
LanguagesChinese, English
License![Apache 2.0][license-apache-2.0]
Zero-Shot-Cloningyes
Dialogueyes

Links: ![GitHub][link-github] ![Website][link-website] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Fish Speech

Fish Speech

Description: State-of-the-art open source TTS and voice cloning model that generates natural, realistic, and emotionally rich speech.

Release Date: May 31, 2025 (v1.5.1)

FeatureValue
Parameters4B (S1), 0.5B (S1-mini)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages8 (EN, JP, KO, ZH, FR, DE, AR, ES)
Streaming
License![Apache 2.0][license-apache-2.0]
Rtf~1:7

Links: ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

Chatterbox

Chatterbox

Description: Family of SOTA open-source TTS models by Resemble AI, covering a single-language English line plus a multilingual V3 release that brings broader language coverage, more consistent speaker similarity, reduced hallucinations, and more natural conversational speech across 23+ languages.

Release Date: April 24, 2025

FeatureValue
Parameters500M (Llama backbone, 0.5B)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages23+
Streaming
License![MIT][license-mit]

Features: First open-source TTS model with explicit emotion exaggeration control, plus an alignment-informed inference pipeline and a watermarked decoder. Multilingual V3 narrows the quality gap to closed systems like ElevenLabs on cross-language voice cloning while staying under 1B parameters.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Orpheus-TTS

Orpheus-TTS

Description: SOTA open-source TTS built on Llama-3b backbone demonstrating emergent capabilities of LLMs for speech synthesis.

Release Date: April 2025

FeatureValue
Parameters3B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMultilingual
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![Website][link-website]

· · · · · · · · · · · · · ·

MegaTTS3

MegaTTS3

Description: Advanced zero-shot speech synthesis with Sparse Alignment Enhanced Latent Diffusion Transformer.

Release Date: March 22, 2025

FeatureValue
Parameters0.45B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Spark-TTS

Spark-TTS

Description: Efficient LLM-Based TTS Model with Single-Stream Decoupled Speech Tokens, built on Qwen2.5.

Release Date: March 2025

FeatureValue
Parameters0.5B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Step-Audio

Step-Audio

Description: Production-ready open-source framework for intelligent speech interaction with unified speech comprehension and generation.

Release Date: February 17, 2025

FeatureValue
Parameters130B (Chat), 3B (TTS)
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesChinese, English, Japanese
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Kokoro-82M

Kokoro-82M

Description: Kokoro is an open-weight Text-to-Speech model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.

Release Date: January 27, 2025 (v1.0)

FeatureValue
Parameters82M
ArchitectureStyleTTS 2, ISTFTNet
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages8 (54 voices)
Streaming
Cost<$0.06 per hour of audio
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

KokoClone

KokoClone

Description: KokoClone is a fast, real-time compatible multilingual voice cloning system built on top of Kokoro-ONNX. It enables users to type text in multiple languages, provide a short 3-10 second reference audio clip, and instantly generate speech in that same voice.

Release Date: 2025

FeatureValue
Parameters82M (Base: Kokoro-ONNX)
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages7 (En, Hi, Fr, Ja, Zh, It, Pt, Es)
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

LuxTTS

LuxTTS

Description: Lightweight ZipVoice-based TTS model for high quality voice cloning at speeds exceeding 150x realtime.

Release Date: 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
Languages-
Streaming
License![Apache 2.0][license-apache-2.0]
Rtf150x
Vram1GB

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

MiMo-Audio

MiMo-Audio

Description: Audio Language Model by Xiaomi functioning as a Few-Shot Learner with SOTA audio understanding.

Release Date: 2025

FeatureValue
Parameters7B
Voice Cloning
Asr
Emotion Control
LanguagesMulti-lingual
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

SoulX-Podcast

SoulX-Podcast

Description: SOTA Multi-Speaker TTS model for generating realistic long-form podcasts with dialectal diversity.

Release Date: 2025

FeatureValue
Parameters-
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesMandarin, English, Cantonese, Sichuanese, Henanese
Streaming
License![Apache 2.0][license-apache-2.0]
Max Duration90+ minutes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

VieNeu-TTS

VieNeu-TTS

Description: Advanced on-device Vietnamese TTS model with instant voice cloning from 3-5 seconds of reference audio.

Release Date: 2025

FeatureValue
Parameters0.3B-0.6B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesVietnamese
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

Dia

Dia

Description: 1.6B parameter TTS model by Nari Labs for generating ultra-realistic dialogue in one pass.

Release Date: June 27, 2024

FeatureValue
Parameters1.6B
Voice Cloning
Asr
Pronunciation
Emotion Control
LanguagesEnglish
Streaming
License![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

MeloTTS

MeloTTS

Description: MeloTTS is a high-quality multi-lingual text-to-speech library from MyShell.ai in collaboration with MIT, supporting English (American, British, Indian, Australian, and a default accent), Spanish, French, Chinese (with mixed Chinese–English capability), Japanese, and Korean. Built on VITS / VITS2 / Bert-VITS2 family work and packaged with both a Python API and a Web UI, it runs fast enough for CPU real-time inference.

Release Date: February 19, 2024

FeatureValue
Voice Cloning
Asr
LanguagesEnglish (American, British, Indian, Australian, Default), Spanish, French, Chinese, Japanese, Korean
Streaming
License![MIT][license-mit]
BaseVITS / VITS2 / Bert-VITS2 family
Mixed Chinese Englishyes

Features: A multi-accent multilingual TTS library that ships both a Python API and a Web UI on top of the VITS-style architecture, with explicit English-accent coverage (American, British, Indian, Australian, Default) and mixed Chinese–English output — designed for fast CPU real-time inference without requiring GPU servers.

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

Kimi-Audio

Kimi-Audio

Description: Open-source audio foundation model by Moonshot AI for audio understanding, generation, and conversation.

Release Date: 2024

FeatureValue
Parameters7B
Voice Cloning
Asr
Emotion Control
LanguagesMulti-lingual
Streaming
License![MIT][license-mit]
![Apache 2.0][license-apache-2.0]

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

eSpeak-NG

eSpeak-NG

Description: eSpeak NG is a compact open-source software text-to-speech synthesizer for Linux, Windows, Android, and other operating systems. It supports more than 100 languages and accents, is a fork of Jonathan Duddington's original eSpeak engine, and uses the formant synthesis method: the model produces speech by explicitly computing the acoustic resonances (formants) of each phoneme, not by concatenating human-speech recordings. The trade is well-known — the speech is clear and usable at high playback speeds, but not as natural or smooth as larger neural or concatenative synthesizers. The compensation is size: the program and its data, including many languages, total a few megabytes. Other synthesis methods supported: Klatt formant synthesis and MBROLA diphone back-end via the documented integration.

Release Date: December 8, 2015

FeatureValue
Parametersn/a (formant-synthesis engine; not a neural model)
Voice Cloning
Asr
Languages100+ languages and accents (see docs/languages.md)
Streaming
License![GPL 3.0][license-gpl-3.0]
Synthesis Methodformant synthesis (primary); Klatt formant synthesis (secondary); MBROLA diphone backend (optional)
Footprinta few MB (program + data + many languages)
Audio OutputWAV file (CLI), direct playback, or shared-library API
Input Formatstext from file / stdin (CLI), SSML (partial), HTML (partial)
PackagesCLI (espeak-ng man page), shared library (libespeak-ng), SAPI5 Windows module
SupersedeseSpeak (Jonathan Duddington's original engine)
Downstream UsageG2P / phonemizer for neural TTS pipelines (e.g. sanoTTS bundles espeak-ng for its duration model)
PlatformsLinux, Windows, Android, Solaris, Mac OS X

Features: eSpeak-NG is the canonical reference implementation of compact multi-language formant synthesis. Its 100+-language coverage in a few megabytes, plus SSML / SAPI5 / MBROLA / shared-library / CLI surfaces, are matched only by neural TTS systems that are orders of magnitude larger. The reason it belongs in a list whose other entries are neural TTS systems is its continued quiet role in the neural stack as a G2P / phonemizer front-end — the phoneme inventory and grapheme-to-phoneme rules that sanoTTS and similar sub-1B neural TTS engines bundle are often just a port of eSpeak-NG's language-data files. So even if the formant-synthesis audio output itself has been surpassed for naturalness, the phoneme infrastructure underneath many of the smaller neural TTS entries on this list still traces back to eSpeak-NG.

Links: ![GitHub][link-github]

· · · · · · · · · · · · · ·


Anything to Audio

Models that can generate audio from multiple input modalities (video, text, image, audio). These are unified frameworks for multimodal audio synthesis.

Anything to Audio Quick Comparison

ModelTextVideoAudioMax DurationSample RateLicense
MiDashengLM-Gen16 kHz![Apache 2.0][license-apache-2.0]
ScenA![Other][license-other]
Nemotron-Labs-Audex-2B![NVIDIA NC][license-nvidia-noncommercial]
Nemotron-Labs-Audex-30B-A3B![NVIDIA NC][license-nvidia-noncommercial]
MOSS-SoundEffect30 s48 kHz![Apache 2.0][license-apache-2.0]
Omni2Sound (Omni2Audio)![CC BY-NC 4.0][license-cc-by-nc-4.0]
ControlFoley44,100 Hz![CC BY-NC 4.0][license-cc-by-nc-4.0]
Woosh![Apache 2.0][license-apache-2.0]
Chroma-4B![Apache 2.0][license-apache-2.0]
Uni-MoE (Audio)![Apache 2.0][license-apache-2.0]
AudioX / Audio-Omni![Apache 2.0][license-apache-2.0]
![CC BY-NC 4.0][license-cc-by-nc-4.0]
HunyuanVideo-Foley48 kHz![Research Only][license-research-only]
PrismAudio![Apache 2.0][license-apache-2.0]
ThinkSound![Apache 2.0][license-apache-2.0]
MMAudio![Apache 2.0][license-apache-2.0]
MiDashengLM-Gen

MiDashengLM-Gen

Description: MiDashengLM-Gen (MiDasheng Language Model for Generation) is an end-to-end framework for unified audio-scene generation from Xiaomi. Built on a pre-trained LLM and the Dasheng audio tokenizer, it couples per-token conditional flow matching with autoregressive generation to produce coherent 16 kHz audio that simultaneously blends speech, music, sound effects and environmental acoustics from a structured text description. It supports 9 languages with emotion control and approaches dedicated TTS intelligibility on speech (Seed-TTS English WER drops from 12.15% to 2.79%) while retaining mixed-audio scene capability, and extends competitively to multilingual settings.

Release Date: August 12, 2026

FeatureValue
Text
Video
Image
Audio
Sample Rate16 kHz
Languages9
Emotion Control
License![Apache 2.0][license-apache-2.0]
Parameters1.7B (Qwen3-1.7B backbone)
ArchitectureDashengTokenizer (768-dim @25Hz) + Qwen3-1.7B + flow-matching DiT (16 layers, hidden 2048)

Features: LLM-conditioned high-dimensional (768-dim @25Hz) audio latents generated without quantization artifacts; audio-text alignment pre-training maps latents into the LLM token space before generation; a learned stop head enables variable-length truncation. First end-to-end trained model for general text-to-audio-scene generation.

Links: ![Demo][link-demo] ![HuggingFace][link-huggingface] ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

ScenA

ScenA

Description: ScenA generates multi-speaker audio scenes — dialogue and conversation with sound effects and ambience — from a text prompt, conditioned on one or more reference-audio clips that set the speakers' voices. Unlike prior multi-speaker dialogue systems it uses no per-turn tags, multi-stream transcripts, or speaker embeddings: a free-form natural-language prompt alone describes the scene. The text prompt determines which reference voice speaks where, allowing overlapping speech, spontaneous paralinguistic events, and scene-level ambient sound — all inherited from the in-the-wild text-to-audio pretraining distribution. The architecture is an audio-only, reference-conditioned flow-matching DiT built on the LTX-2 backbone (~4B parameters, 48 layers). Reference latents are concatenated into the token sequence and distinguished by lightweight identity-aware positional encodings. The training tackles a specifically identified "Reference Shortcut" failure mode — under standard noise schedules the model can identify the matching reference by noisy-target acoustic similarity, bypassing the text prompt — by using a high-noise-biased timestep distribution that forces reliance on the prompt for speaker assignment. Evaluator: CoVoMix2-Dialogue benchmark. Project page, code, paper, and HuggingFace checkpoint are linked below.

Release Date: July 7, 2026

FeatureValue
Parameters~4B (DiT, 48 layers; built on LTX-2 architecture)
Text
Video
Audio
Max Durationnot stated (scene-level generation)
Sample Rate(not stated; inherits LTX-2 audio VAE)
Voice Cloning
Multi Speakeryes
Ambient Soundyes (SFX, room acoustics, overlapping speech)
Architectureflow-matching DiT (LTX-2 backbone, audio-only)
Speaker Assignmentnatural language (no per-turn tags / identity encoders)
Training Fixhigh-noise-biased timestep distribution (defeats Reference Shortcut)
Text Encodergoogle/gemma-3-12b-it
Audio Vaebundled (~365 MB; encodes+decodes so full LTX-2 not needed)
Checkpoint Size~8.2 GB (scena.safetensors) + ~365 MB (audio_vae.safetensors)
License![Other][license-other]
Training Datain-the-wild text-to-audio pretrained, then reference-conditioned fine-tune
EvaluationCoVoMix2-Dialogue (speaker-binding metrics)

Features: The "Reference Shortcut" failure-mode identification is the technical center of the work: under standard diffusion noise schedules, a multi-speaker reference-conditioned model can match each reference to the noisy-target segment by acoustic similarity alone, bypassing the text prompt entirely. ScenA's high-noise-biased timestep distribution forces the model to rely on the prompt for speaker assignment at training time. Combined with the absence of any per-turn speaker structure (tags / transcripts / identity encoders) and the prompt's role as the only speaker-routing signal, this yields multi-speaker conversational scenes with overlapping speech, paralinguistic events, and ambient texture that previous structured-supervision multi-speaker systems filter out by design.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper]

· · · · · · · · · · · · · ·

Nemotron-Labs-Audex-2B

Nemotron-Labs-Audex-2B

Description: Nemotron-Labs-Audex-2B is NVIDIA's smaller sibling of the Audex unified audio-text LLM. Like the 30B-A3B flagship, the 2B is a single model family that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). It is built on the same audio-vocabulary-extended transformer stack as the 30B-A3B but at a densely-parameterized 2B scale (no MoE), so the compute and memory footprint are lowered to a budget tractable on more modest hardware. The 2B checkpoint is the project-tagged SFT variant in the Audex collection — instruction-tuned and ready for inference. Both sizes preserve text reasoning, alignment, knowledge, long-context, and agentic capabilities of the text backbone while adding discrete-token audio I/O.

Release Date: July 6, 2026

FeatureValue
Parameters2B (dense; SFT fine-tune, instruct + reasoning-ready)
Text
Video
Audio
Modalitiestext + audio (input and output)
Max Durationnot stated
Sample Ratenot stated (decoder output)
Voice Cloning
Audio Understandingyes (audio QA, classification)
Asr
Speech Translationyes
Text To Speechyes
Text To Audioyes
Speech To Speech Generationyes
Reasoning Modeyes (thinking + instruct modes inherited from text backbone)
License![NVIDIA NC][license-nvidia-noncommercial]
Pipeline Tagtext-generation
Library Nametransformers
Derived Fromsame family as Nemotron-Labs-Audex-30B-A3B
Companion 30Bnvidia/Nemotron-Labs-Audex-30B-A3B (MoE: 30B total, 3B active)
Spacesnvidia/Nemotron-Labs-Audex, WaveCut/Nemotron-Labs-Audex, hugging-apps/nemotron-labs-audex-2b
Createdat2026-07-06T16:21:07Z
Downloads~2.4k

Features: The 2B sibling matters because it preserves the central thesis of the Audex paper — unified audio-text LLM intelligence without regressing on text intelligence — while dropping the parameter budget substantially. The 30B-A3B MoE hits a 1M-context, agentic flagship tier; the 2B dense version is the same audio-aware architecture extended down to a budget that doesn't require a high-end MoE serving stack. The pair lets users choose on deployment cost rather than on capability sub-selection: the 2B ships the same audio-to-audio + text-to-audio + audio-understanding

  • ASR + speech-translation coverage as the MoE flagship, just at the cost of longer-context / reasoning depth that the MoE was specifically tuned for. Both share the discrete-audio-token vocabulary extension of the text backbone so they can be reasoned about interchangeably.

Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Nemotron-Labs-Audex-30B-A3B

Nemotron-Labs-Audex-30B-A3B

Description: Nemotron-Labs-Audex-30B-A3B is NVIDIA's unified audio-text LLM — a single model that both understands audio (audio QA, speech recognition, speech translation) and generates audio (text-to-speech, text-to-audio, speech-to-speech). Built on Nemotron-Cascade-2-30B-A3B (text-only MoE: 30B parameters, 3B active), Audex extends the vocabulary with discrete audio tokens for speech / general-audio output and adds an audio encoder for speech / general-audio input. Runs in thinking and instruct (non-thinking) modes and supports up to a 1M-token context length — preserving text-reasoning, alignment, knowledge, long-context, and agentic capabilities of the backbone while gaining audio tasks.

Release Date: July 6, 2026

FeatureValue
Parameters30B MoE (3B active)
Modalitiesaudio (input and output)
Audio Understandingyes
Asr
Speech Translationyes
Text To Speechyes
Text To Audioyes
Speech To Speech Generationyes
Voice Cloning
License![NVIDIA NC][license-nvidia-noncommercial]
LanguagesEnglish
Modesthinking, instruct (non-thinking)
Context Length1M tokens
TemplateChatML (with <think>…</think> for thinking mode)
InferencevLLM 0.20.0 (recommended) or transformers >= 4.53.0 (mamba-ssm + causal-conv1d required)

Features: First-class audio I/O for a 30B/3B-active text LLM: extended vocabulary with discrete audio tokens for outputting speech and general audio, plus an audio encoder for input — so the same backbone keeps its strong text reasoning (alignment, knowledge, long-context) and adds ASR + speech translation + TTS + audio generation + S2S without retraining. The MoE form (30B routes, 3B active) keeps inference tractable for a single pipeline that does both.

Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Collection][link-collection]

· · · · · · · · · · · · · ·

MOSS-SoundEffect

MOSS-SoundEffect

Description: MOSS-SoundEffect is the dedicated text-to-sound model in the OpenMOSS / MOSI.AI MOSS-TTS family. It turns natural-language captions into high-fidelity non-speech audio (ambience, urban scenes, creatures, human actions, and short music-like clips).

Release Date: May 25, 2026

FeatureValue
TypeText-to-Sound / SFX generation
ConditioningText
Max Duration30 seconds
Sample Rate48 kHz
License![Apache 2.0][license-apache-2.0]
ArchitectureDiT + Flow Matching + DAC VAE + Qwen3 text encoder
Parameters1.3B (DiT variant 1.3B)
LanguagesEnglish, Chinese
Inference Defaults100 flow-match steps, cfg 4.0, sigma_shift 5.0
Librarydiffusers

Features: Replaces the discrete-token autoregressive v1 (which bottlenecked on vocabulary) with a continuous-latent DiT + Flow Matching paired with a DAC VAE — yielding 30 s stable audio, bilingual English + Chinese prompts, and a clean CFG/sigma-shift inference schedule (cfg 4.0, shift 5.0) that works straight out of the box on the diffusers library.

Links: ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![GitHub][link-github]

· · · · · · · · · · · · · ·

Omni2Sound (Omni2Audio)

Omni2Sound (Omni2Audio)

Description: Omni2Sound — also written Omni2Audio on the project page — is a unified VT2A / V2A / T2A framework and a CVPR 2026 Highlight. A single Diffusion Transformer (DiT) backbone with a decoupled two-branch conditioning design:

Release Date: April 20, 2026

FeatureValue
ConditioningText / Video / Text+Video
ModalitiesVideo, Audio
Asr
Voice Cloning
Text
Video
Image
Audio
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
TasksVT2A, V2A, T2A (single model)
ArchitectureDiT + decoupled Semantic / Temporal branches + 3-stage progressive training
Pipeline Tagtext-to-audio

Features: One single model that is SOTA on three distinct tasks (VT2A, V2A, T2A) without a separate model per mode — decoupled semantic and temporal conditioning let the same DiT backbone handle text-only, video-only, and text+video conditioning by cleanly omitting the missing modality rather than padding it, which is what most prior unified VA models had to do.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Website][link-website] ![Paper][link-paper] ![Benchmark][link-benchmark]

· · · · · · · · · · · · · ·

ControlFoley

ControlFoley

Description: ControlFoley (Xiaomi MiLM Plus) is a unified controllable video-to-audio (foley) generation model. It supports four conditioning combinations under one architecture:

Release Date: April 13, 2026

FeatureValue
ConditioningText / Video / Text + Video / Video + Reference Audio
ModalitiesVideo (visual), Audio (foley)
Asr
Voice Cloning
Text
Video
Image
Audio
Sample Rate44,100 Hz
License![CC BY-NC 4.0][license-cc-by-nc-4.0]
Pipeline Tagtext-to-audio
Librarydiffusers
Cross Modal Conflicthandled via modality-specific control (no explicit router)
Inference SkillClawHub ControlFoley Audio Generator
UpcomingComfyUI nodes (in preparation, expanding to V2A / TV2A / TC-V2A / AC-V2A / T2A)

Features: Modality-specific cross-modal conflict resolution in a single generative stack: text governs semantics, reference audio governs timbre/acoustic style, and video governs temporal synchronization. Rather than routing to a single user-trusted modality, the model decouples control axes so an input disagreement (video shows a dog barking, text asks for a cat) is decomposed into a coherent output that respects each modality's responsibility. Trained with all-modality dropout for modality-robustness, ControlFoley is the first foley system that brings all four conditioning modes — T2A, V2A, TV2A, AC-V2A — under one model.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv] ![Skill][link-skill]

· · · · · · · · · · · · · ·

Woosh

Woosh

Description: Sony AI's sound effect foundation model for text-to-audio and video-to-audio generation. Includes Woosh-AE (audio encoder/decoder), Woosh-Flow/DFlow (T2A), and Woosh-VFlow/DVFlow (V2A) with distilled fast inference variants.

Release Date: 2026

FeatureValue
ArchitectureFlow-based generative models
Text
Video
License![Apache 2.0][license-apache-2.0]
Audio-Encodingyes
Fast-Inferenceyes (Distilled models)

Features: Optimized for sound effects (not general audio) with both public and private model versions. Video-conditioned generation without requiring captions. Competitive with Stable Audio Open and TangoFlux.

Links: ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

Chroma-4B

Chroma-4B

Description: Chroma 1.0 (FlashLabs' Chroma-4B on HuggingFace) is the first open-source, real-time, end-to-end spoken dialogue model that achieves both sub-second end-to-end latency and high-fidelity personalized voice cloning. The pipeline is end-to-end — no separate ASR → LLM → TTS stitch — speech goes in, speech comes out. The architectural centerpiece is an interleaved text-audio token schedule (1 text : 2 audio) that supports streaming generation, so the model can begin emitting audio while the user is still talking (broken-off turns / barge-in handled). Experimental results from the project's paper:

Release Date: November 28, 2025

FeatureValue
Parameters4B
Voice Cloning
Asr
LanguagesEnglish (per benchmark reporting)
Streaming
License![Apache 2.0][license-apache-2.0]
Architectureend-to-end spoken-dialogue LLM; interleaved text-audio token schedule (1:2); custom_code modules
Pipeline Tagany-to-any (HF classification)
Audio Tokenizationchroma tokenizer (RVQ-style per project's tag)
Latency Rtf0.43 (speech out ~2.3× wall-clock)
Speaker Similarity+10.96% relative improvement over human baseline
Inference Librarytransformers (custom_code)
Correlations With Larger Classmatches full dialogue turn at streaming latency
Pretrainedyes (safetensors weights)
Hf Space Demoshysts/Chroma-4B, Pnevka/Chroma-4B
Historypaper arXiv 2601.11141 (2026-01)

Features: Two bets together produce the dual property that no prior open-source spoken-dialogue model has hit simultaneously. First, an interleaved text-audio token schedule (1:2) — text tokens and audio tokens are interleaved at a fixed 1:2 ratio through the sequence, which gives the model a structured place to emit audio while still consuming user audio + text context, supporting sub-second end-to-end latency without a separate ASR / LLM / TTS pipeline. Second, personalized voice cloning baked into the spoke-dialogue model — the cloned voice is not bolted on top by a separate TTS stage (as is the default pattern), it's in-model at the audio-token-generation layer. The empirical payoff is a 10.96% relative speaker-similarity gain over the human baseline (i.e. the cloned voice is closer to the reference speaker than two of the same human speaker's recordings are to each other), while hitting RTF 0.43 — a floor that prior systems exceeded either in latency (no streaming) or in cloning fidelity (parrot the speaker poorly), rarely both.

Links: ![HuggingFace][link-huggingface] ![Paper][link-paper] ![Demo][link-demo]

· · · · · · · · · · · · · ·

Uni-MoE (Audio)

Uni-MoE (Audio)

Description: MoE-based omnimodal model with voice cloning, TTS, T2M (text-to-music), and V2M (video-to-music).

Release Date: October 16, 2025 (Uni-MoE-Audio)

FeatureValue
Parameters-
Voice Cloning
Text
Video
License![Apache 2.0][license-apache-2.0]
Dynamic-Routingyes

Links: ![GitHub][link-github] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

AudioX / Audio-Omni

AudioX / Audio-Omni

Description: Audio-Omni is the first end-to-end framework unifying understanding, generation, and editing across general sound, music, and speech domains. Presented at SIGGRAPH 2026. AudioX is a unified framework integrating text, video, image, and audio conditions.

Release Date: March 2025 (AudioX), 2026 (Audio-Omni)

FeatureValue
Parameters-
Text
Video
Audio
License![Apache 2.0][license-apache-2.0]
![CC BY-NC 4.0][license-cc-by-nc-4.0]

Features: First unified framework covering all three audio domains. Combines frozen multimodal LLM (Qwen2.5-Omni) with trainable Diffusion Transformer for high-fidelity synthesis. Any-to-any audio processing.

Links: ![GitHub][link-github] ![GitHub][link-github] ![HuggingFace][link-huggingface] ![HuggingFace][link-huggingface] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

HunyuanVideo-Foley

HunyuanVideo-Foley

Description: Tencent's end-to-end video sound effect generation model for professional-grade AI Foley sound generation. Analyzes footage and creates immersive audio that matches the visual content perfectly.

Release Date: 2025

FeatureValue
Parameters-
Sample Rate48 kHz
Text
Video
License![Research Only][license-research-only]
High-Quality-Foleyyes
Context-Awareyes

Links: ![GitHub][link-github] ![Demo][link-demo] ![Website][link-website] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

PrismAudio

PrismAudio

Description: Video-to-Audio generation framework with Reinforcement Learning and specialized Chain-of-Thought (CoT) planning. Decomposes reasoning into four specialized modules (Semantic, Temporal, Aesthetic, Spatial CoT) for comprehensive video understanding. Built upon ThinkSound.

Release Date: 2025 (ICLR 2026)

FeatureValue
Parameters518M
Video
License![Apache 2.0][license-apache-2.0]
Cot-Planningyes (4 modules)
Multi-Dimensional-Rlyes
Fast-Grpoyes (Hybrid ODE-SDE)
Inference-Time0.63 seconds

Features: Performance Benchmarks:

MetricVGGSoundAudioCanvas
Semantic (CLAP)0.470.52
Temporal (DeSync↓)0.410.36
Aesthetic (MOS-Q)4.21±0.354.12±0.28

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·

ThinkSound

ThinkSound

Description: Unified Any2Audio generation framework with flow matching guided by Chain-of-Thought (CoT) reasoning. Supports generating or editing audio from video, text, audio, or their combinations. Accepted to NeurIPS 2025.

Release Date: 2025

FeatureValue
Parameters-
Text
Audio
License![Research Only][license-research-only]
![Apache 2.0][license-apache-2.0]
Cot-Driven-Reasoningyes
Interactive-Object-Centric-Editingyes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo]

· · · · · · · · · · · · · ·

MMAudio

MMAudio

Description: Multimodal joint training framework for high-quality synchronized audio generation from video and/or text inputs. State-of-the-art open source model for generating sounds for videos, images, and text prompts.

Release Date: December 2024 (CVPR 2025)

FeatureValue
Parameters-
Text
Video
Image
License![Apache 2.0][license-apache-2.0]
Synchronized-Audioyes
Multimodal-Joint-Trainingyes

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface] ![Demo][link-demo] ![arXiv][link-arxiv]

· · · · · · · · · · · · · ·


Audio Restoration & Enhancement

Audio Restoration & Enhancement Quick Comparison

ModelTypeBandwidth ExtensionInpaintingLicense
RE-USEUniversal Speech Enhancement![NVIDIA NC][license-nvidia-noncommercial]
NovaSRAudio Super-Resolution![Apache 2.0][license-apache-2.0]
QuarkAudio-UniSEUniversal Speech Enhancement![Apache 2.0][license-apache-2.0]
PASESpeech Enhancement![Apache 2.0][license-apache-2.0]
DTT-BSRMusic Source Restoration![MIT][license-mit]
NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges)High-Resolution Audio Restoration![NVIDIA NC][license-nvidia-noncommercial]
ZipEnhancerAcoustic Noise Suppression![Apache 2.0][license-apache-2.0]
AudioSRAudio Super-Resolution![Apache 2.0][license-apache-2.0]
RE-USE

RE-USE

Description: RE-USE (RE-…), NVIDIA's multilingual universal speech enhancement model, targets distortion–perception trade-off by training a single model that balances listening quality against fidelity to the underlying linguistic / speaker / emotional content. Designed to restore diverse degraded speech while leaving everything else (content, identity, prosody, accent, paralinguistic attributes) intact.

Release Date: March 17, 2026

FeatureValue
TypeUniversal Speech Enhancement
Bandwidth Extension
Inpainting
Sample Rate8 / 16 / 22.05 / 24 / 32 / 44.1 / 48 kHz (multi-rate input)
ArchitectureMamba-SSM backbone
Degradation Coverageadditive noise, reverberation, clipping, bandwidth limit, codec artifacts, packet loss, low-quality mics
Language Agnosticyes
License![NVIDIA NC][license-nvidia-noncommercial]

Features: A single Mamba-SSM model that handles seven different input sample rates (no resampling pre-step), covers a broad degradation menu in one checkpoint, stays language-agnostic without per-language training, and explicitly balances distortion reduction against fidelity to the input speech — addressing the universal-SE trade-off that earlier single-purpose enhancers couldn't.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

NovaSR

NovaSR

Description: NovaSR is a tiny audio upsampler (~52 kB parameter count) that bandwidth-extends 16 kHz input up to 48 kHz. Public card on YatharthS/NovaSR advertises realtime factors around 3500× on A100, making it a candidate for real-time on-device super-resolution where model size dominates latency. Inference path is small enough to fit in CPU memory; the use case is speech-bandwidth extension without GPU.

Release Date: January 6, 2026

FeatureValue
TypeAudio Super-Resolution (16 kHz → 48 kHz)
Bandwidth Extension
Inpainting
Channelsmono
License![Apache 2.0][license-apache-2.0]
Parameters52 kB
Streamableyes (low VRAM / runs without GPU)
Realtime Factor~3500× (A100)

Features: A 52 kB-parameter Upsampler that hits ~3500× realtime on GPU and runs on CPU — pushing bandwidth extension below the size / latency envelope where a typical neural upsampler is unacceptable (real-time on-device speech enhancement).

Links: ![GitHub][link-github] ![HuggingFace][link-huggingface]

· · · · · · · · · · · · · ·

QuarkAudio-UniSE

QuarkAudio-UniSE

Description: UniSE is a unified, prompt-free autoregressive speech-enhancement framework built on a decoder-only language model. A single model performs multiple speech-enhancement tasks — speech restoration (SR / denoising), target-speaker extraction (TSE), source separation (SS), and acoustic echo cancellation (AEC, in development) — without explicit task-specific instructions or prompt conditioning; the language model infers the task from the input context. Stack: WavLM as the feature extractor, BiCodec as the discrete codec, and a decoder-only LM as the middle autoregressive backbone. Outputs reconstructed waveform from predicted discrete token sequences.

Release Date: December 22, 2025

FeatureValue
Voice Cloning
Asr
Streaming
LanguagesEnglish (paper demo)
License![Apache 2.0][license-apache-2.0]
TasksSpeech Restoration, Target Speaker Extraction, Source Separation, AEC (developing)
ArchitectureWavLM (feature extractor) + BiCodec (discrete codec) + decoder-only AR-LM
Unifiedyes (single model handles SE, SR, TSE, SS without explicit task prompts)
Prompt Freeyes (LM infers task from input context)
Dataset Signalsnoise + reverb + packet-loss + clean (configurable per task)
TrainingSpeech-enhancement SFT, then multitask joint training

Features: A single decoder-only LM that learns the speech-enhancement task distribution and infers which task to perform from the input context — eliminating the need for task-specific prompts, modules, or fine-tuning when switching between denoising, target-speaker extraction, and separation. Built as an autoregressive discrete-token predictor over a WavLM-extracted / BiCodec-quantised representation, it moves the speech-enhancement workflow from a zoo of specialist models into one generalist.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Paper][link-paper]

· · · · · · · · · · · · · ·

PASE

PASE

Description: PASE (Phonologically Anchored Speech Enhancer) is a generative speech-enhancement model from Cisco Collaboration AI that removes noise and reverberation while preserving linguistic content and speaker identity. It uses two fine-tuned WavLM-derived components:

Release Date: November 8, 2025

FeatureValue
TypeSpeech Enhancement
Bandwidth Extension
Inpainting
Sample Rate16 kHz mono
ArchitectureDenoising WavLM (DRD from WavLM-Large) + Dual-Stream Vocoder (phonetic + acoustic)
Finetuned FromWavLM-Large
Training DataDN5/DNS5 challenge clean + noise, LibriTTS, VCTK, OpenSLR26+28 RIRs
License![Apache 2.0][license-apache-2.0]

Features: Anchors enhancement to phonology instead of spectrum: by reconstructing from a phonetic stream and a separate acoustic stream (per DeWavLM's two representations), PASE keeps the words intact even when the spectrum is severely degraded — substantially lowering hallucinations while still regaining perceptual quality.

Links: ![HuggingFace][link-huggingface] ![GitHub][link-github] ![Demo][link-demo] ![Paper][link-paper]

· · · · · · · · · · · · · ·

DTT-BSR

DTT-BSR

Description: DTT-BSR (DTTNet with BandSequence and RoPE) is a music-source-restoration challenge submission from team AC/DC (Wuhan University) to ICASSP 2026. It is built inside the official MSR-Kit GAN framework, where the baseline generator is replaced by a DTTNet-style time-frequency U-Net and augmented at the bottleneck with:

Release Date: October 16, 2025

FeatureValue
TypeMusic Source Restoration
Bandwidth Extension
Inpainting
ArchitectureDTTNet TFC-TDF U-Net (complex STFT) + Improved Dual-Path BandSplitRNN block + RoPE-Transformer
Inputcomplex STFT (real + imag channels; n_fft=2048, hop=512)
DiscriminatorMulti-Frequency Discriminator (baseline)
FrameworkMSR-Kit GAN (reconstruction + adversarial + feature-matching losses)
License![MIT][license-mit]

Features: Treats music-source restoration as a complex-STFT time-frequency U-Net enhancement at the bottleneck: keep the strong DTTNet dual-path TFC-TDF structure for local spectral patterns, then layer in BandSplitRNN-style sub-band recurrence + RoPE self-attention so the generator can model long-range, cross-band harmonic structure that ordinary GAN baselines miss — critical for restoring non-vocal stems cleanly.

Links: ![GitHub][link-github]

· · · · · · · · · · · · · ·

NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges)

NVIDIA A2SB (Audio-to-Audio Schrodinger Bridges)

Description: A2SB is NVIDIA's audio-to-audio Schrödinger Bridge diffusion model for high-resolution (44.1 kHz) music restoration. It is the first long-audio restoration model that can restore hour-long inputs without boundary artifacts, and it's end-to-end — predicting waveform outputs directly withou

Truncated — view the full README on GitHub.

ai-music
ai-voice
asr
music-generation
tts
voice-cloning

Contributors

wildminder

37 commits