kadirnar/awesome-tts-architectures

A visual catalog of text-to-speech architectures, with model diagrams, concise descriptions, and primary sources.

Python

18

6 commits

updated Sep 13, 2026

See the code

README

Awesome TTS Architectures Awesome

A visual catalog of text-to-speech models, from Tacotron to speech language models. Diagrams, primary sources and short notes for every entry.

380 models and families · Reviewed 2026-09-13

Model list · All diagrams · 2025+ TTS-arxiv-daily collection · Descriptions · Timeline · Methodology · Contribute

The TTS-arxiv-daily collection covers the source list's TTS systems with first paper submissions from January 1, 2025 onward. Every included family has an image, a description, paper links and an explicit GitHub availability status. The complete screening record is available as JSON.

Models

T: text · S: speech or voice reference · A: other audio · I: image · V: video. Scope and labels.

Alphabetical model list · 380 entries
ModelGroupInput → output
A2TTSDiffusion / flowT, S → S
AffectronToken LMT, S → S
AgentSteerTTSToken LMT, S → S
AlignDiTDiffusion / flowT, S, V → S
AMNetParallelT → S
ARCHI-TTSDiffusion / flowT, S → S
ATRIEToken LMT → S
Audiobook-CCToken LMT, S → S
AuEmoChatToken LMT, S, V → S
AuKDiffusion / flowT, S → S, A
Authentic-DubberDiffusion / flowT, S, V → S
AutoSIFTDiffusion / flowT, S → S
AutoStyle-TTSToken LMT, S → S
AVLM (expressive speech)Token LMT, S, V → S
Bagpiper-TTSToken LMT → S
BareWaveDiffusion / flowT, S → S
BarkToken LMT → S, A
BASE TTSToken LMT, S → S
BatonTTS (BatonVoice)Token LMT, S → S
BELLEContinuous LMT, S → S
BitTTSCompactT → S
Block-wise Mimi TTSToken LMT → S
BnTTSToken LMT, S → S
BolboshDiffusion / flowT → S
Borderless Long Speech SynthesisContinuous LMT, S → S
BreezyVoiceToken LMT, S → S
BridgeTTSToken LMT, S → S
BVSToken LMT, V → S, A
CAM-TTSToken LMT, S → S
CapTalkToken LMT, S → S
CAST-TTSDiffusion / flowT, S → S
CaT-TTSToken LMT, S → S
Causal-prosody FastSpeech 2ParallelT → S
CDE-StyleTTSDiffusion / flowT, S → S
Chain-of-Details TTSToken LMT, S → S
Chain-TalkerToken LMT, S → S
ChatterboxToken LMT, S → S
Chatterbox-FlashToken LMT, S → S
ChatTTSToken LMT → S
CLaM-TTSToken LMT, S → S
CLEARContinuous LMT, S → S
Clip-TTSParallelT → S
Compact neural accessibility TTSCompactT → S
Compressed-to-fine speech LMToken LMT, S → S
Confucius4-TTSToken LMT, S → S
Continuous-token diffusion TTSContinuous LMT, S → S
Controllable masked-speech TTSToken LMT, S, A → S
CookVoiceDiffusion / flowT, S → S
CosyEdit2Token LMT, S → S
CoSyncDiTDiffusion / flowT, S, V → S
CosyVoiceToken LMT, S → S
CosyVoice 2Token LMT, S → S
CosyVoice 3Token LMT, S → S
CosyWhisper (WhispSynth)Token LMT, S → S
CoVoMix2Diffusion / flowT, S → S
Cross-Lingual F5-TTSDiffusion / flowT, S → S
CrossAccent-TTSToken LMT, S → S
CSMToken LMT, S → S
CTC-TTSToken LMT, S → S
CtrlSpeechContinuous LMT, S → S
CuteTTSContinuous LMT, S → S
DAIEN-TTSDiffusion / flowT, S, A → S
DARSDiffusion / flowT → S
DCARToken LMT, S → S
Deep VoiceAutoregressiveT → S
Deep Voice 2AutoregressiveT → S
Deep Voice 3AutoregressiveT → S
DeepASMRToken LMT, S → S
DeepDubber-V1Diffusion / flowT, V → S
DeepDubbingToken LMT, S → S
DelightfulTTSParallelT → S
DELTA-TTSToken LMT, S → S
DepFlowDiffusion / flowT, S → S
DiaToken LMT, S → S
Dia2Token LMT, S → S
DialoSpeechToken LMT, S → S
DiEmo-TTSParallelT, S → S
Diff-TTSDiffusion / flowT → S
DiffCSSToken LMT, S → S
DiFlow-TTSToken LMT, S → S
DiFlowDubberToken LMT, S, V → S
DisCo-SpeechToken LMT, S → S
DisSpeechToken LMT → S
DiSTARToken LMT, S → S
DiTARContinuous LMT, S → S
DiTTo-TTSDiffusion / flowT, S → S
DMOSpeech 2Diffusion / flowT, S → S
DMP-TTSDiffusion / flowT, S → S
dots.ttsContinuous LMT, S → S
Dragon-FMToken LMT, S → S
DrawSpeechDiffusion / flowT, I → S
DS-TTSDiffusion / flowT, S → S
DualDubToken LMT, V → S
DualSpeechLMToken LMT, S → S
E2 TTSDiffusion / flowT, S → S
ECTSpeechDiffusion / flowT, S → S
ELLA-VToken LMT, S → S
EME-TTSParallelT → S
EMM-TTSToken LMT, S → S
EmojiVoiceDiffusion / flowT → S
EmoShiftToken LMT → S
EmoSSLSphereToken LMT → S
EmoSteer-TTSDiffusion / flowT, S → S
Emotion-timbre disentangled TTSParallelT, S → S
EmotiVoiceParallelT → S
EmoTra-TTSToken LMT, S → S
EmoVoiceToken LMT → S
End-to-end discrete-token TTSToken LMT, S → S
F5-TTSDiffusion / flowT, S → S
F5R-TTSDiffusion / flowT, S → S
Face-adapted StyleTTS 2Diffusion / flowT, I → S
FaceSpeakDiffusion / flowT, I → S
FacialTalkerToken LMT, S, V → S
FastPitchParallelT → S
FastSpeechParallelT → S
FastSpeech 2ParallelT → S
FC-TTSToken LMT, S → S
FELLEContinuous LMT, S → S
FineCombo-TTSDiffusion / flowT, S → S
FireRedAudioContinuous LMT, S → S, A
FireRedTTSToken LMT, S → S
FireRedTTS-1SToken LMT, S → S
FireRedTTS-2Token LMT, S → S
FireRedTTS3Continuous LMT, S → S
Fish Audio S1 / OpenAudio S1Token LMT, S → S
Fish Audio S2Token LMT, S → S
Fish SpeechToken LMT, S → S
Flamed-TTSDiffusion / flowT, S → S
FlashTTSToken LMT, S → S
FleSpeechToken LMT, S, I → S
FlexiVoiceToken LMT, S → S
FlexSpeechDiffusion / flowT, S → S
FlowtronFlow / VAET, S → S
FNH-TTSFlow / VAET, S → S
Frame-stacked local Transformer TTSToken LMT, S → S
FreyaTTSDiffusion / flowT → S
Gemini 2.5 TTSAPIT → S
Gemini 3.1 Flash TTSAPIT → S
GibbsTTSToken LMT, S → S
GLM-TTSToken LMT, S → S
Glow-TTSFlow / VAET → S
GOAT-TTSToken LMT, S → S
GPAToken LMT, S → S
GPT-4o Mini TTSAPIT → S
GPT-SoVITSToken LMT, S → S
Grad-TTSDiffusion / flowT → S
GRAFTToken LMT, S → S
GSA-TTSParallelT, S → S
GST-TacotronAutoregressiveT, S → S
HabibiDiffusion / flowT, S → S
HD-PPTToken LMT, S → S
Higgs Audio v2Token LMT, S → S
Higgs Audio v2.5Token LMT, S → S
Higgs Audio v3 TTSToken LMT, S → S
HiStyleDiffusion / flowT → S
HoliDubberContinuous LMT, V → S
HoliTok (TTS)Continuous LMT, S → S
Hume Octave TTSAPIT, S → S
ImmersiveTTSDiffusion / flowT, S, A → S, A
IndexTTSToken LMT, S → S
IndexTTS 2.5Token LMT, S → S
IndexTTS2Token LMT, S → S
InstructAudioDiffusion / flowT → S
IntMeanFlowDiffusion / flowT, S → S
Inworld TTS-1Token LMT → S
JaiTTSContinuous LMT, S → S
JAM-FlowDiffusion / flowT, S, V → S
JELLYToken LMT, S → S
JETSParallelT → S
Joint non-autoregressive STT-TTSParallelT → S
JoycentDiffusion / flowT, S → S
JoyVoiceToken LMT, S → S
KABURI-TTSDiffusion / flowT → S
KittenTTSCompactT → S
Koel-TTSToken LMT, S → S
KokoroCompactT → S
Kyutai TTS (DSM)Token LMT, S → S
LanStyleTTSParallelT, S → S
LatinXToken LMT, S → S
LE2E-TTSCompactT → S
LightSpeechParallelT → S
LLaDA-TTSToken LMT, S → S
LlasaToken LMT, S → S
Llasa+Token LMT, S → S
LLMVoXToken LMT → S
Lombard Matcha-TTSDiffusion / flowT → S
LongCat-AudioDiTDiffusion / flowT, S → S
LoRP-TTSDiffusion / flowT, S → S
Luna-TTSToken LMT, S → S
M3-TTSDiffusion / flowT, S → S
MAGIC-TTSToken LMT, S → S
MagpieTTS-LFToken LMT, S → S
MambaVoiceCloningDiffusion / flowT, S → S
MamTraDiffusion / flowT, S → S
ManchuTTSDiffusion / flowT → S
Marco-VoiceToken LMT, S → S
MARS6Token LMT, S → S
Masked-style TTSToken LMT, S → S
MaskGCTToken LMT, S → S
Matcha-TTSDiffusion / flowT → S
MAVEToken LMT, S → S
Mega-TTSToken LMT, S → S
Mega-TTS 2Token LMT, S → S
MegaTTS 3Diffusion / flowT, S → S
Meitei Mayek TTSAutoregressiveT → S
Mel-LLM (TTS)Continuous LMT → S
MELA-TTSContinuous LMT, S → S
MELDToken LMT, S → S
MellotronAutoregressiveT, S → S
MeloTTSFlow / VAET → S
MetisToken LMT, S → S
MFCIG-CSSToken LMT, S, V → S
MiDashengLM-GenContinuous LMT → S, A
MiniMax-SpeechToken LMT, S → S
MixedG2P-T5Token LMT, S → S
MM-MovieDubberDiffusion / flowT, V → S
MoE-TTSToken LMT → S
MoonCastToken LMT, S → S
MOSS-TTSToken LMT, S → S
MOSS-TTS-NanoToken LMT, S → S
MOSS-TTS-RealtimeToken LMT, S → S
MOSS-TTSDToken LMT, S → S
MOSS-VoiceGeneratorToken LMT → S
MP-ELDContinuous LMT, S → S
MPE-TTSToken LMT, S → S
Multistage multimodal TTSDiffusion / flowT, I → S
Muyan-TTSToken LMT, S → S
NaturalSpeechFlow / VAET → S
NaturalSpeech 2Diffusion / flowT, S → S
NaturalSpeech 3Diffusion / flowT, S → S
NeuTTS AirToken LMT, S → S
NeuTTS NanoToken LMT, S → S
NeuTTS-2EToken LMT → S
NR-LauraTTSToken LMT, S → S
NVSpeech TTSToken LMT, S → S
Nüshu-PitchVITSFlow / VAET → S
Ojibwe-Mi'kmaq-Maliseet TTSDiffusion / flowT → S
OmniVoiceToken LMT, S → S
OpusLMToken LMT, S → S
Orpheus TTSToken LMT, S → S
OscillaTTSDiffusion / flowT, S → S
OuteTTSToken LMT, S → S
OV-InstructTTSToken LMT → S
OZSpeechToken LMT, S → S
PALLEToken LMT, S → S
Parallel GPTToken LMT, S → S
Parallel TacotronParallelT → S
Parallel Tacotron 2ParallelT → S
ParaStyleTTSParallelT → S
Parler-TTSToken LMT → S
Parler-TTS Hinglish adaptationToken LMT → S
PFluxTTSDiffusion / flowT, S → S
Phoenix TTSToken LMT, S → S
Phoneme-tone adaptive Thai TTSParallelT, S → S
PilotTTSToken LMT, S → S
Piper (VITS voices)Flow / VAET → S
Pocket TTSContinuous LMT, S → S
PortaSpeechFlow / VAET → S
PROEMOParallelT → S
Progressive face-conditioned TTSFlow / VAET, I → S
Prompt-Unseen-EmotionToken LMT → S
PromptTTSParallelT → S
PromptTTS 2Diffusion / flowT → S
ProtoDisent-TTSFlow / VAET, S → S
PS-TTSToken LMT, S → S
QTTSToken LMT, S → S
Qwen-Audio-3.0-TTSToken LMT, S → S
Qwen3-TTSToken LMT, S → S
RADKA-CSSToken LMT, S → S
RALL-EToken LMT, S → S
Raon-OpenTTSDiffusion / flowT, S → S
RapFlow-TTSDiffusion / flowT, S → S
ReGenVoiceDiffusion / flowT, S → S
ReStyle-TTSDiffusion / flowT, S → S
RTFree-F5Diffusion / flowT, S → S
RV-TTSDiffusion / flowT, I → S
RWKVTTSToken LMT, S → S
S5-TTSToken LMT, S → S
Sarashina2.2-TTSToken LMT, S → S
SASLMContinuous LMT, S → S
Seed-TTSToken LMT, S → S
Self-distilled zero-shot TTSCompactT, S → S
SelfTTSFlow / VAET, S → S
SemaVoiceContinuous LMT, S → S
SemBridgeContinuous LMT, S → S
Shallow Flow Matching TTSDiffusion / flowT, S → S
SLEDContinuous LMT, S → S
SlimSpeechCompactT, S → S
SMLLEToken LMT, S → S
SoulX-PodcastToken LMT, S → S
Spark-TTSToken LMT, S → S
SpeakStreamContinuous LMT, S → S
SPEAR-TTSToken LMT, S → S
SpeechAccentLLMToken LMT, S → S
SpeechEditToken LMT, S → S
SpeechT5AutoregressiveT, S → S
SpeechXToken LMT, S → S
SpeedySpeechParallelT → S
Spotlight-TTSDiffusion / flowT, S → S
StellarTTSCompactT, S → S
Step-Audio-EditXToken LMT, S → S
Step-Audio-TTSToken LMT, S → S
StepAudio 2.5 TTSToken LMT, S → S
Stochastic-alignment continuous TTSContinuous LMT, S → S
StreamMelContinuous LMT, S → S
StyleTTSParallelT, S → S
StyleTTS 2Diffusion / flowT, S → S
SupertonicCompactT, S → S
Supertonic 2CompactT → S
Supertonic 3CompactT → S
SwanVoiceDiffusion / flowT, S → S
SyncSpeechToken LMT, S → S
TacotronAutoregressiveT → S
Tacotron 2AutoregressiveT → S
TADAContinuous LMT, S → S
TED-TTSToken LMT, S → S
Tibetan-TTSToken LMT, S → S
TinyWaveToken LMT, S → S
TLDR (TTS)Token LMT, S → S
TMD-TTS (formerly FMSD-TTS)Diffusion / flowT → S
TontaubeV1Token LMT, S → S
Tortoise TTSToken LMT, S → S
Transformer TTSAutoregressiveT → S
TTS-CtrlNetDiffusion / flowT, S → S
TTS-TransducerToken LMT, S → S
TTSYorubaConcatenativeT → S
UDDETTSToken LMT → S
UmbraTTSDiffusion / flowT, S, A → S
UniFlow-AudioDiffusion / flowT, S, A, I, V → S, A
UNISONDiffusion / flowT, S, A → S, A
UniSonateDiffusion / flowT → S
UniSpeakerDiffusion / flowT, S, I → S
UniTAFToken LMT → S
UniTalkerToken LMT, S, V → S
UniTTSToken LMT, S → S
UniVocalToken LMT, S → S
UniVoice (ASR and TTS)Diffusion / flowT, S → S
UniVoice (speech and singing)Diffusion / flowT, S → S
UniWav (TTS)Diffusion / flowT, S → S, A
USCF-conditioned TTSDiffusion / flowT, S → S
V-CASSToken LMT, V → S
VALL-EToken LMT, S → S
VALL-E 2Token LMT, S → S
VALL-E XToken LMT, S → S
VALL-TToken LMT, S → S
VclipFlow / VAET, I → S
VevoToken LMT, S → S
VibeVoiceContinuous LMT, S → S
VibeVoice-RealtimeContinuous LMT → S
VisualSpeechParallelT, V → S
VITSFlow / VAET → S
VITS2Flow / VAET → S
VividVoiceDiffusion / flowT, I, V → S
VocalNet-M2Token LMT, S → S
VoiceboxDiffusion / flowT, S → S
VoiceChat-TTSContinuous LMT → S
VoiceCraftToken LMT, S → S
VoiceCraft-DubToken LMT, S, V → S
VoiceDesignerDiffusion / flowT, S → S
VoiceSculptorToken LMT, S → S
VoxCPMContinuous LMT, S → S
VoxCPM2Continuous LMT, S → S
Voxtral TTSToken LMT, S → S
VoXtreamToken LMT, S → S
VoXtream2Token LMT, S → S
VSpeechLMToken LMT, V → S
Wave-TacotronAutoregressiveT → S
WavTTSDiffusion / flowT, S → S
WenetSpeech-Wu TTSToken LMT, S → S
WeSConToken LMT, S → S
WhisperSpeechToken LMT, S → S
WordVoiceToken LMT, S → S
X-VoiceDiffusion / flowT, S → S
X2Streaming-TTSToken LMT, S → S
XEmoRAGToken LMT, S → S
XTTSToken LMT, S → S
YourTTSFlow / VAET, S → S
ZipVoiceDiffusion / flowT, S → S
ZipVoice-DialogDiffusion / flowT, S → S
ZonosToken LMT, S → S

Model figures

Figures are credited to their sources. Editorial input/output diagrams are labeled. Credits.

A2TTS

A2TTS extracts a voice embedding from a short recording and conditions a diffusion acoustic decoder on it. Reference-aware duration prediction improves timing consistency for multilingual synthesis in low-resource Indian languages.

Paper · GitHub: no author-linked repository found · Details

A2TTS — Figure 1

Figure 1 · Source

Affectron

Affectron extends a verbal-speech backbone to place nonverbal vocalizations in emotionally and contextually appropriate positions. Augmented training examples and structural masking enable expressive utterances containing events such as laughter while preserving the spoken content.

Paper · GitHub · Project · Details

Affectron — Figure 2

Figure 2 · Source

AgentSteerTTS

AgentSteerTTS separates identity and emotional-prosodic representations, then grounds compound instructions in retrieved acoustic examples. A controller combines these conditions and uses feedback to refine expressive speech while preserving the intended speaker.

Paper · GitHub: no author-linked repository found · Details

AgentSteerTTS — Figure 4

Figure 4 · Source

AlignDiT

AlignDiT aligns text, visual information and acoustic conditions before diffusion-based speech generation. Modality-specific guidance balances these inputs, targeting synchronized, intelligible speech that follows the timing and expression of the supplied scene.

Paper · GitHub · Details

AlignDiT — Figure 1

Figure 1 · Source

AMNet

AMNet adds phrase-structure information and local convolutional modeling to a parallel Mandarin acoustic model. These changes help capture contextual pauses, emphasis and intonation before the accompanying vocoder reconstructs the waveform.

Paper · GitHub: no author-linked repository found · Details

AMNet — Paper figure

Paper figure · Source

ARCHI-TTS

ARCHI-TTS uses a dedicated semantic alignment module to coordinate text and reference acoustic features. Reusing encoder features across denoising steps reduces repeated computation while the flow model generates the target speech.

Paper · Project · GitHub: no author-linked repository found · Details

ARCHI-TTS — Figure 1

Figure 1 · Source

ATRIE

ATRIE converts character descriptions into separate voice-identity and dynamic prosody conditions. A compact adapter learns from a larger language-model teacher and modulates a GPT-SoVITS-based synthesizer, supporting expressive persona-driven speech generation.

Paper · GitHub: no author-linked repository found · Details

ATRIE — Figure 1

Figure 1 · Source

Audiobook-CC

Audiobook-CC models context beyond individual sentences and separates style instructions from voice prompts. Distillation strengthens emotional expression, supporting multi-character narration with more consistent voices and performance across longer passages.

Paper · GitHub: no author-linked repository found · Details

Audiobook-CC — Figure 1

Figure 1 · Source

AuEmoChat

AuEmoChat learns a discrete emotion representation from speech and compresses dialogue history while retaining emotionally relevant information. Its language model predicts emotion and speech tokens, which a context-conditioned flow decoder renders into expressive conversational audio.

Paper · GitHub · Details

AuEmoChat — Figure 2

Figure 2 · Source

AuK

AuK combines language-model conditioning, an audio VAE and successive multimodal and single-stream diffusion blocks. One model handles reference-based speech synthesis and instruction-guided editing; its distilled AuK-Flash variant reduces the number of generation steps.

Paper · GitHub · Details

AuK — Figure 4

Figure 4 · Source

Authentic-Dubber

Authentic-Dubber retrieves emotionally relevant audiovisual examples and progressively incorporates them into speech generation. Its director-actor formulation connects a scene's visual context and reference delivery to the target transcript for expressive movie dubbing.

Paper · GitHub · Details

Authentic-Dubber — Figure 2

Figure 2 · Source

AutoSIFT

AutoSIFT divides a reference voice's style into attribute-specific components and a residual representation. Text instructions replace selected attributes while unmentioned characteristics remain conditioned on the reference, allowing partial style editing during speech synthesis.

Paper · GitHub: no author-linked repository found · Details

AutoSIFT — Figure 1

Figure 1 · Source

AutoStyle-TTS

AutoStyle-TTS matches the target text against a collection of expressive speech examples using learned textual embeddings. The selected recording provides style conditioning for synthesis, automatically adapting delivery to the content.

Paper · GitHub · Project · Details

AutoStyle-TTS — Paper figure

Paper figure · Source

AVLM (expressive speech)

This audio-visual language model adds full-face information to an expressive speech backbone. Training on emotion and dialogue tasks connects facial cues with spoken delivery, enabling speech generation that uses visual as well as acoustic conversational context.

Paper · GitHub · Details

AVLM (expressive speech) — Figure 2

Figure 2 · Source

Bagpiper-TTS

Bagpiper-TTS converts a natural-language request into a detailed speech plan containing words and delivery information. The generator follows that plan for tasks ranging from ordinary narration to multi-speaker rendering, role-play and singing.

Paper · Project · GitHub: no author-linked repository found · Details

Bagpiper-TTS — Figure 1

Figure 1 · Source

BareWave

BareWave generates speech in waveform space using a single inference path. Representation alignment, staged noise scheduling and perceptual objectives guide training, replacing the separate acoustic-feature and waveform-reconstruction stages common in other TTS systems.

Paper · Project · GitHub: no author-linked repository found · Details

BareWave — Figure 2

Figure 2 · Source

Bark

Hierarchical autoregressive audio tokens.

GitHub · Details

Bark — Editorial input/output diagram

Editorial input/output diagram · Source

BASE TTS

Autoregressive speechcodes and convolutional decoder.

Paper · Details

BASE TTS — Figure 1

Figure 1 · Source

BatonTTS (BatonVoice)

BatonVoice interprets a user's expressive request and translates it into controls for its dedicated BatonTTS generator. Separating instruction interpretation from acoustic rendering enables more explicit feature control, including transfer to languages outside the control-training data.

Paper · GitHub: no author-linked repository found · Details

BatonTTS (BatonVoice) — Figure 1

Figure 1 · Source

BELLE

BELLE predicts both speech values and their uncertainty in a continuous autoregressive synthesizer. Multiple synthetic renditions of the same text provide training support for the variance estimate, enabling richer acoustic distributions without adding an iterative inference stage.

Paper · GitHub · Project · Details

BELLE — Figure 1

Figure 1 · Source

BitTTS

BitTTS reduces storage and computation through extremely low-bit trained weights and indexed parameter sharing. It targets speech generation on constrained devices, preserving a full synthesis path while shrinking the model representation.

Paper · GitHub: no author-linked repository found · Details

BitTTS — Figure 1

Figure 1 · Source

Block-wise Mimi TTS

This streaming system replaces continuous acoustic regression with direct prediction of Mimi codec layers. A modified FastSpeech 2 backbone supplies aligned features and a depth-wise decoder fills residual codebooks, producing successive speech blocks without temporal autoregression.

Paper · GitHub: no author-linked repository found · Details

Block-wise Mimi TTS — Figure 1 (paper page 10)

Figure 1 (paper page 10) · Source

BnTTS

BnTTS extends an XTTS-based multilingual pipeline to Bangla using language-specific phonetic adaptations. A small amount of target-speaker audio supports personalization, with the model designed for limited-resource speech synthesis.

Paper · GitHub: no author-linked repository found · Details

BnTTS — Figure 1

Figure 1 · Source

Bolbosh

Bolbosh adapts Matcha-TTS to Kashmiri with language-aware text processing and cross-language training. Its design targets the pronunciation and script challenges of a low-resource language while retaining efficient non-autoregressive acoustic generation.

Paper · GitHub · Details

Bolbosh — Figure 1

Figure 1 · Source

Borderless Long Speech Synthesis

This system organizes speech instructions at global, sentence and token levels to guide extended recordings. A continuous-token backbone uses explicit planning and condition dropout to combine voice design, multi-speaker rendering and changing acoustic or emotional context.

Paper · GitHub: no author-linked repository found · Details

Borderless Long Speech Synthesis — Editorial input/output diagram

Editorial input/output diagram · Source

BreezyVoice

BreezyVoice combines supervised speech tokens, a language model and flow-based acoustics with a pronunciation frontend. Its Taiwanese Mandarin adaptation provides explicit phonetic control for characters with multiple readings while retaining reference-based voice synthesis.

Paper · GitHub · Details

BreezyVoice — Figure 1

Figure 1 · Source

BridgeTTS

BridgeTTS uses the BridgeCode dual representation to shorten the sequence predicted by its autoregressive language model. Bridging modules reconstruct more detailed continuous acoustic features from those sparse tokens, balancing generation speed with voice fidelity.

Paper · GitHub: no author-linked repository found · Details

BridgeTTS — Figure 2

Figure 2 · Source

BVS

Beyond Video-to-SFX predicts audio semantic tokens from visual information and phonetic cues, then refines them into acoustic tokens. The two-stage generator produces intelligible speech whose timing and environmental sound fit the supplied video.

Paper · GitHub: no author-linked repository found · Details

BVS — Figure 1

Figure 1 · Source

CAM-TTS

CAM-TTS retains global narrative information and retrieves local details through an updatable memory block. Prefix attention combines those memories with preceding context, guiding sentence-level synthesis across longer paragraphs.

Paper · GitHub: no author-linked repository found · Details

CAM-TTS — Figure 2

Figure 2 · Source

CapTalk

CapTalk designs voices from descriptions for individual utterances and multi-speaker dialogue. Hierarchical conditioning separates stable speaker identity from changing turn-level delivery, while explicit planning tokens control dynamic expressive attributes.

Paper · GitHub: no author-linked repository found · Details

CapTalk — Figure 1

Figure 1 · Source

CAST-TTS

CAST-TTS maps a voice description or a reference recording into a common timbre-conditioning interface. Cross-attention delivers that information to the synthesizer, enabling voice design and reference-based cloning within the same generation framework.

Paper · GitHub · Project · Details

CAST-TTS — Figure 1

Figure 1 · Source

CaT-TTS

CaT-TTS separates textual understanding from acoustic generation in a two-Transformer architecture. During decoding, a masked parallel inference procedure guides speech-token predictions to reduce local errors in zero-shot voice synthesis.

Paper · GitHub: no author-linked repository found · Details

CaT-TTS — Figure 2

Figure 2 · Source

Causal-prosody FastSpeech 2

This FastSpeech 2 extension explicitly models emotion alongside duration, pitch and energy. Counterfactual training separates emotional changes from linguistic content, allowing users to modify prosody while preserving the intended words.

Paper · GitHub: no author-linked repository found · Details

Causal-prosody FastSpeech 2 — Figure 1 (paper page 3)

Figure 1 (paper page 3) · Source

CDE-StyleTTS

CDE-StyleTTS lets acoustic states evolve continuously along a phoneme sequence whose timing comes from durations. Sampling this trajectory supplies the acoustic decoder with timing-sensitive representations, providing a way to transfer changing expressive style instead of merely repeating phoneme embeddings.

Paper · GitHub · Details

CDE-StyleTTS — Paper figure

Paper figure · Source

Chain-of-Details TTS

Chain-of-Details TTS progressively predicts speech at increasing temporal resolutions using a shared decoder. The coarsest stage provides an implicit phonetic plan, while subsequent stages recover timing detail without a separate phoneme-duration predictor.

Paper · GitHub: no author-linked repository found · Details

Chain-of-Details TTS — Paper figure

Paper figure · Source

Chain-Talker

Chain-Talker first derives an emotional description from dialogue history, then predicts semantic speech codes. A final rendering stage combines these plans to synthesize expressive responses whose delivery fits the conversational context.

Paper · GitHub · Details

Chain-Talker — Figure 2

Figure 2 · Source

Chatterbox

Chatterbox synthesizes speech from text and a voice reference, with controls for expressive delivery. Its multilingual models focus on cross-language voice consistency, while Turbo and Nano use a smaller backbone and a single-step acoustic decoder. These variants serve different narration, conversational playback and local-device requirements; their capabilities are not interchangeable.

GitHub · Details

Chatterbox — Editorial input/output diagram

Editorial input/output diagram · Source

Chatterbox-Flash

Chatterbox-Flash generates speech-token blocks in parallel while keeping block-by-block streaming. Calibration against common-token priors and confidence-based stopping improve its discrete diffusion decoding after adaptation from a pretrained autoregressive synthesizer.

Paper · GitHub · Details

Chatterbox-Flash — Figure 3

Figure 3 · Source

ChatTTS

Autoregressive speech-token generation.

GitHub · Details

ChatTTS — Editorial input/output diagram

Editorial input/output diagram · Source

CLaM-TTS

Probabilistic residual quantization and multi-token LM.

Paper · Details

CLaM-TTS — Figure 1

Figure 1 · Source

CLEAR

CLEAR models speech directly in a continuous latent space, avoiding discrete codec-token prediction. Its zero-shot generator combines reference-voice conditioning with incremental audio production, targeting a balance between naturalness and response latency.

Paper · GitHub: no author-linked repository found · Details

CLEAR — Figure 1

Figure 1 · Source

Clip-TTS

Clip-TTS trains its textual representation against corresponding mel-spectrogram information through a contrastive objective. The acoustic Transformer uses the resulting context-aware features to improve prosodic interpretation during speech generation.

Paper · GitHub: no author-linked repository found · Details

Clip-TTS — Figure 3

Figure 3 · Source

Compact neural accessibility TTS

This compact synthesis system combines a shared-parameter text frontend with an efficient recurrent waveform generator. It targets responsive accessibility voices on low-power devices, where small storage requirements and immediate playback matter alongside naturalness.

Paper · GitHub: no author-linked repository found · Details

Compact neural accessibility TTS — Figure 1

Figure 1 · Source

Compressed-to-fine speech LM

This speech-language-model design keeps recent acoustic tokens and voice prompts at full detail while compressing distant context. The asymmetric representation reduces redundant long-sequence processing without discarding the local cues needed for pronunciation and vocal consistency.

Paper · GitHub: no author-linked repository found · Details

Compressed-to-fine speech LM — Figure 1

Figure 1 · Source

Confucius4-TTS

Confucius4-TTS extracts voice characteristics from self-supervised speech features without requiring a transcript of the reference recording. A language model predicts semantic tokens and a flow decoder generates mel-spectrograms, enabling voice cloning within and across fourteen languages.

Paper · GitHub · Details

Confucius4-TTS — Figure 1

Figure 1 · Source

Continuous-token diffusion TTS

This model combines a language head that predicts boundaries with a diffusion head that generates continuous acoustic frames. Masked and staged training stabilize speaker-reference conditioning, providing a text-to-speech path within a multimodal language-model architecture.

Paper · GitHub: no author-linked repository found · Details

Continuous-token diffusion TTS — Figure 2

Figure 2 · Source

Controllable masked-speech TTS

This synthesizer separates reference voice information from acoustic background conditions. An explicit task control selects whether background sound is retained or removed, allowing personalized speech generation under different environmental requirements.

Paper · GitHub: no author-linked repository found · Details

Controllable masked-speech TTS — Paper figure

Paper figure · Source

CookVoice

CookVoice aligns textual content, style and prosodic controls to acoustic frames before speech generation. The same compact model supports spoken and sung voices, reference imitation and editing, allowing individual voice attributes to be controlled within a shared synthesis pipeline.

Paper · Project · GitHub: no author-linked repository found · Details

CookVoice — Figure 1

Figure 1 · Source

CosyEdit2

CosyEdit2 adapts a text-speech language model and acoustic decoder for consistent speech editing, then refines them with editing-specific rewards. The paper also evaluates the resulting improvement in zero-shot text-to-speech, connecting local editing consistency with reference-based synthesis.

Paper · Project · GitHub: no author-linked repository found · Details

CosyEdit2 — Figure 1

Figure 1 · Source

CoSyncDiT

CoSyncDiT guides flow-based speech synthesis through acoustic-style adaptation, visual calibration and timed context alignment. These stages connect the supplied transcript and scene information to expressive, synchronized movie dubbing.

Paper · GitHub · Details

CoSyncDiT — Figure 2

Figure 2 · Source

CosyVoice

Supervised semantic tokens and flow decoder.

Paper · Details

CosyVoice — Figure 1

Figure 1 · Source

CosyVoice 2

Text/speech LM and chunk-aware flow matching.

Paper · Details

CosyVoice 2 — Figure 1

Figure 1 · Source

CosyVoice 3

CosyVoice 3 extends streaming, reference-conditioned synthesis with a tokenizer trained on several speech-understanding tasks and a reward model for post-training. It targets reliable pronunciation, speaker identity and prosody across languages, dialects and less controlled text. The research scaling experiments and the downloadable Fun-CosyVoice3 checkpoint represent different model configurations.

Paper · GitHub · Project · Details

CosyVoice 3 — Figure 2

Figure 2 · Source

CosyWhisper (WhispSynth)

The WhispSynth generation pipeline combines a CosyVoice synthesizer with pitch-free digital signal processing to produce whispered speech. It supports multilingual whisper generation while avoiding the voiced pitch patterns of ordinary speech synthesis.

Paper · GitHub · Details

CosyWhisper (WhispSynth) — Figure 2

Figure 2 · Source

CoVoMix2

CoVoMix2 generates scripted dialogue directly with a flow-matching model, using reference voices without their transcripts. Speaker-disentangled text, sentence alignment and prompt masking support controlled timing and overlapping turns.

Paper · GitHub · Details

CoVoMix2 — Figure 1

Figure 1 · Source

Cross-Lingual F5-TTS

Cross-Lingual F5-TTS changes reference preparation and training so the generated text need not be paired with a reference transcript. Word-aligned acoustic prompts support cross-language voice cloning while reusing the flow-matching synthesis backbone.

Paper · Project · GitHub: no author-linked repository found · Details

Cross-Lingual F5-TTS — Figure 1

Figure 1 · Source

CrossAccent-TTS

CrossAccent-TTS separates speaker identity from accent-related information in a speech synthesis model. Weighted language embeddings control the accent subspace, allowing gradual accent changes and cross-language synthesis while retaining the reference speaker's timbre.

Paper · GitHub: no author-linked repository found · Details

CrossAccent-TTS — Figure 1

Figure 1 · Source

CSM

CSM uses the text and audio of preceding speaker turns to shape the delivery of the next utterance. A Llama backbone predicts speech representations, and a smaller decoder completes Mimi audio codes. It is a contextual speech renderer: an application supplies the words to say, including any responses written by a separate language model.

GitHub · Details

CSM — Editorial input/output diagram

Editorial input/output diagram · Source

CTC-TTS

CTC-TTS uses automatically derived alignment and two-word interleaving to train an incremental speech language model. Its length-concatenated and feature-stacked variants make different trade-offs between speech quality and generation latency.

Paper · GitHub · Details

CTC-TTS — Figure 2

Figure 2 · Source

CtrlSpeech

CtrlSpeech adds local pitch, loudness and duration conditioning to a patch-autoregressive diffusion synthesizer. A separate global speaker condition preserves the reference voice while users modify the delivery of individual words or phonemes.

Paper · GitHub · Details

CtrlSpeech — Figure 2

Figure 2 · Source

CuteTTS

CuteTTS combines a causal audio VAE, an autoregressive patch model and an explicitly speaker-conditioned flow head. Generating patches rather than individual frames reduces sequential work; a distilled variant targets faster streaming while preserving reference-voice synthesis.

Paper · GitHub · Details

CuteTTS — Figure 1

Figure 1 · Source

DAIEN-TTS

DAIEN-TTS separates a reference recording into speech and environmental components, then conditions acoustic generation on them independently. Its extended formulation additionally models reverberation and uses separate guidance controls for speech, noise and room acoustics, enabling voice cloning into a chosen environment.

Paper 1 · Project 1 · Paper 2 · GitHub · Project 2 · Details

DAIEN-TTS — Paper figure

Paper figure · Source

DARS

DARS separately models pathological timing and acoustic style to synthesize dysarthric speech. A multistage rhythm predictor and conditioned flow model generate targeted examples for improving recognition under limited real-world speech data.

Paper · GitHub: no author-linked repository found · Details

DARS — Figure 1

Figure 1 · Source

DCAR

DCAR changes the number of acoustic tokens predicted at each autoregressive step. Adapting the chunk size to the generation state shortens sequential processing while maintaining content alignment and reference-conditioned speech quality.

Paper · GitHub: no author-linked repository found · Details

DCAR — Figure 1

Figure 1 · Source

Deep Voice

Neural TTS pipeline with autoregressive WaveNet synthesis.

Paper · Details

Deep Voice — Figure 1

Figure 1 · Source

Deep Voice 2

Speaker-conditioned neural pipeline and WaveNet synthesis.

Paper · Details

Deep Voice 2 — Figure 1

Figure 1 · Source

Deep Voice 3

Convolutional attention encoder-decoder and converter.

Paper · Details

Deep Voice 3 — Figure 1

Figure 1 · Source

DeepASMR

DeepASMR separates ASMR delivery from the reference speaker's identity using discrete speech representations. A language model predicts content and style, while a flow-based acoustic decoder transfers the requested performance into the target voice.

Paper · GitHub: no author-linked repository found · Details

DeepASMR — Paper figure

Paper figure · Source

DeepDubber-V1

DeepDubber-V1 interprets visual scenes and dubbing requirements before generating the target speech. Its multimodal conditions distinguish narration, monologue and dialogue, guiding both expressive delivery and synchronization.

Paper · GitHub · Details

DeepDubber-V1 — Figure 1

Figure 1 · Source

DeepDubbing

DeepDubbing assigns voices to characters and conditions speech rendering on the surrounding script. Its voice-design and instruction-synthesis stages support multi-participant audiobook production while maintaining character identity and context-sensitive expression.

Paper · GitHub · Details

DeepDubbing — Figure 1

Figure 1 · Source

DelightfulTTS

Conformer acoustic model with explicit and implicit prosody.

Paper · Details

DelightfulTTS — Figure 1

Figure 1 · Source

DELTA-TTS

DELTA-TTS converts a pretrained speech language model to parallel discrete diffusion using lightweight adaptation. Local convolution and confidence-based decoding help retain acoustic structure while the model fills speech-token positions in an order determined by prediction confidence.

Paper · GitHub: no author-linked repository found · Details

DELTA-TTS — Figure 1

Figure 1 · Source

DepFlow

DepFlow separates a depression-related acoustic representation from speaker identity and spoken content. The representation conditions a flow-matching synthesizer, enabling controlled synthetic speech for investigating acoustic cues and data augmentation.

Paper · GitHub: no author-linked repository found · Details

DepFlow — Figure 1

Figure 1 · Source

Dia

Dia turns a speaker-tagged transcript into conversational audio, including supported nonverbal events such as laughter and coughing. Reference audio and its transcript can establish speaker identity and delivery. Its English checkpoint is useful for scripted exchanges and dialogue narration, with the conversation content supplied by the user rather than generated by the model.

GitHub · Details

Dia — Editorial input/output diagram

Editorial input/output diagram · Source

Dia2

Dia2 begins synthesizing before the complete script is available, allowing an application to feed words incrementally. Audio prefixes provide speaker and conversational context, while a streaming codec path reconstructs the output. The released 1B and 2B checkpoints focus on English dialogue; prefix conditioning helps keep voices consistent between generations.

GitHub · Details

Dia2 — Editorial input/output diagram

Editorial input/output diagram · Source

DialoSpeech

DialoSpeech models two speakers on separate tracks and uses chunked flow matching for acoustic rendering. The design supports expressive scripted dialogue, including interactions whose timing is difficult to reproduce by joining independent utterances.

Paper · Project · GitHub: no author-linked repository found · Details

DialoSpeech — Figure 2

Figure 2 · Source

DiEmo-TTS

DiEmo-TTS distills emotion information from speech while suppressing unrelated speaker characteristics. Cluster-based sampling and representation perturbation improve cross-speaker emotion transfer, including situations where extensive emotion labels are unavailable.

Paper · GitHub · Details

DiEmo-TTS — Figure 1

Figure 1 · Source

Diff-TTS

Text-conditioned denoising diffusion acoustic model.

Paper · Details

Diff-TTS — Figure 3

Figure 3 · Source

DiffCSS

DiffCSS samples prosody representations from multimodal conversational context using a diffusion model. A prosody-conditioned speech language model renders those samples, allowing several expressive deliveries that remain consistent with the same dialogue.

Paper · GitHub: no author-linked repository found · Details

DiffCSS — Paper figure

Paper figure · Source

DiFlow-TTS

DiFlow-TTS maps phonemes into linguistic content and generates separate prosody and acoustic token streams through discrete flow matching. Factorizing these responsibilities supports compact zero-shot synthesis with fewer sequential generation steps.

Paper · GitHub · Details

DiFlow-TTS — Figure 2

Figure 2 · Source

DiFlowDubber

DiFlowDubber first learns linguistic content and separate prosodic-acoustic tokens through a discrete-flow TTS model. A subsequent video-dubbing stage aligns those representations to visual timing, connecting voice generation with synchronized lip movements.

Paper · GitHub · Project · Details

DiFlowDubber — Figure 2

Figure 2 · Source

DisCo-Speech

DisCo-Speech learns a codec that separates content, delivery and speaker identity. A language model predicts combined content-prosody tokens while the decoder receives a separate timbre representation, enabling reference cloning with independently controllable vocal attributes.

Paper · GitHub · Project · Details

DisCo-Speech — Figure 1

Figure 1 · Source

DisSpeech

DisSpeech maps Mandarin text and marked stuttering events to semantic speech tokens without temporal autoregression. Pitch and energy modeling guide acoustic reconstruction, enabling controlled repetitions and other disfluencies for speech synthesis and recognition-data augmentation.

Paper · GitHub: no author-linked repository found · Details

DisSpeech — Figure 2

Figure 2 · Source

DiSTAR

DiSTAR first drafts blocks of residual-quantized speech tokens with a language model. A masked diffusion decoder then fills acoustic detail within each block, combining temporal planning and parallel refinement entirely in discrete codec space.

Paper · GitHub: no author-linked repository found · Details

DiSTAR — Figure 1

Figure 1 · Source

DiTAR

DiTAR predicts a sequence of compressed acoustic patches using a language model, then generates each patch's detail through diffusion. Separating global temporal planning from local reconstruction supports zero-shot speech synthesis with controllable sampling diversity.

Paper · Project · GitHub: no author-linked repository found · Details

DiTAR — Figure 1

Figure 1 · Source

DiTTo-TTS

Latent diffusion Transformer with speech-length prediction.

Paper · Details

DiTTo-TTS — Figure 1

Figure 1 · Source

DMOSpeech 2

DMOSpeech 2 adds reinforcement learning to duration prediction in an already metric-optimized synthesizer. Rewards derived from speaker similarity and transcription accuracy guide timing choices, linking prosodic planning with reference-voice and content objectives.

Paper · GitHub · Project · Details

DMOSpeech 2 — Figure 1

Figure 1 · Source

DMP-TTS

DMP-TTS maps style descriptions and reference recordings into a shared conditioning space. Chained guidance controls content, timbre and style separately, allowing detailed synthesis adjustments within a latent diffusion Transformer.

Paper · GitHub: no author-linked repository found · Details

DMP-TTS — Figure 1

Figure 1 · Source

dots.tts

dots.tts predicts continuous acoustic representations from multilingual text and voice references. Its flow-matching output head supports both audio streaming and streaming text input; guidance-aware distillation reduces the work needed to generate each audio packet.

Paper · GitHub · Project · Details

dots.tts — Figure 1

Figure 1 · Source

Dragon-FM

Dragon-FM predicts successive speech chunks autoregressively while refining the tokens inside each chunk with bidirectional flow matching. Compact acoustic codes and cross-chunk caching reduce generation overhead and support longer content such as podcasts.

Paper · Project · GitHub: no author-linked repository found · Details

Dragon-FM — Figure 1

Figure 1 · Source

DrawSpeech

DrawSpeech turns user-drawn prosodic curves into detailed pitch and energy conditions for synthesis. A diffusion model fills in the acoustic detail, enabling localized expressive control beyond an overall style description.

Paper · GitHub · Details

DrawSpeech — Paper figure

Paper figure · Source

DS-TTS

DS-TTS extracts complementary voice characteristics through two style encoders. Dynamic modulation conditions the acoustic generator on these representations, supporting unseen speakers and adapting synthesis across different sentence lengths.

Paper · GitHub: no author-linked repository found · Details

DS-TTS — Paper figure

Paper figure · Source

DualDub

DualDub generates spoken dialogue and background sound together from video and textual conditions. A cross-modal aligner coordinates the two decoding heads, targeting temporally synchronized soundtracks rather than speech rendered in isolation.

Paper · GitHub: no author-linked repository found · Details

DualDub — Figure 2

Figure 2 · Source

DualSpeechLM

DualSpeechLM uses understanding-oriented speech tokens as input and acoustic codec tokens for generation. Semantic supervision and staged conditioning coordinate the two representations, supporting speech synthesis within a unified understanding-and-generation architecture.

Paper · GitHub: no author-linked repository found · Details

DualSpeechLM — Figure 3

Figure 3 · Source

E2 TTS

Flow matching with filler-token text conditioning.

Paper · Details

E2 TTS — Figure 1

Figure 1 · Source

ECTSpeech

ECTSpeech gradually tightens consistency constraints on a pretrained diffusion synthesizer. The resulting generator maps a noise state to speech in one step, reducing repeated denoising while retaining the conditioning of the original TTS model.

Paper · GitHub: no author-linked repository found · Details

ECTSpeech — Figure 2

Figure 2 · Source

ELLA-V

Alignment-guided token reordering.

Paper · Details

ELLA-V — Figure 1

Figure 1 · Source

EME-TTS

EME-TTS jointly models emotional delivery and local emphasis instead of treating them as independent effects. Automatically derived emphasis labels and variance-related features help users stress selected material while retaining a recognizable target emotion.

Paper · GitHub: no author-linked repository found · Details

EME-TTS — Figure 1

Figure 1 · Source

EMM-TTS

EMM-TTS separates emotional content modeling from speaker-specific acoustic generation. Speaker-consistency objectives and adaptive normalization help transfer emotion across languages while retaining the reference speaker's timbre.

Paper · GitHub: no author-linked repository found · Details

EMM-TTS — Figure 2

Figure 2 · Source

EmojiVoice

EmojiVoice adds interpretable emoji prompts to the text encoder and flow predictor of Matcha-TTS. Changing prompts across phrases varies expression during longer robot utterances, providing a lightweight control interface for expressive synthesis.

Paper · GitHub · Details

EmojiVoice — Paper figure

Paper figure · Source

EmoShift

EmoShift adds a lightweight layer that learns emotion-dependent changes to a TTS model's hidden representation. The controls adjust expressive delivery and emotion intensity while leaving most of the underlying synthesis backbone intact.

Paper · GitHub: no author-linked repository found · Details

EmoShift — Figure 2

Figure 2 · Source

EmoSSLSphere

EmoSSLSphere combines an emotion representation constrained to a sphere with discrete units derived from self-supervised speech features. The model synthesizes emotional speech across languages while organizing expressive controls independently of the textual content.

Paper · GitHub: no author-linked repository found · Details

EmoSSLSphere — Figure 2

Figure 2 · Source

EmoSteer-TTS

EmoSteer-TTS extracts emotion-related directions from a pretrained synthesizer's internal activations. Applying those directions during inference changes, mixes or removes emotional expression without retraining; the paper evaluates the method across several different TTS backbones.

Paper · GitHub: no author-linked repository found · Details

EmoSteer-TTS — Figure 3

Figure 3 · Source

Emotion-timbre disentangled TTS

This emotional synthesizer learns separate reference encoders for timbre and emotion. A mutual-information objective reduces their overlap, while phoneme-level emotion prediction carries changing expression into the generated acoustic sequence.

Paper · GitHub · Project · Details

Emotion-timbre disentangled TTS — Figure 1

Figure 1 · Source

EmotiVoice

PromptTTS-derived style and content conditioning.

GitHub · Details

EmotiVoice — Editorial input/output diagram

Editorial input/output diagram · Source

EmoTra-TTS

EmoTra-TTS introduces frame-level valence, arousal and dominance controls into both prosodic planning and acoustic decoding. Synthetic transition examples teach the model to move between emotions within an utterance while keeping its wording and speaker identity consistent.

Paper · GitHub · Project · Details

EmoTra-TTS — Figure 2

Figure 2 · Source

EmoVoice

EmoVoice interprets free-form textual descriptions of emotional delivery. Its phoneme-boost variant predicts phonetic and acoustic information together to improve content consistency while retaining expressive style control.

Paper · GitHub · Details

EmoVoice — Figure 1

Figure 1 · Source

End-to-end discrete-token TTS

This system trains the discrete speech representation together with the language and acoustic models that consume it. Reconstruction and recognition feedback also update token prediction, reducing mismatches between separately trained components in reference-conditioned speech synthesis.

Paper · GitHub: no author-linked repository found · Details

End-to-end discrete-token TTS — Figure 1

Figure 1 · Source

F5-TTS

F5-TTS learns text-guided speech infilling with flow matching, refining character representations before a Transformer predicts the speech trajectory. A reference clip supplies voice context, and Sway Sampling controls how inference steps are distributed. The 2025 v1 Base release refines training and inference within the same general architecture for zero-shot speech synthesis.

Paper · GitHub · Details

F5-TTS — Figure 1

Figure 1 · Source

F5R-TTS

F5R-TTS adapts a flow-based synthesizer to reinforcement learning through a probabilistic formulation of generation. Transcription accuracy and speaker-similarity rewards refine the model's ability to preserve both target content and the reference voice.

Paper · GitHub · Details

F5R-TTS — Figure 2

Figure 2 · Source

Face-adapted StyleTTS 2

This model maps facial features into the style space of StyleTTS 2 through a lightweight learned adapter. It synthesizes text in a plausible face-conditioned voice without an audio reference; the paper evaluates transfer to unseen identities and another synthesis language.

Paper · GitHub: no author-linked repository found · Details

Face-adapted StyleTTS 2 — Figure 1

Figure 1 · Source

FaceSpeak

FaceSpeak extracts speaker-related and expressive information from real or stylized portraits. These visual conditions guide text-to-speech generation, allowing the requested words to be rendered in a plausible voice and emotion associated with the image.

Paper · GitHub: no author-linked repository found · Details

FaceSpeak — Figure 3

Figure 3 · Source

FacialTalker

FacialTalker quantizes facial action information and combines it with text and speech history in a conversational synthesizer. Joint preference training over visual and speech tokens helps the generated delivery reflect the interlocutor's facial expression and dialogue context.

Paper · GitHub · Details

FacialTalker — Figure 2

Figure 2 · Source

FastPitch

Parallel Transformer with explicit pitch prediction.

Paper · Details

FastPitch — Figure 1

Figure 1 · Source

FastSpeech

Feed-forward Transformer and duration-based length regulator.

Paper · Details

FastSpeech — Figure 1, PDF p. 4

Figure 1, PDF p. 4 · Source

FastSpeech 2

Parallel Transformer with duration, pitch and energy prediction.

Paper · Details

FastSpeech 2 — Figure 1, PDF p. 3

Figure 1, PDF p. 3 · Source

FC-TTS

FC-TTS conditions generation on two recordings, one supplying delivery style and the other speaker identity. Specialized representation processing and auxiliary training objectives aim to prevent either condition from leaking unwanted attributes into the synthesized voice.

Paper · GitHub: no author-linked repository found · Details

FC-TTS — Figure 1

Figure 1 · Source

FELLE

FELLE generates continuous acoustic frames sequentially, using the preceding frame to shape the next flow-matching prior. A coarse-to-fine acoustic head refines each prediction, supporting reference-conditioned speech without discrete speech-token classification.

Paper · GitHub: no author-linked repository found · Details

FELLE — Figure 1

Figure 1 · Source

FineCombo-TTS

FineCombo-TTS interprets style descriptions relative to a supplied speech reference. A flow-based variance predictor models how acoustic attributes should change, enabling precise relative edits without requiring an explicit independent embedding for every voice attribute.

Paper · Project · GitHub: no author-linked repository found · Details

FineCombo-TTS — Figure 2

Figure 2 · Source

FireRedAudio

FireRedAudio uses different acoustic encoders for understanding audio and conditioning speech generation. Its shared language model drives a flow-matching decoder over continuous RedAE latents, supporting voice cloning, instruction-controlled synthesis and speech editing within the broader audio model.

Paper · GitHub · Details

FireRedAudio — Figure 1

Figure 1 · Source

FireRedTTS

Text-to-semantic LM and speech decoder.

Paper · Details

FireRedTTS — Figure 3

Figure 3 · Source

FireRedTTS-1S

FireRedTTS-1S extends the FireRed synthesis line with incremental acoustic decoding. Its chunked flow-matching and frame-autoregressive multi-stream decoder options provide different trade-offs between initial latency and sustained generation speed.

Paper · GitHub: no author-linked repository found · Details

FireRedTTS-1S — Figure 1

Figure 1 · Source

FireRedTTS-2

FireRedTTS-2 models chronological sequences of speaker-labeled text and speech using a large Transformer plus a smaller codebook decoder. A low-rate streaming tokenizer reduces the number of audio steps. It targets long conversations and podcasts where speaker changes, turn-specific delivery and continuity across utterances matter.

Paper · GitHub · Details

FireRedTTS-2 — Figure 1

Figure 1 · Source

FireRedTTS3

FireRedTTS3 uses a semantically supervised audio autoencoder to make continuous speech representations easier to predict. The Base variant provides multilingual reference cloning; the Instruct variant adds natural-language voice design and editing of spoken content or acoustic attributes.

Paper · GitHub · Details

FireRedTTS3 — Figure 1

Figure 1 · Source

Fish Audio S1 / OpenAudio S1

The S1 family combines multilingual voice conditioning with explicit markers for emotion, tone and nonverbal sounds. Its full model and distilled S1-mini offer different deployment sizes, with reinforcement learning used to refine generation. It is suited to expressive narration and character dialogue, although access and capabilities depend on the selected release.

Model card · Details

Fish Audio S1 / OpenAudio S1 — Editorial input/output diagram

Editorial input/output diagram · Source

Fish Audio S2

Fish Audio S2 extends the Fish speech-model line with natural-language delivery instructions and multi-speaker, multi-turn synthesis. Its training pipeline uses speech descriptions, quality assessment and reward modeling to improve controllability. The released inference stack supports streaming, making the model relevant to both scripted audio production and incremental spoken responses.

Paper · GitHub · Details

Fish Audio S2 — Figure 2

Figure 2 · Source

Fish Speech

Dual-autoregressive slow/fast transformers.

Paper · GitHub · Details

Fish Speech — Figure 2

Figure 2 · Source

Flamed-TTS

Flamed-TTS combines representations of different speech attributes in an attention-free generator. Its reformulated flow-matching process targets efficient zero-shot synthesis with flexible pacing and reduced sequential computation.

Paper · Project · GitHub: no author-linked repository found · Details

Flamed-TTS — Figure 1

Figure 1 · Source

FlashTTS

FlashTTS processes incoming text and speech context on staggered tracks so synthesis can begin before sentence completion. Multi-token prediction accelerates the language model, while a distilled acoustic decoder reduces waveform-generation delay.

Paper · GitHub · Project · Details

FlashTTS — Figure 1

Figure 1 · Source

FleSpeech

FleSpeech unifies text, voice recordings and visual prompts into a common conditioning representation. Its multistage generator uses those controls to manipulate voice and delivery attributes flexibly rather than requiring one fixed prompt modality.

Paper · Project · GitHub: no author-linked repository found · Details

FleSpeech — Figure 2

Figure 2 · Source

FlexiVoice

FlexiVoice accepts an optional style instruction and an optional reference voice alongside the target text. Progressive preference training teaches the model to follow both conditions while reducing unwanted coupling between wording, speaker identity and delivery.

Paper · Project · GitHub: no author-linked repository found · Details

FlexiVoice — Figure 1

Figure 1 · Source

FlexSpeech

FlexSpeech separates timing control from the acoustic synthesis component to balance stable pronunciation and natural expression. A small set of style examples can adapt the duration module without retraining the full generator, supporting efficient delivery customization.

Paper · Project · GitHub: no author-linked repository found · Details

FlexSpeech — Figure 1

Figure 1 · Source

Flowtron

Autoregressive normalizing flows over mel spectrograms.

Paper · Details

Flowtron — Figure 1

Figure 1 · Source

FNH-TTS

FNH-TTS routes linguistic and speaker information through several duration experts to model varied timing patterns. It combines this predictor with changes to waveform generation, aiming for robust end-to-end speech synthesis across voices and prosodic conditions.

Paper · GitHub: no author-linked repository found · Details

FNH-TTS — Figure 1

Figure 1 · Source

Frame-stacked local Transformer TTS

This architecture lets a global language model predict several speech frames at a time and delegates their codec entries to a smaller local Transformer. The paper compares sequential local decoding with iterative masked prediction, showing how frame stacking changes the trade-off between synthesis throughput and acoustic fidelity.

Paper · GitHub: no author-linked repository found · Details

Frame-stacked local Transformer TTS — Figure 1

Figure 1 · Source

FreyaTTS

FreyaTTS is a Turkish-focused non-autoregressive synthesizer that maps character sequences to continuous audio latents. A frozen waveform autoencoder reconstructs the speech, while duration prediction and voice-focused post-training support efficient conversational playback without phoneme or discrete speech tokenization.

Paper · GitHub: no author-linked repository found · Details

FreyaTTS — Figure 1

Figure 1 · Source

Gemini 2.5 TTS

Gemini 2.5 TTS converts supplied text into speech with prompt-based control of accent, pace, style and emotion. The Flash and Pro interfaces support single-speaker narration and two-speaker scripts with separately assigned voices. These are dedicated speech-generation endpoints; their internal acoustic architecture is not fully disclosed in the public documentation.

Docs 1 · Docs 2 · Docs 3 · Details

Gemini 2.5 TTS — Editorial input/output diagram

Editorial input/output diagram · Source

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS adds expressive audio tags to prompt-steered speech generation, giving authors more local control over narration and delivery. It targets natural, responsive multilingual synthesis through a managed API. The public preview documentation describes the interface and controls, without enough architectural detail to reconstruct the underlying speech generator.

Docs 1 · Docs 2 · Details

Gemini 3.1 Flash TTS — Editorial input/output diagram

Editorial input/output diagram · Source

GibbsTTS

GibbsTTS generates discrete speech tokens through a continuous-time jump process. Metric-aware transition scheduling and a finite-step correction improve how token states evolve, supporting zero-shot voice synthesis with discrete flow matching.

Paper · GitHub · Project · Details

GibbsTTS — Figure 1

Figure 1 · Source

GLM-TTS

GLM-TTS first predicts speech tokens autoregressively, then converts them into audio with a diffusion decoder. Pitch-aware tokenization and multi-reward reinforcement learning target pronunciation, speaker similarity and expression. Hybrid phoneme/text input and LoRA voice adaptation provide controls for applications that need repeatable pronunciation and customized voices.

Paper · GitHub · Details

GLM-TTS — Figure 1

Figure 1 · Source

Glow-TTS

Normalizing-flow acoustic model and monotonic alignment search.

Paper · Details

Glow-TTS — Figure 1, PDF p. 3

Figure 1, PDF p. 3 · Source

GOAT-TTS

GOAT-TTS encodes continuous voice information in one branch and predicts speech tokens in another. Partial language-model adaptation preserves textual knowledge, while multi-token prediction supports streaming synthesis with reference-based paralinguistic conditioning.

Paper · GitHub: no author-linked repository found · Details

GOAT-TTS — Figure 1

Figure 1 · Source

GPA

General-Purpose Audio uses one autoregressive backbone to predict discrete speech tokens across synthesis, recognition and conversion tasks. Its TTS path combines target text with voice context, with shared multitask training and scalable inference.

Paper · GitHub · Details

GPA — Figure 1

Figure 1 · Source

GPT-4o Mini TTS

GPT-4o Mini TTS combines the text to be spoken with instructions that steer accent, speed, tone and emotional delivery. The Speech API can stream audio before the full result is complete and supports several output formats. It provides a managed synthesis component for narration and voice applications; public documentation does not disclose the complete acoustic architecture.

Docs 1 · Docs 2 · Details

GPT-4o Mini TTS — Editorial input/output diagram

Editorial input/output diagram · Source

GPT-SoVITS

GPT-SoVITS couples text-to-semantic token prediction with a reference-conditioned speech decoder. Its 2025 V3, V4 and V2 Pro releases extend a workflow that supports both zero-shot synthesis and voice adaptation from a small training set. The surrounding WebUI helps prepare data, while recognition and source-separation utilities remain separate components.

GitHub · Docs · Details

GPT-SoVITS — Editorial input/output diagram

Editorial input/output diagram · Source

Grad-TTS

Score-based diffusion decoder and monotonic alignment search.

Paper · Details

Grad-TTS — Figure 2

Figure 2 · Source

GRAFT

GRAFT attaches codec tokens from a spoken word example to that word's location in the text prompt. Separate target-speaker conditioning allows the pronunciation hint to come from another voice while the synthesized sentence retains the desired speaker.

Paper · GitHub: no author-linked repository found · Details

GRAFT — Paper figure

Paper figure · Source

GSA-TTS

GSA-TTS extracts local style information at successive levels and combines it through attention into a global reference condition. This richer style representation guides the acoustic model when synthesizing an unseen speaker's voice.

Paper · GitHub: no author-linked repository found · Details

GSA-TTS — Figure 1

Figure 1 · Source

GST-Tacotron

Tacotron with a reference encoder and global style tokens.

Paper · Details

GST-Tacotron — Figure 1

Figure 1 · Source

Habibi

Habibi trains an Arabic synthesizer progressively from standard language to regional dialects using curated public speech. It targets zero-shot voice cloning across dialects and reading without mandatory diacritic marks.

Paper · GitHub · Project · Details

Habibi — Figure 1

Figure 1 · Source

HD-PPT

HD-PPT learns speech codes that distinguish spoken content from instruction-related preferences. A language model predicts semantic information, expressive style and acoustic detail in sequence, improving the mapping from natural-language requests to controllable speech.

Paper · GitHub: no author-linked repository found · Details

HD-PPT — Figure 1

Figure 1 · Source

Higgs Audio v2

Higgs Audio v2 combines interleaved text/audio modeling with a unified speech tokenizer and DualFFN layers for acoustic prediction. Reference clips and scene context influence voices, while text context shapes prosody across narration and multi-speaker scripts. The generation checkpoint covers synthesis; the separate understanding branch in the family diagram is not another output mode of this checkpoint.

Model card · Details

Higgs Audio v2 — Official architecture diagram

Official architecture diagram · Source

Higgs Audio v2.5

Higgs Audio v2.5, now documented as Higgs TTS 2.5, reduces the autoregressive audio Transformer to 1B parameters. GRPO-based alignment and a curated voice dataset refine pronunciation, cloning and expressive control tags. It targets production narration and conversational speech with lower computational requirements than the preceding 3B generation model.

Announcement · Details

Higgs Audio v2.5 — Editorial input/output diagram

Editorial input/output diagram · Source

Higgs Audio v3 TTS

Higgs TTS 3 uses an autoregressive decoder over interleaved text and eight speech codebooks, with a delay pattern and fused input/output projections. Reference audio establishes a voice, while inline tokens control emotion, style, pauses and sound effects. The 4B release targets multilingual conversational speech and expressive response rendering.

Model card · Details

Higgs Audio v3 TTS — Official architecture diagram

Official architecture diagram · Source

HiStyle

HiStyle predicts a voice's timbre first and finer delivery attributes afterward from textual descriptions. Contrastive text-audio alignment organizes these style representations before they condition a speech synthesizer.

Paper · GitHub: no author-linked repository found · Details

HiStyle — Figure 2

Figure 2 · Source

HoliDubber

HoliDubber conditions audio generation on video and a text prompt describing speech and sound effects. A causal model plans successive latent patches and a local diffusion Transformer generates their detail, supporting synchronized dubbing within complex acoustic scenes.

Paper · Project · GitHub: no author-linked repository found · Details

HoliDubber — Figure 2

Figure 2 · Source

HoliTok (TTS)

HoliTok combines linguistic and acoustic information in a continuous representation designed for both understanding and generation. Its downstream autoregressive model and diffusion decoder demonstrate a text-to-speech path using the same latents employed for recognition.

Paper · GitHub · Details

HoliTok (TTS) — Figure 1

Figure 1 · Source

Hume Octave TTS

Octave uses text context and acting instructions to adjust pronunciation, emphasis, tempo and emotional delivery. Its API supports voice creation from descriptions, voice cloning and continuation across longer passages. Octave 1 and the Octave 2 preview have different feature coverage; the public interface is documented more fully than the internal speech-model architecture.

Docs · Details

Hume Octave TTS — Editorial input/output diagram

Editorial input/output diagram · Source

ImmersiveTTS

ImmersiveTTS jointly models spoken content and its surrounding acoustic scene in a multimodal diffusion Transformer. Speech and general-audio representations provide complementary training signals, helping generated wording remain intelligible within the requested environmental context.

Paper · GitHub · Project · Details

ImmersiveTTS — Figure 1

Figure 1 · Source

IndexTTS

IndexTTS adapts the XTTS/Tortoise approach with a Conformer reference encoder and a BigVGAN2 speech decoder. Hybrid character/pinyin input gives explicit control over difficult Chinese pronunciations. Its central use case is zero-shot voice cloning with predictable text rendering, including content that benefits from pronunciation correction.

Paper · GitHub · Project · Details

IndexTTS — Figure 1

Figure 1 · Source

IndexTTS 2.5

IndexTTS 2.5 shortens semantic sequences with a lower-rate codec and replaces the acoustic module's backbone with a Zipformer design. Multilingual training strategies and reinforcement learning extend pronunciation and emotion transfer across languages. It retains reference-based voice conditioning while reducing the cost of semantic and acoustic generation.

Paper · GitHub · Details

IndexTTS 2.5 — Figure 1

Figure 1 · Source

IndexTTS2

IndexTTS2 separates speaker identity from emotional style so that different references can control timbre and delivery. Its autoregressive formulation also supports explicit output-token budgeting for duration control, alongside unconstrained generation. These mechanisms are intended for expressive speech and timing-sensitive work such as dubbing; availability of controls should be checked in the chosen implementation.

Paper · GitHub · Details

IndexTTS2 — Figure 1

Figure 1 · Source

InstructAudio

InstructAudio combines natural-language instructions with phonemes or lyrics in a common generation format. Joint and single-stream diffusion layers synthesize speech or music while controlling attributes such as voice, emotion, accent or musical character.

Paper · Project · GitHub: no author-linked repository found · Details

InstructAudio — Figure 2

Figure 2 · Source

IntMeanFlow

IntMeanFlow distills a flow-based speech generator to predict integrated acoustic updates over larger intervals. A search for effective sampling steps further reduces decoding work, enabling reference-conditioned synthesis with fewer iterative refinements.

Paper · Project · GitHub: no author-linked repository found · Details

IntMeanFlow — Figure 1

Figure 1 · Source

Inworld TTS-1

Inworld TTS-1 and TTS-1-Max are multilingual autoregressive synthesizers designed for low-latency speech output. Textual audio markup controls emotions and nonverbal vocalizations, with model variants offering different capacity and inference-cost trade-offs.

Paper · GitHub · Details

Inworld TTS-1 — Figure 1

Figure 1 · Source

JaiTTS

JaiTTS continually trains a VoxCPM-derived synthesizer on Thai-centered speech data. Its semantic planning, residual acoustic modeling and local diffusion decoding retain reference-based voice cloning while targeting fluent Thai pronunciation and delivery.

Paper · GitHub · Details

JaiTTS — Figure 1

Figure 1 · Source

JAM-Flow

JAM-Flow couples audio and motion diffusion modules within a shared model. Text, voice references and optional motion conditions support synchronized speech and facial movement, with an infilling objective allowing several conditioning combinations.

Paper · GitHub: no author-linked repository found · Details

JAM-Flow — Figure 3

Figure 3 · Source

JELLY

JELLY combines an emotion-aware Q-Former with several partially adapted language-model modules. Joint emotion recognition and contextual reasoning guide a speech synthesizer toward responses whose delivery matches the conversation.

Paper · GitHub · Project · Details

JELLY — Paper figure

Paper figure · Source

JETS

Joint FastSpeech 2 and HiFi-GAN with learned alignment.

Paper · Details

JETS — Figure 1, PDF p. 2

Figure 1, PDF p. 2 · Source

Joint non-autoregressive STT-TTS

This model handles text and speech within one non-autoregressive architecture. Its TTS path predicts acoustic output from text, and feeding partial predictions back into the model improves generation through iterative refinement while also supporting recognition training.

Paper · GitHub: no author-linked repository found · Details

Joint non-autoregressive STT-TTS — Paper figure

Paper figure · Source

Joycent

Joycent separates accent information from speaker identity using an adversarially trained accent encoder. It injects accent and speaker features at different text-encoder layers, supporting accent-conditioned synthesis without requiring a separate accented phoneme prediction stage.

Paper · GitHub · Project · Details

Joycent — Figure 1

Figure 1 · Source

JoyVoice

JoyVoice conditions long-form speech on speaker-labeled text and shared conversational context. Autoregressive hidden states feed an acoustic diffusion decoder, allowing multiple speakers, changing expression and more flexible turn boundaries within one synthesis system.

Paper · Project · GitHub: no author-linked repository found · Details

JoyVoice — Figure 2

Figure 2 · Source

KABURI-TTS

KABURI-TTS renders each participant on a separate audio channel from time-aligned phonemes and speaker activity. Supplying the timing layout explicitly lets the system synthesize overlapping speech, backchannels and interruptions for two-speaker conversations.

Paper · GitHub: no author-linked repository found · Details

KABURI-TTS — Figure 1

Figure 1 · Source

KittenTTS

KittenTTS provides small ONNX speech models with built-in voices and adjustable playback speed. The Mini, Micro and Nano releases offer different size and inference tradeoffs for CPU-oriented applications. The official README documents usage more fully than internal acoustic design, so the catalog presents it as a compact synthesis family without asserting an undisclosed architecture.

GitHub · Details

KittenTTS — Editorial input/output diagram

Editorial input/output diagram · Source

Koel-TTS

Koel-TTS explores several ways to condition a Transformer synthesizer on text and reference audio. Automatic speech-recognition and speaker-verification feedback, together with classifier-free guidance, improve adherence to the requested words and voice.

Paper · Project · GitHub: no author-linked repository found · Details

Koel-TTS — Figure 1

Figure 1 · Source

Kokoro

Kokoro's 2025 v1.0 release uses a compact StyleTTS 2-derived decoder with an iSTFTNet waveform generator. The released model relies on preset voice representations and omits style diffusion and a reference encoder. It is suited to lightweight narration and application speech where a small synthesis model and ready-made voices are useful.

Model card · Details

Kokoro — Editorial input/output diagram

Editorial input/output diagram · Source

Kyutai TTS (DSM)

Kyutai TTS treats text and speech as aligned streams separated by a controlled delay. A decoder-only language model can therefore emit audio as text arrives, instead of waiting for a complete utterance. This formulation supports incremental synthesis for voice interfaces and long streams, with speaker conditioning supplied through the TTS implementation.

Paper · GitHub · Details

Kyutai TTS (DSM) — Figure 1

Figure 1 · Source

LanStyleTTS

LanStyleTTS standardizes phonetic inputs and introduces local style conditioning across languages. The framework can augment several parallel acoustic backbones, allowing one multilingual model to vary delivery at phoneme level.

Paper · GitHub: no author-linked repository found · Details

LanStyleTTS — Paper figure

Paper figure · Source

LatinX

LatinX uses staged text-to-audio training, voice-cloning adaptation and automatic preference alignment. The resulting multilingual Transformer renders text in the source speaker's voice, supporting the synthesis stage of cross-language speech translation.

Paper · Project · GitHub: no author-linked repository found · Details

LatinX — Figure 1

Figure 1 · Source

LE2E-TTS

LE2E-TTS trains a compact text-to-waveform pipeline end to end rather than separately optimizing acoustic and waveform stages. It targets local devices where model size, response time and compute cost constrain deployment.

Paper · GitHub: no author-linked repository found · Details

LE2E-TTS — Figure 1

Figure 1 · Source

LightSpeech

FastSpeech-derived architecture found by neural architecture search.

Paper · Details

LightSpeech — Editorial input/output diagram

Editorial input/output diagram · Source

LLaDA-TTS

LLaDA-TTS adapts a speech language model to fill masked token sequences in parallel. Its bidirectional generation also supports inserting, replacing or deleting spoken words, combining reference-based TTS and speech editing through the same model.

Paper · GitHub: no author-linked repository found · Details

LLaDA-TTS — Figure 1

Figure 1 · Source

Llasa

Llasa maps text and an optional speech prompt to a single stream of codec tokens using a Llama-style Transformer. The simple representation makes standard language-model scaling and sampling techniques applicable to synthesis. Its research also explores speech-model verifiers that select samples for content accuracy, voice consistency or emotional expression.

Paper · GitHub · Project · Details

Llasa — Figure 2

Figure 2 · Source

Llasa+

Llasa+ adds multi-token prediction modules to a frozen Llasa backbone and checks their proposals with that backbone. A causal codec decoder turns accepted tokens into streaming audio. The resulting design addresses autoregressive latency while retaining the original speech model, making it relevant to systems that need incremental playback without retraining an entire backbone.

Paper · GitHub · Details

Llasa+ — Figure 1

Figure 1 · Source

LLMVoX

LLMVoX connects to an upstream language model through a streaming queue interface. Its small speech generator renders incoming text incrementally, allowing long conversations without tightly coupling the synthesizer to one particular language-model backbone.

Paper · GitHub · Project · Details

LLMVoX — Figure 2

Figure 2 · Source

Lombard Matcha-TTS

This Matcha-TTS extension learns vocal effort and articulation from automatically derived labels. It provides continuous controls over speech clarity and loudness-related effort, together with word-level emphasis, to synthesize the clearer delivery used in noisy listening conditions.

Paper · GitHub: no author-linked repository found · Details

Lombard Matcha-TTS — Figure 1

Figure 1 · Source

LongCat-AudioDiT

LongCat-AudioDiT maps text and reference speech into continuous waveform latents using a diffusion model. A jointly considered waveform autoencoder and adapted inference guidance address the interaction between acoustic reconstruction and zero-shot synthesis quality.

Paper · GitHub · Details

LongCat-AudioDiT — Figure 1

Figure 1 · Source

LoRP-TTS

LoRP-TTS adapts a pretrained zero-shot synthesizer using small low-rank parameter updates. It focuses on preserving a target speaker from limited, potentially noisy or spontaneous recordings whose acoustic conditions differ from the original training data.

Paper · GitHub: no author-linked repository found · Details

LoRP-TTS — Figure 3

Figure 3 · Source

Luna-TTS

Luna-TTS adapts an autoregressive text backbone into a speech diffusion language model. Its parallel and Realtime variants share a tokenizer and training lineage; Realtime predicts successive codec blocks while denoising each block in parallel for incremental audio delivery.

Paper · Project · GitHub: no author-linked repository found · Details

Luna-TTS — Figure 1 (paper page 4)

Figure 1 (paper page 4) · Source

M3-TTS

M3-TTS uses joint text-audio diffusion layers to learn alignment without first stretching text into a guessed acoustic timeline. Additional single-stream layers refine acoustic details, producing reference-conditioned speech through a compressed mel representation.

Paper · GitHub: no author-linked repository found · Details

M3-TTS — Figure 1

Figure 1 · Source

MAGIC-TTS

MAGIC-TTS exposes timing controls for selected speech tokens and pauses. Training includes incomplete control signals so the model can follow local edits where provided and infer natural timing elsewhere, supporting precise pacing without requiring every segment to be specified.

Paper · GitHub · Details

MAGIC-TTS — Figure 1

Figure 1 · Source

MagpieTTS-LF

MagpieTTS-LF extends MagpieTTS at inference time by retaining acoustic and textual context across sentence boundaries. Soft alignment priors and history-aware encoding support coherent longer narration without retraining the synthesizer on long recordings.

Paper · GitHub: no author-linked repository found · Details

MagpieTTS-LF — Figure 1

Figure 1 · Source

MambaVoiceCloning

MambaVoiceCloning uses state-space modules to encode phonemes, learn their timing and condition expressive synthesis. A training-only alignment teacher supplies timing supervision, while the generation path targets efficient long sequences and limited-lookahead speech streaming.

Paper · GitHub · Details

MambaVoiceCloning — Figure 1

Figure 1 · Source

MamTra

MamTra mixes state-space and attention layers to retain global context while reducing the cost of long sequences. Knowledge transfer from a pretrained Transformer initializes the hybrid synthesizer, combining efficient local processing with expressive acoustic modeling.

Paper · GitHub · Project · Details

MamTra — Figure 1

Figure 1 · Source

ManchuTTS

ManchuTTS builds multilevel text representations suited to Manchu and feeds them into a convolutional diffusion Transformer. Its non-autoregressive generator and augmented training data target speech synthesis where naturally recorded material is scarce.

Paper · GitHub: no author-linked repository found · Details

ManchuTTS — Paper figure

Paper figure · Source

Marco-Voice

Marco-Voice learns separate speaker and emotion representations using contrastive training. Rotating the emotional representation provides smooth expressive control while preserving the reference voice across different delivery styles.

Paper · GitHub · Details

Marco-Voice — Figure 1

Figure 1 · Source

MARS6

MARS6 encodes text and a speaker representation before generating hierarchical acoustic codes. Its compact encoder-decoder design targets expressive speech and reference-voice cloning with reduced inference cost.

Paper · GitHub · Project · Details

MARS6 — Paper figure

Paper figure · Source

Masked-style TTS

This controllable system first predicts a masked-autoencoder-derived speech style representation from text and controls. A second model generates codec tokens, allowing speaker characteristics and expressive attributes to be specified separately during synthesis.

Paper · GitHub: no author-linked repository found · Details

Masked-style TTS — Paper figure

Paper figure · Source

MaskGCT

Masked generative codec transformers.

Paper · Details

MaskGCT — Figure 1

Figure 1 · Source

Matcha-TTS

Conditional flow matching with an encoder-decoder acoustic model.

Paper · Details

Matcha-TTS — Figure 1

Figure 1 · Source

MAVE

MAVE combines a state-space backbone with cross-attention to generate speech conditioned on text and acoustic context. It supports zero-shot voice synthesis and editing while reducing the attention-memory requirements of a comparable Transformer-based codec model.

Paper · GitHub: no author-linked repository found · Details

MAVE — Figure 1

Figure 1 · Source

Mega-TTS

Disentangled speech factors and prosody LM.

Paper · Details

Mega-TTS — Figure 1

Figure 1 · Source

Mega-TTS 2

Prosody language model and multi-sentence prompting.

Paper · Details

Mega-TTS 2 — Figure 1

Figure 1 · Source

MegaTTS 3

MegaTTS 3 guides a latent diffusion Transformer with sparse text-speech alignment boundaries, leaving the model room to learn finer timing. Classifier-free guidance controls accent strength, while piecewise rectified flow reduces sampling work. The design targets robust zero-shot voice synthesis with more flexible alignment than a fully fixed duration sequence.

Paper · GitHub · Details

MegaTTS 3 — Figure 1

Figure 1 · Source

Meitei Mayek TTS

This Manipuri speech synthesizer maps Meitei Mayek writing to an ARPAbet-based phoneme representation before acoustic generation with Tacotron 2. A HiFi-GAN vocoder reconstructs the waveform. The paper develops a single-speaker system for a language with limited training resources and tonal pronunciation requirements.

Paper · GitHub: no author-linked repository found · Details

Meitei Mayek TTS — Paper figure

Paper figure · Source

Mel-LLM (TTS)

The synthesis experiment in Mel-LLM extends a language model to predict mel-based acoustic information directly. Its next-token VAE decoder demonstrates a text-to-speech path within an otherwise understanding-focused model; the paper presents this as a proof of concept with quality limitations.

Paper · GitHub: no author-linked repository found · Details

Mel-LLM (TTS) — Fig. 1 (paper page 2)

Fig. 1 (paper page 2) · Source

MELA-TTS

MELA-TTS predicts continuous mel-spectrogram frames from text and speaker conditions. A training-time alignment module connects the decoder to recognition-derived semantic features, helping the joint Transformer-diffusion model retain linguistic structure without discrete speech tokenization.

Paper · GitHub: no author-linked repository found · Details

MELA-TTS — Figure 1

Figure 1 · Source

MELD

MELD learns discrete latent variables from mel-spectrograms jointly with its speech language model. This shared optimization supports zero-shot synthesis and recognition while addressing omissions and excessive silence associated with less coordinated acoustic representations.

Paper · GitHub: no author-linked repository found · Details

MELD — Figures 1–2 (paper page 2; panels assembled)

Figures 1–2 (paper page 2; panels assembled) · Source

Mellotron

Tacotron 2 with global style tokens, pitch and rhythm conditioning.

Paper · Details

Mellotron — Editorial input/output diagram

Editorial input/output diagram · Source

MeloTTS

VITS-family multilingual speech synthesis.

GitHub · Details

MeloTTS — Editorial input/output diagram

Editorial input/output diagram · Source

Metis

Metis pretrains on unlabeled speech before adapting to task-specific conditions such as text. Self-supervised semantic tokens and acoustic codes support a shared foundation for reference-based TTS, conversion and other speech-generation tasks.

Paper · GitHub · Project · Details

Metis — Paper figure

Paper figure · Source

MFCIG-CSS

MFCIG-CSS represents dialogue history through separate graphs of meaning and vocal expression. Fine-grained multimodal interactions condition the speech synthesizer, helping each scripted response fit the surrounding conversation.

Paper · GitHub · Details

MFCIG-CSS — Figure 1

Figure 1 · Source

MiDashengLM-Gen

MiDashengLM-Gen trains a language model together with a conditional flow-matching output head to generate variable-length audio. Its text conditioning supports scenes containing intelligible speech alongside music or other sounds, with generation performed over continuous audio representations.

Paper · GitHub · Project · Details

MiDashengLM-Gen — Paper figure

Paper figure · Source

MiniMax-Speech

MiniMax-Speech extracts speaker characteristics directly from reference audio without requiring its transcript, then generates speech with an autoregressive Transformer and Flow-VAE. The research emphasizes multilingual zero-shot cloning and expressive delivery. Additional adaptation mechanisms support emotion control, description-based voice creation and more specialized voice cloning without replacing the base model.

Paper · GitHub: no author-linked repository found · Details

MiniMax-Speech — Figure 1

Figure 1 · Source

MixedG2P-T5

MixedG2P-T5 learns acoustic units from speech and uses a language-model synthesis path for text containing mixed scripts. It reduces dependence on manually designed grapheme-to-phoneme rules while retaining accent and intonation information in the speech representation.

Paper · GitHub: no author-linked repository found · Details

MixedG2P-T5 — Figure 3

Figure 3 · Source

MM-MovieDubber

MM-MovieDubber interprets scene information to distinguish dialogue, narration and monologue delivery. A speech generator then uses the resulting multimodal conditions with the target content to render expressive movie dubbing.

Paper · GitHub: no author-linked repository found · Details

MM-MovieDubber — Figure 2

Figure 2 · Source

MoE-TTS

MoE-TTS augments a frozen text language model with speech-specific expert parameters. Retaining the original language knowledge helps the synthesizer interpret unfamiliar style descriptions while learning the acoustic generation task.

Paper · GitHub: no author-linked repository found · Details

MoE-TTS — Figure 1

Figure 1 · Source

MoonCast

MoonCast combines podcast script preparation with a synthesizer trained for longer, spontaneous-sounding delivery. Voice references allow unseen speakers to render the resulting conversation, while discourse-level context supports more natural transitions than isolated sentence synthesis.

Paper · GitHub · Project · Details

MoonCast — Figure 1

Figure 1 · Source

MOSS-TTS

MOSS-TTS offers two generators over a shared discrete audio representation: a delay-pattern model and a model with a frame-local Transformer. They balance long-context control against efficient codebook prediction and speaker preservation. The family supports reference-conditioned synthesis, pronunciation and duration controls, with later checkpoints extending language handling and explicit pauses.

Paper · GitHub · Details

MOSS-TTS — Figure 2, PDF p. 8

Figure 2, PDF p. 8 · Source

MOSS-TTS-Nano

MOSS-TTS-Nano packages multilingual voice cloning into a roughly 100M-parameter speech generator with a compact audio tokenizer. Streaming output and an ONNX inference path make it relevant to CPU-based readers and local applications. Its published performance depends on the runtime and hardware, and the tokenizer is a separate part of the deployment footprint.

GitHub 1 · GitHub 2 · Details

MOSS-TTS-Nano — Official architecture diagram

Official architecture diagram · Source

MOSS-TTS-Realtime

MOSS-TTS-Realtime uses a Qwen3-derived backbone for linguistic context and a smaller local Transformer to predict audio codebooks. Text and speech tokens are handled at different levels so the system can accept text and emit audio incrementally. It targets low-latency spoken responses while preserving context across the generated utterance.

GitHub · Docs · Details

MOSS-TTS-Realtime — Editorial input/output diagram

Editorial input/output diagram · Source

MOSS-TTSD

MOSS-TTSD turns a dialogue script with explicit speaker tags into a continuous multi-party recording. Long-context modeling helps maintain speaker identity, turn assignment and acoustic continuity, while short references can define voices. It is designed for podcasts, commentary and other scripted conversations; the source paper evaluates dialogue-specific consistency as well as intelligibility.

Paper · GitHub · Details

MOSS-TTSD — Figure 2, PDF p. 4

Figure 2, PDF p. 4 · Source

MOSS-VoiceGenerator

MOSS-VoiceGenerator creates a speaking voice from a natural-language description rather than requiring an example speaker recording. Training on expressive cinematic speech exposes it to varied delivery and acoustic conditions. The model is intended for character design, storytelling and role-based narration where the desired voice must be specified in words.

Paper · GitHub · Details

MOSS-VoiceGenerator — Figure 1

Figure 1 · Source

MP-ELD

MP-ELD predicts low-rate continuous speech tokens through several information paths with separate local encoders. A flow decoder combines their predictions, while the accompanying Locodec representation is designed to limit accumulated errors during long speech generation.

Paper · GitHub: no author-linked repository found · Details

MP-ELD — Figure 2

Figure 2 · Source

MPE-TTS

MPE-TTS combines reference speech and textual prompts to specify an unseen speaker and the desired emotion. A prosody predictor and emotion-consistency objective carry those controls into the synthesized acoustic performance.

Paper · GitHub: no author-linked repository found · Details

MPE-TTS — Figure 1

Figure 1 · Source

Multistage multimodal TTS

This framework learns face and text conditioning in separate stages before using them for voice synthesis. Visual knowledge distillation and training across text-face and text-speech pairs reduce reliance on fully matched multimodal recordings.

Paper · GitHub: no author-linked repository found · Details

Multistage multimodal TTS — Paper figure

Paper figure · Source

Muyan-TTS

Muyan-TTS trains a speech language model on a large podcast collection for expressive reference-based synthesis. The release documents data preparation, training and optimized inference, with an emphasis on reproducible podcast-style voice generation.

Paper · GitHub · Details

Muyan-TTS — Figure 1

Figure 1 · Source

NaturalSpeech

Text-to-waveform VAE with enhanced prior and duration modeling.

Paper · Details

NaturalSpeech — Figure 1

Figure 1 · Source

NaturalSpeech 2

Latent diffusion over neural-codec representations.

Paper · Details

NaturalSpeech 2 — Figure 1

Figure 1 · Source

NaturalSpeech 3

Factorized speech codec and attribute-wise diffusion.

Paper · Details

NaturalSpeech 3 — Figure 3

Figure 3 · Source

NeuTTS Air

NeuTTS Air pairs a phoneme-conditioned language model with NeuCodec to synthesize a reference voice locally. Quantized GGUF backbones support incremental generation through the documented streaming backend. It targets embedded and desktop voice applications, with the codec's compute and memory requirements considered alongside those of the language-model backbone.

GitHub · Details

NeuTTS Air — Editorial input/output diagram

Editorial input/output diagram · Source

NeuTTS Nano

NeuTTS Nano reduces the speech-model backbone while retaining phoneme conditioning, reference-based cloning and NeuCodec reconstruction. Its English, German, French and Spanish models share a design but use separate language-specific checkpoints. Quantized variants support compact local deployments, and streaming depends on the selected inference backend.

GitHub · Details

NeuTTS Nano — Editorial input/output diagram

Editorial input/output diagram · Source

NeuTTS-2E

NeuTTS-2E accepts text directly and adds explicit emotional delivery to the NeuTTS language-model-and-codec pipeline. The released configuration supplies four fixed speaker presets instead of arbitrary reference-based cloning. It is intended for compact expressive speech applications, with streaming available through the documented GGUF inference path.

GitHub · Details

NeuTTS-2E — Editorial input/output diagram

Editorial input/output diagram · Source

NR-LauraTTS

NR-LauraTTS cleans the discrete representation of a noisy voice prompt before passing it to LauraTTS. Token prediction and embedding refinement reduce background contamination, supporting reference cloning when the available recording is acoustically imperfect.

Paper · GitHub · Project · Details

NR-LauraTTS — Figure 1

Figure 1 · Source

NVSpeech TTS

The NVSpeech pipeline includes a TTS model that renders text with explicitly marked nonverbal events. Word-level annotations connect ordinary speech with vocalizations such as laughter, providing a shared representation for recognition and controllable audio generation.

Paper · Project · GitHub: no author-linked repository found · Details

NVSpeech TTS — Figure 2

Figure 2 · Source

Nüshu-PitchVITS

Nüshu-PitchVITS uses pitch annotations from Nüshu's writing system to guide acoustic generation under very limited data. A frame-level pitch predictor conditions the VITS waveform path, allowing syllable recordings and linguistic tone knowledge to support sentence synthesis.

Paper · GitHub: no author-linked repository found · Details

Nüshu-PitchVITS — Figure 3

Figure 3 · Source

Ojibwe-Mi'kmaq-Maliseet TTS

This model family shares speech-synthesis training across three related Indigenous languages. The paper compares attention-based and attention-free flow architectures, demonstrating how joint linguistic coverage can help languages with limited recordings.

Paper · GitHub · Details

Ojibwe-Mi'kmaq-Maliseet TTS — Figure 1

Figure 1 · Source

OmniVoice

OmniVoice predicts multiple acoustic codebooks directly from text using a masked, non-autoregressive diffusion language model. Random masking across codebooks and initialization from a pretrained language model support multilingual generation without a separate text-to-semantic stage. It focuses on broad-language zero-shot synthesis and voice conditioning, rather than visual or general-purpose omni interaction.

Paper · GitHub · Details

OmniVoice — Figure 1

Figure 1 · Source

OpusLM

OpusLM extends text language models through speech-text pretraining on public data. Its interleaved representation supports speech recognition, text-conditioned synthesis and textual continuation within a transparent family of shared backbones.

Paper · GitHub: no author-linked repository found · Details

OpusLM — Figure 1

Figure 1 · Source

Orpheus TTS

Orpheus TTS repurposes a Llama-family language model to generate speech codec tokens from text. Emotion tags and speaker conditioning guide expressive delivery, while the project supplies inference and adaptation workflows. English releases and multilingual previews have different coverage, so a checkpoint's documented capabilities matter when selecting it for narration or a voice application.

GitHub · Details

Orpheus TTS — Editorial input/output diagram

Editorial input/output diagram · Source

OscillaTTS

OscillaTTS changes the periodic nonlinearities used in a style-diffusion synthesis backbone. Adjustable oscillatory modulation is designed to capture rapid pitch and amplitude changes while a linear bypass stabilizes the acoustic signal, targeting sharper expressive prosody.

Paper · GitHub: no author-linked repository found · Details

OscillaTTS — Figure 1

Figure 1 · Source

OuteTTS

OuteTTS represents speech in a form that can be generated by a decoder-only language model and reconstructed by an audio decoder. A speaker reference guides vocal identity, style and accent. Its 1.0 line supports standard LLM serving backends, but generation settings and repetition handling need to follow the matching model implementation.

GitHub · Details

OuteTTS — Editorial input/output diagram

Editorial input/output diagram · Source

OV-InstructTTS

OV-InstructTTS interprets voice and delivery descriptions beyond a fixed inventory of style labels. Its reasoning-based conditioning connects broader textual requests with expressive speech generation, supported by a dedicated instruction-speech dataset.

Paper · GitHub · Details

OV-InstructTTS — Figure 2

Figure 2 · Source

OZSpeech

OZSpeech generates disentangled speech components with a flow model conditioned on a learned prior. The design targets single-step zero-shot synthesis while separately modeling content, prosody and speaker-related information from the voice prompt.

Paper · GitHub · Project · Details

OZSpeech — Figure 1

Figure 1 · Source

PALLE

PALLE generates variable-length speech spans at fixed decoding steps, combining temporal planning with parallel token prediction. A second non-autoregressive stage refines the initial sequence, supporting efficient zero-shot synthesis.

Paper · GitHub: no author-linked repository found · Details

PALLE — Figure 3

Figure 3 · Source

Parallel GPT

Parallel GPT divides speech generation between a general autoregressive predictor and a non-autoregressive detail model. The parallel refinement stage conditions on the initial tokens, balancing independence and interaction between semantic and acoustic information.

Paper · GitHub: no author-linked repository found · Details

Parallel GPT — Paper figure

Paper figure · Source

Parallel Tacotron

Parallel acoustic model with a variational residual encoder.

Paper · Details

Parallel Tacotron — Figure 1

Figure 1 · Source

Parallel Tacotron 2

Parallel synthesis with differentiable duration modeling.

Paper · Details

Parallel Tacotron 2 — Figure 1

Figure 1 · Source

ParaStyleTTS

ParaStyleTTS converts textual style prompts into separate controls for prosody and broader paralinguistic characteristics. The lightweight adaptation design targets expressive speech from descriptions while making the roles of the two conditioning levels explicit.

Paper · GitHub · Project · Details

ParaStyleTTS — Figure 1

Figure 1 · Source

Parler-TTS

Description-conditioned codec language model.

GitHub · Paper · Details

Parler-TTS — Figure 1

Figure 1 · Source

Parler-TTS Hinglish adaptation

This Parler-TTS extension introduces language-specific phonetic alignment and emotion embeddings for Hindi and Indian English. Its conditioning targets code-switched utterances whose accent and emotional delivery change coherently across language boundaries.

Paper · GitHub · Details

Parler-TTS Hinglish adaptation — Figure 2

Figure 2 · Source

PFluxTTS

PFluxTTS combines two acoustic-generation paths by fusing their predicted vector fields at inference. Sequential reference embeddings support transcript-free cross-language voice cloning, and a super-resolution vocoder reconstructs high-rate output audio.

Paper · Project · GitHub: no author-linked repository found · Details

PFluxTTS — Figure 1

Figure 1 · Source

Phoenix TTS

Phoenix TTS aligns its speech tokenizer with the downstream flow-matching acoustic decoder during training. An autoregressive language model predicts the resulting discrete representation, supporting text-to-speech and voice conversion without separating token design from acoustic reconstruction.

Paper · GitHub: no author-linked repository found · Details

Phoenix TTS — Figure 2

Figure 2 · Source

Phoneme-tone adaptive Thai TTS

This Thai speech synthesizer encodes phonemes and tones with a language-specific BERT model, then predicts duration, pitch and energy for a GAN-trained waveform decoder. A reference-derived style vector supports voice cloning. Multilingual pretraining of acoustic feature extractors and Thai adaptation address limited language-specific data.

Paper · GitHub: no author-linked repository found · Details

Phoneme-tone adaptive Thai TTS — Figure 2

Figure 2 · Source

PilotTTS

PilotTTS uses paired recordings and Q-Former conditioning to separate a speaker's identity from delivery style. The model supports reference cloning, emotional and nonverbal expression, and Chinese dialect synthesis within a shared autoregressive pipeline.

Paper · GitHub · Details

PilotTTS — Figure 3

Figure 3 · Source

Piper (VITS voices)

VITS voice models exported for local inference.

GitHub · Docs · Details

Piper (VITS voices) — Editorial input/output diagram

Editorial input/output diagram · Source

Pocket TTS

Pocket TTS uses continuous autoregressive speech modeling with a flow-based output mechanism, avoiding long sequences of discrete acoustic codebooks. It combines a small language-model backbone with streaming audio reconstruction and reusable voice conditioning. The project targets CPU-based speech synthesis; language-specific models and runtime choices affect its speed and voice behavior.

GitHub · Paper · Details

Pocket TTS — Figure 1

Figure 1 · Source

PortaSpeech

Variational acoustic model and flow-based post-net.

Paper · Details

PortaSpeech — Figure 1, PDF p. 4

Figure 1, PDF p. 4 · Source

PROEMO

PROEMO combines emotional prompts with an explicit intensity control in a multi-speaker synthesizer. The conditioning adjusts delivery strength and prosodic variation, allowing the same spoken text to be rendered with different emotional performances.

Paper · GitHub: no author-linked repository found · Details

PROEMO — Figure 1

Figure 1 · Source

Progressive face-conditioned TTS

This face-conditioned synthesizer combines local facial regions into progressively broader visual representations. Joint visual and acoustic attribute learning and multiple photographs of each training speaker align the face representation with voice characteristics, conditioning speech generation on text and a face image.

Paper · GitHub: no author-linked repository found · Details

Progressive face-conditioned TTS — Figure 1

Figure 1 · Source

Prompt-Unseen-Emotion

Prompt-Unseen-Emotion learns the relationship between emotion descriptions and speech using a language-model synthesis backbone. Weighted combinations of known emotions and contextual language knowledge allow expressive delivery outside the original categorical training labels.

Paper · GitHub: no author-linked repository found · Details

Prompt-Unseen-Emotion — Paper figure

Paper figure · Source

PromptTTS

Style and content text encoders with a speech decoder.

Paper · Details

PromptTTS — Figure 1

Figure 1 · Source

PromptTTS 2

Prompt-conditioned TTS with a diffusion variation network.

Paper · Details

PromptTTS 2 — Figure 1

Figure 1 · Source

ProtoDisent-TTS

ProtoDisent-TTS learns a codebook of healthy and dysarthric articulation patterns separately from speaker identity. Adversarial constraints reduce pathological information in the speaker representation, enabling controlled synthesis of articulation characteristics in a target voice.

Paper · Project · GitHub: no author-linked repository found · Details

ProtoDisent-TTS — Figure 1

Figure 1 · Source

PS-TTS

PS-TTS uses vowel-based alignment to coordinate the timing and phonetic structure of dubbed speech. Its PS-Comet variant also considers semantic preservation when choosing translated text, connecting translation choices with a TTS rendering stage.

Paper · GitHub: no author-linked repository found · Details

PS-TTS — Fig. 1 (paper page 3)

Fig. 1 (paper page 3) · Source

QTTS

QTTS predicts residual speech codes produced by its QDAC tokenizer. Hierarchical parallel and delayed multihead variants organize codebook dependencies differently, offering alternative balances between acoustic detail and sequential decoding cost.

Paper · GitHub: no author-linked repository found · Details

QTTS — Figure 2

Figure 2 · Source

Qwen-Audio-3.0-TTS

Qwen-Audio-3.0-TTS combines compact semantic speech tokens with progressively trained language and acoustic models. Natural-language instructions and inline tags control delivery, while multilingual reference conditioning supports voice cloning and longer speech generation under varied recording conditions.

Paper · Project · GitHub: no author-linked repository found · Details

Qwen-Audio-3.0-TTS — Figure 3

Figure 3 · Source

Qwen3-TTS

Qwen3-TTS combines a dual-track speech language model with tokenizers designed for compact streaming audio. The released 12Hz line separates Base voice cloning, CustomVoice preset-speaker control and VoiceDesign creation from descriptions. These variants support different conditioning interfaces, allowing applications to choose between reproducing a reference voice and directing a new voice through text.

Paper · GitHub · Details

Qwen3-TTS — Figure 3

Figure 3 · Source

RADKA-CSS

RADKA-CSS retrieves dialogue examples related to the current conversation in both meaning and delivery. A graph-based aggregation mechanism combines their style information with current context, conditioning expressive conversational speech synthesis.

Paper · GitHub · Project · Details

RADKA-CSS — Figure 2

Figure 2 · Source

RALL-E

Prosody-guided codec language modeling.

Paper · Details

RALL-E — Figure 1

Figure 1 · Source

Raon-OpenTTS

Raon-OpenTTS is a family of reference-conditioned diffusion synthesizers trained on a large, documented English speech collection. The release pairs its models with data processing and evaluation resources, allowing robustness across varied acoustic conditions to be examined alongside clean-speech quality.

Paper · GitHub · Details

Raon-OpenTTS — Figure 1

Figure 1 · Source

RapFlow-TTS

RapFlow-TTS regularizes the acoustic velocity field so longer generation steps remain consistent. Time-interval scheduling and adversarial objectives improve the resulting few-step synthesizer, reducing the iterations needed to render reference-conditioned speech.

Paper · GitHub · Details

RapFlow-TTS — Figure 1

Figure 1 · Source

ReGenVoice

ReGenVoice applies the ReGen representation-and-waveform modeling approach to text-to-speech. Multiple levels of generated conditioning help reconstruct detailed waveforms from compressed latents, linking efficient acoustic representation with reference-conditioned speech synthesis.

Paper · Project · GitHub: no author-linked repository found · Details

ReGenVoice — Figure 1

Figure 1 · Source

ReStyle-TTS

ReStyle-TTS changes vocal attributes relative to a reference recording. Independent text and reference guidance, composable style adapters and timbre-consistency optimization allow continuous expressive edits while limiting changes to the speaker's identity.

Paper · GitHub: no author-linked repository found · Details

ReStyle-TTS — Figure 1

Figure 1 · Source

RTFree-F5

RTFree-F5 replaces the transcript normally associated with an F5-TTS reference recording with projected speech features. A lightweight adapter reuses the pretrained generator, enabling transcript-free voice conditioning, including references whose pronunciation makes transcription unreliable.

Paper · GitHub: no author-linked repository found · Details

RTFree-F5 — Figure 1

Figure 1 · Source

RV-TTS

Revival with Voice learns voice identity from face images and delivery attributes from descriptions. Audio-only training data and stylized portrait augmentation broaden its input coverage, enabling controlled speech from real faces or artistic portraits.

Paper · GitHub: no author-linked repository found · Details

RV-TTS — Figure 1

Figure 1 · Source

RWKVTTS

RWKVTTS uses the recurrent RWKV-7 architecture for speech synthesis in place of a conventional Transformer backbone. Its token-generation path targets efficient streaming and reduced state-management cost while conditioning audio on the supplied text.

Paper · GitHub · Details

RWKVTTS — Figure 2

Figure 2 · Source

S5-TTS

S5-TTS adapts T5-TTS for word-by-word synthesis using limited future text. Lookahead-aware masks, convolutional auxiliary attention and distillation let the model start speaking before the complete sentence is available while retaining voice conditioning and alignment.

Paper · GitHub: no author-linked repository found · Details

S5-TTS — Figure 1

Figure 1 · Source

Sarashina2.2-TTS

Sarashina2.2-TTS emphasizes reliable Japanese pronunciation, including characters with several possible readings. Its semantic language model and acoustic flow decoder use reference speech for voice conditioning, with balanced multilingual training to reduce dependence on the reference language.

Paper · GitHub · Details

Sarashina2.2-TTS — Figure 1

Figure 1 · Source

SASLM

SASLM derives expressive intent from its own evolving semantic states through an information bottleneck. Acoustic feedback aligns generated speech with that intent, reducing the need for externally supplied emotion labels in context-sensitive speech rendering.

Paper · GitHub · Project · Details

SASLM — Figure 3

Figure 3 · Source

Seed-TTS

Autoregressive speech foundation model.

Paper · Details

Seed-TTS — Figure 1

Figure 1 · Source

Self-distilled zero-shot TTS

This zero-shot synthesizer learns linguistic content and reference-speaker attributes through separate representations. Two-stage self-distillation creates aligned examples that strengthen their separation, targeting stable voice cloning with a small inference footprint.

Paper · GitHub: no author-linked repository found · Details

Self-distilled zero-shot TTS — Figure 1

Figure 1 · Source

SelfTTS

SelfTTS learns separate representations of a speaker's identity and emotional delivery using contrastive and adversarial objectives. It then improves synthesis through self-generated training examples, allowing emotion transfer to speakers originally recorded with neutral expression.

Paper · GitHub · Project · Details

SelfTTS — Figure 1

Figure 1 · Source

SemaVoice

SemaVoice organizes its audio VAE latents using guidance from speech foundation-model representations. A continuous autoregressive backbone and patch-level diffusion head then synthesize reference-conditioned speech with greater emphasis on linguistic coherence.

Paper · GitHub: no author-linked repository found · Details

SemaVoice — Figure 1

Figure 1 · Source

SemBridge

SemBridge uses discrete semantic targets during training to organize both acoustic latents and language-model hidden states. The resulting continuous generator supports zero-shot speech synthesis with stronger content alignment, without needing to generate the auxiliary semantic tokens during inference.

Paper · GitHub · Details

SemBridge — Figure 1

Figure 1 · Source

Shallow Flow Matching TTS

Shallow Flow Matching adds a lightweight head that predicts an intermediate acoustic state for a flow-based synthesizer. Starting refinement closer to the target reduces the remaining generation path, allowing coarse-to-fine speech synthesis with less iterative work.

Paper · GitHub · Details

Shallow Flow Matching TTS — Figure 2

Figure 2 · Source

SLED

SLED learns the conditional distribution of acoustic latents using an energy-distance objective rather than discrete token classification. Its autoregressive generator samples continuous speech representations, simplifying synthesis while retaining acoustic detail.

Paper · GitHub · Details

SLED — Figure 2

Figure 2 · Source

SlimSpeech

SlimSpeech reduces the parameter count of a rectified-flow TTS model and transfers knowledge into a lightweight generator. It targets efficient reference-conditioned speech synthesis while retaining the acoustic quality of a larger teacher.

Paper · GitHub: no author-linked repository found · Details

SlimSpeech — Paper figure

Paper figure · Source

SMLLE

SMLLE uses a transducer to align incoming text with semantic speech tokens and duration information. A separate autoregressive stage generates acoustic frames, with controlled access to future text stabilizing incremental synthesis.

Paper · Project · GitHub: no author-linked repository found · Details

SMLLE — Figure 1

Figure 1 · Source

SoulX-Podcast

SoulX-Podcast synthesizes conversational scripts with reference voices, dialect choices and nonverbal expression. Its long-form training targets consistent speaker identity and natural transitions across turns, while also supporting ordinary single-speaker TTS.

Paper · GitHub · Project · Details

SoulX-Podcast — Figure 3

Figure 3 · Source

Spark-TTS

Spark-TTS uses BiCodec to separate changing linguistic content from global speaker attributes, then predicts these tokens with a Qwen2.5 backbone. That separation supports both reference-based cloning and direct control of attributes such as speaking rate and pitch. It is useful for controllable speech generation where a reference recording alone is insufficient to specify the desired delivery.

Paper · GitHub · Details

Spark-TTS — Figure 3

Figure 3 · Source

SpeakStream

SpeakStream trains on text interleaved with corresponding speech and generates audio as new text becomes available. The synthesis module remains compatible with an upstream text-streaming language model, supporting responsive conversational playback.

Paper · Project · GitHub: no author-linked repository found · Details

SpeakStream — Paper figure

Paper figure · Source

SPEAR-TTS

Text-to-semantic and semantic-to-acoustic LMs.

Paper · Details

SPEAR-TTS — Figure 1

Figure 1 · Source

SpeechAccentLLM

SpeechAccentLLM uses a content tokenizer trained with transcription alignment and jointly learns accent conversion and synthesis. A reconstruction refinement stage improves generated speech while separating accent-related changes from the target speaker's identity.

Paper · GitHub: no author-linked repository found · Details

SpeechAccentLLM — Figure 3

Figure 3 · Source

SpeechEdit

SpeechEdit combines text, instruction tokens and reference audio in a shared codec-language-model sequence. Paired examples that differ in selected attributes teach localized changes while retaining other aspects of the reference voice and delivery.

Paper · Project · GitHub: no author-linked repository found · Details

SpeechEdit — Figure 1

Figure 1 · Source

SpeechT5

Shared encoder–decoder with modality interfaces.

Paper · Details

SpeechT5 — Figure 2

Figure 2 · Source

SpeechX

Prompted neural codec language model.

Paper · Details

SpeechX — Figure 1

Figure 1 · Source

SpeedySpeech

Residual convolutional acoustic model with duration expansion.

Paper · Details

SpeedySpeech — Figure 3

Figure 3 · Source

Spotlight-TTS

Spotlight-TTS extracts expressive reference information primarily from voiced speech and adjusts the resulting style direction before acoustic generation. The method targets more faithful expression transfer; the catalog combines its related paper records under one model.

Paper 1 · Paper 2 · GitHub: no author-linked repository found · Details

Spotlight-TTS — Figure 1

Figure 1 · Source

StellarTTS

StellarTTS encodes phoneme timing sparsely and uses a lightweight masked Transformer to generate speech tokens in parallel. A semantic-aware codec supports waveform reconstruction, while explicit timing representations provide control over pronunciation, duration and prosody.

Paper · Project · GitHub: no author-linked repository found · Details

StellarTTS — Paper figure

Paper figure · Source

Step-Audio-EditX

Step-Audio-EditX supports reference-based synthesis and repeated edits to emotion, speaking style or nonverbal delivery. Training on deliberately contrasting synthetic examples teaches the model to follow expressive changes without a separate attribute-embedding module.

Paper · GitHub · Details

Step-Audio-EditX — Figure 2

Figure 2 · Source

Step-Audio-TTS

Step-Audio-TTS is the compact synthesis component produced through the broader Step-Audio speech-data and distillation pipeline. Text and voice conditioning drive speech-token generation, while the released workflow exposes delivery controls. The TTS checkpoint is used to render supplied content; reasoning, tool use and dialogue management belong to other components of the system.

Paper · GitHub · Details

Step-Audio-TTS — Figure 2

Figure 2 · Source

StepAudio 2.5 TTS

The TTS mode of StepAudio 2.5 uses a shared speech-language foundation with synthesis-specific decoding and preference training. Rich contextual supervision and feedback target controllable expression; this card describes its speech-rendering path within the broader system.

Paper · GitHub: no author-linked repository found · Details

StepAudio 2.5 TTS — Figure 1

Figure 1 · Source

Stochastic-alignment continuous TTS

This synthesizer predicts continuous speech latents using a Gaussian-mixture conditional distribution. A stochastic monotonic alignment mechanism keeps the acoustic sequence ordered against the text, offering an alternative to autoregressive discrete-codec modeling.

Paper · GitHub: no author-linked repository found · Details

Stochastic-alignment continuous TTS — Figure 1

Figure 1 · Source

StreamMel

StreamMel alternates text tokens with continuous acoustic frames in one streaming synthesis model. This organization lets incoming text guide speech immediately while retaining reference-voice information across the generated audio stream.

Paper · GitHub: no author-linked repository found · Details

StreamMel — Paper figure

Paper figure · Source

StyleTTS

Style-conditioned parallel synthesis with a transferable aligner.

Paper · Details

StyleTTS — Figure 1

Figure 1 · Source

StyleTTS 2

Style diffusion and adversarial training with speech-model discriminators.

Paper · Details

StyleTTS 2 — Figure 1, PDF p. 4

Figure 1, PDF p. 4 · Source

Supertonic

The SupertonicTTS research system compresses speech into continuous latents and predicts them from character-level text with flow matching. ConvNeXt blocks, temporal compression and a separate duration predictor keep synthesis compact. The later Supertonic ONNX release exposes preset voice-style assets; its packaged configurations should not be equated with the paper's 44M-parameter research model.

Model card · Paper · GitHub · Project · Details

Supertonic — Figure 1

Figure 1 · Source

Supertonic 2

Supertonic 2 extends the local ONNX synthesis line to five languages while retaining voice-style conditioning and a compact model. It provides a practical path to multilingual narration on devices that can run the supplied inference stack. Creating a new voice-style asset is a separate workflow from generating speech with an existing asset.

Model card · Details

Supertonic 2 — Editorial input/output diagram

Editorial input/output diagram · Source

Supertonic 3

Supertonic 3 expands language coverage and adds expression tags while keeping local ONNX inference and preset voice styles. The release targets more reliable reading across short and long text, with controls for events such as breaths or laughter. Custom voice-style creation is offered through a separate service; downloaded styles can then condition local synthesis.

Model card · Details

Supertonic 3 — Editorial input/output diagram

Editorial input/output diagram · Source

SwanVoice

SwanVoice generates monologues or dialogues with up to four speakers using raw text, voice references and speaker-turn conditions. Pause markers and optional pronunciation substitutions provide textual control, while staged dialogue training supports longer expressive speech.

Paper · Project · GitHub: no author-linked repository found · Details

SwanVoice — Figure 2

Figure 2 · Source

SyncSpeech

SyncSpeech uses a temporal masking scheme to coordinate sequential speech structure with parallel token decoding. This hybrid organization targets faster first audio and higher throughput while retaining reference-conditioned synthesis quality.

Paper · Project · GitHub: no author-linked repository found · Details

SyncSpeech — Figure 1

Figure 1 · Source

Tacotron

Attention-based recurrent spectrogram synthesis.

Paper · Details

Tacotron — Figure 1

Figure 1 · Source

Tacotron 2

Recurrent attention model and WaveNet vocoder.

Paper · Details

Tacotron 2 — Figure 1

Figure 1 · Source

TADA

TADA aligns text tokens one-to-one with continuous acoustic units. A language model with a flow-matching head predicts these synchronized representations, reducing the ambiguity of text-speech alignment during reference-conditioned synthesis.

Paper · GitHub · Details

TADA — Figure 2

Figure 2 · Source

TED-TTS

TED-TTS modifies conditioning and decoding in a pretrained zero-shot synthesizer to control different parts of an utterance. Segment-specific emotion masks and duration steering support local changes while coordinating transitions and the overall stopping point.

Paper · GitHub · Details

TED-TTS — Figure 1

Figure 1 · Source

Tibetan-TTS

Tibetan-TTS adapts a large speech generator through language-specific text representation, tokenizer changes and cross-language training. Its data preparation and quality enhancement target scarce Tibetan recordings and the differences between written forms and spoken pronunciation.

Paper · GitHub: no author-linked repository found · Details

Tibetan-TTS — Figure 2

Figure 2 · Source

<a

Truncated — view the full README on GitHub.

Contributors

kadirnar

6 commits

kadirnar/awesome-tts-architectures

A visual catalog of text-to-speech architectures, with model diagrams, concise descriptions, and primary sources.

Python

18

6 commits

updated Sep 13, 2026

See the code

README

Awesome TTS Architectures Awesome

A visual catalog of text-to-speech models, from Tacotron to speech language models. Diagrams, primary sources and short notes for every entry.

380 models and families · Reviewed 2026-09-13

Model list · All diagrams · 2025+ TTS-arxiv-daily collection · Descriptions · Timeline · Methodology · Contribute

The TTS-arxiv-daily collection covers the source list's TTS systems with first paper submissions from January 1, 2025 onward. Every included family has an image, a description, paper links and an explicit GitHub availability status. The complete screening record is available as JSON.

Models

T: text · S: speech or voice reference · A: other audio · I: image · V: video. Scope and labels.

Alphabetical model list · 380 entries
ModelGroupInput → output
A2TTSDiffusion / flowT, S → S
AffectronToken LMT, S → S
AgentSteerTTSToken LMT, S → S
AlignDiTDiffusion / flowT, S, V → S
AMNetParallelT → S
ARCHI-TTSDiffusion / flowT, S → S
ATRIEToken LMT → S
Audiobook-CCToken LMT, S → S
AuEmoChatToken LMT, S, V → S
AuKDiffusion / flowT, S → S, A
Authentic-DubberDiffusion / flowT, S, V → S
AutoSIFTDiffusion / flowT, S → S
AutoStyle-TTSToken LMT, S → S
AVLM (expressive speech)Token LMT, S, V → S
Bagpiper-TTSToken LMT → S
BareWaveDiffusion / flowT, S → S
BarkToken LMT → S, A
BASE TTSToken LMT, S → S
BatonTTS (BatonVoice)Token LMT, S → S
BELLEContinuous LMT, S → S
BitTTSCompactT → S
Block-wise Mimi TTSToken LMT → S
BnTTSToken LMT, S → S
BolboshDiffusion / flowT → S
Borderless Long Speech SynthesisContinuous LMT, S → S
BreezyVoiceToken LMT, S → S
BridgeTTSToken LMT, S → S
BVSToken LMT, V → S, A
CAM-TTSToken LMT, S → S
CapTalkToken LMT, S → S
CAST-TTSDiffusion / flowT, S → S
CaT-TTSToken LMT, S → S
Causal-prosody FastSpeech 2ParallelT → S
CDE-StyleTTSDiffusion / flowT, S → S
Chain-of-Details TTSToken LMT, S → S
Chain-TalkerToken LMT, S → S
ChatterboxToken LMT, S → S
Chatterbox-FlashToken LMT, S → S
ChatTTSToken LMT → S
CLaM-TTSToken LMT, S → S
CLEARContinuous LMT, S → S
Clip-TTSParallelT → S
Compact neural accessibility TTSCompactT → S
Compressed-to-fine speech LMToken LMT, S → S
Confucius4-TTSToken LMT, S → S
Continuous-token diffusion TTSContinuous LMT, S → S
Controllable masked-speech TTSToken LMT, S, A → S
CookVoiceDiffusion / flowT, S → S
CosyEdit2Token LMT, S → S
CoSyncDiTDiffusion / flowT, S, V → S
CosyVoiceToken LMT, S → S
CosyVoice 2Token LMT, S → S
CosyVoice 3Token LMT, S → S
CosyWhisper (WhispSynth)Token LMT, S → S
CoVoMix2Diffusion / flowT, S → S
Cross-Lingual F5-TTSDiffusion / flowT, S → S
CrossAccent-TTSToken LMT, S → S
CSMToken LMT, S → S
CTC-TTSToken LMT, S → S
CtrlSpeechContinuous LMT, S → S
CuteTTSContinuous LMT, S → S
DAIEN-TTSDiffusion / flowT, S, A → S
DARSDiffusion / flowT → S
DCARToken LMT, S → S
Deep VoiceAutoregressiveT → S
Deep Voice 2AutoregressiveT → S
Deep Voice 3AutoregressiveT → S
DeepASMRToken LMT, S → S
DeepDubber-V1Diffusion / flowT, V → S
DeepDubbingToken LMT, S → S
DelightfulTTSParallelT → S
DELTA-TTSToken LMT, S → S
DepFlowDiffusion / flowT, S → S
DiaToken LMT, S → S
Dia2Token LMT, S → S
DialoSpeechToken LMT, S → S
DiEmo-TTSParallelT, S → S
Diff-TTSDiffusion / flowT → S
DiffCSSToken LMT, S → S
DiFlow-TTSToken LMT, S → S
DiFlowDubberToken LMT, S, V → S
DisCo-SpeechToken LMT, S → S
DisSpeechToken LMT → S
DiSTARToken LMT, S → S
DiTARContinuous LMT, S → S
DiTTo-TTSDiffusion / flowT, S → S
DMOSpeech 2Diffusion / flowT, S → S
DMP-TTSDiffusion / flowT, S → S
dots.ttsContinuous LMT, S → S
Dragon-FMToken LMT, S → S
DrawSpeechDiffusion / flowT, I → S
DS-TTSDiffusion / flowT, S → S
DualDubToken LMT, V → S
DualSpeechLMToken LMT, S → S
E2 TTSDiffusion / flowT, S → S
ECTSpeechDiffusion / flowT, S → S
ELLA-VToken LMT, S → S
EME-TTSParallelT → S
EMM-TTSToken LMT, S → S
EmojiVoiceDiffusion / flowT → S
EmoShiftToken LMT → S
EmoSSLSphereToken LMT → S
EmoSteer-TTSDiffusion / flowT, S → S
Emotion-timbre disentangled TTSParallelT, S → S
EmotiVoiceParallelT → S
EmoTra-TTSToken LMT, S → S
EmoVoiceToken LMT → S
End-to-end discrete-token TTSToken LMT, S → S
F5-TTSDiffusion / flowT, S → S
F5R-TTSDiffusion / flowT, S → S
Face-adapted StyleTTS 2Diffusion / flowT, I → S
FaceSpeakDiffusion / flowT, I → S
FacialTalkerToken LMT, S, V → S
FastPitchParallelT → S
FastSpeechParallelT → S
FastSpeech 2ParallelT → S
FC-TTSToken LMT, S → S
FELLEContinuous LMT, S → S
FineCombo-TTSDiffusion / flowT, S → S
FireRedAudioContinuous LMT, S → S, A
FireRedTTSToken LMT, S → S
FireRedTTS-1SToken LMT, S → S
FireRedTTS-2Token LMT, S → S
FireRedTTS3Continuous LMT, S → S
Fish Audio S1 / OpenAudio S1Token LMT, S → S
Fish Audio S2Token LMT, S → S
Fish SpeechToken LMT, S → S
Flamed-TTSDiffusion / flowT, S → S
FlashTTSToken LMT, S → S
FleSpeechToken LMT, S, I → S
FlexiVoiceToken LMT, S → S
FlexSpeechDiffusion / flowT, S → S
FlowtronFlow / VAET, S → S
FNH-TTSFlow / VAET, S → S
Frame-stacked local Transformer TTSToken LMT, S → S
FreyaTTSDiffusion / flowT → S
Gemini 2.5 TTSAPIT → S
Gemini 3.1 Flash TTSAPIT → S
GibbsTTSToken LMT, S → S
GLM-TTSToken LMT, S → S
Glow-TTSFlow / VAET → S
GOAT-TTSToken LMT, S → S
GPAToken LMT, S → S
GPT-4o Mini TTSAPIT → S
GPT-SoVITSToken LMT, S → S
Grad-TTSDiffusion / flowT → S
GRAFTToken LMT, S → S
GSA-TTSParallelT, S → S
GST-TacotronAutoregressiveT, S → S
HabibiDiffusion / flowT, S → S
HD-PPTToken LMT, S → S
Higgs Audio v2Token LMT, S → S
Higgs Audio v2.5Token LMT, S → S
Higgs Audio v3 TTSToken LMT, S → S
HiStyleDiffusion / flowT → S
HoliDubberContinuous LMT, V → S
HoliTok (TTS)Continuous LMT, S → S
Hume Octave TTSAPIT, S → S
ImmersiveTTSDiffusion / flowT, S, A → S, A
IndexTTSToken LMT, S → S
IndexTTS 2.5Token LMT, S → S
IndexTTS2Token LMT, S → S
InstructAudioDiffusion / flowT → S
IntMeanFlowDiffusion / flowT, S → S
Inworld TTS-1Token LMT → S
JaiTTSContinuous LMT, S → S
JAM-FlowDiffusion / flowT, S, V → S
JELLYToken LMT, S → S
JETSParallelT → S
Joint non-autoregressive STT-TTSParallelT → S
JoycentDiffusion / flowT, S → S
JoyVoiceToken LMT, S → S
KABURI-TTSDiffusion / flowT → S
KittenTTSCompactT → S
Koel-TTSToken LMT, S → S
KokoroCompactT → S
Kyutai TTS (DSM)Token LMT, S → S
LanStyleTTSParallelT, S → S
LatinXToken LMT, S → S
LE2E-TTSCompactT → S
LightSpeechParallelT → S
LLaDA-TTSToken LMT, S → S
LlasaToken LMT, S → S
Llasa+Token LMT, S → S
LLMVoXToken LMT → S
Lombard Matcha-TTSDiffusion / flowT → S
LongCat-AudioDiTDiffusion / flowT, S → S
LoRP-TTSDiffusion / flowT, S → S
Luna-TTSToken LMT, S → S
M3-TTSDiffusion / flowT, S → S
MAGIC-TTSToken LMT, S → S
MagpieTTS-LFToken LMT, S → S
MambaVoiceCloningDiffusion / flowT, S → S
MamTraDiffusion / flowT, S → S
ManchuTTSDiffusion / flowT → S
Marco-VoiceToken LMT, S → S
MARS6Token LMT, S → S
Masked-style TTSToken LMT, S → S
MaskGCTToken LMT, S → S
Matcha-TTSDiffusion / flowT → S
MAVEToken LMT, S → S
Mega-TTSToken LMT, S → S
Mega-TTS 2Token LMT, S → S
MegaTTS 3Diffusion / flowT, S → S
Meitei Mayek TTSAutoregressiveT → S
Mel-LLM (TTS)Continuous LMT → S
MELA-TTSContinuous LMT, S → S
MELDToken LMT, S → S
MellotronAutoregressiveT, S → S
MeloTTSFlow / VAET → S
MetisToken LMT, S → S
MFCIG-CSSToken LMT, S, V → S
MiDashengLM-GenContinuous LMT → S, A
MiniMax-SpeechToken LMT, S → S
MixedG2P-T5Token LMT, S → S
MM-MovieDubberDiffusion / flowT, V → S
MoE-TTSToken LMT → S
MoonCastToken LMT, S → S
MOSS-TTSToken LMT, S → S
MOSS-TTS-NanoToken LMT, S → S
MOSS-TTS-RealtimeToken LMT, S → S
MOSS-TTSDToken LMT, S → S
MOSS-VoiceGeneratorToken LMT → S
MP-ELDContinuous LMT, S → S
MPE-TTSToken LMT, S → S
Multistage multimodal TTSDiffusion / flowT, I → S
Muyan-TTSToken LMT, S → S
NaturalSpeechFlow / VAET → S
NaturalSpeech 2Diffusion / flowT, S → S
NaturalSpeech 3Diffusion / flowT, S → S
NeuTTS AirToken LMT, S → S
NeuTTS NanoToken LMT, S → S
NeuTTS-2EToken LMT → S
NR-LauraTTSToken LMT, S → S
NVSpeech TTSToken LMT, S → S
Nüshu-PitchVITSFlow / VAET → S
Ojibwe-Mi'kmaq-Maliseet TTSDiffusion / flowT → S
OmniVoiceToken LMT, S → S
OpusLMToken LMT, S → S
Orpheus TTSToken LMT, S → S
OscillaTTSDiffusion / flowT, S → S
OuteTTSToken LMT, S → S
OV-InstructTTSToken LMT → S
OZSpeechToken LMT, S → S
PALLEToken LMT, S → S
Parallel GPTToken LMT, S → S
Parallel TacotronParallelT → S
Parallel Tacotron 2ParallelT → S
ParaStyleTTSParallelT → S
Parler-TTSToken LMT → S
Parler-TTS Hinglish adaptationToken LMT → S
PFluxTTSDiffusion / flowT, S → S
Phoenix TTSToken LMT, S → S
Phoneme-tone adaptive Thai TTSParallelT, S → S
PilotTTSToken LMT, S → S
Piper (VITS voices)Flow / VAET → S
Pocket TTSContinuous LMT, S → S
PortaSpeechFlow / VAET → S
PROEMOParallelT → S
Progressive face-conditioned TTSFlow / VAET, I → S
Prompt-Unseen-EmotionToken LMT → S
PromptTTSParallelT → S
PromptTTS 2Diffusion / flowT → S
ProtoDisent-TTSFlow / VAET, S → S
PS-TTSToken LMT, S → S
QTTSToken LMT, S → S
Qwen-Audio-3.0-TTSToken LMT, S → S
Qwen3-TTSToken LMT, S → S
RADKA-CSSToken LMT, S → S
RALL-EToken LMT, S → S
Raon-OpenTTSDiffusion / flowT, S → S
RapFlow-TTSDiffusion / flowT, S → S
ReGenVoiceDiffusion / flowT, S → S
ReStyle-TTSDiffusion / flowT, S → S
RTFree-F5Diffusion / flowT, S → S
RV-TTSDiffusion / flowT, I → S
RWKVTTSToken LMT, S → S
S5-TTSToken LMT, S → S
Sarashina2.2-TTSToken LMT, S → S
SASLMContinuous LMT, S → S
Seed-TTSToken LMT, S → S
Self-distilled zero-shot TTSCompactT, S → S
SelfTTSFlow / VAET, S → S
SemaVoiceContinuous LMT, S → S
SemBridgeContinuous LMT, S → S
Shallow Flow Matching TTSDiffusion / flowT, S → S
SLEDContinuous LMT, S → S
SlimSpeechCompactT, S → S
SMLLEToken LMT, S → S
SoulX-PodcastToken LMT, S → S
Spark-TTSToken LMT, S → S
SpeakStreamContinuous LMT, S → S
SPEAR-TTSToken LMT, S → S
SpeechAccentLLMToken LMT, S → S
SpeechEditToken LMT, S → S
SpeechT5AutoregressiveT, S → S
SpeechXToken LMT, S → S
SpeedySpeechParallelT → S
Spotlight-TTSDiffusion / flowT, S → S
StellarTTSCompactT, S → S
Step-Audio-EditXToken LMT, S → S
Step-Audio-TTSToken LMT, S → S
StepAudio 2.5 TTSToken LMT, S → S
Stochastic-alignment continuous TTSContinuous LMT, S → S
StreamMelContinuous LMT, S → S
StyleTTSParallelT, S → S
StyleTTS 2Diffusion / flowT, S → S
SupertonicCompactT, S → S
Supertonic 2CompactT → S
Supertonic 3CompactT → S
SwanVoiceDiffusion / flowT, S → S
SyncSpeechToken LMT, S → S
TacotronAutoregressiveT → S
Tacotron 2AutoregressiveT → S
TADAContinuous LMT, S → S
TED-TTSToken LMT, S → S
Tibetan-TTSToken LMT, S → S
TinyWaveToken LMT, S → S
TLDR (TTS)Token LMT, S → S
TMD-TTS (formerly FMSD-TTS)Diffusion / flowT → S
TontaubeV1Token LMT, S → S
Tortoise TTSToken LMT, S → S
Transformer TTSAutoregressiveT → S
TTS-CtrlNetDiffusion / flowT, S → S
TTS-TransducerToken LMT, S → S
TTSYorubaConcatenativeT → S
UDDETTSToken LMT → S
UmbraTTSDiffusion / flowT, S, A → S
UniFlow-AudioDiffusion / flowT, S, A, I, V → S, A
UNISONDiffusion / flowT, S, A → S, A
UniSonateDiffusion / flowT → S
UniSpeakerDiffusion / flowT, S, I → S
UniTAFToken LMT → S
UniTalkerToken LMT, S, V → S
UniTTSToken LMT, S → S
UniVocalToken LMT, S → S
UniVoice (ASR and TTS)Diffusion / flowT, S → S
UniVoice (speech and singing)Diffusion / flowT, S → S
UniWav (TTS)Diffusion / flowT, S → S, A
USCF-conditioned TTSDiffusion / flowT, S → S
V-CASSToken LMT, V → S
VALL-EToken LMT, S → S
VALL-E 2Token LMT, S → S
VALL-E XToken LMT, S → S
VALL-TToken LMT, S → S
VclipFlow / VAET, I → S
VevoToken LMT, S → S
VibeVoiceContinuous LMT, S → S
VibeVoice-RealtimeContinuous LMT → S
VisualSpeechParallelT, V → S
VITSFlow / VAET → S
VITS2Flow / VAET → S
VividVoiceDiffusion / flowT, I, V → S
VocalNet-M2Token LMT, S → S
VoiceboxDiffusion / flowT, S → S
VoiceChat-TTSContinuous LMT → S
VoiceCraftToken LMT, S → S
VoiceCraft-DubToken LMT, S, V → S
VoiceDesignerDiffusion / flowT, S → S
VoiceSculptorToken LMT, S → S
VoxCPMContinuous LMT, S → S
VoxCPM2Continuous LMT, S → S
Voxtral TTSToken LMT, S → S
VoXtreamToken LMT, S → S
VoXtream2Token LMT, S → S
VSpeechLMToken LMT, V → S
Wave-TacotronAutoregressiveT → S
WavTTSDiffusion / flowT, S → S
WenetSpeech-Wu TTSToken LMT, S → S
WeSConToken LMT, S → S
WhisperSpeechToken LMT, S → S
WordVoiceToken LMT, S → S
X-VoiceDiffusion / flowT, S → S
X2Streaming-TTSToken LMT, S → S
XEmoRAGToken LMT, S → S
XTTSToken LMT, S → S
YourTTSFlow / VAET, S → S
ZipVoiceDiffusion / flowT, S → S
ZipVoice-DialogDiffusion / flowT, S → S
ZonosToken LMT, S → S

Model figures

Figures are credited to their sources. Editorial input/output diagrams are labeled. Credits.

A2TTS

A2TTS extracts a voice embedding from a short recording and conditions a diffusion acoustic decoder on it. Reference-aware duration prediction improves timing consistency for multilingual synthesis in low-resource Indian languages.

Paper · GitHub: no author-linked repository found · Details

A2TTS — Figure 1

Figure 1 · Source

Affectron

Affectron extends a verbal-speech backbone to place nonverbal vocalizations in emotionally and contextually appropriate positions. Augmented training examples and structural masking enable expressive utterances containing events such as laughter while preserving the spoken content.

Paper · GitHub · Project · Details

Affectron — Figure 2

Figure 2 · Source

AgentSteerTTS

AgentSteerTTS separates identity and emotional-prosodic representations, then grounds compound instructions in retrieved acoustic examples. A controller combines these conditions and uses feedback to refine expressive speech while preserving the intended speaker.

Paper · GitHub: no author-linked repository found · Details

AgentSteerTTS — Figure 4

Figure 4 · Source

AlignDiT

AlignDiT aligns text, visual information and acoustic conditions before diffusion-based speech generation. Modality-specific guidance balances these inputs, targeting synchronized, intelligible speech that follows the timing and expression of the supplied scene.

Paper · GitHub · Details

AlignDiT — Figure 1

Figure 1 · Source

AMNet

AMNet adds phrase-structure information and local convolutional modeling to a parallel Mandarin acoustic model. These changes help capture contextual pauses, emphasis and intonation before the accompanying vocoder reconstructs the waveform.

Paper · GitHub: no author-linked repository found · Details

AMNet — Paper figure

Paper figure · Source

ARCHI-TTS

ARCHI-TTS uses a dedicated semantic alignment module to coordinate text and reference acoustic features. Reusing encoder features across denoising steps reduces repeated computation while the flow model generates the target speech.

Paper · Project · GitHub: no author-linked repository found · Details

ARCHI-TTS — Figure 1

Figure 1 · Source

ATRIE

ATRIE converts character descriptions into separate voice-identity and dynamic prosody conditions. A compact adapter learns from a larger language-model teacher and modulates a GPT-SoVITS-based synthesizer, supporting expressive persona-driven speech generation.

Paper · GitHub: no author-linked repository found · Details

ATRIE — Figure 1

Figure 1 · Source

Audiobook-CC

Audiobook-CC models context beyond individual sentences and separates style instructions from voice prompts. Distillation strengthens emotional expression, supporting multi-character narration with more consistent voices and performance across longer passages.

Paper · GitHub: no author-linked repository found · Details

Audiobook-CC — Figure 1

Figure 1 · Source

AuEmoChat

AuEmoChat learns a discrete emotion representation from speech and compresses dialogue history while retaining emotionally relevant information. Its language model predicts emotion and speech tokens, which a context-conditioned flow decoder renders into expressive conversational audio.

Paper · GitHub · Details

AuEmoChat — Figure 2

Figure 2 · Source

AuK

AuK combines language-model conditioning, an audio VAE and successive multimodal and single-stream diffusion blocks. One model handles reference-based speech synthesis and instruction-guided editing; its distilled AuK-Flash variant reduces the number of generation steps.

Paper · GitHub · Details

AuK — Figure 4

Figure 4 · Source

Authentic-Dubber

Authentic-Dubber retrieves emotionally relevant audiovisual examples and progressively incorporates them into speech generation. Its director-actor formulation connects a scene's visual context and reference delivery to the target transcript for expressive movie dubbing.

Paper · GitHub · Details

Authentic-Dubber — Figure 2

Figure 2 · Source

AutoSIFT

AutoSIFT divides a reference voice's style into attribute-specific components and a residual representation. Text instructions replace selected attributes while unmentioned characteristics remain conditioned on the reference, allowing partial style editing during speech synthesis.

Paper · GitHub: no author-linked repository found · Details

AutoSIFT — Figure 1

Figure 1 · Source

AutoStyle-TTS

AutoStyle-TTS matches the target text against a collection of expressive speech examples using learned textual embeddings. The selected recording provides style conditioning for synthesis, automatically adapting delivery to the content.

Paper · GitHub · Project · Details

AutoStyle-TTS — Paper figure

Paper figure · Source

AVLM (expressive speech)

This audio-visual language model adds full-face information to an expressive speech backbone. Training on emotion and dialogue tasks connects facial cues with spoken delivery, enabling speech generation that uses visual as well as acoustic conversational context.

Paper · GitHub · Details

AVLM (expressive speech) — Figure 2

Figure 2 · Source

Bagpiper-TTS

Bagpiper-TTS converts a natural-language request into a detailed speech plan containing words and delivery information. The generator follows that plan for tasks ranging from ordinary narration to multi-speaker rendering, role-play and singing.

Paper · Project · GitHub: no author-linked repository found · Details

Bagpiper-TTS — Figure 1

Figure 1 · Source

BareWave

BareWave generates speech in waveform space using a single inference path. Representation alignment, staged noise scheduling and perceptual objectives guide training, replacing the separate acoustic-feature and waveform-reconstruction stages common in other TTS systems.

Paper · Project · GitHub: no author-linked repository found · Details

BareWave — Figure 2

Figure 2 · Source

Bark

Hierarchical autoregressive audio tokens.

GitHub · Details

Bark — Editorial input/output diagram

Editorial input/output diagram · Source

BASE TTS

Autoregressive speechcodes and convolutional decoder.

Paper · Details

BASE TTS — Figure 1

Figure 1 · Source

BatonTTS (BatonVoice)

BatonVoice interprets a user's expressive request and translates it into controls for its dedicated BatonTTS generator. Separating instruction interpretation from acoustic rendering enables more explicit feature control, including transfer to languages outside the control-training data.

Paper · GitHub: no author-linked repository found · Details

BatonTTS (BatonVoice) — Figure 1

Figure 1 · Source

BELLE

BELLE predicts both speech values and their uncertainty in a continuous autoregressive synthesizer. Multiple synthetic renditions of the same text provide training support for the variance estimate, enabling richer acoustic distributions without adding an iterative inference stage.

Paper · GitHub · Project · Details

BELLE — Figure 1

Figure 1 · Source

BitTTS

BitTTS reduces storage and computation through extremely low-bit trained weights and indexed parameter sharing. It targets speech generation on constrained devices, preserving a full synthesis path while shrinking the model representation.

Paper · GitHub: no author-linked repository found · Details

BitTTS — Figure 1

Figure 1 · Source

Block-wise Mimi TTS

This streaming system replaces continuous acoustic regression with direct prediction of Mimi codec layers. A modified FastSpeech 2 backbone supplies aligned features and a depth-wise decoder fills residual codebooks, producing successive speech blocks without temporal autoregression.

Paper · GitHub: no author-linked repository found · Details

Block-wise Mimi TTS — Figure 1 (paper page 10)

Figure 1 (paper page 10) · Source

BnTTS

BnTTS extends an XTTS-based multilingual pipeline to Bangla using language-specific phonetic adaptations. A small amount of target-speaker audio supports personalization, with the model designed for limited-resource speech synthesis.

Paper · GitHub: no author-linked repository found · Details

BnTTS — Figure 1

Figure 1 · Source

Bolbosh

Bolbosh adapts Matcha-TTS to Kashmiri with language-aware text processing and cross-language training. Its design targets the pronunciation and script challenges of a low-resource language while retaining efficient non-autoregressive acoustic generation.

Paper · GitHub · Details

Bolbosh — Figure 1

Figure 1 · Source

Borderless Long Speech Synthesis

This system organizes speech instructions at global, sentence and token levels to guide extended recordings. A continuous-token backbone uses explicit planning and condition dropout to combine voice design, multi-speaker rendering and changing acoustic or emotional context.

Paper · GitHub: no author-linked repository found · Details

Borderless Long Speech Synthesis — Editorial input/output diagram

Editorial input/output diagram · Source

BreezyVoice

BreezyVoice combines supervised speech tokens, a language model and flow-based acoustics with a pronunciation frontend. Its Taiwanese Mandarin adaptation provides explicit phonetic control for characters with multiple readings while retaining reference-based voice synthesis.

Paper · GitHub · Details

BreezyVoice — Figure 1

Figure 1 · Source

BridgeTTS

BridgeTTS uses the BridgeCode dual representation to shorten the sequence predicted by its autoregressive language model. Bridging modules reconstruct more detailed continuous acoustic features from those sparse tokens, balancing generation speed with voice fidelity.

Paper · GitHub: no author-linked repository found · Details

BridgeTTS — Figure 2

Figure 2 · Source

BVS

Beyond Video-to-SFX predicts audio semantic tokens from visual information and phonetic cues, then refines them into acoustic tokens. The two-stage generator produces intelligible speech whose timing and environmental sound fit the supplied video.

Paper · GitHub: no author-linked repository found · Details

BVS — Figure 1

Figure 1 · Source

CAM-TTS

CAM-TTS retains global narrative information and retrieves local details through an updatable memory block. Prefix attention combines those memories with preceding context, guiding sentence-level synthesis across longer paragraphs.

Paper · GitHub: no author-linked repository found · Details

CAM-TTS — Figure 2

Figure 2 · Source

CapTalk

CapTalk designs voices from descriptions for individual utterances and multi-speaker dialogue. Hierarchical conditioning separates stable speaker identity from changing turn-level delivery, while explicit planning tokens control dynamic expressive attributes.

Paper · GitHub: no author-linked repository found · Details

CapTalk — Figure 1

Figure 1 · Source

CAST-TTS

CAST-TTS maps a voice description or a reference recording into a common timbre-conditioning interface. Cross-attention delivers that information to the synthesizer, enabling voice design and reference-based cloning within the same generation framework.

Paper · GitHub · Project · Details

CAST-TTS — Figure 1

Figure 1 · Source

CaT-TTS

CaT-TTS separates textual understanding from acoustic generation in a two-Transformer architecture. During decoding, a masked parallel inference procedure guides speech-token predictions to reduce local errors in zero-shot voice synthesis.

Paper · GitHub: no author-linked repository found · Details

CaT-TTS — Figure 2

Figure 2 · Source

Causal-prosody FastSpeech 2

This FastSpeech 2 extension explicitly models emotion alongside duration, pitch and energy. Counterfactual training separates emotional changes from linguistic content, allowing users to modify prosody while preserving the intended words.

Paper · GitHub: no author-linked repository found · Details

Causal-prosody FastSpeech 2 — Figure 1 (paper page 3)

Figure 1 (paper page 3) · Source

CDE-StyleTTS

CDE-StyleTTS lets acoustic states evolve continuously along a phoneme sequence whose timing comes from durations. Sampling this trajectory supplies the acoustic decoder with timing-sensitive representations, providing a way to transfer changing expressive style instead of merely repeating phoneme embeddings.

Paper · GitHub · Details

CDE-StyleTTS — Paper figure

Paper figure · Source

Chain-of-Details TTS

Chain-of-Details TTS progressively predicts speech at increasing temporal resolutions using a shared decoder. The coarsest stage provides an implicit phonetic plan, while subsequent stages recover timing detail without a separate phoneme-duration predictor.

Paper · GitHub: no author-linked repository found · Details

Chain-of-Details TTS — Paper figure

Paper figure · Source

Chain-Talker

Chain-Talker first derives an emotional description from dialogue history, then predicts semantic speech codes. A final rendering stage combines these plans to synthesize expressive responses whose delivery fits the conversational context.

Paper · GitHub · Details

Chain-Talker — Figure 2

Figure 2 · Source

Chatterbox

Chatterbox synthesizes speech from text and a voice reference, with controls for expressive delivery. Its multilingual models focus on cross-language voice consistency, while Turbo and Nano use a smaller backbone and a single-step acoustic decoder. These variants serve different narration, conversational playback and local-device requirements; their capabilities are not interchangeable.

GitHub · Details

Chatterbox — Editorial input/output diagram

Editorial input/output diagram · Source

Chatterbox-Flash

Chatterbox-Flash generates speech-token blocks in parallel while keeping block-by-block streaming. Calibration against common-token priors and confidence-based stopping improve its discrete diffusion decoding after adaptation from a pretrained autoregressive synthesizer.

Paper · GitHub · Details

Chatterbox-Flash — Figure 3

Figure 3 · Source

ChatTTS

Autoregressive speech-token generation.

GitHub · Details

ChatTTS — Editorial input/output diagram

Editorial input/output diagram · Source

CLaM-TTS

Probabilistic residual quantization and multi-token LM.

Paper · Details

CLaM-TTS — Figure 1

Figure 1 · Source

CLEAR

CLEAR models speech directly in a continuous latent space, avoiding discrete codec-token prediction. Its zero-shot generator combines reference-voice conditioning with incremental audio production, targeting a balance between naturalness and response latency.

Paper · GitHub: no author-linked repository found · Details

CLEAR — Figure 1

Figure 1 · Source

Clip-TTS

Clip-TTS trains its textual representation against corresponding mel-spectrogram information through a contrastive objective. The acoustic Transformer uses the resulting context-aware features to improve prosodic interpretation during speech generation.

Paper · GitHub: no author-linked repository found · Details

Clip-TTS — Figure 3

Figure 3 · Source

Compact neural accessibility TTS

This compact synthesis system combines a shared-parameter text frontend with an efficient recurrent waveform generator. It targets responsive accessibility voices on low-power devices, where small storage requirements and immediate playback matter alongside naturalness.

Paper · GitHub: no author-linked repository found · Details

Compact neural accessibility TTS — Figure 1

Figure 1 · Source

Compressed-to-fine speech LM

This speech-language-model design keeps recent acoustic tokens and voice prompts at full detail while compressing distant context. The asymmetric representation reduces redundant long-sequence processing without discarding the local cues needed for pronunciation and vocal consistency.

Paper · GitHub: no author-linked repository found · Details

Compressed-to-fine speech LM — Figure 1

Figure 1 · Source

Confucius4-TTS

Confucius4-TTS extracts voice characteristics from self-supervised speech features without requiring a transcript of the reference recording. A language model predicts semantic tokens and a flow decoder generates mel-spectrograms, enabling voice cloning within and across fourteen languages.

Paper · GitHub · Details

Confucius4-TTS — Figure 1

Figure 1 · Source

Continuous-token diffusion TTS

This model combines a language head that predicts boundaries with a diffusion head that generates continuous acoustic frames. Masked and staged training stabilize speaker-reference conditioning, providing a text-to-speech path within a multimodal language-model architecture.

Paper · GitHub: no author-linked repository found · Details

Continuous-token diffusion TTS — Figure 2

Figure 2 · Source

Controllable masked-speech TTS

This synthesizer separates reference voice information from acoustic background conditions. An explicit task control selects whether background sound is retained or removed, allowing personalized speech generation under different environmental requirements.

Paper · GitHub: no author-linked repository found · Details

Controllable masked-speech TTS — Paper figure

Paper figure · Source

CookVoice

CookVoice aligns textual content, style and prosodic controls to acoustic frames before speech generation. The same compact model supports spoken and sung voices, reference imitation and editing, allowing individual voice attributes to be controlled within a shared synthesis pipeline.

Paper · Project · GitHub: no author-linked repository found · Details

CookVoice — Figure 1

Figure 1 · Source

CosyEdit2

CosyEdit2 adapts a text-speech language model and acoustic decoder for consistent speech editing, then refines them with editing-specific rewards. The paper also evaluates the resulting improvement in zero-shot text-to-speech, connecting local editing consistency with reference-based synthesis.

Paper · Project · GitHub: no author-linked repository found · Details

CosyEdit2 — Figure 1

Figure 1 · Source

CoSyncDiT

CoSyncDiT guides flow-based speech synthesis through acoustic-style adaptation, visual calibration and timed context alignment. These stages connect the supplied transcript and scene information to expressive, synchronized movie dubbing.

Paper · GitHub · Details

CoSyncDiT — Figure 2

Figure 2 · Source

CosyVoice

Supervised semantic tokens and flow decoder.

Paper · Details

CosyVoice — Figure 1

Figure 1 · Source

CosyVoice 2

Text/speech LM and chunk-aware flow matching.

Paper · Details

CosyVoice 2 — Figure 1

Figure 1 · Source

CosyVoice 3

CosyVoice 3 extends streaming, reference-conditioned synthesis with a tokenizer trained on several speech-understanding tasks and a reward model for post-training. It targets reliable pronunciation, speaker identity and prosody across languages, dialects and less controlled text. The research scaling experiments and the downloadable Fun-CosyVoice3 checkpoint represent different model configurations.

Paper · GitHub · Project · Details

CosyVoice 3 — Figure 2

Figure 2 · Source

CosyWhisper (WhispSynth)

The WhispSynth generation pipeline combines a CosyVoice synthesizer with pitch-free digital signal processing to produce whispered speech. It supports multilingual whisper generation while avoiding the voiced pitch patterns of ordinary speech synthesis.

Paper · GitHub · Details

CosyWhisper (WhispSynth) — Figure 2

Figure 2 · Source

CoVoMix2

CoVoMix2 generates scripted dialogue directly with a flow-matching model, using reference voices without their transcripts. Speaker-disentangled text, sentence alignment and prompt masking support controlled timing and overlapping turns.

Paper · GitHub · Details

CoVoMix2 — Figure 1

Figure 1 · Source

Cross-Lingual F5-TTS

Cross-Lingual F5-TTS changes reference preparation and training so the generated text need not be paired with a reference transcript. Word-aligned acoustic prompts support cross-language voice cloning while reusing the flow-matching synthesis backbone.

Paper · Project · GitHub: no author-linked repository found · Details

Cross-Lingual F5-TTS — Figure 1

Figure 1 · Source

CrossAccent-TTS

CrossAccent-TTS separates speaker identity from accent-related information in a speech synthesis model. Weighted language embeddings control the accent subspace, allowing gradual accent changes and cross-language synthesis while retaining the reference speaker's timbre.

Paper · GitHub: no author-linked repository found · Details

CrossAccent-TTS — Figure 1

Figure 1 · Source

CSM

CSM uses the text and audio of preceding speaker turns to shape the delivery of the next utterance. A Llama backbone predicts speech representations, and a smaller decoder completes Mimi audio codes. It is a contextual speech renderer: an application supplies the words to say, including any responses written by a separate language model.

GitHub · Details

CSM — Editorial input/output diagram

Editorial input/output diagram · Source

CTC-TTS

CTC-TTS uses automatically derived alignment and two-word interleaving to train an incremental speech language model. Its length-concatenated and feature-stacked variants make different trade-offs between speech quality and generation latency.

Paper · GitHub · Details

CTC-TTS — Figure 2

Figure 2 · Source

CtrlSpeech

CtrlSpeech adds local pitch, loudness and duration conditioning to a patch-autoregressive diffusion synthesizer. A separate global speaker condition preserves the reference voice while users modify the delivery of individual words or phonemes.

Paper · GitHub · Details

CtrlSpeech — Figure 2

Figure 2 · Source

CuteTTS

CuteTTS combines a causal audio VAE, an autoregressive patch model and an explicitly speaker-conditioned flow head. Generating patches rather than individual frames reduces sequential work; a distilled variant targets faster streaming while preserving reference-voice synthesis.

Paper · GitHub · Details

CuteTTS — Figure 1

Figure 1 · Source

DAIEN-TTS

DAIEN-TTS separates a reference recording into speech and environmental components, then conditions acoustic generation on them independently. Its extended formulation additionally models reverberation and uses separate guidance controls for speech, noise and room acoustics, enabling voice cloning into a chosen environment.

Paper 1 · Project 1 · Paper 2 · GitHub · Project 2 · Details

DAIEN-TTS — Paper figure

Paper figure · Source

DARS

DARS separately models pathological timing and acoustic style to synthesize dysarthric speech. A multistage rhythm predictor and conditioned flow model generate targeted examples for improving recognition under limited real-world speech data.

Paper · GitHub: no author-linked repository found · Details

DARS — Figure 1

Figure 1 · Source

DCAR

DCAR changes the number of acoustic tokens predicted at each autoregressive step. Adapting the chunk size to the generation state shortens sequential processing while maintaining content alignment and reference-conditioned speech quality.

Paper · GitHub: no author-linked repository found · Details

DCAR — Figure 1

Figure 1 · Source

Deep Voice

Neural TTS pipeline with autoregressive WaveNet synthesis.

Paper · Details

Deep Voice — Figure 1

Figure 1 · Source

Deep Voice 2

Speaker-conditioned neural pipeline and WaveNet synthesis.

Paper · Details

Deep Voice 2 — Figure 1

Figure 1 · Source

Deep Voice 3

Convolutional attention encoder-decoder and converter.

Paper · Details

Deep Voice 3 — Figure 1

Figure 1 · Source

DeepASMR

DeepASMR separates ASMR delivery from the reference speaker's identity using discrete speech representations. A language model predicts content and style, while a flow-based acoustic decoder transfers the requested performance into the target voice.

Paper · GitHub: no author-linked repository found · Details

DeepASMR — Paper figure

Paper figure · Source

DeepDubber-V1

DeepDubber-V1 interprets visual scenes and dubbing requirements before generating the target speech. Its multimodal conditions distinguish narration, monologue and dialogue, guiding both expressive delivery and synchronization.

Paper · GitHub · Details

DeepDubber-V1 — Figure 1

Figure 1 · Source

DeepDubbing

DeepDubbing assigns voices to characters and conditions speech rendering on the surrounding script. Its voice-design and instruction-synthesis stages support multi-participant audiobook production while maintaining character identity and context-sensitive expression.

Paper · GitHub · Details

DeepDubbing — Figure 1

Figure 1 · Source

DelightfulTTS

Conformer acoustic model with explicit and implicit prosody.

Paper · Details

DelightfulTTS — Figure 1

Figure 1 · Source

DELTA-TTS

DELTA-TTS converts a pretrained speech language model to parallel discrete diffusion using lightweight adaptation. Local convolution and confidence-based decoding help retain acoustic structure while the model fills speech-token positions in an order determined by prediction confidence.

Paper · GitHub: no author-linked repository found · Details

DELTA-TTS — Figure 1

Figure 1 · Source

DepFlow

DepFlow separates a depression-related acoustic representation from speaker identity and spoken content. The representation conditions a flow-matching synthesizer, enabling controlled synthetic speech for investigating acoustic cues and data augmentation.

Paper · GitHub: no author-linked repository found · Details

DepFlow — Figure 1

Figure 1 · Source

Dia

Dia turns a speaker-tagged transcript into conversational audio, including supported nonverbal events such as laughter and coughing. Reference audio and its transcript can establish speaker identity and delivery. Its English checkpoint is useful for scripted exchanges and dialogue narration, with the conversation content supplied by the user rather than generated by the model.

GitHub · Details

Dia — Editorial input/output diagram

Editorial input/output diagram · Source

Dia2

Dia2 begins synthesizing before the complete script is available, allowing an application to feed words incrementally. Audio prefixes provide speaker and conversational context, while a streaming codec path reconstructs the output. The released 1B and 2B checkpoints focus on English dialogue; prefix conditioning helps keep voices consistent between generations.

GitHub · Details

Dia2 — Editorial input/output diagram

Editorial input/output diagram · Source

DialoSpeech

DialoSpeech models two speakers on separate tracks and uses chunked flow matching for acoustic rendering. The design supports expressive scripted dialogue, including interactions whose timing is difficult to reproduce by joining independent utterances.

Paper · Project · GitHub: no author-linked repository found · Details

DialoSpeech — Figure 2

Figure 2 · Source

DiEmo-TTS

DiEmo-TTS distills emotion information from speech while suppressing unrelated speaker characteristics. Cluster-based sampling and representation perturbation improve cross-speaker emotion transfer, including situations where extensive emotion labels are unavailable.

Paper · GitHub · Details

DiEmo-TTS — Figure 1

Figure 1 · Source

Diff-TTS

Text-conditioned denoising diffusion acoustic model.

Paper · Details

Diff-TTS — Figure 3

Figure 3 · Source

DiffCSS

DiffCSS samples prosody representations from multimodal conversational context using a diffusion model. A prosody-conditioned speech language model renders those samples, allowing several expressive deliveries that remain consistent with the same dialogue.

Paper · GitHub: no author-linked repository found · Details

DiffCSS — Paper figure

Paper figure · Source

DiFlow-TTS

DiFlow-TTS maps phonemes into linguistic content and generates separate prosody and acoustic token streams through discrete flow matching. Factorizing these responsibilities supports compact zero-shot synthesis with fewer sequential generation steps.

Paper · GitHub · Details

DiFlow-TTS — Figure 2

Figure 2 · Source

DiFlowDubber

DiFlowDubber first learns linguistic content and separate prosodic-acoustic tokens through a discrete-flow TTS model. A subsequent video-dubbing stage aligns those representations to visual timing, connecting voice generation with synchronized lip movements.

Paper · GitHub · Project · Details

DiFlowDubber — Figure 2

Figure 2 · Source

DisCo-Speech

DisCo-Speech learns a codec that separates content, delivery and speaker identity. A language model predicts combined content-prosody tokens while the decoder receives a separate timbre representation, enabling reference cloning with independently controllable vocal attributes.

Paper · GitHub · Project · Details

DisCo-Speech — Figure 1

Figure 1 · Source

DisSpeech

DisSpeech maps Mandarin text and marked stuttering events to semantic speech tokens without temporal autoregression. Pitch and energy modeling guide acoustic reconstruction, enabling controlled repetitions and other disfluencies for speech synthesis and recognition-data augmentation.

Paper · GitHub: no author-linked repository found · Details

DisSpeech — Figure 2

Figure 2 · Source

DiSTAR

DiSTAR first drafts blocks of residual-quantized speech tokens with a language model. A masked diffusion decoder then fills acoustic detail within each block, combining temporal planning and parallel refinement entirely in discrete codec space.

Paper · GitHub: no author-linked repository found · Details

DiSTAR — Figure 1

Figure 1 · Source

DiTAR

DiTAR predicts a sequence of compressed acoustic patches using a language model, then generates each patch's detail through diffusion. Separating global temporal planning from local reconstruction supports zero-shot speech synthesis with controllable sampling diversity.

Paper · Project · GitHub: no author-linked repository found · Details

DiTAR — Figure 1

Figure 1 · Source

DiTTo-TTS

Latent diffusion Transformer with speech-length prediction.

Paper · Details

DiTTo-TTS — Figure 1

Figure 1 · Source

DMOSpeech 2

DMOSpeech 2 adds reinforcement learning to duration prediction in an already metric-optimized synthesizer. Rewards derived from speaker similarity and transcription accuracy guide timing choices, linking prosodic planning with reference-voice and content objectives.

Paper · GitHub · Project · Details

DMOSpeech 2 — Figure 1

Figure 1 · Source

DMP-TTS

DMP-TTS maps style descriptions and reference recordings into a shared conditioning space. Chained guidance controls content, timbre and style separately, allowing detailed synthesis adjustments within a latent diffusion Transformer.

Paper · GitHub: no author-linked repository found · Details

DMP-TTS — Figure 1

Figure 1 · Source

dots.tts

dots.tts predicts continuous acoustic representations from multilingual text and voice references. Its flow-matching output head supports both audio streaming and streaming text input; guidance-aware distillation reduces the work needed to generate each audio packet.

Paper · GitHub · Project · Details

dots.tts — Figure 1

Figure 1 · Source

Dragon-FM

Dragon-FM predicts successive speech chunks autoregressively while refining the tokens inside each chunk with bidirectional flow matching. Compact acoustic codes and cross-chunk caching reduce generation overhead and support longer content such as podcasts.

Paper · Project · GitHub: no author-linked repository found · Details

Dragon-FM — Figure 1

Figure 1 · Source

DrawSpeech

DrawSpeech turns user-drawn prosodic curves into detailed pitch and energy conditions for synthesis. A diffusion model fills in the acoustic detail, enabling localized expressive control beyond an overall style description.

Paper · GitHub · Details

DrawSpeech — Paper figure

Paper figure · Source

DS-TTS

DS-TTS extracts complementary voice characteristics through two style encoders. Dynamic modulation conditions the acoustic generator on these representations, supporting unseen speakers and adapting synthesis across different sentence lengths.

Paper · GitHub: no author-linked repository found · Details

DS-TTS — Paper figure

Paper figure · Source

DualDub

DualDub generates spoken dialogue and background sound together from video and textual conditions. A cross-modal aligner coordinates the two decoding heads, targeting temporally synchronized soundtracks rather than speech rendered in isolation.

Paper · GitHub: no author-linked repository found · Details

DualDub — Figure 2

Figure 2 · Source

DualSpeechLM

DualSpeechLM uses understanding-oriented speech tokens as input and acoustic codec tokens for generation. Semantic supervision and staged conditioning coordinate the two representations, supporting speech synthesis within a unified understanding-and-generation architecture.

Paper · GitHub: no author-linked repository found · Details

DualSpeechLM — Figure 3

Figure 3 · Source

E2 TTS

Flow matching with filler-token text conditioning.

Paper · Details

E2 TTS — Figure 1

Figure 1 · Source

ECTSpeech

ECTSpeech gradually tightens consistency constraints on a pretrained diffusion synthesizer. The resulting generator maps a noise state to speech in one step, reducing repeated denoising while retaining the conditioning of the original TTS model.

Paper · GitHub: no author-linked repository found · Details

ECTSpeech — Figure 2

Figure 2 · Source

ELLA-V

Alignment-guided token reordering.

Paper · Details

ELLA-V — Figure 1

Figure 1 · Source

EME-TTS

EME-TTS jointly models emotional delivery and local emphasis instead of treating them as independent effects. Automatically derived emphasis labels and variance-related features help users stress selected material while retaining a recognizable target emotion.

Paper · GitHub: no author-linked repository found · Details

EME-TTS — Figure 1

Figure 1 · Source

EMM-TTS

EMM-TTS separates emotional content modeling from speaker-specific acoustic generation. Speaker-consistency objectives and adaptive normalization help transfer emotion across languages while retaining the reference speaker's timbre.

Paper · GitHub: no author-linked repository found · Details

EMM-TTS — Figure 2

Figure 2 · Source

EmojiVoice

EmojiVoice adds interpretable emoji prompts to the text encoder and flow predictor of Matcha-TTS. Changing prompts across phrases varies expression during longer robot utterances, providing a lightweight control interface for expressive synthesis.

Paper · GitHub · Details

EmojiVoice — Paper figure

Paper figure · Source

EmoShift

EmoShift adds a lightweight layer that learns emotion-dependent changes to a TTS model's hidden representation. The controls adjust expressive delivery and emotion intensity while leaving most of the underlying synthesis backbone intact.

Paper · GitHub: no author-linked repository found · Details

EmoShift — Figure 2

Figure 2 · Source

EmoSSLSphere

EmoSSLSphere combines an emotion representation constrained to a sphere with discrete units derived from self-supervised speech features. The model synthesizes emotional speech across languages while organizing expressive controls independently of the textual content.

Paper · GitHub: no author-linked repository found · Details

EmoSSLSphere — Figure 2

Figure 2 · Source

EmoSteer-TTS

EmoSteer-TTS extracts emotion-related directions from a pretrained synthesizer's internal activations. Applying those directions during inference changes, mixes or removes emotional expression without retraining; the paper evaluates the method across several different TTS backbones.

Paper · GitHub: no author-linked repository found · Details

EmoSteer-TTS — Figure 3

Figure 3 · Source

Emotion-timbre disentangled TTS

This emotional synthesizer learns separate reference encoders for timbre and emotion. A mutual-information objective reduces their overlap, while phoneme-level emotion prediction carries changing expression into the generated acoustic sequence.

Paper · GitHub · Project · Details

Emotion-timbre disentangled TTS — Figure 1

Figure 1 · Source

EmotiVoice

PromptTTS-derived style and content conditioning.

GitHub · Details

EmotiVoice — Editorial input/output diagram

Editorial input/output diagram · Source

EmoTra-TTS

EmoTra-TTS introduces frame-level valence, arousal and dominance controls into both prosodic planning and acoustic decoding. Synthetic transition examples teach the model to move between emotions within an utterance while keeping its wording and speaker identity consistent.

Paper · GitHub · Project · Details

EmoTra-TTS — Figure 2

Figure 2 · Source

EmoVoice

EmoVoice interprets free-form textual descriptions of emotional delivery. Its phoneme-boost variant predicts phonetic and acoustic information together to improve content consistency while retaining expressive style control.

Paper · GitHub · Details

EmoVoice — Figure 1

Figure 1 · Source

End-to-end discrete-token TTS

This system trains the discrete speech representation together with the language and acoustic models that consume it. Reconstruction and recognition feedback also update token prediction, reducing mismatches between separately trained components in reference-conditioned speech synthesis.

Paper · GitHub: no author-linked repository found · Details

End-to-end discrete-token TTS — Figure 1

Figure 1 · Source

F5-TTS

F5-TTS learns text-guided speech infilling with flow matching, refining character representations before a Transformer predicts the speech trajectory. A reference clip supplies voice context, and Sway Sampling controls how inference steps are distributed. The 2025 v1 Base release refines training and inference within the same general architecture for zero-shot speech synthesis.

Paper · GitHub · Details

F5-TTS — Figure 1

Figure 1 · Source

F5R-TTS

F5R-TTS adapts a flow-based synthesizer to reinforcement learning through a probabilistic formulation of generation. Transcription accuracy and speaker-similarity rewards refine the model's ability to preserve both target content and the reference voice.

Paper · GitHub · Details

F5R-TTS — Figure 2

Figure 2 · Source

Face-adapted StyleTTS 2

This model maps facial features into the style space of StyleTTS 2 through a lightweight learned adapter. It synthesizes text in a plausible face-conditioned voice without an audio reference; the paper evaluates transfer to unseen identities and another synthesis language.

Paper · GitHub: no author-linked repository found · Details

Face-adapted StyleTTS 2 — Figure 1

Figure 1 · Source

FaceSpeak

FaceSpeak extracts speaker-related and expressive information from real or stylized portraits. These visual conditions guide text-to-speech generation, allowing the requested words to be rendered in a plausible voice and emotion associated with the image.

Paper · GitHub: no author-linked repository found · Details

FaceSpeak — Figure 3

Figure 3 · Source

FacialTalker

FacialTalker quantizes facial action information and combines it with text and speech history in a conversational synthesizer. Joint preference training over visual and speech tokens helps the generated delivery reflect the interlocutor's facial expression and dialogue context.

Paper · GitHub · Details

FacialTalker — Figure 2

Figure 2 · Source

FastPitch

Parallel Transformer with explicit pitch prediction.

Paper · Details

FastPitch — Figure 1

Figure 1 · Source

FastSpeech

Feed-forward Transformer and duration-based length regulator.

Paper · Details

FastSpeech — Figure 1, PDF p. 4

Figure 1, PDF p. 4 · Source

FastSpeech 2

Parallel Transformer with duration, pitch and energy prediction.

Paper · Details

FastSpeech 2 — Figure 1, PDF p. 3

Figure 1, PDF p. 3 · Source

FC-TTS

FC-TTS conditions generation on two recordings, one supplying delivery style and the other speaker identity. Specialized representation processing and auxiliary training objectives aim to prevent either condition from leaking unwanted attributes into the synthesized voice.

Paper · GitHub: no author-linked repository found · Details

FC-TTS — Figure 1

Figure 1 · Source

FELLE

FELLE generates continuous acoustic frames sequentially, using the preceding frame to shape the next flow-matching prior. A coarse-to-fine acoustic head refines each prediction, supporting reference-conditioned speech without discrete speech-token classification.

Paper · GitHub: no author-linked repository found · Details

FELLE — Figure 1

Figure 1 · Source

FineCombo-TTS

FineCombo-TTS interprets style descriptions relative to a supplied speech reference. A flow-based variance predictor models how acoustic attributes should change, enabling precise relative edits without requiring an explicit independent embedding for every voice attribute.

Paper · Project · GitHub: no author-linked repository found · Details

FineCombo-TTS — Figure 2

Figure 2 · Source

FireRedAudio

FireRedAudio uses different acoustic encoders for understanding audio and conditioning speech generation. Its shared language model drives a flow-matching decoder over continuous RedAE latents, supporting voice cloning, instruction-controlled synthesis and speech editing within the broader audio model.

Paper · GitHub · Details

FireRedAudio — Figure 1

Figure 1 · Source

FireRedTTS

Text-to-semantic LM and speech decoder.

Paper · Details

FireRedTTS — Figure 3

Figure 3 · Source

FireRedTTS-1S

FireRedTTS-1S extends the FireRed synthesis line with incremental acoustic decoding. Its chunked flow-matching and frame-autoregressive multi-stream decoder options provide different trade-offs between initial latency and sustained generation speed.

Paper · GitHub: no author-linked repository found · Details

FireRedTTS-1S — Figure 1

Figure 1 · Source

FireRedTTS-2

FireRedTTS-2 models chronological sequences of speaker-labeled text and speech using a large Transformer plus a smaller codebook decoder. A low-rate streaming tokenizer reduces the number of audio steps. It targets long conversations and podcasts where speaker changes, turn-specific delivery and continuity across utterances matter.

Paper · GitHub · Details

FireRedTTS-2 — Figure 1

Figure 1 · Source

FireRedTTS3

FireRedTTS3 uses a semantically supervised audio autoencoder to make continuous speech representations easier to predict. The Base variant provides multilingual reference cloning; the Instruct variant adds natural-language voice design and editing of spoken content or acoustic attributes.

Paper · GitHub · Details

FireRedTTS3 — Figure 1

Figure 1 · Source

Fish Audio S1 / OpenAudio S1

The S1 family combines multilingual voice conditioning with explicit markers for emotion, tone and nonverbal sounds. Its full model and distilled S1-mini offer different deployment sizes, with reinforcement learning used to refine generation. It is suited to expressive narration and character dialogue, although access and capabilities depend on the selected release.

Model card · Details

Fish Audio S1 / OpenAudio S1 — Editorial input/output diagram

Editorial input/output diagram · Source

Fish Audio S2

Fish Audio S2 extends the Fish speech-model line with natural-language delivery instructions and multi-speaker, multi-turn synthesis. Its training pipeline uses speech descriptions, quality assessment and reward modeling to improve controllability. The released inference stack supports streaming, making the model relevant to both scripted audio production and incremental spoken responses.

Paper · GitHub · Details

Fish Audio S2 — Figure 2

Figure 2 · Source

Fish Speech

Dual-autoregressive slow/fast transformers.

Paper · GitHub · Details

Fish Speech — Figure 2

Figure 2 · Source

Flamed-TTS

Flamed-TTS combines representations of different speech attributes in an attention-free generator. Its reformulated flow-matching process targets efficient zero-shot synthesis with flexible pacing and reduced sequential computation.

Paper · Project · GitHub: no author-linked repository found · Details

Flamed-TTS — Figure 1

Figure 1 · Source

FlashTTS

FlashTTS processes incoming text and speech context on staggered tracks so synthesis can begin before sentence completion. Multi-token prediction accelerates the language model, while a distilled acoustic decoder reduces waveform-generation delay.

Paper · GitHub · Project · Details

FlashTTS — Figure 1

Figure 1 · Source

FleSpeech

FleSpeech unifies text, voice recordings and visual prompts into a common conditioning representation. Its multistage generator uses those controls to manipulate voice and delivery attributes flexibly rather than requiring one fixed prompt modality.

Paper · Project · GitHub: no author-linked repository found · Details

FleSpeech — Figure 2

Figure 2 · Source

FlexiVoice

FlexiVoice accepts an optional style instruction and an optional reference voice alongside the target text. Progressive preference training teaches the model to follow both conditions while reducing unwanted coupling between wording, speaker identity and delivery.

Paper · Project · GitHub: no author-linked repository found · Details

FlexiVoice — Figure 1

Figure 1 · Source

FlexSpeech

FlexSpeech separates timing control from the acoustic synthesis component to balance stable pronunciation and natural expression. A small set of style examples can adapt the duration module without retraining the full generator, supporting efficient delivery customization.

Paper · Project · GitHub: no author-linked repository found · Details

FlexSpeech — Figure 1

Figure 1 · Source

Flowtron

Autoregressive normalizing flows over mel spectrograms.

Paper · Details

Flowtron — Figure 1

Figure 1 · Source

FNH-TTS

FNH-TTS routes linguistic and speaker information through several duration experts to model varied timing patterns. It combines this predictor with changes to waveform generation, aiming for robust end-to-end speech synthesis across voices and prosodic conditions.

Paper · GitHub: no author-linked repository found · Details

FNH-TTS — Figure 1

Figure 1 · Source

Frame-stacked local Transformer TTS

This architecture lets a global language model predict several speech frames at a time and delegates their codec entries to a smaller local Transformer. The paper compares sequential local decoding with iterative masked prediction, showing how frame stacking changes the trade-off between synthesis throughput and acoustic fidelity.

Paper · GitHub: no author-linked repository found · Details

Frame-stacked local Transformer TTS — Figure 1

Figure 1 · Source

FreyaTTS

FreyaTTS is a Turkish-focused non-autoregressive synthesizer that maps character sequences to continuous audio latents. A frozen waveform autoencoder reconstructs the speech, while duration prediction and voice-focused post-training support efficient conversational playback without phoneme or discrete speech tokenization.

Paper · GitHub: no author-linked repository found · Details

FreyaTTS — Figure 1

Figure 1 · Source

Gemini 2.5 TTS

Gemini 2.5 TTS converts supplied text into speech with prompt-based control of accent, pace, style and emotion. The Flash and Pro interfaces support single-speaker narration and two-speaker scripts with separately assigned voices. These are dedicated speech-generation endpoints; their internal acoustic architecture is not fully disclosed in the public documentation.

Docs 1 · Docs 2 · Docs 3 · Details

Gemini 2.5 TTS — Editorial input/output diagram

Editorial input/output diagram · Source

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS adds expressive audio tags to prompt-steered speech generation, giving authors more local control over narration and delivery. It targets natural, responsive multilingual synthesis through a managed API. The public preview documentation describes the interface and controls, without enough architectural detail to reconstruct the underlying speech generator.

Docs 1 · Docs 2 · Details

Gemini 3.1 Flash TTS — Editorial input/output diagram

Editorial input/output diagram · Source

GibbsTTS

GibbsTTS generates discrete speech tokens through a continuous-time jump process. Metric-aware transition scheduling and a finite-step correction improve how token states evolve, supporting zero-shot voice synthesis with discrete flow matching.

Paper · GitHub · Project · Details

GibbsTTS — Figure 1

Figure 1 · Source

GLM-TTS

GLM-TTS first predicts speech tokens autoregressively, then converts them into audio with a diffusion decoder. Pitch-aware tokenization and multi-reward reinforcement learning target pronunciation, speaker similarity and expression. Hybrid phoneme/text input and LoRA voice adaptation provide controls for applications that need repeatable pronunciation and customized voices.

Paper · GitHub · Details

GLM-TTS — Figure 1

Figure 1 · Source

Glow-TTS

Normalizing-flow acoustic model and monotonic alignment search.

Paper · Details

Glow-TTS — Figure 1, PDF p. 3

Figure 1, PDF p. 3 · Source

GOAT-TTS

GOAT-TTS encodes continuous voice information in one branch and predicts speech tokens in another. Partial language-model adaptation preserves textual knowledge, while multi-token prediction supports streaming synthesis with reference-based paralinguistic conditioning.

Paper · GitHub: no author-linked repository found · Details

GOAT-TTS — Figure 1

Figure 1 · Source

GPA

General-Purpose Audio uses one autoregressive backbone to predict discrete speech tokens across synthesis, recognition and conversion tasks. Its TTS path combines target text with voice context, with shared multitask training and scalable inference.

Paper · GitHub · Details

GPA — Figure 1

Figure 1 · Source

GPT-4o Mini TTS

GPT-4o Mini TTS combines the text to be spoken with instructions that steer accent, speed, tone and emotional delivery. The Speech API can stream audio before the full result is complete and supports several output formats. It provides a managed synthesis component for narration and voice applications; public documentation does not disclose the complete acoustic architecture.

Docs 1 · Docs 2 · Details

GPT-4o Mini TTS — Editorial input/output diagram

Editorial input/output diagram · Source

GPT-SoVITS

GPT-SoVITS couples text-to-semantic token prediction with a reference-conditioned speech decoder. Its 2025 V3, V4 and V2 Pro releases extend a workflow that supports both zero-shot synthesis and voice adaptation from a small training set. The surrounding WebUI helps prepare data, while recognition and source-separation utilities remain separate components.

GitHub · Docs · Details

GPT-SoVITS — Editorial input/output diagram

Editorial input/output diagram · Source

Grad-TTS

Score-based diffusion decoder and monotonic alignment search.

Paper · Details

Grad-TTS — Figure 2

Figure 2 · Source

GRAFT

GRAFT attaches codec tokens from a spoken word example to that word's location in the text prompt. Separate target-speaker conditioning allows the pronunciation hint to come from another voice while the synthesized sentence retains the desired speaker.

Paper · GitHub: no author-linked repository found · Details

GRAFT — Paper figure

Paper figure · Source

GSA-TTS

GSA-TTS extracts local style information at successive levels and combines it through attention into a global reference condition. This richer style representation guides the acoustic model when synthesizing an unseen speaker's voice.

Paper · GitHub: no author-linked repository found · Details

GSA-TTS — Figure 1

Figure 1 · Source

GST-Tacotron

Tacotron with a reference encoder and global style tokens.

Paper · Details

GST-Tacotron — Figure 1

Figure 1 · Source

Habibi

Habibi trains an Arabic synthesizer progressively from standard language to regional dialects using curated public speech. It targets zero-shot voice cloning across dialects and reading without mandatory diacritic marks.

Paper · GitHub · Project · Details

Habibi — Figure 1

Figure 1 · Source

HD-PPT

HD-PPT learns speech codes that distinguish spoken content from instruction-related preferences. A language model predicts semantic information, expressive style and acoustic detail in sequence, improving the mapping from natural-language requests to controllable speech.

Paper · GitHub: no author-linked repository found · Details

HD-PPT — Figure 1

Figure 1 · Source

Higgs Audio v2

Higgs Audio v2 combines interleaved text/audio modeling with a unified speech tokenizer and DualFFN layers for acoustic prediction. Reference clips and scene context influence voices, while text context shapes prosody across narration and multi-speaker scripts. The generation checkpoint covers synthesis; the separate understanding branch in the family diagram is not another output mode of this checkpoint.

Model card · Details

Higgs Audio v2 — Official architecture diagram

Official architecture diagram · Source

Higgs Audio v2.5

Higgs Audio v2.5, now documented as Higgs TTS 2.5, reduces the autoregressive audio Transformer to 1B parameters. GRPO-based alignment and a curated voice dataset refine pronunciation, cloning and expressive control tags. It targets production narration and conversational speech with lower computational requirements than the preceding 3B generation model.

Announcement · Details

Higgs Audio v2.5 — Editorial input/output diagram

Editorial input/output diagram · Source

Higgs Audio v3 TTS

Higgs TTS 3 uses an autoregressive decoder over interleaved text and eight speech codebooks, with a delay pattern and fused input/output projections. Reference audio establishes a voice, while inline tokens control emotion, style, pauses and sound effects. The 4B release targets multilingual conversational speech and expressive response rendering.

Model card · Details

Higgs Audio v3 TTS — Official architecture diagram

Official architecture diagram · Source

HiStyle

HiStyle predicts a voice's timbre first and finer delivery attributes afterward from textual descriptions. Contrastive text-audio alignment organizes these style representations before they condition a speech synthesizer.

Paper · GitHub: no author-linked repository found · Details

HiStyle — Figure 2

Figure 2 · Source

HoliDubber

HoliDubber conditions audio generation on video and a text prompt describing speech and sound effects. A causal model plans successive latent patches and a local diffusion Transformer generates their detail, supporting synchronized dubbing within complex acoustic scenes.

Paper · Project · GitHub: no author-linked repository found · Details

HoliDubber — Figure 2

Figure 2 · Source

HoliTok (TTS)

HoliTok combines linguistic and acoustic information in a continuous representation designed for both understanding and generation. Its downstream autoregressive model and diffusion decoder demonstrate a text-to-speech path using the same latents employed for recognition.

Paper · GitHub · Details

HoliTok (TTS) — Figure 1

Figure 1 · Source

Hume Octave TTS

Octave uses text context and acting instructions to adjust pronunciation, emphasis, tempo and emotional delivery. Its API supports voice creation from descriptions, voice cloning and continuation across longer passages. Octave 1 and the Octave 2 preview have different feature coverage; the public interface is documented more fully than the internal speech-model architecture.

Docs · Details

Hume Octave TTS — Editorial input/output diagram

Editorial input/output diagram · Source

ImmersiveTTS

ImmersiveTTS jointly models spoken content and its surrounding acoustic scene in a multimodal diffusion Transformer. Speech and general-audio representations provide complementary training signals, helping generated wording remain intelligible within the requested environmental context.

Paper · GitHub · Project · Details

ImmersiveTTS — Figure 1

Figure 1 · Source

IndexTTS

IndexTTS adapts the XTTS/Tortoise approach with a Conformer reference encoder and a BigVGAN2 speech decoder. Hybrid character/pinyin input gives explicit control over difficult Chinese pronunciations. Its central use case is zero-shot voice cloning with predictable text rendering, including content that benefits from pronunciation correction.

Paper · GitHub · Project · Details

IndexTTS — Figure 1

Figure 1 · Source

IndexTTS 2.5

IndexTTS 2.5 shortens semantic sequences with a lower-rate codec and replaces the acoustic module's backbone with a Zipformer design. Multilingual training strategies and reinforcement learning extend pronunciation and emotion transfer across languages. It retains reference-based voice conditioning while reducing the cost of semantic and acoustic generation.

Paper · GitHub · Details

IndexTTS 2.5 — Figure 1

Figure 1 · Source

IndexTTS2

IndexTTS2 separates speaker identity from emotional style so that different references can control timbre and delivery. Its autoregressive formulation also supports explicit output-token budgeting for duration control, alongside unconstrained generation. These mechanisms are intended for expressive speech and timing-sensitive work such as dubbing; availability of controls should be checked in the chosen implementation.

Paper · GitHub · Details

IndexTTS2 — Figure 1

Figure 1 · Source

InstructAudio

InstructAudio combines natural-language instructions with phonemes or lyrics in a common generation format. Joint and single-stream diffusion layers synthesize speech or music while controlling attributes such as voice, emotion, accent or musical character.

Paper · Project · GitHub: no author-linked repository found · Details

InstructAudio — Figure 2

Figure 2 · Source

IntMeanFlow

IntMeanFlow distills a flow-based speech generator to predict integrated acoustic updates over larger intervals. A search for effective sampling steps further reduces decoding work, enabling reference-conditioned synthesis with fewer iterative refinements.

Paper · Project · GitHub: no author-linked repository found · Details

IntMeanFlow — Figure 1

Figure 1 · Source

Inworld TTS-1

Inworld TTS-1 and TTS-1-Max are multilingual autoregressive synthesizers designed for low-latency speech output. Textual audio markup controls emotions and nonverbal vocalizations, with model variants offering different capacity and inference-cost trade-offs.

Paper · GitHub · Details

Inworld TTS-1 — Figure 1

Figure 1 · Source

JaiTTS

JaiTTS continually trains a VoxCPM-derived synthesizer on Thai-centered speech data. Its semantic planning, residual acoustic modeling and local diffusion decoding retain reference-based voice cloning while targeting fluent Thai pronunciation and delivery.

Paper · GitHub · Details

JaiTTS — Figure 1

Figure 1 · Source

JAM-Flow

JAM-Flow couples audio and motion diffusion modules within a shared model. Text, voice references and optional motion conditions support synchronized speech and facial movement, with an infilling objective allowing several conditioning combinations.

Paper · GitHub: no author-linked repository found · Details

JAM-Flow — Figure 3

Figure 3 · Source

JELLY

JELLY combines an emotion-aware Q-Former with several partially adapted language-model modules. Joint emotion recognition and contextual reasoning guide a speech synthesizer toward responses whose delivery matches the conversation.

Paper · GitHub · Project · Details

JELLY — Paper figure

Paper figure · Source

JETS

Joint FastSpeech 2 and HiFi-GAN with learned alignment.

Paper · Details

JETS — Figure 1, PDF p. 2

Figure 1, PDF p. 2 · Source

Joint non-autoregressive STT-TTS

This model handles text and speech within one non-autoregressive architecture. Its TTS path predicts acoustic output from text, and feeding partial predictions back into the model improves generation through iterative refinement while also supporting recognition training.

Paper · GitHub: no author-linked repository found · Details

Joint non-autoregressive STT-TTS — Paper figure

Paper figure · Source

Joycent

Joycent separates accent information from speaker identity using an adversarially trained accent encoder. It injects accent and speaker features at different text-encoder layers, supporting accent-conditioned synthesis without requiring a separate accented phoneme prediction stage.

Paper · GitHub · Project · Details

Joycent — Figure 1

Figure 1 · Source

JoyVoice

JoyVoice conditions long-form speech on speaker-labeled text and shared conversational context. Autoregressive hidden states feed an acoustic diffusion decoder, allowing multiple speakers, changing expression and more flexible turn boundaries within one synthesis system.

Paper · Project · GitHub: no author-linked repository found · Details

JoyVoice — Figure 2

Figure 2 · Source

KABURI-TTS

KABURI-TTS renders each participant on a separate audio channel from time-aligned phonemes and speaker activity. Supplying the timing layout explicitly lets the system synthesize overlapping speech, backchannels and interruptions for two-speaker conversations.

Paper · GitHub: no author-linked repository found · Details

KABURI-TTS — Figure 1

Figure 1 · Source

KittenTTS

KittenTTS provides small ONNX speech models with built-in voices and adjustable playback speed. The Mini, Micro and Nano releases offer different size and inference tradeoffs for CPU-oriented applications. The official README documents usage more fully than internal acoustic design, so the catalog presents it as a compact synthesis family without asserting an undisclosed architecture.

GitHub · Details

KittenTTS — Editorial input/output diagram

Editorial input/output diagram · Source

Koel-TTS

Koel-TTS explores several ways to condition a Transformer synthesizer on text and reference audio. Automatic speech-recognition and speaker-verification feedback, together with classifier-free guidance, improve adherence to the requested words and voice.

Paper · Project · GitHub: no author-linked repository found · Details

Koel-TTS — Figure 1

Figure 1 · Source

Kokoro

Kokoro's 2025 v1.0 release uses a compact StyleTTS 2-derived decoder with an iSTFTNet waveform generator. The released model relies on preset voice representations and omits style diffusion and a reference encoder. It is suited to lightweight narration and application speech where a small synthesis model and ready-made voices are useful.

Model card · Details

Kokoro — Editorial input/output diagram

Editorial input/output diagram · Source

Kyutai TTS (DSM)

Kyutai TTS treats text and speech as aligned streams separated by a controlled delay. A decoder-only language model can therefore emit audio as text arrives, instead of waiting for a complete utterance. This formulation supports incremental synthesis for voice interfaces and long streams, with speaker conditioning supplied through the TTS implementation.

Paper · GitHub · Details

Kyutai TTS (DSM) — Figure 1

Figure 1 · Source

LanStyleTTS

LanStyleTTS standardizes phonetic inputs and introduces local style conditioning across languages. The framework can augment several parallel acoustic backbones, allowing one multilingual model to vary delivery at phoneme level.

Paper · GitHub: no author-linked repository found · Details

LanStyleTTS — Paper figure

Paper figure · Source

LatinX

LatinX uses staged text-to-audio training, voice-cloning adaptation and automatic preference alignment. The resulting multilingual Transformer renders text in the source speaker's voice, supporting the synthesis stage of cross-language speech translation.

Paper · Project · GitHub: no author-linked repository found · Details

LatinX — Figure 1

Figure 1 · Source

LE2E-TTS

LE2E-TTS trains a compact text-to-waveform pipeline end to end rather than separately optimizing acoustic and waveform stages. It targets local devices where model size, response time and compute cost constrain deployment.

Paper · GitHub: no author-linked repository found · Details

LE2E-TTS — Figure 1

Figure 1 · Source

LightSpeech

FastSpeech-derived architecture found by neural architecture search.

Paper · Details

LightSpeech — Editorial input/output diagram

Editorial input/output diagram · Source

LLaDA-TTS

LLaDA-TTS adapts a speech language model to fill masked token sequences in parallel. Its bidirectional generation also supports inserting, replacing or deleting spoken words, combining reference-based TTS and speech editing through the same model.

Paper · GitHub: no author-linked repository found · Details

LLaDA-TTS — Figure 1

Figure 1 · Source

Llasa

Llasa maps text and an optional speech prompt to a single stream of codec tokens using a Llama-style Transformer. The simple representation makes standard language-model scaling and sampling techniques applicable to synthesis. Its research also explores speech-model verifiers that select samples for content accuracy, voice consistency or emotional expression.

Paper · GitHub · Project · Details

Llasa — Figure 2

Figure 2 · Source

Llasa+

Llasa+ adds multi-token prediction modules to a frozen Llasa backbone and checks their proposals with that backbone. A causal codec decoder turns accepted tokens into streaming audio. The resulting design addresses autoregressive latency while retaining the original speech model, making it relevant to systems that need incremental playback without retraining an entire backbone.

Paper · GitHub · Details

Llasa+ — Figure 1

Figure 1 · Source

LLMVoX

LLMVoX connects to an upstream language model through a streaming queue interface. Its small speech generator renders incoming text incrementally, allowing long conversations without tightly coupling the synthesizer to one particular language-model backbone.

Paper · GitHub · Project · Details

LLMVoX — Figure 2

Figure 2 · Source

Lombard Matcha-TTS

This Matcha-TTS extension learns vocal effort and articulation from automatically derived labels. It provides continuous controls over speech clarity and loudness-related effort, together with word-level emphasis, to synthesize the clearer delivery used in noisy listening conditions.

Paper · GitHub: no author-linked repository found · Details

Lombard Matcha-TTS — Figure 1

Figure 1 · Source

LongCat-AudioDiT

LongCat-AudioDiT maps text and reference speech into continuous waveform latents using a diffusion model. A jointly considered waveform autoencoder and adapted inference guidance address the interaction between acoustic reconstruction and zero-shot synthesis quality.

Paper · GitHub · Details

LongCat-AudioDiT — Figure 1

Figure 1 · Source

LoRP-TTS

LoRP-TTS adapts a pretrained zero-shot synthesizer using small low-rank parameter updates. It focuses on preserving a target speaker from limited, potentially noisy or spontaneous recordings whose acoustic conditions differ from the original training data.

Paper · GitHub: no author-linked repository found · Details

LoRP-TTS — Figure 3

Figure 3 · Source

Luna-TTS

Luna-TTS adapts an autoregressive text backbone into a speech diffusion language model. Its parallel and Realtime variants share a tokenizer and training lineage; Realtime predicts successive codec blocks while denoising each block in parallel for incremental audio delivery.

Paper · Project · GitHub: no author-linked repository found · Details

Luna-TTS — Figure 1 (paper page 4)

Figure 1 (paper page 4) · Source

M3-TTS

M3-TTS uses joint text-audio diffusion layers to learn alignment without first stretching text into a guessed acoustic timeline. Additional single-stream layers refine acoustic details, producing reference-conditioned speech through a compressed mel representation.

Paper · GitHub: no author-linked repository found · Details

M3-TTS — Figure 1

Figure 1 · Source

MAGIC-TTS

MAGIC-TTS exposes timing controls for selected speech tokens and pauses. Training includes incomplete control signals so the model can follow local edits where provided and infer natural timing elsewhere, supporting precise pacing without requiring every segment to be specified.

Paper · GitHub · Details

MAGIC-TTS — Figure 1

Figure 1 · Source

MagpieTTS-LF

MagpieTTS-LF extends MagpieTTS at inference time by retaining acoustic and textual context across sentence boundaries. Soft alignment priors and history-aware encoding support coherent longer narration without retraining the synthesizer on long recordings.

Paper · GitHub: no author-linked repository found · Details

MagpieTTS-LF — Figure 1

Figure 1 · Source

MambaVoiceCloning

MambaVoiceCloning uses state-space modules to encode phonemes, learn their timing and condition expressive synthesis. A training-only alignment teacher supplies timing supervision, while the generation path targets efficient long sequences and limited-lookahead speech streaming.

Paper · GitHub · Details

MambaVoiceCloning — Figure 1

Figure 1 · Source

MamTra

MamTra mixes state-space and attention layers to retain global context while reducing the cost of long sequences. Knowledge transfer from a pretrained Transformer initializes the hybrid synthesizer, combining efficient local processing with expressive acoustic modeling.

Paper · GitHub · Project · Details

MamTra — Figure 1

Figure 1 · Source

ManchuTTS

ManchuTTS builds multilevel text representations suited to Manchu and feeds them into a convolutional diffusion Transformer. Its non-autoregressive generator and augmented training data target speech synthesis where naturally recorded material is scarce.

Paper · GitHub: no author-linked repository found · Details

ManchuTTS — Paper figure

Paper figure · Source

Marco-Voice

Marco-Voice learns separate speaker and emotion representations using contrastive training. Rotating the emotional representation provides smooth expressive control while preserving the reference voice across different delivery styles.

Paper · GitHub · Details

Marco-Voice — Figure 1

Figure 1 · Source

MARS6

MARS6 encodes text and a speaker representation before generating hierarchical acoustic codes. Its compact encoder-decoder design targets expressive speech and reference-voice cloning with reduced inference cost.

Paper · GitHub · Project · Details

MARS6 — Paper figure

Paper figure · Source

Masked-style TTS

This controllable system first predicts a masked-autoencoder-derived speech style representation from text and controls. A second model generates codec tokens, allowing speaker characteristics and expressive attributes to be specified separately during synthesis.

Paper · GitHub: no author-linked repository found · Details

Masked-style TTS — Paper figure

Paper figure · Source

MaskGCT

Masked generative codec transformers.

Paper · Details

MaskGCT — Figure 1

Figure 1 · Source

Matcha-TTS

Conditional flow matching with an encoder-decoder acoustic model.

Paper · Details

Matcha-TTS — Figure 1

Figure 1 · Source

MAVE

MAVE combines a state-space backbone with cross-attention to generate speech conditioned on text and acoustic context. It supports zero-shot voice synthesis and editing while reducing the attention-memory requirements of a comparable Transformer-based codec model.

Paper · GitHub: no author-linked repository found · Details

MAVE — Figure 1

Figure 1 · Source

Mega-TTS

Disentangled speech factors and prosody LM.

Paper · Details

Mega-TTS — Figure 1

Figure 1 · Source

Mega-TTS 2

Prosody language model and multi-sentence prompting.

Paper · Details

Mega-TTS 2 — Figure 1

Figure 1 · Source

MegaTTS 3

MegaTTS 3 guides a latent diffusion Transformer with sparse text-speech alignment boundaries, leaving the model room to learn finer timing. Classifier-free guidance controls accent strength, while piecewise rectified flow reduces sampling work. The design targets robust zero-shot voice synthesis with more flexible alignment than a fully fixed duration sequence.

Paper · GitHub · Details

MegaTTS 3 — Figure 1

Figure 1 · Source

Meitei Mayek TTS

This Manipuri speech synthesizer maps Meitei Mayek writing to an ARPAbet-based phoneme representation before acoustic generation with Tacotron 2. A HiFi-GAN vocoder reconstructs the waveform. The paper develops a single-speaker system for a language with limited training resources and tonal pronunciation requirements.

Paper · GitHub: no author-linked repository found · Details

Meitei Mayek TTS — Paper figure

Paper figure · Source

Mel-LLM (TTS)

The synthesis experiment in Mel-LLM extends a language model to predict mel-based acoustic information directly. Its next-token VAE decoder demonstrates a text-to-speech path within an otherwise understanding-focused model; the paper presents this as a proof of concept with quality limitations.

Paper · GitHub: no author-linked repository found · Details

Mel-LLM (TTS) — Fig. 1 (paper page 2)

Fig. 1 (paper page 2) · Source

MELA-TTS

MELA-TTS predicts continuous mel-spectrogram frames from text and speaker conditions. A training-time alignment module connects the decoder to recognition-derived semantic features, helping the joint Transformer-diffusion model retain linguistic structure without discrete speech tokenization.

Paper · GitHub: no author-linked repository found · Details

MELA-TTS — Figure 1

Figure 1 · Source

MELD

MELD learns discrete latent variables from mel-spectrograms jointly with its speech language model. This shared optimization supports zero-shot synthesis and recognition while addressing omissions and excessive silence associated with less coordinated acoustic representations.

Paper · GitHub: no author-linked repository found · Details

MELD — Figures 1–2 (paper page 2; panels assembled)

Figures 1–2 (paper page 2; panels assembled) · Source

Mellotron

Tacotron 2 with global style tokens, pitch and rhythm conditioning.

Paper · Details

Mellotron — Editorial input/output diagram

Editorial input/output diagram · Source

MeloTTS

VITS-family multilingual speech synthesis.

GitHub · Details

MeloTTS — Editorial input/output diagram

Editorial input/output diagram · Source

Metis

Metis pretrains on unlabeled speech before adapting to task-specific conditions such as text. Self-supervised semantic tokens and acoustic codes support a shared foundation for reference-based TTS, conversion and other speech-generation tasks.

Paper · GitHub · Project · Details

Metis — Paper figure

Paper figure · Source

MFCIG-CSS

MFCIG-CSS represents dialogue history through separate graphs of meaning and vocal expression. Fine-grained multimodal interactions condition the speech synthesizer, helping each scripted response fit the surrounding conversation.

Paper · GitHub · Details

MFCIG-CSS — Figure 1

Figure 1 · Source

MiDashengLM-Gen

MiDashengLM-Gen trains a language model together with a conditional flow-matching output head to generate variable-length audio. Its text conditioning supports scenes containing intelligible speech alongside music or other sounds, with generation performed over continuous audio representations.

Paper · GitHub · Project · Details

MiDashengLM-Gen — Paper figure

Paper figure · Source

MiniMax-Speech

MiniMax-Speech extracts speaker characteristics directly from reference audio without requiring its transcript, then generates speech with an autoregressive Transformer and Flow-VAE. The research emphasizes multilingual zero-shot cloning and expressive delivery. Additional adaptation mechanisms support emotion control, description-based voice creation and more specialized voice cloning without replacing the base model.

Paper · GitHub: no author-linked repository found · Details

MiniMax-Speech — Figure 1

Figure 1 · Source

MixedG2P-T5

MixedG2P-T5 learns acoustic units from speech and uses a language-model synthesis path for text containing mixed scripts. It reduces dependence on manually designed grapheme-to-phoneme rules while retaining accent and intonation information in the speech representation.

Paper · GitHub: no author-linked repository found · Details

MixedG2P-T5 — Figure 3

Figure 3 · Source

MM-MovieDubber

MM-MovieDubber interprets scene information to distinguish dialogue, narration and monologue delivery. A speech generator then uses the resulting multimodal conditions with the target content to render expressive movie dubbing.

Paper · GitHub: no author-linked repository found · Details

MM-MovieDubber — Figure 2

Figure 2 · Source

MoE-TTS

MoE-TTS augments a frozen text language model with speech-specific expert parameters. Retaining the original language knowledge helps the synthesizer interpret unfamiliar style descriptions while learning the acoustic generation task.

Paper · GitHub: no author-linked repository found · Details

MoE-TTS — Figure 1

Figure 1 · Source

MoonCast

MoonCast combines podcast script preparation with a synthesizer trained for longer, spontaneous-sounding delivery. Voice references allow unseen speakers to render the resulting conversation, while discourse-level context supports more natural transitions than isolated sentence synthesis.

Paper · GitHub · Project · Details

MoonCast — Figure 1

Figure 1 · Source

MOSS-TTS

MOSS-TTS offers two generators over a shared discrete audio representation: a delay-pattern model and a model with a frame-local Transformer. They balance long-context control against efficient codebook prediction and speaker preservation. The family supports reference-conditioned synthesis, pronunciation and duration controls, with later checkpoints extending language handling and explicit pauses.

Paper · GitHub · Details

MOSS-TTS — Figure 2, PDF p. 8

Figure 2, PDF p. 8 · Source

MOSS-TTS-Nano

MOSS-TTS-Nano packages multilingual voice cloning into a roughly 100M-parameter speech generator with a compact audio tokenizer. Streaming output and an ONNX inference path make it relevant to CPU-based readers and local applications. Its published performance depends on the runtime and hardware, and the tokenizer is a separate part of the deployment footprint.

GitHub 1 · GitHub 2 · Details

MOSS-TTS-Nano — Official architecture diagram

Official architecture diagram · Source

MOSS-TTS-Realtime

MOSS-TTS-Realtime uses a Qwen3-derived backbone for linguistic context and a smaller local Transformer to predict audio codebooks. Text and speech tokens are handled at different levels so the system can accept text and emit audio incrementally. It targets low-latency spoken responses while preserving context across the generated utterance.

GitHub · Docs · Details

MOSS-TTS-Realtime — Editorial input/output diagram

Editorial input/output diagram · Source

MOSS-TTSD

MOSS-TTSD turns a dialogue script with explicit speaker tags into a continuous multi-party recording. Long-context modeling helps maintain speaker identity, turn assignment and acoustic continuity, while short references can define voices. It is designed for podcasts, commentary and other scripted conversations; the source paper evaluates dialogue-specific consistency as well as intelligibility.

Paper · GitHub · Details

MOSS-TTSD — Figure 2, PDF p. 4

Figure 2, PDF p. 4 · Source

MOSS-VoiceGenerator

MOSS-VoiceGenerator creates a speaking voice from a natural-language description rather than requiring an example speaker recording. Training on expressive cinematic speech exposes it to varied delivery and acoustic conditions. The model is intended for character design, storytelling and role-based narration where the desired voice must be specified in words.

Paper · GitHub · Details

MOSS-VoiceGenerator — Figure 1

Figure 1 · Source

MP-ELD

MP-ELD predicts low-rate continuous speech tokens through several information paths with separate local encoders. A flow decoder combines their predictions, while the accompanying Locodec representation is designed to limit accumulated errors during long speech generation.

Paper · GitHub: no author-linked repository found · Details

MP-ELD — Figure 2

Figure 2 · Source

MPE-TTS

MPE-TTS combines reference speech and textual prompts to specify an unseen speaker and the desired emotion. A prosody predictor and emotion-consistency objective carry those controls into the synthesized acoustic performance.

Paper · GitHub: no author-linked repository found · Details

MPE-TTS — Figure 1

Figure 1 · Source

Multistage multimodal TTS

This framework learns face and text conditioning in separate stages before using them for voice synthesis. Visual knowledge distillation and training across text-face and text-speech pairs reduce reliance on fully matched multimodal recordings.

Paper · GitHub: no author-linked repository found · Details

Multistage multimodal TTS — Paper figure

Paper figure · Source

Muyan-TTS

Muyan-TTS trains a speech language model on a large podcast collection for expressive reference-based synthesis. The release documents data preparation, training and optimized inference, with an emphasis on reproducible podcast-style voice generation.

Paper · GitHub · Details

Muyan-TTS — Figure 1

Figure 1 · Source

NaturalSpeech

Text-to-waveform VAE with enhanced prior and duration modeling.

Paper · Details

NaturalSpeech — Figure 1

Figure 1 · Source

NaturalSpeech 2

Latent diffusion over neural-codec representations.

Paper · Details

NaturalSpeech 2 — Figure 1

Figure 1 · Source

NaturalSpeech 3

Factorized speech codec and attribute-wise diffusion.

Paper · Details

NaturalSpeech 3 — Figure 3

Figure 3 · Source

NeuTTS Air

NeuTTS Air pairs a phoneme-conditioned language model with NeuCodec to synthesize a reference voice locally. Quantized GGUF backbones support incremental generation through the documented streaming backend. It targets embedded and desktop voice applications, with the codec's compute and memory requirements considered alongside those of the language-model backbone.

GitHub · Details

NeuTTS Air — Editorial input/output diagram

Editorial input/output diagram · Source

NeuTTS Nano

NeuTTS Nano reduces the speech-model backbone while retaining phoneme conditioning, reference-based cloning and NeuCodec reconstruction. Its English, German, French and Spanish models share a design but use separate language-specific checkpoints. Quantized variants support compact local deployments, and streaming depends on the selected inference backend.

GitHub · Details

NeuTTS Nano — Editorial input/output diagram

Editorial input/output diagram · Source

NeuTTS-2E

NeuTTS-2E accepts text directly and adds explicit emotional delivery to the NeuTTS language-model-and-codec pipeline. The released configuration supplies four fixed speaker presets instead of arbitrary reference-based cloning. It is intended for compact expressive speech applications, with streaming available through the documented GGUF inference path.

GitHub · Details

NeuTTS-2E — Editorial input/output diagram

Editorial input/output diagram · Source

NR-LauraTTS

NR-LauraTTS cleans the discrete representation of a noisy voice prompt before passing it to LauraTTS. Token prediction and embedding refinement reduce background contamination, supporting reference cloning when the available recording is acoustically imperfect.

Paper · GitHub · Project · Details

NR-LauraTTS — Figure 1

Figure 1 · Source

NVSpeech TTS

The NVSpeech pipeline includes a TTS model that renders text with explicitly marked nonverbal events. Word-level annotations connect ordinary speech with vocalizations such as laughter, providing a shared representation for recognition and controllable audio generation.

Paper · Project · GitHub: no author-linked repository found · Details

NVSpeech TTS — Figure 2

Figure 2 · Source

Nüshu-PitchVITS

Nüshu-PitchVITS uses pitch annotations from Nüshu's writing system to guide acoustic generation under very limited data. A frame-level pitch predictor conditions the VITS waveform path, allowing syllable recordings and linguistic tone knowledge to support sentence synthesis.

Paper · GitHub: no author-linked repository found · Details

Nüshu-PitchVITS — Figure 3

Figure 3 · Source

Ojibwe-Mi'kmaq-Maliseet TTS

This model family shares speech-synthesis training across three related Indigenous languages. The paper compares attention-based and attention-free flow architectures, demonstrating how joint linguistic coverage can help languages with limited recordings.

Paper · GitHub · Details

Ojibwe-Mi'kmaq-Maliseet TTS — Figure 1

Figure 1 · Source

OmniVoice

OmniVoice predicts multiple acoustic codebooks directly from text using a masked, non-autoregressive diffusion language model. Random masking across codebooks and initialization from a pretrained language model support multilingual generation without a separate text-to-semantic stage. It focuses on broad-language zero-shot synthesis and voice conditioning, rather than visual or general-purpose omni interaction.

Paper · GitHub · Details

OmniVoice — Figure 1

Figure 1 · Source

OpusLM

OpusLM extends text language models through speech-text pretraining on public data. Its interleaved representation supports speech recognition, text-conditioned synthesis and textual continuation within a transparent family of shared backbones.

Paper · GitHub: no author-linked repository found · Details

OpusLM — Figure 1

Figure 1 · Source

Orpheus TTS

Orpheus TTS repurposes a Llama-family language model to generate speech codec tokens from text. Emotion tags and speaker conditioning guide expressive delivery, while the project supplies inference and adaptation workflows. English releases and multilingual previews have different coverage, so a checkpoint's documented capabilities matter when selecting it for narration or a voice application.

GitHub · Details

Orpheus TTS — Editorial input/output diagram

Editorial input/output diagram · Source

OscillaTTS

OscillaTTS changes the periodic nonlinearities used in a style-diffusion synthesis backbone. Adjustable oscillatory modulation is designed to capture rapid pitch and amplitude changes while a linear bypass stabilizes the acoustic signal, targeting sharper expressive prosody.

Paper · GitHub: no author-linked repository found · Details

OscillaTTS — Figure 1

Figure 1 · Source

OuteTTS

OuteTTS represents speech in a form that can be generated by a decoder-only language model and reconstructed by an audio decoder. A speaker reference guides vocal identity, style and accent. Its 1.0 line supports standard LLM serving backends, but generation settings and repetition handling need to follow the matching model implementation.

GitHub · Details

OuteTTS — Editorial input/output diagram

Editorial input/output diagram · Source

OV-InstructTTS

OV-InstructTTS interprets voice and delivery descriptions beyond a fixed inventory of style labels. Its reasoning-based conditioning connects broader textual requests with expressive speech generation, supported by a dedicated instruction-speech dataset.

Paper · GitHub · Details

OV-InstructTTS — Figure 2

Figure 2 · Source

OZSpeech

OZSpeech generates disentangled speech components with a flow model conditioned on a learned prior. The design targets single-step zero-shot synthesis while separately modeling content, prosody and speaker-related information from the voice prompt.

Paper · GitHub · Project · Details

OZSpeech — Figure 1

Figure 1 · Source

PALLE

PALLE generates variable-length speech spans at fixed decoding steps, combining temporal planning with parallel token prediction. A second non-autoregressive stage refines the initial sequence, supporting efficient zero-shot synthesis.

Paper · GitHub: no author-linked repository found · Details

PALLE — Figure 3

Figure 3 · Source

Parallel GPT

Parallel GPT divides speech generation between a general autoregressive predictor and a non-autoregressive detail model. The parallel refinement stage conditions on the initial tokens, balancing independence and interaction between semantic and acoustic information.

Paper · GitHub: no author-linked repository found · Details

Parallel GPT — Paper figure

Paper figure · Source

Parallel Tacotron

Parallel acoustic model with a variational residual encoder.

Paper · Details

Parallel Tacotron — Figure 1

Figure 1 · Source

Parallel Tacotron 2

Parallel synthesis with differentiable duration modeling.

Paper · Details

Parallel Tacotron 2 — Figure 1

Figure 1 · Source

ParaStyleTTS

ParaStyleTTS converts textual style prompts into separate controls for prosody and broader paralinguistic characteristics. The lightweight adaptation design targets expressive speech from descriptions while making the roles of the two conditioning levels explicit.

Paper · GitHub · Project · Details

ParaStyleTTS — Figure 1

Figure 1 · Source

Parler-TTS

Description-conditioned codec language model.

GitHub · Paper · Details

Parler-TTS — Figure 1

Figure 1 · Source

Parler-TTS Hinglish adaptation

This Parler-TTS extension introduces language-specific phonetic alignment and emotion embeddings for Hindi and Indian English. Its conditioning targets code-switched utterances whose accent and emotional delivery change coherently across language boundaries.

Paper · GitHub · Details

Parler-TTS Hinglish adaptation — Figure 2

Figure 2 · Source

PFluxTTS

PFluxTTS combines two acoustic-generation paths by fusing their predicted vector fields at inference. Sequential reference embeddings support transcript-free cross-language voice cloning, and a super-resolution vocoder reconstructs high-rate output audio.

Paper · Project · GitHub: no author-linked repository found · Details

PFluxTTS — Figure 1

Figure 1 · Source

Phoenix TTS

Phoenix TTS aligns its speech tokenizer with the downstream flow-matching acoustic decoder during training. An autoregressive language model predicts the resulting discrete representation, supporting text-to-speech and voice conversion without separating token design from acoustic reconstruction.

Paper · GitHub: no author-linked repository found · Details

Phoenix TTS — Figure 2

Figure 2 · Source

Phoneme-tone adaptive Thai TTS

This Thai speech synthesizer encodes phonemes and tones with a language-specific BERT model, then predicts duration, pitch and energy for a GAN-trained waveform decoder. A reference-derived style vector supports voice cloning. Multilingual pretraining of acoustic feature extractors and Thai adaptation address limited language-specific data.

Paper · GitHub: no author-linked repository found · Details

Phoneme-tone adaptive Thai TTS — Figure 2

Figure 2 · Source

PilotTTS

PilotTTS uses paired recordings and Q-Former conditioning to separate a speaker's identity from delivery style. The model supports reference cloning, emotional and nonverbal expression, and Chinese dialect synthesis within a shared autoregressive pipeline.

Paper · GitHub · Details

PilotTTS — Figure 3

Figure 3 · Source

Piper (VITS voices)

VITS voice models exported for local inference.

GitHub · Docs · Details

Piper (VITS voices) — Editorial input/output diagram

Editorial input/output diagram · Source

Pocket TTS

Pocket TTS uses continuous autoregressive speech modeling with a flow-based output mechanism, avoiding long sequences of discrete acoustic codebooks. It combines a small language-model backbone with streaming audio reconstruction and reusable voice conditioning. The project targets CPU-based speech synthesis; language-specific models and runtime choices affect its speed and voice behavior.

GitHub · Paper · Details

Pocket TTS — Figure 1

Figure 1 · Source

PortaSpeech

Variational acoustic model and flow-based post-net.

Paper · Details

PortaSpeech — Figure 1, PDF p. 4

Figure 1, PDF p. 4 · Source

PROEMO

PROEMO combines emotional prompts with an explicit intensity control in a multi-speaker synthesizer. The conditioning adjusts delivery strength and prosodic variation, allowing the same spoken text to be rendered with different emotional performances.

Paper · GitHub: no author-linked repository found · Details

PROEMO — Figure 1

Figure 1 · Source

Progressive face-conditioned TTS

This face-conditioned synthesizer combines local facial regions into progressively broader visual representations. Joint visual and acoustic attribute learning and multiple photographs of each training speaker align the face representation with voice characteristics, conditioning speech generation on text and a face image.

Paper · GitHub: no author-linked repository found · Details

Progressive face-conditioned TTS — Figure 1

Figure 1 · Source

Prompt-Unseen-Emotion

Prompt-Unseen-Emotion learns the relationship between emotion descriptions and speech using a language-model synthesis backbone. Weighted combinations of known emotions and contextual language knowledge allow expressive delivery outside the original categorical training labels.

Paper · GitHub: no author-linked repository found · Details

Prompt-Unseen-Emotion — Paper figure

Paper figure · Source

PromptTTS

Style and content text encoders with a speech decoder.

Paper · Details

PromptTTS — Figure 1

Figure 1 · Source

PromptTTS 2

Prompt-conditioned TTS with a diffusion variation network.

Paper · Details

PromptTTS 2 — Figure 1

Figure 1 · Source

ProtoDisent-TTS

ProtoDisent-TTS learns a codebook of healthy and dysarthric articulation patterns separately from speaker identity. Adversarial constraints reduce pathological information in the speaker representation, enabling controlled synthesis of articulation characteristics in a target voice.

Paper · Project · GitHub: no author-linked repository found · Details

ProtoDisent-TTS — Figure 1

Figure 1 · Source

PS-TTS

PS-TTS uses vowel-based alignment to coordinate the timing and phonetic structure of dubbed speech. Its PS-Comet variant also considers semantic preservation when choosing translated text, connecting translation choices with a TTS rendering stage.

Paper · GitHub: no author-linked repository found · Details

PS-TTS — Fig. 1 (paper page 3)

Fig. 1 (paper page 3) · Source

QTTS

QTTS predicts residual speech codes produced by its QDAC tokenizer. Hierarchical parallel and delayed multihead variants organize codebook dependencies differently, offering alternative balances between acoustic detail and sequential decoding cost.

Paper · GitHub: no author-linked repository found · Details

QTTS — Figure 2

Figure 2 · Source

Qwen-Audio-3.0-TTS

Qwen-Audio-3.0-TTS combines compact semantic speech tokens with progressively trained language and acoustic models. Natural-language instructions and inline tags control delivery, while multilingual reference conditioning supports voice cloning and longer speech generation under varied recording conditions.

Paper · Project · GitHub: no author-linked repository found · Details

Qwen-Audio-3.0-TTS — Figure 3

Figure 3 · Source

Qwen3-TTS

Qwen3-TTS combines a dual-track speech language model with tokenizers designed for compact streaming audio. The released 12Hz line separates Base voice cloning, CustomVoice preset-speaker control and VoiceDesign creation from descriptions. These variants support different conditioning interfaces, allowing applications to choose between reproducing a reference voice and directing a new voice through text.

Paper · GitHub · Details

Qwen3-TTS — Figure 3

Figure 3 · Source

RADKA-CSS

RADKA-CSS retrieves dialogue examples related to the current conversation in both meaning and delivery. A graph-based aggregation mechanism combines their style information with current context, conditioning expressive conversational speech synthesis.

Paper · GitHub · Project · Details

RADKA-CSS — Figure 2

Figure 2 · Source

RALL-E

Prosody-guided codec language modeling.

Paper · Details

RALL-E — Figure 1

Figure 1 · Source

Raon-OpenTTS

Raon-OpenTTS is a family of reference-conditioned diffusion synthesizers trained on a large, documented English speech collection. The release pairs its models with data processing and evaluation resources, allowing robustness across varied acoustic conditions to be examined alongside clean-speech quality.

Paper · GitHub · Details

Raon-OpenTTS — Figure 1

Figure 1 · Source

RapFlow-TTS

RapFlow-TTS regularizes the acoustic velocity field so longer generation steps remain consistent. Time-interval scheduling and adversarial objectives improve the resulting few-step synthesizer, reducing the iterations needed to render reference-conditioned speech.

Paper · GitHub · Details

RapFlow-TTS — Figure 1

Figure 1 · Source

ReGenVoice

ReGenVoice applies the ReGen representation-and-waveform modeling approach to text-to-speech. Multiple levels of generated conditioning help reconstruct detailed waveforms from compressed latents, linking efficient acoustic representation with reference-conditioned speech synthesis.

Paper · Project · GitHub: no author-linked repository found · Details

ReGenVoice — Figure 1

Figure 1 · Source

ReStyle-TTS

ReStyle-TTS changes vocal attributes relative to a reference recording. Independent text and reference guidance, composable style adapters and timbre-consistency optimization allow continuous expressive edits while limiting changes to the speaker's identity.

Paper · GitHub: no author-linked repository found · Details

ReStyle-TTS — Figure 1

Figure 1 · Source

RTFree-F5

RTFree-F5 replaces the transcript normally associated with an F5-TTS reference recording with projected speech features. A lightweight adapter reuses the pretrained generator, enabling transcript-free voice conditioning, including references whose pronunciation makes transcription unreliable.

Paper · GitHub: no author-linked repository found · Details

RTFree-F5 — Figure 1

Figure 1 · Source

RV-TTS

Revival with Voice learns voice identity from face images and delivery attributes from descriptions. Audio-only training data and stylized portrait augmentation broaden its input coverage, enabling controlled speech from real faces or artistic portraits.

Paper · GitHub: no author-linked repository found · Details

RV-TTS — Figure 1

Figure 1 · Source

RWKVTTS

RWKVTTS uses the recurrent RWKV-7 architecture for speech synthesis in place of a conventional Transformer backbone. Its token-generation path targets efficient streaming and reduced state-management cost while conditioning audio on the supplied text.

Paper · GitHub · Details

RWKVTTS — Figure 2

Figure 2 · Source

S5-TTS

S5-TTS adapts T5-TTS for word-by-word synthesis using limited future text. Lookahead-aware masks, convolutional auxiliary attention and distillation let the model start speaking before the complete sentence is available while retaining voice conditioning and alignment.

Paper · GitHub: no author-linked repository found · Details

S5-TTS — Figure 1

Figure 1 · Source

Sarashina2.2-TTS

Sarashina2.2-TTS emphasizes reliable Japanese pronunciation, including characters with several possible readings. Its semantic language model and acoustic flow decoder use reference speech for voice conditioning, with balanced multilingual training to reduce dependence on the reference language.

Paper · GitHub · Details

Sarashina2.2-TTS — Figure 1

Figure 1 · Source

SASLM

SASLM derives expressive intent from its own evolving semantic states through an information bottleneck. Acoustic feedback aligns generated speech with that intent, reducing the need for externally supplied emotion labels in context-sensitive speech rendering.

Paper · GitHub · Project · Details

SASLM — Figure 3

Figure 3 · Source

Seed-TTS

Autoregressive speech foundation model.

Paper · Details

Seed-TTS — Figure 1

Figure 1 · Source

Self-distilled zero-shot TTS

This zero-shot synthesizer learns linguistic content and reference-speaker attributes through separate representations. Two-stage self-distillation creates aligned examples that strengthen their separation, targeting stable voice cloning with a small inference footprint.

Paper · GitHub: no author-linked repository found · Details

Self-distilled zero-shot TTS — Figure 1

Figure 1 · Source

SelfTTS

SelfTTS learns separate representations of a speaker's identity and emotional delivery using contrastive and adversarial objectives. It then improves synthesis through self-generated training examples, allowing emotion transfer to speakers originally recorded with neutral expression.

Paper · GitHub · Project · Details

SelfTTS — Figure 1

Figure 1 · Source

SemaVoice

SemaVoice organizes its audio VAE latents using guidance from speech foundation-model representations. A continuous autoregressive backbone and patch-level diffusion head then synthesize reference-conditioned speech with greater emphasis on linguistic coherence.

Paper · GitHub: no author-linked repository found · Details

SemaVoice — Figure 1

Figure 1 · Source

SemBridge

SemBridge uses discrete semantic targets during training to organize both acoustic latents and language-model hidden states. The resulting continuous generator supports zero-shot speech synthesis with stronger content alignment, without needing to generate the auxiliary semantic tokens during inference.

Paper · GitHub · Details

SemBridge — Figure 1

Figure 1 · Source

Shallow Flow Matching TTS

Shallow Flow Matching adds a lightweight head that predicts an intermediate acoustic state for a flow-based synthesizer. Starting refinement closer to the target reduces the remaining generation path, allowing coarse-to-fine speech synthesis with less iterative work.

Paper · GitHub · Details

Shallow Flow Matching TTS — Figure 2

Figure 2 · Source

SLED

SLED learns the conditional distribution of acoustic latents using an energy-distance objective rather than discrete token classification. Its autoregressive generator samples continuous speech representations, simplifying synthesis while retaining acoustic detail.

Paper · GitHub · Details

SLED — Figure 2

Figure 2 · Source

SlimSpeech

SlimSpeech reduces the parameter count of a rectified-flow TTS model and transfers knowledge into a lightweight generator. It targets efficient reference-conditioned speech synthesis while retaining the acoustic quality of a larger teacher.

Paper · GitHub: no author-linked repository found · Details

SlimSpeech — Paper figure

Paper figure · Source

SMLLE

SMLLE uses a transducer to align incoming text with semantic speech tokens and duration information. A separate autoregressive stage generates acoustic frames, with controlled access to future text stabilizing incremental synthesis.

Paper · Project · GitHub: no author-linked repository found · Details

SMLLE — Figure 1

Figure 1 · Source

SoulX-Podcast

SoulX-Podcast synthesizes conversational scripts with reference voices, dialect choices and nonverbal expression. Its long-form training targets consistent speaker identity and natural transitions across turns, while also supporting ordinary single-speaker TTS.

Paper · GitHub · Project · Details

SoulX-Podcast — Figure 3

Figure 3 · Source

Spark-TTS

Spark-TTS uses BiCodec to separate changing linguistic content from global speaker attributes, then predicts these tokens with a Qwen2.5 backbone. That separation supports both reference-based cloning and direct control of attributes such as speaking rate and pitch. It is useful for controllable speech generation where a reference recording alone is insufficient to specify the desired delivery.

Paper · GitHub · Details

Spark-TTS — Figure 3

Figure 3 · Source

SpeakStream

SpeakStream trains on text interleaved with corresponding speech and generates audio as new text becomes available. The synthesis module remains compatible with an upstream text-streaming language model, supporting responsive conversational playback.

Paper · Project · GitHub: no author-linked repository found · Details

SpeakStream — Paper figure

Paper figure · Source

SPEAR-TTS

Text-to-semantic and semantic-to-acoustic LMs.

Paper · Details

SPEAR-TTS — Figure 1

Figure 1 · Source

SpeechAccentLLM

SpeechAccentLLM uses a content tokenizer trained with transcription alignment and jointly learns accent conversion and synthesis. A reconstruction refinement stage improves generated speech while separating accent-related changes from the target speaker's identity.

Paper · GitHub: no author-linked repository found · Details

SpeechAccentLLM — Figure 3

Figure 3 · Source

SpeechEdit

SpeechEdit combines text, instruction tokens and reference audio in a shared codec-language-model sequence. Paired examples that differ in selected attributes teach localized changes while retaining other aspects of the reference voice and delivery.

Paper · Project · GitHub: no author-linked repository found · Details

SpeechEdit — Figure 1

Figure 1 · Source

SpeechT5

Shared encoder–decoder with modality interfaces.

Paper · Details

SpeechT5 — Figure 2

Figure 2 · Source

SpeechX

Prompted neural codec language model.

Paper · Details

SpeechX — Figure 1

Figure 1 · Source

SpeedySpeech

Residual convolutional acoustic model with duration expansion.

Paper · Details

SpeedySpeech — Figure 3

Figure 3 · Source

Spotlight-TTS

Spotlight-TTS extracts expressive reference information primarily from voiced speech and adjusts the resulting style direction before acoustic generation. The method targets more faithful expression transfer; the catalog combines its related paper records under one model.

Paper 1 · Paper 2 · GitHub: no author-linked repository found · Details

Spotlight-TTS — Figure 1

Figure 1 · Source

StellarTTS

StellarTTS encodes phoneme timing sparsely and uses a lightweight masked Transformer to generate speech tokens in parallel. A semantic-aware codec supports waveform reconstruction, while explicit timing representations provide control over pronunciation, duration and prosody.

Paper · Project · GitHub: no author-linked repository found · Details

StellarTTS — Paper figure

Paper figure · Source

Step-Audio-EditX

Step-Audio-EditX supports reference-based synthesis and repeated edits to emotion, speaking style or nonverbal delivery. Training on deliberately contrasting synthetic examples teaches the model to follow expressive changes without a separate attribute-embedding module.

Paper · GitHub · Details

Step-Audio-EditX — Figure 2

Figure 2 · Source

Step-Audio-TTS

Step-Audio-TTS is the compact synthesis component produced through the broader Step-Audio speech-data and distillation pipeline. Text and voice conditioning drive speech-token generation, while the released workflow exposes delivery controls. The TTS checkpoint is used to render supplied content; reasoning, tool use and dialogue management belong to other components of the system.

Paper · GitHub · Details

Step-Audio-TTS — Figure 2

Figure 2 · Source

StepAudio 2.5 TTS

The TTS mode of StepAudio 2.5 uses a shared speech-language foundation with synthesis-specific decoding and preference training. Rich contextual supervision and feedback target controllable expression; this card describes its speech-rendering path within the broader system.

Paper · GitHub: no author-linked repository found · Details

StepAudio 2.5 TTS — Figure 1

Figure 1 · Source

Stochastic-alignment continuous TTS

This synthesizer predicts continuous speech latents using a Gaussian-mixture conditional distribution. A stochastic monotonic alignment mechanism keeps the acoustic sequence ordered against the text, offering an alternative to autoregressive discrete-codec modeling.

Paper · GitHub: no author-linked repository found · Details

Stochastic-alignment continuous TTS — Figure 1

Figure 1 · Source

StreamMel

StreamMel alternates text tokens with continuous acoustic frames in one streaming synthesis model. This organization lets incoming text guide speech immediately while retaining reference-voice information across the generated audio stream.

Paper · GitHub: no author-linked repository found · Details

StreamMel — Paper figure

Paper figure · Source

StyleTTS

Style-conditioned parallel synthesis with a transferable aligner.

Paper · Details

StyleTTS — Figure 1

Figure 1 · Source

StyleTTS 2

Style diffusion and adversarial training with speech-model discriminators.

Paper · Details

StyleTTS 2 — Figure 1, PDF p. 4

Figure 1, PDF p. 4 · Source

Supertonic

The SupertonicTTS research system compresses speech into continuous latents and predicts them from character-level text with flow matching. ConvNeXt blocks, temporal compression and a separate duration predictor keep synthesis compact. The later Supertonic ONNX release exposes preset voice-style assets; its packaged configurations should not be equated with the paper's 44M-parameter research model.

Model card · Paper · GitHub · Project · Details

Supertonic — Figure 1

Figure 1 · Source

Supertonic 2

Supertonic 2 extends the local ONNX synthesis line to five languages while retaining voice-style conditioning and a compact model. It provides a practical path to multilingual narration on devices that can run the supplied inference stack. Creating a new voice-style asset is a separate workflow from generating speech with an existing asset.

Model card · Details

Supertonic 2 — Editorial input/output diagram

Editorial input/output diagram · Source

Supertonic 3

Supertonic 3 expands language coverage and adds expression tags while keeping local ONNX inference and preset voice styles. The release targets more reliable reading across short and long text, with controls for events such as breaths or laughter. Custom voice-style creation is offered through a separate service; downloaded styles can then condition local synthesis.

Model card · Details

Supertonic 3 — Editorial input/output diagram

Editorial input/output diagram · Source

SwanVoice

SwanVoice generates monologues or dialogues with up to four speakers using raw text, voice references and speaker-turn conditions. Pause markers and optional pronunciation substitutions provide textual control, while staged dialogue training supports longer expressive speech.

Paper · Project · GitHub: no author-linked repository found · Details

SwanVoice — Figure 2

Figure 2 · Source

SyncSpeech

SyncSpeech uses a temporal masking scheme to coordinate sequential speech structure with parallel token decoding. This hybrid organization targets faster first audio and higher throughput while retaining reference-conditioned synthesis quality.

Paper · Project · GitHub: no author-linked repository found · Details

SyncSpeech — Figure 1

Figure 1 · Source

Tacotron

Attention-based recurrent spectrogram synthesis.

Paper · Details

Tacotron — Figure 1

Figure 1 · Source

Tacotron 2

Recurrent attention model and WaveNet vocoder.

Paper · Details

Tacotron 2 — Figure 1

Figure 1 · Source

TADA

TADA aligns text tokens one-to-one with continuous acoustic units. A language model with a flow-matching head predicts these synchronized representations, reducing the ambiguity of text-speech alignment during reference-conditioned synthesis.

Paper · GitHub · Details

TADA — Figure 2

Figure 2 · Source

TED-TTS

TED-TTS modifies conditioning and decoding in a pretrained zero-shot synthesizer to control different parts of an utterance. Segment-specific emotion masks and duration steering support local changes while coordinating transitions and the overall stopping point.

Paper · GitHub · Details

TED-TTS — Figure 1

Figure 1 · Source

Tibetan-TTS

Tibetan-TTS adapts a large speech generator through language-specific text representation, tokenizer changes and cross-language training. Its data preparation and quality enhancement target scarce Tibetan recordings and the differences between written forms and spoken pronunciation.

Paper · GitHub: no author-linked repository found · Details

Tibetan-TTS — Figure 2

Figure 2 · Source

<a

Truncated — view the full README on GitHub.

Contributors

kadirnar

6 commits

Languages

Python

100.0%