A visual catalog of text-to-speech architectures, with model diagrams, concise descriptions, and primary sources.
Python
18
6 commits
updated Sep 13, 2026
A visual catalog of text-to-speech models, from Tacotron to speech language models. Diagrams, primary sources and short notes for every entry.
380 models and families · Reviewed 2026-09-13
Model list · All diagrams · 2025+ TTS-arxiv-daily collection · Descriptions · Timeline · Methodology · Contribute
The TTS-arxiv-daily collection covers the source list's TTS systems with first paper submissions from January 1, 2025 onward. Every included family has an image, a description, paper links and an explicit GitHub availability status. The complete screening record is available as JSON.
T: text · S: speech or voice reference · A: other audio · I: image · V: video. Scope and labels.
| Model | Group | Input → output |
|---|---|---|
| A2TTS | Diffusion / flow | T, S → S |
| Affectron | Token LM | T, S → S |
| AgentSteerTTS | Token LM | T, S → S |
| AlignDiT | Diffusion / flow | T, S, V → S |
| AMNet | Parallel | T → S |
| ARCHI-TTS | Diffusion / flow | T, S → S |
| ATRIE | Token LM | T → S |
| Audiobook-CC | Token LM | T, S → S |
| AuEmoChat | Token LM | T, S, V → S |
| AuK | Diffusion / flow | T, S → S, A |
| Authentic-Dubber | Diffusion / flow | T, S, V → S |
| AutoSIFT | Diffusion / flow | T, S → S |
| AutoStyle-TTS | Token LM | T, S → S |
| AVLM (expressive speech) | Token LM | T, S, V → S |
| Bagpiper-TTS | Token LM | T → S |
| BareWave | Diffusion / flow | T, S → S |
| Bark | Token LM | T → S, A |
| BASE TTS | Token LM | T, S → S |
| BatonTTS (BatonVoice) | Token LM | T, S → S |
| BELLE | Continuous LM | T, S → S |
| BitTTS | Compact | T → S |
| Block-wise Mimi TTS | Token LM | T → S |
| BnTTS | Token LM | T, S → S |
| Bolbosh | Diffusion / flow | T → S |
| Borderless Long Speech Synthesis | Continuous LM | T, S → S |
| BreezyVoice | Token LM | T, S → S |
| BridgeTTS | Token LM | T, S → S |
| BVS | Token LM | T, V → S, A |
| CAM-TTS | Token LM | T, S → S |
| CapTalk | Token LM | T, S → S |
| CAST-TTS | Diffusion / flow | T, S → S |
| CaT-TTS | Token LM | T, S → S |
| Causal-prosody FastSpeech 2 | Parallel | T → S |
| CDE-StyleTTS | Diffusion / flow | T, S → S |
| Chain-of-Details TTS | Token LM | T, S → S |
| Chain-Talker | Token LM | T, S → S |
| Chatterbox | Token LM | T, S → S |
| Chatterbox-Flash | Token LM | T, S → S |
| ChatTTS | Token LM | T → S |
| CLaM-TTS | Token LM | T, S → S |
| CLEAR | Continuous LM | T, S → S |
| Clip-TTS | Parallel | T → S |
| Compact neural accessibility TTS | Compact | T → S |
| Compressed-to-fine speech LM | Token LM | T, S → S |
| Confucius4-TTS | Token LM | T, S → S |
| Continuous-token diffusion TTS | Continuous LM | T, S → S |
| Controllable masked-speech TTS | Token LM | T, S, A → S |
| CookVoice | Diffusion / flow | T, S → S |
| CosyEdit2 | Token LM | T, S → S |
| CoSyncDiT | Diffusion / flow | T, S, V → S |
| CosyVoice | Token LM | T, S → S |
| CosyVoice 2 | Token LM | T, S → S |
| CosyVoice 3 | Token LM | T, S → S |
| CosyWhisper (WhispSynth) | Token LM | T, S → S |
| CoVoMix2 | Diffusion / flow | T, S → S |
| Cross-Lingual F5-TTS | Diffusion / flow | T, S → S |
| CrossAccent-TTS | Token LM | T, S → S |
| CSM | Token LM | T, S → S |
| CTC-TTS | Token LM | T, S → S |
| CtrlSpeech | Continuous LM | T, S → S |
| CuteTTS | Continuous LM | T, S → S |
| DAIEN-TTS | Diffusion / flow | T, S, A → S |
| DARS | Diffusion / flow | T → S |
| DCAR | Token LM | T, S → S |
| Deep Voice | Autoregressive | T → S |
| Deep Voice 2 | Autoregressive | T → S |
| Deep Voice 3 | Autoregressive | T → S |
| DeepASMR | Token LM | T, S → S |
| DeepDubber-V1 | Diffusion / flow | T, V → S |
| DeepDubbing | Token LM | T, S → S |
| DelightfulTTS | Parallel | T → S |
| DELTA-TTS | Token LM | T, S → S |
| DepFlow | Diffusion / flow | T, S → S |
| Dia | Token LM | T, S → S |
| Dia2 | Token LM | T, S → S |
| DialoSpeech | Token LM | T, S → S |
| DiEmo-TTS | Parallel | T, S → S |
| Diff-TTS | Diffusion / flow | T → S |
| DiffCSS | Token LM | T, S → S |
| DiFlow-TTS | Token LM | T, S → S |
| DiFlowDubber | Token LM | T, S, V → S |
| DisCo-Speech | Token LM | T, S → S |
| DisSpeech | Token LM | T → S |
| DiSTAR | Token LM | T, S → S |
| DiTAR | Continuous LM | T, S → S |
| DiTTo-TTS | Diffusion / flow | T, S → S |
| DMOSpeech 2 | Diffusion / flow | T, S → S |
| DMP-TTS | Diffusion / flow | T, S → S |
| dots.tts | Continuous LM | T, S → S |
| Dragon-FM | Token LM | T, S → S |
| DrawSpeech | Diffusion / flow | T, I → S |
| DS-TTS | Diffusion / flow | T, S → S |
| DualDub | Token LM | T, V → S |
| DualSpeechLM | Token LM | T, S → S |
| E2 TTS | Diffusion / flow | T, S → S |
| ECTSpeech | Diffusion / flow | T, S → S |
| ELLA-V | Token LM | T, S → S |
| EME-TTS | Parallel | T → S |
| EMM-TTS | Token LM | T, S → S |
| EmojiVoice | Diffusion / flow | T → S |
| EmoShift | Token LM | T → S |
| EmoSSLSphere | Token LM | T → S |
| EmoSteer-TTS | Diffusion / flow | T, S → S |
| Emotion-timbre disentangled TTS | Parallel | T, S → S |
| EmotiVoice | Parallel | T → S |
| EmoTra-TTS | Token LM | T, S → S |
| EmoVoice | Token LM | T → S |
| End-to-end discrete-token TTS | Token LM | T, S → S |
| F5-TTS | Diffusion / flow | T, S → S |
| F5R-TTS | Diffusion / flow | T, S → S |
| Face-adapted StyleTTS 2 | Diffusion / flow | T, I → S |
| FaceSpeak | Diffusion / flow | T, I → S |
| FacialTalker | Token LM | T, S, V → S |
| FastPitch | Parallel | T → S |
| FastSpeech | Parallel | T → S |
| FastSpeech 2 | Parallel | T → S |
| FC-TTS | Token LM | T, S → S |
| FELLE | Continuous LM | T, S → S |
| FineCombo-TTS | Diffusion / flow | T, S → S |
| FireRedAudio | Continuous LM | T, S → S, A |
| FireRedTTS | Token LM | T, S → S |
| FireRedTTS-1S | Token LM | T, S → S |
| FireRedTTS-2 | Token LM | T, S → S |
| FireRedTTS3 | Continuous LM | T, S → S |
| Fish Audio S1 / OpenAudio S1 | Token LM | T, S → S |
| Fish Audio S2 | Token LM | T, S → S |
| Fish Speech | Token LM | T, S → S |
| Flamed-TTS | Diffusion / flow | T, S → S |
| FlashTTS | Token LM | T, S → S |
| FleSpeech | Token LM | T, S, I → S |
| FlexiVoice | Token LM | T, S → S |
| FlexSpeech | Diffusion / flow | T, S → S |
| Flowtron | Flow / VAE | T, S → S |
| FNH-TTS | Flow / VAE | T, S → S |
| Frame-stacked local Transformer TTS | Token LM | T, S → S |
| FreyaTTS | Diffusion / flow | T → S |
| Gemini 2.5 TTS | API | T → S |
| Gemini 3.1 Flash TTS | API | T → S |
| GibbsTTS | Token LM | T, S → S |
| GLM-TTS | Token LM | T, S → S |
| Glow-TTS | Flow / VAE | T → S |
| GOAT-TTS | Token LM | T, S → S |
| GPA | Token LM | T, S → S |
| GPT-4o Mini TTS | API | T → S |
| GPT-SoVITS | Token LM | T, S → S |
| Grad-TTS | Diffusion / flow | T → S |
| GRAFT | Token LM | T, S → S |
| GSA-TTS | Parallel | T, S → S |
| GST-Tacotron | Autoregressive | T, S → S |
| Habibi | Diffusion / flow | T, S → S |
| HD-PPT | Token LM | T, S → S |
| Higgs Audio v2 | Token LM | T, S → S |
| Higgs Audio v2.5 | Token LM | T, S → S |
| Higgs Audio v3 TTS | Token LM | T, S → S |
| HiStyle | Diffusion / flow | T → S |
| HoliDubber | Continuous LM | T, V → S |
| HoliTok (TTS) | Continuous LM | T, S → S |
| Hume Octave TTS | API | T, S → S |
| ImmersiveTTS | Diffusion / flow | T, S, A → S, A |
| IndexTTS | Token LM | T, S → S |
| IndexTTS 2.5 | Token LM | T, S → S |
| IndexTTS2 | Token LM | T, S → S |
| InstructAudio | Diffusion / flow | T → S |
| IntMeanFlow | Diffusion / flow | T, S → S |
| Inworld TTS-1 | Token LM | T → S |
| JaiTTS | Continuous LM | T, S → S |
| JAM-Flow | Diffusion / flow | T, S, V → S |
| JELLY | Token LM | T, S → S |
| JETS | Parallel | T → S |
| Joint non-autoregressive STT-TTS | Parallel | T → S |
| Joycent | Diffusion / flow | T, S → S |
| JoyVoice | Token LM | T, S → S |
| KABURI-TTS | Diffusion / flow | T → S |
| KittenTTS | Compact | T → S |
| Koel-TTS | Token LM | T, S → S |
| Kokoro | Compact | T → S |
| Kyutai TTS (DSM) | Token LM | T, S → S |
| LanStyleTTS | Parallel | T, S → S |
| LatinX | Token LM | T, S → S |
| LE2E-TTS | Compact | T → S |
| LightSpeech | Parallel | T → S |
| LLaDA-TTS | Token LM | T, S → S |
| Llasa | Token LM | T, S → S |
| Llasa+ | Token LM | T, S → S |
| LLMVoX | Token LM | T → S |
| Lombard Matcha-TTS | Diffusion / flow | T → S |
| LongCat-AudioDiT | Diffusion / flow | T, S → S |
| LoRP-TTS | Diffusion / flow | T, S → S |
| Luna-TTS | Token LM | T, S → S |
| M3-TTS | Diffusion / flow | T, S → S |
| MAGIC-TTS | Token LM | T, S → S |
| MagpieTTS-LF | Token LM | T, S → S |
| MambaVoiceCloning | Diffusion / flow | T, S → S |
| MamTra | Diffusion / flow | T, S → S |
| ManchuTTS | Diffusion / flow | T → S |
| Marco-Voice | Token LM | T, S → S |
| MARS6 | Token LM | T, S → S |
| Masked-style TTS | Token LM | T, S → S |
| MaskGCT | Token LM | T, S → S |
| Matcha-TTS | Diffusion / flow | T → S |
| MAVE | Token LM | T, S → S |
| Mega-TTS | Token LM | T, S → S |
| Mega-TTS 2 | Token LM | T, S → S |
| MegaTTS 3 | Diffusion / flow | T, S → S |
| Meitei Mayek TTS | Autoregressive | T → S |
| Mel-LLM (TTS) | Continuous LM | T → S |
| MELA-TTS | Continuous LM | T, S → S |
| MELD | Token LM | T, S → S |
| Mellotron | Autoregressive | T, S → S |
| MeloTTS | Flow / VAE | T → S |
| Metis | Token LM | T, S → S |
| MFCIG-CSS | Token LM | T, S, V → S |
| MiDashengLM-Gen | Continuous LM | T → S, A |
| MiniMax-Speech | Token LM | T, S → S |
| MixedG2P-T5 | Token LM | T, S → S |
| MM-MovieDubber | Diffusion / flow | T, V → S |
| MoE-TTS | Token LM | T → S |
| MoonCast | Token LM | T, S → S |
| MOSS-TTS | Token LM | T, S → S |
| MOSS-TTS-Nano | Token LM | T, S → S |
| MOSS-TTS-Realtime | Token LM | T, S → S |
| MOSS-TTSD | Token LM | T, S → S |
| MOSS-VoiceGenerator | Token LM | T → S |
| MP-ELD | Continuous LM | T, S → S |
| MPE-TTS | Token LM | T, S → S |
| Multistage multimodal TTS | Diffusion / flow | T, I → S |
| Muyan-TTS | Token LM | T, S → S |
| NaturalSpeech | Flow / VAE | T → S |
| NaturalSpeech 2 | Diffusion / flow | T, S → S |
| NaturalSpeech 3 | Diffusion / flow | T, S → S |
| NeuTTS Air | Token LM | T, S → S |
| NeuTTS Nano | Token LM | T, S → S |
| NeuTTS-2E | Token LM | T → S |
| NR-LauraTTS | Token LM | T, S → S |
| NVSpeech TTS | Token LM | T, S → S |
| Nüshu-PitchVITS | Flow / VAE | T → S |
| Ojibwe-Mi'kmaq-Maliseet TTS | Diffusion / flow | T → S |
| OmniVoice | Token LM | T, S → S |
| OpusLM | Token LM | T, S → S |
| Orpheus TTS | Token LM | T, S → S |
| OscillaTTS | Diffusion / flow | T, S → S |
| OuteTTS | Token LM | T, S → S |
| OV-InstructTTS | Token LM | T → S |
| OZSpeech | Token LM | T, S → S |
| PALLE | Token LM | T, S → S |
| Parallel GPT | Token LM | T, S → S |
| Parallel Tacotron | Parallel | T → S |
| Parallel Tacotron 2 | Parallel | T → S |
| ParaStyleTTS | Parallel | T → S |
| Parler-TTS | Token LM | T → S |
| Parler-TTS Hinglish adaptation | Token LM | T → S |
| PFluxTTS | Diffusion / flow | T, S → S |
| Phoenix TTS | Token LM | T, S → S |
| Phoneme-tone adaptive Thai TTS | Parallel | T, S → S |
| PilotTTS | Token LM | T, S → S |
| Piper (VITS voices) | Flow / VAE | T → S |
| Pocket TTS | Continuous LM | T, S → S |
| PortaSpeech | Flow / VAE | T → S |
| PROEMO | Parallel | T → S |
| Progressive face-conditioned TTS | Flow / VAE | T, I → S |
| Prompt-Unseen-Emotion | Token LM | T → S |
| PromptTTS | Parallel | T → S |
| PromptTTS 2 | Diffusion / flow | T → S |
| ProtoDisent-TTS | Flow / VAE | T, S → S |
| PS-TTS | Token LM | T, S → S |
| QTTS | Token LM | T, S → S |
| Qwen-Audio-3.0-TTS | Token LM | T, S → S |
| Qwen3-TTS | Token LM | T, S → S |
| RADKA-CSS | Token LM | T, S → S |
| RALL-E | Token LM | T, S → S |
| Raon-OpenTTS | Diffusion / flow | T, S → S |
| RapFlow-TTS | Diffusion / flow | T, S → S |
| ReGenVoice | Diffusion / flow | T, S → S |
| ReStyle-TTS | Diffusion / flow | T, S → S |
| RTFree-F5 | Diffusion / flow | T, S → S |
| RV-TTS | Diffusion / flow | T, I → S |
| RWKVTTS | Token LM | T, S → S |
| S5-TTS | Token LM | T, S → S |
| Sarashina2.2-TTS | Token LM | T, S → S |
| SASLM | Continuous LM | T, S → S |
| Seed-TTS | Token LM | T, S → S |
| Self-distilled zero-shot TTS | Compact | T, S → S |
| SelfTTS | Flow / VAE | T, S → S |
| SemaVoice | Continuous LM | T, S → S |
| SemBridge | Continuous LM | T, S → S |
| Shallow Flow Matching TTS | Diffusion / flow | T, S → S |
| SLED | Continuous LM | T, S → S |
| SlimSpeech | Compact | T, S → S |
| SMLLE | Token LM | T, S → S |
| SoulX-Podcast | Token LM | T, S → S |
| Spark-TTS | Token LM | T, S → S |
| SpeakStream | Continuous LM | T, S → S |
| SPEAR-TTS | Token LM | T, S → S |
| SpeechAccentLLM | Token LM | T, S → S |
| SpeechEdit | Token LM | T, S → S |
| SpeechT5 | Autoregressive | T, S → S |
| SpeechX | Token LM | T, S → S |
| SpeedySpeech | Parallel | T → S |
| Spotlight-TTS | Diffusion / flow | T, S → S |
| StellarTTS | Compact | T, S → S |
| Step-Audio-EditX | Token LM | T, S → S |
| Step-Audio-TTS | Token LM | T, S → S |
| StepAudio 2.5 TTS | Token LM | T, S → S |
| Stochastic-alignment continuous TTS | Continuous LM | T, S → S |
| StreamMel | Continuous LM | T, S → S |
| StyleTTS | Parallel | T, S → S |
| StyleTTS 2 | Diffusion / flow | T, S → S |
| Supertonic | Compact | T, S → S |
| Supertonic 2 | Compact | T → S |
| Supertonic 3 | Compact | T → S |
| SwanVoice | Diffusion / flow | T, S → S |
| SyncSpeech | Token LM | T, S → S |
| Tacotron | Autoregressive | T → S |
| Tacotron 2 | Autoregressive | T → S |
| TADA | Continuous LM | T, S → S |
| TED-TTS | Token LM | T, S → S |
| Tibetan-TTS | Token LM | T, S → S |
| TinyWave | Token LM | T, S → S |
| TLDR (TTS) | Token LM | T, S → S |
| TMD-TTS (formerly FMSD-TTS) | Diffusion / flow | T → S |
| TontaubeV1 | Token LM | T, S → S |
| Tortoise TTS | Token LM | T, S → S |
| Transformer TTS | Autoregressive | T → S |
| TTS-CtrlNet | Diffusion / flow | T, S → S |
| TTS-Transducer | Token LM | T, S → S |
| TTSYoruba | Concatenative | T → S |
| UDDETTS | Token LM | T → S |
| UmbraTTS | Diffusion / flow | T, S, A → S |
| UniFlow-Audio | Diffusion / flow | T, S, A, I, V → S, A |
| UNISON | Diffusion / flow | T, S, A → S, A |
| UniSonate | Diffusion / flow | T → S |
| UniSpeaker | Diffusion / flow | T, S, I → S |
| UniTAF | Token LM | T → S |
| UniTalker | Token LM | T, S, V → S |
| UniTTS | Token LM | T, S → S |
| UniVocal | Token LM | T, S → S |
| UniVoice (ASR and TTS) | Diffusion / flow | T, S → S |
| UniVoice (speech and singing) | Diffusion / flow | T, S → S |
| UniWav (TTS) | Diffusion / flow | T, S → S, A |
| USCF-conditioned TTS | Diffusion / flow | T, S → S |
| V-CASS | Token LM | T, V → S |
| VALL-E | Token LM | T, S → S |
| VALL-E 2 | Token LM | T, S → S |
| VALL-E X | Token LM | T, S → S |
| VALL-T | Token LM | T, S → S |
| Vclip | Flow / VAE | T, I → S |
| Vevo | Token LM | T, S → S |
| VibeVoice | Continuous LM | T, S → S |
| VibeVoice-Realtime | Continuous LM | T → S |
| VisualSpeech | Parallel | T, V → S |
| VITS | Flow / VAE | T → S |
| VITS2 | Flow / VAE | T → S |
| VividVoice | Diffusion / flow | T, I, V → S |
| VocalNet-M2 | Token LM | T, S → S |
| Voicebox | Diffusion / flow | T, S → S |
| VoiceChat-TTS | Continuous LM | T → S |
| VoiceCraft | Token LM | T, S → S |
| VoiceCraft-Dub | Token LM | T, S, V → S |
| VoiceDesigner | Diffusion / flow | T, S → S |
| VoiceSculptor | Token LM | T, S → S |
| VoxCPM | Continuous LM | T, S → S |
| VoxCPM2 | Continuous LM | T, S → S |
| Voxtral TTS | Token LM | T, S → S |
| VoXtream | Token LM | T, S → S |
| VoXtream2 | Token LM | T, S → S |
| VSpeechLM | Token LM | T, V → S |
| Wave-Tacotron | Autoregressive | T → S |
| WavTTS | Diffusion / flow | T, S → S |
| WenetSpeech-Wu TTS | Token LM | T, S → S |
| WeSCon | Token LM | T, S → S |
| WhisperSpeech | Token LM | T, S → S |
| WordVoice | Token LM | T, S → S |
| X-Voice | Diffusion / flow | T, S → S |
| X2Streaming-TTS | Token LM | T, S → S |
| XEmoRAG | Token LM | T, S → S |
| XTTS | Token LM | T, S → S |
| YourTTS | Flow / VAE | T, S → S |
| ZipVoice | Diffusion / flow | T, S → S |
| ZipVoice-Dialog | Diffusion / flow | T, S → S |
| Zonos | Token LM | T, S → S |
Figures are credited to their sources. Editorial input/output diagrams are labeled. Credits.
A2TTS extracts a voice embedding from a short recording and conditions a diffusion acoustic decoder on it. Reference-aware duration prediction improves timing consistency for multilingual synthesis in low-resource Indian languages.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Affectron extends a verbal-speech backbone to place nonverbal vocalizations in emotionally and contextually appropriate positions. Augmented training examples and structural masking enable expressive utterances containing events such as laughter while preserving the spoken content.
Paper · GitHub · Project · Details

Figure 2 · Source
AgentSteerTTS separates identity and emotional-prosodic representations, then grounds compound instructions in retrieved acoustic examples. A controller combines these conditions and uses feedback to refine expressive speech while preserving the intended speaker.
Paper · GitHub: no author-linked repository found · Details

Figure 4 · Source
AlignDiT aligns text, visual information and acoustic conditions before diffusion-based speech generation. Modality-specific guidance balances these inputs, targeting synchronized, intelligible speech that follows the timing and expression of the supplied scene.

Figure 1 · Source
AMNet adds phrase-structure information and local convolutional modeling to a parallel Mandarin acoustic model. These changes help capture contextual pauses, emphasis and intonation before the accompanying vocoder reconstructs the waveform.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
ARCHI-TTS uses a dedicated semantic alignment module to coordinate text and reference acoustic features. Reusing encoder features across denoising steps reduces repeated computation while the flow model generates the target speech.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
ATRIE converts character descriptions into separate voice-identity and dynamic prosody conditions. A compact adapter learns from a larger language-model teacher and modulates a GPT-SoVITS-based synthesizer, supporting expressive persona-driven speech generation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Audiobook-CC models context beyond individual sentences and separates style instructions from voice prompts. Distillation strengthens emotional expression, supporting multi-character narration with more consistent voices and performance across longer passages.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
AuEmoChat learns a discrete emotion representation from speech and compresses dialogue history while retaining emotionally relevant information. Its language model predicts emotion and speech tokens, which a context-conditioned flow decoder renders into expressive conversational audio.

Figure 2 · Source
AuK combines language-model conditioning, an audio VAE and successive multimodal and single-stream diffusion blocks. One model handles reference-based speech synthesis and instruction-guided editing; its distilled AuK-Flash variant reduces the number of generation steps.

Figure 4 · Source
Authentic-Dubber retrieves emotionally relevant audiovisual examples and progressively incorporates them into speech generation. Its director-actor formulation connects a scene's visual context and reference delivery to the target transcript for expressive movie dubbing.

Figure 2 · Source
AutoSIFT divides a reference voice's style into attribute-specific components and a residual representation. Text instructions replace selected attributes while unmentioned characteristics remain conditioned on the reference, allowing partial style editing during speech synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
AutoStyle-TTS matches the target text against a collection of expressive speech examples using learned textual embeddings. The selected recording provides style conditioning for synthesis, automatically adapting delivery to the content.
Paper · GitHub · Project · Details

Paper figure · Source
This audio-visual language model adds full-face information to an expressive speech backbone. Training on emotion and dialogue tasks connects facial cues with spoken delivery, enabling speech generation that uses visual as well as acoustic conversational context.

Figure 2 · Source
Bagpiper-TTS converts a natural-language request into a detailed speech plan containing words and delivery information. The generator follows that plan for tasks ranging from ordinary narration to multi-speaker rendering, role-play and singing.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
BareWave generates speech in waveform space using a single inference path. Representation alignment, staged noise scheduling and perceptual objectives guide training, replacing the separate acoustic-feature and waveform-reconstruction stages common in other TTS systems.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
Hierarchical autoregressive audio tokens.
Editorial input/output diagram · Source
Autoregressive speechcodes and convolutional decoder.

Figure 1 · Source
BatonVoice interprets a user's expressive request and translates it into controls for its dedicated BatonTTS generator. Separating instruction interpretation from acoustic rendering enables more explicit feature control, including transfer to languages outside the control-training data.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
BELLE predicts both speech values and their uncertainty in a continuous autoregressive synthesizer. Multiple synthetic renditions of the same text provide training support for the variance estimate, enabling richer acoustic distributions without adding an iterative inference stage.
Paper · GitHub · Project · Details

Figure 1 · Source
BitTTS reduces storage and computation through extremely low-bit trained weights and indexed parameter sharing. It targets speech generation on constrained devices, preserving a full synthesis path while shrinking the model representation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This streaming system replaces continuous acoustic regression with direct prediction of Mimi codec layers. A modified FastSpeech 2 backbone supplies aligned features and a depth-wise decoder fills residual codebooks, producing successive speech blocks without temporal autoregression.
Paper · GitHub: no author-linked repository found · Details

Figure 1 (paper page 10) · Source
BnTTS extends an XTTS-based multilingual pipeline to Bangla using language-specific phonetic adaptations. A small amount of target-speaker audio supports personalization, with the model designed for limited-resource speech synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Bolbosh adapts Matcha-TTS to Kashmiri with language-aware text processing and cross-language training. Its design targets the pronunciation and script challenges of a low-resource language while retaining efficient non-autoregressive acoustic generation.

Figure 1 · Source
This system organizes speech instructions at global, sentence and token levels to guide extended recordings. A continuous-token backbone uses explicit planning and condition dropout to combine voice design, multi-speaker rendering and changing acoustic or emotional context.
Paper · GitHub: no author-linked repository found · Details
Editorial input/output diagram · Source
BreezyVoice combines supervised speech tokens, a language model and flow-based acoustics with a pronunciation frontend. Its Taiwanese Mandarin adaptation provides explicit phonetic control for characters with multiple readings while retaining reference-based voice synthesis.

Figure 1 · Source
BridgeTTS uses the BridgeCode dual representation to shorten the sequence predicted by its autoregressive language model. Bridging modules reconstruct more detailed continuous acoustic features from those sparse tokens, balancing generation speed with voice fidelity.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
Beyond Video-to-SFX predicts audio semantic tokens from visual information and phonetic cues, then refines them into acoustic tokens. The two-stage generator produces intelligible speech whose timing and environmental sound fit the supplied video.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
CAM-TTS retains global narrative information and retrieves local details through an updatable memory block. Prefix attention combines those memories with preceding context, guiding sentence-level synthesis across longer paragraphs.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
CapTalk designs voices from descriptions for individual utterances and multi-speaker dialogue. Hierarchical conditioning separates stable speaker identity from changing turn-level delivery, while explicit planning tokens control dynamic expressive attributes.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
CAST-TTS maps a voice description or a reference recording into a common timbre-conditioning interface. Cross-attention delivers that information to the synthesizer, enabling voice design and reference-based cloning within the same generation framework.
Paper · GitHub · Project · Details

Figure 1 · Source
CaT-TTS separates textual understanding from acoustic generation in a two-Transformer architecture. During decoding, a masked parallel inference procedure guides speech-token predictions to reduce local errors in zero-shot voice synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
This FastSpeech 2 extension explicitly models emotion alongside duration, pitch and energy. Counterfactual training separates emotional changes from linguistic content, allowing users to modify prosody while preserving the intended words.
Paper · GitHub: no author-linked repository found · Details

Figure 1 (paper page 3) · Source
CDE-StyleTTS lets acoustic states evolve continuously along a phoneme sequence whose timing comes from durations. Sampling this trajectory supplies the acoustic decoder with timing-sensitive representations, providing a way to transfer changing expressive style instead of merely repeating phoneme embeddings.

Paper figure · Source
Chain-of-Details TTS progressively predicts speech at increasing temporal resolutions using a shared decoder. The coarsest stage provides an implicit phonetic plan, while subsequent stages recover timing detail without a separate phoneme-duration predictor.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Chain-Talker first derives an emotional description from dialogue history, then predicts semantic speech codes. A final rendering stage combines these plans to synthesize expressive responses whose delivery fits the conversational context.

Figure 2 · Source
Chatterbox synthesizes speech from text and a voice reference, with controls for expressive delivery. Its multilingual models focus on cross-language voice consistency, while Turbo and Nano use a smaller backbone and a single-step acoustic decoder. These variants serve different narration, conversational playback and local-device requirements; their capabilities are not interchangeable.
Editorial input/output diagram · Source
Chatterbox-Flash generates speech-token blocks in parallel while keeping block-by-block streaming. Calibration against common-token priors and confidence-based stopping improve its discrete diffusion decoding after adaptation from a pretrained autoregressive synthesizer.

Figure 3 · Source
Autoregressive speech-token generation.
Editorial input/output diagram · Source
Probabilistic residual quantization and multi-token LM.

Figure 1 · Source
CLEAR models speech directly in a continuous latent space, avoiding discrete codec-token prediction. Its zero-shot generator combines reference-voice conditioning with incremental audio production, targeting a balance between naturalness and response latency.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Clip-TTS trains its textual representation against corresponding mel-spectrogram information through a contrastive objective. The acoustic Transformer uses the resulting context-aware features to improve prosodic interpretation during speech generation.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
This compact synthesis system combines a shared-parameter text frontend with an efficient recurrent waveform generator. It targets responsive accessibility voices on low-power devices, where small storage requirements and immediate playback matter alongside naturalness.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This speech-language-model design keeps recent acoustic tokens and voice prompts at full detail while compressing distant context. The asymmetric representation reduces redundant long-sequence processing without discarding the local cues needed for pronunciation and vocal consistency.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Confucius4-TTS extracts voice characteristics from self-supervised speech features without requiring a transcript of the reference recording. A language model predicts semantic tokens and a flow decoder generates mel-spectrograms, enabling voice cloning within and across fourteen languages.

Figure 1 · Source
This model combines a language head that predicts boundaries with a diffusion head that generates continuous acoustic frames. Masked and staged training stabilize speaker-reference conditioning, providing a text-to-speech path within a multimodal language-model architecture.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
This synthesizer separates reference voice information from acoustic background conditions. An explicit task control selects whether background sound is retained or removed, allowing personalized speech generation under different environmental requirements.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
CookVoice aligns textual content, style and prosodic controls to acoustic frames before speech generation. The same compact model supports spoken and sung voices, reference imitation and editing, allowing individual voice attributes to be controlled within a shared synthesis pipeline.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
CosyEdit2 adapts a text-speech language model and acoustic decoder for consistent speech editing, then refines them with editing-specific rewards. The paper also evaluates the resulting improvement in zero-shot text-to-speech, connecting local editing consistency with reference-based synthesis.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
CoSyncDiT guides flow-based speech synthesis through acoustic-style adaptation, visual calibration and timed context alignment. These stages connect the supplied transcript and scene information to expressive, synchronized movie dubbing.

Figure 2 · Source
Supervised semantic tokens and flow decoder.

Figure 1 · Source
Text/speech LM and chunk-aware flow matching.

Figure 1 · Source
CosyVoice 3 extends streaming, reference-conditioned synthesis with a tokenizer trained on several speech-understanding tasks and a reward model for post-training. It targets reliable pronunciation, speaker identity and prosody across languages, dialects and less controlled text. The research scaling experiments and the downloadable Fun-CosyVoice3 checkpoint represent different model configurations.
Paper · GitHub · Project · Details

Figure 2 · Source
The WhispSynth generation pipeline combines a CosyVoice synthesizer with pitch-free digital signal processing to produce whispered speech. It supports multilingual whisper generation while avoiding the voiced pitch patterns of ordinary speech synthesis.

Figure 2 · Source
CoVoMix2 generates scripted dialogue directly with a flow-matching model, using reference voices without their transcripts. Speaker-disentangled text, sentence alignment and prompt masking support controlled timing and overlapping turns.

Figure 1 · Source
Cross-Lingual F5-TTS changes reference preparation and training so the generated text need not be paired with a reference transcript. Word-aligned acoustic prompts support cross-language voice cloning while reusing the flow-matching synthesis backbone.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
CrossAccent-TTS separates speaker identity from accent-related information in a speech synthesis model. Weighted language embeddings control the accent subspace, allowing gradual accent changes and cross-language synthesis while retaining the reference speaker's timbre.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
CSM uses the text and audio of preceding speaker turns to shape the delivery of the next utterance. A Llama backbone predicts speech representations, and a smaller decoder completes Mimi audio codes. It is a contextual speech renderer: an application supplies the words to say, including any responses written by a separate language model.
Editorial input/output diagram · Source
CTC-TTS uses automatically derived alignment and two-word interleaving to train an incremental speech language model. Its length-concatenated and feature-stacked variants make different trade-offs between speech quality and generation latency.

Figure 2 · Source
CtrlSpeech adds local pitch, loudness and duration conditioning to a patch-autoregressive diffusion synthesizer. A separate global speaker condition preserves the reference voice while users modify the delivery of individual words or phonemes.

Figure 2 · Source
CuteTTS combines a causal audio VAE, an autoregressive patch model and an explicitly speaker-conditioned flow head. Generating patches rather than individual frames reduces sequential work; a distilled variant targets faster streaming while preserving reference-voice synthesis.

Figure 1 · Source
DAIEN-TTS separates a reference recording into speech and environmental components, then conditions acoustic generation on them independently. Its extended formulation additionally models reverberation and uses separate guidance controls for speech, noise and room acoustics, enabling voice cloning into a chosen environment.
Paper 1 · Project 1 · Paper 2 · GitHub · Project 2 · Details

Paper figure · Source
DARS separately models pathological timing and acoustic style to synthesize dysarthric speech. A multistage rhythm predictor and conditioned flow model generate targeted examples for improving recognition under limited real-world speech data.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
DCAR changes the number of acoustic tokens predicted at each autoregressive step. Adapting the chunk size to the generation state shortens sequential processing while maintaining content alignment and reference-conditioned speech quality.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Neural TTS pipeline with autoregressive WaveNet synthesis.

Figure 1 · Source
Speaker-conditioned neural pipeline and WaveNet synthesis.

Figure 1 · Source
Convolutional attention encoder-decoder and converter.

Figure 1 · Source
DeepASMR separates ASMR delivery from the reference speaker's identity using discrete speech representations. A language model predicts content and style, while a flow-based acoustic decoder transfers the requested performance into the target voice.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
DeepDubber-V1 interprets visual scenes and dubbing requirements before generating the target speech. Its multimodal conditions distinguish narration, monologue and dialogue, guiding both expressive delivery and synchronization.

Figure 1 · Source
DeepDubbing assigns voices to characters and conditions speech rendering on the surrounding script. Its voice-design and instruction-synthesis stages support multi-participant audiobook production while maintaining character identity and context-sensitive expression.

Figure 1 · Source
Conformer acoustic model with explicit and implicit prosody.

Figure 1 · Source
DELTA-TTS converts a pretrained speech language model to parallel discrete diffusion using lightweight adaptation. Local convolution and confidence-based decoding help retain acoustic structure while the model fills speech-token positions in an order determined by prediction confidence.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
DepFlow separates a depression-related acoustic representation from speaker identity and spoken content. The representation conditions a flow-matching synthesizer, enabling controlled synthetic speech for investigating acoustic cues and data augmentation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Dia turns a speaker-tagged transcript into conversational audio, including supported nonverbal events such as laughter and coughing. Reference audio and its transcript can establish speaker identity and delivery. Its English checkpoint is useful for scripted exchanges and dialogue narration, with the conversation content supplied by the user rather than generated by the model.
Editorial input/output diagram · Source
Dia2 begins synthesizing before the complete script is available, allowing an application to feed words incrementally. Audio prefixes provide speaker and conversational context, while a streaming codec path reconstructs the output. The released 1B and 2B checkpoints focus on English dialogue; prefix conditioning helps keep voices consistent between generations.
Editorial input/output diagram · Source
DialoSpeech models two speakers on separate tracks and uses chunked flow matching for acoustic rendering. The design supports expressive scripted dialogue, including interactions whose timing is difficult to reproduce by joining independent utterances.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
DiEmo-TTS distills emotion information from speech while suppressing unrelated speaker characteristics. Cluster-based sampling and representation perturbation improve cross-speaker emotion transfer, including situations where extensive emotion labels are unavailable.

Figure 1 · Source
Text-conditioned denoising diffusion acoustic model.

Figure 3 · Source
DiffCSS samples prosody representations from multimodal conversational context using a diffusion model. A prosody-conditioned speech language model renders those samples, allowing several expressive deliveries that remain consistent with the same dialogue.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
DiFlow-TTS maps phonemes into linguistic content and generates separate prosody and acoustic token streams through discrete flow matching. Factorizing these responsibilities supports compact zero-shot synthesis with fewer sequential generation steps.

Figure 2 · Source
DiFlowDubber first learns linguistic content and separate prosodic-acoustic tokens through a discrete-flow TTS model. A subsequent video-dubbing stage aligns those representations to visual timing, connecting voice generation with synchronized lip movements.
Paper · GitHub · Project · Details

Figure 2 · Source
DisCo-Speech learns a codec that separates content, delivery and speaker identity. A language model predicts combined content-prosody tokens while the decoder receives a separate timbre representation, enabling reference cloning with independently controllable vocal attributes.
Paper · GitHub · Project · Details

Figure 1 · Source
DisSpeech maps Mandarin text and marked stuttering events to semantic speech tokens without temporal autoregression. Pitch and energy modeling guide acoustic reconstruction, enabling controlled repetitions and other disfluencies for speech synthesis and recognition-data augmentation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
DiSTAR first drafts blocks of residual-quantized speech tokens with a language model. A masked diffusion decoder then fills acoustic detail within each block, combining temporal planning and parallel refinement entirely in discrete codec space.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
DiTAR predicts a sequence of compressed acoustic patches using a language model, then generates each patch's detail through diffusion. Separating global temporal planning from local reconstruction supports zero-shot speech synthesis with controllable sampling diversity.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Latent diffusion Transformer with speech-length prediction.

Figure 1 · Source
DMOSpeech 2 adds reinforcement learning to duration prediction in an already metric-optimized synthesizer. Rewards derived from speaker similarity and transcription accuracy guide timing choices, linking prosodic planning with reference-voice and content objectives.
Paper · GitHub · Project · Details

Figure 1 · Source
DMP-TTS maps style descriptions and reference recordings into a shared conditioning space. Chained guidance controls content, timbre and style separately, allowing detailed synthesis adjustments within a latent diffusion Transformer.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
dots.tts predicts continuous acoustic representations from multilingual text and voice references. Its flow-matching output head supports both audio streaming and streaming text input; guidance-aware distillation reduces the work needed to generate each audio packet.
Paper · GitHub · Project · Details

Figure 1 · Source
Dragon-FM predicts successive speech chunks autoregressively while refining the tokens inside each chunk with bidirectional flow matching. Compact acoustic codes and cross-chunk caching reduce generation overhead and support longer content such as podcasts.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
DrawSpeech turns user-drawn prosodic curves into detailed pitch and energy conditions for synthesis. A diffusion model fills in the acoustic detail, enabling localized expressive control beyond an overall style description.

Paper figure · Source
DS-TTS extracts complementary voice characteristics through two style encoders. Dynamic modulation conditions the acoustic generator on these representations, supporting unseen speakers and adapting synthesis across different sentence lengths.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
DualDub generates spoken dialogue and background sound together from video and textual conditions. A cross-modal aligner coordinates the two decoding heads, targeting temporally synchronized soundtracks rather than speech rendered in isolation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
DualSpeechLM uses understanding-oriented speech tokens as input and acoustic codec tokens for generation. Semantic supervision and staged conditioning coordinate the two representations, supporting speech synthesis within a unified understanding-and-generation architecture.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
Flow matching with filler-token text conditioning.

Figure 1 · Source
ECTSpeech gradually tightens consistency constraints on a pretrained diffusion synthesizer. The resulting generator maps a noise state to speech in one step, reducing repeated denoising while retaining the conditioning of the original TTS model.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
Alignment-guided token reordering.

Figure 1 · Source
EME-TTS jointly models emotional delivery and local emphasis instead of treating them as independent effects. Automatically derived emphasis labels and variance-related features help users stress selected material while retaining a recognizable target emotion.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
EMM-TTS separates emotional content modeling from speaker-specific acoustic generation. Speaker-consistency objectives and adaptive normalization help transfer emotion across languages while retaining the reference speaker's timbre.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
EmojiVoice adds interpretable emoji prompts to the text encoder and flow predictor of Matcha-TTS. Changing prompts across phrases varies expression during longer robot utterances, providing a lightweight control interface for expressive synthesis.

Paper figure · Source
EmoShift adds a lightweight layer that learns emotion-dependent changes to a TTS model's hidden representation. The controls adjust expressive delivery and emotion intensity while leaving most of the underlying synthesis backbone intact.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
EmoSSLSphere combines an emotion representation constrained to a sphere with discrete units derived from self-supervised speech features. The model synthesizes emotional speech across languages while organizing expressive controls independently of the textual content.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
EmoSteer-TTS extracts emotion-related directions from a pretrained synthesizer's internal activations. Applying those directions during inference changes, mixes or removes emotional expression without retraining; the paper evaluates the method across several different TTS backbones.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
This emotional synthesizer learns separate reference encoders for timbre and emotion. A mutual-information objective reduces their overlap, while phoneme-level emotion prediction carries changing expression into the generated acoustic sequence.
Paper · GitHub · Project · Details

Figure 1 · Source
PromptTTS-derived style and content conditioning.
Editorial input/output diagram · Source
EmoTra-TTS introduces frame-level valence, arousal and dominance controls into both prosodic planning and acoustic decoding. Synthetic transition examples teach the model to move between emotions within an utterance while keeping its wording and speaker identity consistent.
Paper · GitHub · Project · Details

Figure 2 · Source
EmoVoice interprets free-form textual descriptions of emotional delivery. Its phoneme-boost variant predicts phonetic and acoustic information together to improve content consistency while retaining expressive style control.

Figure 1 · Source
This system trains the discrete speech representation together with the language and acoustic models that consume it. Reconstruction and recognition feedback also update token prediction, reducing mismatches between separately trained components in reference-conditioned speech synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
F5-TTS learns text-guided speech infilling with flow matching, refining character representations before a Transformer predicts the speech trajectory. A reference clip supplies voice context, and Sway Sampling controls how inference steps are distributed. The 2025 v1 Base release refines training and inference within the same general architecture for zero-shot speech synthesis.

Figure 1 · Source
F5R-TTS adapts a flow-based synthesizer to reinforcement learning through a probabilistic formulation of generation. Transcription accuracy and speaker-similarity rewards refine the model's ability to preserve both target content and the reference voice.

Figure 2 · Source
This model maps facial features into the style space of StyleTTS 2 through a lightweight learned adapter. It synthesizes text in a plausible face-conditioned voice without an audio reference; the paper evaluates transfer to unseen identities and another synthesis language.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FaceSpeak extracts speaker-related and expressive information from real or stylized portraits. These visual conditions guide text-to-speech generation, allowing the requested words to be rendered in a plausible voice and emotion associated with the image.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
FacialTalker quantizes facial action information and combines it with text and speech history in a conversational synthesizer. Joint preference training over visual and speech tokens helps the generated delivery reflect the interlocutor's facial expression and dialogue context.

Figure 2 · Source
Parallel Transformer with explicit pitch prediction.

Figure 1 · Source
Feed-forward Transformer and duration-based length regulator.

Figure 1, PDF p. 4 · Source
Parallel Transformer with duration, pitch and energy prediction.

Figure 1, PDF p. 3 · Source
FC-TTS conditions generation on two recordings, one supplying delivery style and the other speaker identity. Specialized representation processing and auxiliary training objectives aim to prevent either condition from leaking unwanted attributes into the synthesized voice.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FELLE generates continuous acoustic frames sequentially, using the preceding frame to shape the next flow-matching prior. A coarse-to-fine acoustic head refines each prediction, supporting reference-conditioned speech without discrete speech-token classification.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FineCombo-TTS interprets style descriptions relative to a supplied speech reference. A flow-based variance predictor models how acoustic attributes should change, enabling precise relative edits without requiring an explicit independent embedding for every voice attribute.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
FireRedAudio uses different acoustic encoders for understanding audio and conditioning speech generation. Its shared language model drives a flow-matching decoder over continuous RedAE latents, supporting voice cloning, instruction-controlled synthesis and speech editing within the broader audio model.

Figure 1 · Source
Text-to-semantic LM and speech decoder.

Figure 3 · Source
FireRedTTS-1S extends the FireRed synthesis line with incremental acoustic decoding. Its chunked flow-matching and frame-autoregressive multi-stream decoder options provide different trade-offs between initial latency and sustained generation speed.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FireRedTTS-2 models chronological sequences of speaker-labeled text and speech using a large Transformer plus a smaller codebook decoder. A low-rate streaming tokenizer reduces the number of audio steps. It targets long conversations and podcasts where speaker changes, turn-specific delivery and continuity across utterances matter.

Figure 1 · Source
FireRedTTS3 uses a semantically supervised audio autoencoder to make continuous speech representations easier to predict. The Base variant provides multilingual reference cloning; the Instruct variant adds natural-language voice design and editing of spoken content or acoustic attributes.

Figure 1 · Source
The S1 family combines multilingual voice conditioning with explicit markers for emotion, tone and nonverbal sounds. Its full model and distilled S1-mini offer different deployment sizes, with reinforcement learning used to refine generation. It is suited to expressive narration and character dialogue, although access and capabilities depend on the selected release.
Editorial input/output diagram · Source
Fish Audio S2 extends the Fish speech-model line with natural-language delivery instructions and multi-speaker, multi-turn synthesis. Its training pipeline uses speech descriptions, quality assessment and reward modeling to improve controllability. The released inference stack supports streaming, making the model relevant to both scripted audio production and incremental spoken responses.

Figure 2 · Source
Dual-autoregressive slow/fast transformers.

Figure 2 · Source
Flamed-TTS combines representations of different speech attributes in an attention-free generator. Its reformulated flow-matching process targets efficient zero-shot synthesis with flexible pacing and reduced sequential computation.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
FlashTTS processes incoming text and speech context on staggered tracks so synthesis can begin before sentence completion. Multi-token prediction accelerates the language model, while a distilled acoustic decoder reduces waveform-generation delay.
Paper · GitHub · Project · Details

Figure 1 · Source
FleSpeech unifies text, voice recordings and visual prompts into a common conditioning representation. Its multistage generator uses those controls to manipulate voice and delivery attributes flexibly rather than requiring one fixed prompt modality.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
FlexiVoice accepts an optional style instruction and an optional reference voice alongside the target text. Progressive preference training teaches the model to follow both conditions while reducing unwanted coupling between wording, speaker identity and delivery.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
FlexSpeech separates timing control from the acoustic synthesis component to balance stable pronunciation and natural expression. A small set of style examples can adapt the duration module without retraining the full generator, supporting efficient delivery customization.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Autoregressive normalizing flows over mel spectrograms.

Figure 1 · Source
FNH-TTS routes linguistic and speaker information through several duration experts to model varied timing patterns. It combines this predictor with changes to waveform generation, aiming for robust end-to-end speech synthesis across voices and prosodic conditions.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This architecture lets a global language model predict several speech frames at a time and delegates their codec entries to a smaller local Transformer. The paper compares sequential local decoding with iterative masked prediction, showing how frame stacking changes the trade-off between synthesis throughput and acoustic fidelity.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FreyaTTS is a Turkish-focused non-autoregressive synthesizer that maps character sequences to continuous audio latents. A frozen waveform autoencoder reconstructs the speech, while duration prediction and voice-focused post-training support efficient conversational playback without phoneme or discrete speech tokenization.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Gemini 2.5 TTS converts supplied text into speech with prompt-based control of accent, pace, style and emotion. The Flash and Pro interfaces support single-speaker narration and two-speaker scripts with separately assigned voices. These are dedicated speech-generation endpoints; their internal acoustic architecture is not fully disclosed in the public documentation.
Docs 1 · Docs 2 · Docs 3 · Details
Editorial input/output diagram · Source
Gemini 3.1 Flash TTS adds expressive audio tags to prompt-steered speech generation, giving authors more local control over narration and delivery. It targets natural, responsive multilingual synthesis through a managed API. The public preview documentation describes the interface and controls, without enough architectural detail to reconstruct the underlying speech generator.
Editorial input/output diagram · Source
GibbsTTS generates discrete speech tokens through a continuous-time jump process. Metric-aware transition scheduling and a finite-step correction improve how token states evolve, supporting zero-shot voice synthesis with discrete flow matching.
Paper · GitHub · Project · Details

Figure 1 · Source
GLM-TTS first predicts speech tokens autoregressively, then converts them into audio with a diffusion decoder. Pitch-aware tokenization and multi-reward reinforcement learning target pronunciation, speaker similarity and expression. Hybrid phoneme/text input and LoRA voice adaptation provide controls for applications that need repeatable pronunciation and customized voices.

Figure 1 · Source
Normalizing-flow acoustic model and monotonic alignment search.

Figure 1, PDF p. 3 · Source
GOAT-TTS encodes continuous voice information in one branch and predicts speech tokens in another. Partial language-model adaptation preserves textual knowledge, while multi-token prediction supports streaming synthesis with reference-based paralinguistic conditioning.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
General-Purpose Audio uses one autoregressive backbone to predict discrete speech tokens across synthesis, recognition and conversion tasks. Its TTS path combines target text with voice context, with shared multitask training and scalable inference.

Figure 1 · Source
GPT-4o Mini TTS combines the text to be spoken with instructions that steer accent, speed, tone and emotional delivery. The Speech API can stream audio before the full result is complete and supports several output formats. It provides a managed synthesis component for narration and voice applications; public documentation does not disclose the complete acoustic architecture.
Editorial input/output diagram · Source
GPT-SoVITS couples text-to-semantic token prediction with a reference-conditioned speech decoder. Its 2025 V3, V4 and V2 Pro releases extend a workflow that supports both zero-shot synthesis and voice adaptation from a small training set. The surrounding WebUI helps prepare data, while recognition and source-separation utilities remain separate components.
Editorial input/output diagram · Source
Score-based diffusion decoder and monotonic alignment search.

Figure 2 · Source
GRAFT attaches codec tokens from a spoken word example to that word's location in the text prompt. Separate target-speaker conditioning allows the pronunciation hint to come from another voice while the synthesized sentence retains the desired speaker.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
GSA-TTS extracts local style information at successive levels and combines it through attention into a global reference condition. This richer style representation guides the acoustic model when synthesizing an unseen speaker's voice.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Tacotron with a reference encoder and global style tokens.

Figure 1 · Source
Habibi trains an Arabic synthesizer progressively from standard language to regional dialects using curated public speech. It targets zero-shot voice cloning across dialects and reading without mandatory diacritic marks.
Paper · GitHub · Project · Details

Figure 1 · Source
HD-PPT learns speech codes that distinguish spoken content from instruction-related preferences. A language model predicts semantic information, expressive style and acoustic detail in sequence, improving the mapping from natural-language requests to controllable speech.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Higgs Audio v2 combines interleaved text/audio modeling with a unified speech tokenizer and DualFFN layers for acoustic prediction. Reference clips and scene context influence voices, while text context shapes prosody across narration and multi-speaker scripts. The generation checkpoint covers synthesis; the separate understanding branch in the family diagram is not another output mode of this checkpoint.

Official architecture diagram · Source
Higgs Audio v2.5, now documented as Higgs TTS 2.5, reduces the autoregressive audio Transformer to 1B parameters. GRPO-based alignment and a curated voice dataset refine pronunciation, cloning and expressive control tags. It targets production narration and conversational speech with lower computational requirements than the preceding 3B generation model.
Editorial input/output diagram · Source
Higgs TTS 3 uses an autoregressive decoder over interleaved text and eight speech codebooks, with a delay pattern and fused input/output projections. Reference audio establishes a voice, while inline tokens control emotion, style, pauses and sound effects. The 4B release targets multilingual conversational speech and expressive response rendering.

Official architecture diagram · Source
HiStyle predicts a voice's timbre first and finer delivery attributes afterward from textual descriptions. Contrastive text-audio alignment organizes these style representations before they condition a speech synthesizer.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
HoliDubber conditions audio generation on video and a text prompt describing speech and sound effects. A causal model plans successive latent patches and a local diffusion Transformer generates their detail, supporting synchronized dubbing within complex acoustic scenes.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
HoliTok combines linguistic and acoustic information in a continuous representation designed for both understanding and generation. Its downstream autoregressive model and diffusion decoder demonstrate a text-to-speech path using the same latents employed for recognition.

Figure 1 · Source
Octave uses text context and acting instructions to adjust pronunciation, emphasis, tempo and emotional delivery. Its API supports voice creation from descriptions, voice cloning and continuation across longer passages. Octave 1 and the Octave 2 preview have different feature coverage; the public interface is documented more fully than the internal speech-model architecture.
Editorial input/output diagram · Source
ImmersiveTTS jointly models spoken content and its surrounding acoustic scene in a multimodal diffusion Transformer. Speech and general-audio representations provide complementary training signals, helping generated wording remain intelligible within the requested environmental context.
Paper · GitHub · Project · Details

Figure 1 · Source
IndexTTS adapts the XTTS/Tortoise approach with a Conformer reference encoder and a BigVGAN2 speech decoder. Hybrid character/pinyin input gives explicit control over difficult Chinese pronunciations. Its central use case is zero-shot voice cloning with predictable text rendering, including content that benefits from pronunciation correction.
Paper · GitHub · Project · Details

Figure 1 · Source
IndexTTS 2.5 shortens semantic sequences with a lower-rate codec and replaces the acoustic module's backbone with a Zipformer design. Multilingual training strategies and reinforcement learning extend pronunciation and emotion transfer across languages. It retains reference-based voice conditioning while reducing the cost of semantic and acoustic generation.

Figure 1 · Source
IndexTTS2 separates speaker identity from emotional style so that different references can control timbre and delivery. Its autoregressive formulation also supports explicit output-token budgeting for duration control, alongside unconstrained generation. These mechanisms are intended for expressive speech and timing-sensitive work such as dubbing; availability of controls should be checked in the chosen implementation.

Figure 1 · Source
InstructAudio combines natural-language instructions with phonemes or lyrics in a common generation format. Joint and single-stream diffusion layers synthesize speech or music while controlling attributes such as voice, emotion, accent or musical character.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
IntMeanFlow distills a flow-based speech generator to predict integrated acoustic updates over larger intervals. A search for effective sampling steps further reduces decoding work, enabling reference-conditioned synthesis with fewer iterative refinements.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Inworld TTS-1 and TTS-1-Max are multilingual autoregressive synthesizers designed for low-latency speech output. Textual audio markup controls emotions and nonverbal vocalizations, with model variants offering different capacity and inference-cost trade-offs.

Figure 1 · Source
JaiTTS continually trains a VoxCPM-derived synthesizer on Thai-centered speech data. Its semantic planning, residual acoustic modeling and local diffusion decoding retain reference-based voice cloning while targeting fluent Thai pronunciation and delivery.

Figure 1 · Source
JAM-Flow couples audio and motion diffusion modules within a shared model. Text, voice references and optional motion conditions support synchronized speech and facial movement, with an infilling objective allowing several conditioning combinations.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
JELLY combines an emotion-aware Q-Former with several partially adapted language-model modules. Joint emotion recognition and contextual reasoning guide a speech synthesizer toward responses whose delivery matches the conversation.
Paper · GitHub · Project · Details

Paper figure · Source
Joint FastSpeech 2 and HiFi-GAN with learned alignment.

Figure 1, PDF p. 2 · Source
This model handles text and speech within one non-autoregressive architecture. Its TTS path predicts acoustic output from text, and feeding partial predictions back into the model improves generation through iterative refinement while also supporting recognition training.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Joycent separates accent information from speaker identity using an adversarially trained accent encoder. It injects accent and speaker features at different text-encoder layers, supporting accent-conditioned synthesis without requiring a separate accented phoneme prediction stage.
Paper · GitHub · Project · Details

Figure 1 · Source
JoyVoice conditions long-form speech on speaker-labeled text and shared conversational context. Autoregressive hidden states feed an acoustic diffusion decoder, allowing multiple speakers, changing expression and more flexible turn boundaries within one synthesis system.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
KABURI-TTS renders each participant on a separate audio channel from time-aligned phonemes and speaker activity. Supplying the timing layout explicitly lets the system synthesize overlapping speech, backchannels and interruptions for two-speaker conversations.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
KittenTTS provides small ONNX speech models with built-in voices and adjustable playback speed. The Mini, Micro and Nano releases offer different size and inference tradeoffs for CPU-oriented applications. The official README documents usage more fully than internal acoustic design, so the catalog presents it as a compact synthesis family without asserting an undisclosed architecture.
Editorial input/output diagram · Source
Koel-TTS explores several ways to condition a Transformer synthesizer on text and reference audio. Automatic speech-recognition and speaker-verification feedback, together with classifier-free guidance, improve adherence to the requested words and voice.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Kokoro's 2025 v1.0 release uses a compact StyleTTS 2-derived decoder with an iSTFTNet waveform generator. The released model relies on preset voice representations and omits style diffusion and a reference encoder. It is suited to lightweight narration and application speech where a small synthesis model and ready-made voices are useful.
Editorial input/output diagram · Source
Kyutai TTS treats text and speech as aligned streams separated by a controlled delay. A decoder-only language model can therefore emit audio as text arrives, instead of waiting for a complete utterance. This formulation supports incremental synthesis for voice interfaces and long streams, with speaker conditioning supplied through the TTS implementation.

Figure 1 · Source
LanStyleTTS standardizes phonetic inputs and introduces local style conditioning across languages. The framework can augment several parallel acoustic backbones, allowing one multilingual model to vary delivery at phoneme level.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
LatinX uses staged text-to-audio training, voice-cloning adaptation and automatic preference alignment. The resulting multilingual Transformer renders text in the source speaker's voice, supporting the synthesis stage of cross-language speech translation.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
LE2E-TTS trains a compact text-to-waveform pipeline end to end rather than separately optimizing acoustic and waveform stages. It targets local devices where model size, response time and compute cost constrain deployment.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FastSpeech-derived architecture found by neural architecture search.
Editorial input/output diagram · Source
LLaDA-TTS adapts a speech language model to fill masked token sequences in parallel. Its bidirectional generation also supports inserting, replacing or deleting spoken words, combining reference-based TTS and speech editing through the same model.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Llasa maps text and an optional speech prompt to a single stream of codec tokens using a Llama-style Transformer. The simple representation makes standard language-model scaling and sampling techniques applicable to synthesis. Its research also explores speech-model verifiers that select samples for content accuracy, voice consistency or emotional expression.
Paper · GitHub · Project · Details

Figure 2 · Source
Llasa+ adds multi-token prediction modules to a frozen Llasa backbone and checks their proposals with that backbone. A causal codec decoder turns accepted tokens into streaming audio. The resulting design addresses autoregressive latency while retaining the original speech model, making it relevant to systems that need incremental playback without retraining an entire backbone.

Figure 1 · Source
LLMVoX connects to an upstream language model through a streaming queue interface. Its small speech generator renders incoming text incrementally, allowing long conversations without tightly coupling the synthesizer to one particular language-model backbone.
Paper · GitHub · Project · Details

Figure 2 · Source
This Matcha-TTS extension learns vocal effort and articulation from automatically derived labels. It provides continuous controls over speech clarity and loudness-related effort, together with word-level emphasis, to synthesize the clearer delivery used in noisy listening conditions.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
LongCat-AudioDiT maps text and reference speech into continuous waveform latents using a diffusion model. A jointly considered waveform autoencoder and adapted inference guidance address the interaction between acoustic reconstruction and zero-shot synthesis quality.

Figure 1 · Source
LoRP-TTS adapts a pretrained zero-shot synthesizer using small low-rank parameter updates. It focuses on preserving a target speaker from limited, potentially noisy or spontaneous recordings whose acoustic conditions differ from the original training data.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
Luna-TTS adapts an autoregressive text backbone into a speech diffusion language model. Its parallel and Realtime variants share a tokenizer and training lineage; Realtime predicts successive codec blocks while denoising each block in parallel for incremental audio delivery.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 (paper page 4) · Source
M3-TTS uses joint text-audio diffusion layers to learn alignment without first stretching text into a guessed acoustic timeline. Additional single-stream layers refine acoustic details, producing reference-conditioned speech through a compressed mel representation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MAGIC-TTS exposes timing controls for selected speech tokens and pauses. Training includes incomplete control signals so the model can follow local edits where provided and infer natural timing elsewhere, supporting precise pacing without requiring every segment to be specified.

Figure 1 · Source
MagpieTTS-LF extends MagpieTTS at inference time by retaining acoustic and textual context across sentence boundaries. Soft alignment priors and history-aware encoding support coherent longer narration without retraining the synthesizer on long recordings.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MambaVoiceCloning uses state-space modules to encode phonemes, learn their timing and condition expressive synthesis. A training-only alignment teacher supplies timing supervision, while the generation path targets efficient long sequences and limited-lookahead speech streaming.

Figure 1 · Source
MamTra mixes state-space and attention layers to retain global context while reducing the cost of long sequences. Knowledge transfer from a pretrained Transformer initializes the hybrid synthesizer, combining efficient local processing with expressive acoustic modeling.
Paper · GitHub · Project · Details

Figure 1 · Source
ManchuTTS builds multilevel text representations suited to Manchu and feeds them into a convolutional diffusion Transformer. Its non-autoregressive generator and augmented training data target speech synthesis where naturally recorded material is scarce.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Marco-Voice learns separate speaker and emotion representations using contrastive training. Rotating the emotional representation provides smooth expressive control while preserving the reference voice across different delivery styles.

Figure 1 · Source
MARS6 encodes text and a speaker representation before generating hierarchical acoustic codes. Its compact encoder-decoder design targets expressive speech and reference-voice cloning with reduced inference cost.
Paper · GitHub · Project · Details

Paper figure · Source
This controllable system first predicts a masked-autoencoder-derived speech style representation from text and controls. A second model generates codec tokens, allowing speaker characteristics and expressive attributes to be specified separately during synthesis.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Masked generative codec transformers.

Figure 1 · Source
Conditional flow matching with an encoder-decoder acoustic model.

Figure 1 · Source
MAVE combines a state-space backbone with cross-attention to generate speech conditioned on text and acoustic context. It supports zero-shot voice synthesis and editing while reducing the attention-memory requirements of a comparable Transformer-based codec model.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Disentangled speech factors and prosody LM.

Figure 1 · Source
Prosody language model and multi-sentence prompting.

Figure 1 · Source
MegaTTS 3 guides a latent diffusion Transformer with sparse text-speech alignment boundaries, leaving the model room to learn finer timing. Classifier-free guidance controls accent strength, while piecewise rectified flow reduces sampling work. The design targets robust zero-shot voice synthesis with more flexible alignment than a fully fixed duration sequence.

Figure 1 · Source
This Manipuri speech synthesizer maps Meitei Mayek writing to an ARPAbet-based phoneme representation before acoustic generation with Tacotron 2. A HiFi-GAN vocoder reconstructs the waveform. The paper develops a single-speaker system for a language with limited training resources and tonal pronunciation requirements.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
The synthesis experiment in Mel-LLM extends a language model to predict mel-based acoustic information directly. Its next-token VAE decoder demonstrates a text-to-speech path within an otherwise understanding-focused model; the paper presents this as a proof of concept with quality limitations.
Paper · GitHub: no author-linked repository found · Details

Fig. 1 (paper page 2) · Source
MELA-TTS predicts continuous mel-spectrogram frames from text and speaker conditions. A training-time alignment module connects the decoder to recognition-derived semantic features, helping the joint Transformer-diffusion model retain linguistic structure without discrete speech tokenization.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MELD learns discrete latent variables from mel-spectrograms jointly with its speech language model. This shared optimization supports zero-shot synthesis and recognition while addressing omissions and excessive silence associated with less coordinated acoustic representations.
Paper · GitHub: no author-linked repository found · Details

Figures 1–2 (paper page 2; panels assembled) · Source
Tacotron 2 with global style tokens, pitch and rhythm conditioning.
Editorial input/output diagram · Source
VITS-family multilingual speech synthesis.
Editorial input/output diagram · Source
Metis pretrains on unlabeled speech before adapting to task-specific conditions such as text. Self-supervised semantic tokens and acoustic codes support a shared foundation for reference-based TTS, conversion and other speech-generation tasks.
Paper · GitHub · Project · Details

Paper figure · Source
MFCIG-CSS represents dialogue history through separate graphs of meaning and vocal expression. Fine-grained multimodal interactions condition the speech synthesizer, helping each scripted response fit the surrounding conversation.

Figure 1 · Source
MiDashengLM-Gen trains a language model together with a conditional flow-matching output head to generate variable-length audio. Its text conditioning supports scenes containing intelligible speech alongside music or other sounds, with generation performed over continuous audio representations.
Paper · GitHub · Project · Details

Paper figure · Source
MiniMax-Speech extracts speaker characteristics directly from reference audio without requiring its transcript, then generates speech with an autoregressive Transformer and Flow-VAE. The research emphasizes multilingual zero-shot cloning and expressive delivery. Additional adaptation mechanisms support emotion control, description-based voice creation and more specialized voice cloning without replacing the base model.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MixedG2P-T5 learns acoustic units from speech and uses a language-model synthesis path for text containing mixed scripts. It reduces dependence on manually designed grapheme-to-phoneme rules while retaining accent and intonation information in the speech representation.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
MM-MovieDubber interprets scene information to distinguish dialogue, narration and monologue delivery. A speech generator then uses the resulting multimodal conditions with the target content to render expressive movie dubbing.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
MoE-TTS augments a frozen text language model with speech-specific expert parameters. Retaining the original language knowledge helps the synthesizer interpret unfamiliar style descriptions while learning the acoustic generation task.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MoonCast combines podcast script preparation with a synthesizer trained for longer, spontaneous-sounding delivery. Voice references allow unseen speakers to render the resulting conversation, while discourse-level context supports more natural transitions than isolated sentence synthesis.
Paper · GitHub · Project · Details

Figure 1 · Source
MOSS-TTS offers two generators over a shared discrete audio representation: a delay-pattern model and a model with a frame-local Transformer. They balance long-context control against efficient codebook prediction and speaker preservation. The family supports reference-conditioned synthesis, pronunciation and duration controls, with later checkpoints extending language handling and explicit pauses.

Figure 2, PDF p. 8 · Source
MOSS-TTS-Nano packages multilingual voice cloning into a roughly 100M-parameter speech generator with a compact audio tokenizer. Streaming output and an ONNX inference path make it relevant to CPU-based readers and local applications. Its published performance depends on the runtime and hardware, and the tokenizer is a separate part of the deployment footprint.

Official architecture diagram · Source
MOSS-TTS-Realtime uses a Qwen3-derived backbone for linguistic context and a smaller local Transformer to predict audio codebooks. Text and speech tokens are handled at different levels so the system can accept text and emit audio incrementally. It targets low-latency spoken responses while preserving context across the generated utterance.
Editorial input/output diagram · Source
MOSS-TTSD turns a dialogue script with explicit speaker tags into a continuous multi-party recording. Long-context modeling helps maintain speaker identity, turn assignment and acoustic continuity, while short references can define voices. It is designed for podcasts, commentary and other scripted conversations; the source paper evaluates dialogue-specific consistency as well as intelligibility.

Figure 2, PDF p. 4 · Source
MOSS-VoiceGenerator creates a speaking voice from a natural-language description rather than requiring an example speaker recording. Training on expressive cinematic speech exposes it to varied delivery and acoustic conditions. The model is intended for character design, storytelling and role-based narration where the desired voice must be specified in words.

Figure 1 · Source
MP-ELD predicts low-rate continuous speech tokens through several information paths with separate local encoders. A flow decoder combines their predictions, while the accompanying Locodec representation is designed to limit accumulated errors during long speech generation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
MPE-TTS combines reference speech and textual prompts to specify an unseen speaker and the desired emotion. A prosody predictor and emotion-consistency objective carry those controls into the synthesized acoustic performance.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This framework learns face and text conditioning in separate stages before using them for voice synthesis. Visual knowledge distillation and training across text-face and text-speech pairs reduce reliance on fully matched multimodal recordings.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Muyan-TTS trains a speech language model on a large podcast collection for expressive reference-based synthesis. The release documents data preparation, training and optimized inference, with an emphasis on reproducible podcast-style voice generation.

Figure 1 · Source
Text-to-waveform VAE with enhanced prior and duration modeling.

Figure 1 · Source
Latent diffusion over neural-codec representations.

Figure 1 · Source
Factorized speech codec and attribute-wise diffusion.

Figure 3 · Source
NeuTTS Air pairs a phoneme-conditioned language model with NeuCodec to synthesize a reference voice locally. Quantized GGUF backbones support incremental generation through the documented streaming backend. It targets embedded and desktop voice applications, with the codec's compute and memory requirements considered alongside those of the language-model backbone.
Editorial input/output diagram · Source
NeuTTS Nano reduces the speech-model backbone while retaining phoneme conditioning, reference-based cloning and NeuCodec reconstruction. Its English, German, French and Spanish models share a design but use separate language-specific checkpoints. Quantized variants support compact local deployments, and streaming depends on the selected inference backend.
Editorial input/output diagram · Source
NeuTTS-2E accepts text directly and adds explicit emotional delivery to the NeuTTS language-model-and-codec pipeline. The released configuration supplies four fixed speaker presets instead of arbitrary reference-based cloning. It is intended for compact expressive speech applications, with streaming available through the documented GGUF inference path.
Editorial input/output diagram · Source
NR-LauraTTS cleans the discrete representation of a noisy voice prompt before passing it to LauraTTS. Token prediction and embedding refinement reduce background contamination, supporting reference cloning when the available recording is acoustically imperfect.
Paper · GitHub · Project · Details

Figure 1 · Source
The NVSpeech pipeline includes a TTS model that renders text with explicitly marked nonverbal events. Word-level annotations connect ordinary speech with vocalizations such as laughter, providing a shared representation for recognition and controllable audio generation.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
Nüshu-PitchVITS uses pitch annotations from Nüshu's writing system to guide acoustic generation under very limited data. A frame-level pitch predictor conditions the VITS waveform path, allowing syllable recordings and linguistic tone knowledge to support sentence synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
This model family shares speech-synthesis training across three related Indigenous languages. The paper compares attention-based and attention-free flow architectures, demonstrating how joint linguistic coverage can help languages with limited recordings.

Figure 1 · Source
OmniVoice predicts multiple acoustic codebooks directly from text using a masked, non-autoregressive diffusion language model. Random masking across codebooks and initialization from a pretrained language model support multilingual generation without a separate text-to-semantic stage. It focuses on broad-language zero-shot synthesis and voice conditioning, rather than visual or general-purpose omni interaction.

Figure 1 · Source
OpusLM extends text language models through speech-text pretraining on public data. Its interleaved representation supports speech recognition, text-conditioned synthesis and textual continuation within a transparent family of shared backbones.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Orpheus TTS repurposes a Llama-family language model to generate speech codec tokens from text. Emotion tags and speaker conditioning guide expressive delivery, while the project supplies inference and adaptation workflows. English releases and multilingual previews have different coverage, so a checkpoint's documented capabilities matter when selecting it for narration or a voice application.
Editorial input/output diagram · Source
OscillaTTS changes the periodic nonlinearities used in a style-diffusion synthesis backbone. Adjustable oscillatory modulation is designed to capture rapid pitch and amplitude changes while a linear bypass stabilizes the acoustic signal, targeting sharper expressive prosody.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
OuteTTS represents speech in a form that can be generated by a decoder-only language model and reconstructed by an audio decoder. A speaker reference guides vocal identity, style and accent. Its 1.0 line supports standard LLM serving backends, but generation settings and repetition handling need to follow the matching model implementation.
Editorial input/output diagram · Source
OV-InstructTTS interprets voice and delivery descriptions beyond a fixed inventory of style labels. Its reasoning-based conditioning connects broader textual requests with expressive speech generation, supported by a dedicated instruction-speech dataset.

Figure 2 · Source
OZSpeech generates disentangled speech components with a flow model conditioned on a learned prior. The design targets single-step zero-shot synthesis while separately modeling content, prosody and speaker-related information from the voice prompt.
Paper · GitHub · Project · Details

Figure 1 · Source
PALLE generates variable-length speech spans at fixed decoding steps, combining temporal planning with parallel token prediction. A second non-autoregressive stage refines the initial sequence, supporting efficient zero-shot synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
Parallel GPT divides speech generation between a general autoregressive predictor and a non-autoregressive detail model. The parallel refinement stage conditions on the initial tokens, balancing independence and interaction between semantic and acoustic information.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Parallel acoustic model with a variational residual encoder.

Figure 1 · Source
Parallel synthesis with differentiable duration modeling.

Figure 1 · Source
ParaStyleTTS converts textual style prompts into separate controls for prosody and broader paralinguistic characteristics. The lightweight adaptation design targets expressive speech from descriptions while making the roles of the two conditioning levels explicit.
Paper · GitHub · Project · Details

Figure 1 · Source
Description-conditioned codec language model.

Figure 1 · Source
This Parler-TTS extension introduces language-specific phonetic alignment and emotion embeddings for Hindi and Indian English. Its conditioning targets code-switched utterances whose accent and emotional delivery change coherently across language boundaries.

Figure 2 · Source
PFluxTTS combines two acoustic-generation paths by fusing their predicted vector fields at inference. Sequential reference embeddings support transcript-free cross-language voice cloning, and a super-resolution vocoder reconstructs high-rate output audio.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Phoenix TTS aligns its speech tokenizer with the downstream flow-matching acoustic decoder during training. An autoregressive language model predicts the resulting discrete representation, supporting text-to-speech and voice conversion without separating token design from acoustic reconstruction.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
This Thai speech synthesizer encodes phonemes and tones with a language-specific BERT model, then predicts duration, pitch and energy for a GAN-trained waveform decoder. A reference-derived style vector supports voice cloning. Multilingual pretraining of acoustic feature extractors and Thai adaptation address limited language-specific data.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
PilotTTS uses paired recordings and Q-Former conditioning to separate a speaker's identity from delivery style. The model supports reference cloning, emotional and nonverbal expression, and Chinese dialect synthesis within a shared autoregressive pipeline.

Figure 3 · Source
VITS voice models exported for local inference.
Editorial input/output diagram · Source
Pocket TTS uses continuous autoregressive speech modeling with a flow-based output mechanism, avoiding long sequences of discrete acoustic codebooks. It combines a small language-model backbone with streaming audio reconstruction and reusable voice conditioning. The project targets CPU-based speech synthesis; language-specific models and runtime choices affect its speed and voice behavior.

Figure 1 · Source
Variational acoustic model and flow-based post-net.

Figure 1, PDF p. 4 · Source
PROEMO combines emotional prompts with an explicit intensity control in a multi-speaker synthesizer. The conditioning adjusts delivery strength and prosodic variation, allowing the same spoken text to be rendered with different emotional performances.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This face-conditioned synthesizer combines local facial regions into progressively broader visual representations. Joint visual and acoustic attribute learning and multiple photographs of each training speaker align the face representation with voice characteristics, conditioning speech generation on text and a face image.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Prompt-Unseen-Emotion learns the relationship between emotion descriptions and speech using a language-model synthesis backbone. Weighted combinations of known emotions and contextual language knowledge allow expressive delivery outside the original categorical training labels.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Style and content text encoders with a speech decoder.

Figure 1 · Source
Prompt-conditioned TTS with a diffusion variation network.

Figure 1 · Source
ProtoDisent-TTS learns a codebook of healthy and dysarthric articulation patterns separately from speaker identity. Adversarial constraints reduce pathological information in the speaker representation, enabling controlled synthesis of articulation characteristics in a target voice.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
PS-TTS uses vowel-based alignment to coordinate the timing and phonetic structure of dubbed speech. Its PS-Comet variant also considers semantic preservation when choosing translated text, connecting translation choices with a TTS rendering stage.
Paper · GitHub: no author-linked repository found · Details

Fig. 1 (paper page 3) · Source
QTTS predicts residual speech codes produced by its QDAC tokenizer. Hierarchical parallel and delayed multihead variants organize codebook dependencies differently, offering alternative balances between acoustic detail and sequential decoding cost.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
Qwen-Audio-3.0-TTS combines compact semantic speech tokens with progressively trained language and acoustic models. Natural-language instructions and inline tags control delivery, while multilingual reference conditioning supports voice cloning and longer speech generation under varied recording conditions.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 3 · Source
Qwen3-TTS combines a dual-track speech language model with tokenizers designed for compact streaming audio. The released 12Hz line separates Base voice cloning, CustomVoice preset-speaker control and VoiceDesign creation from descriptions. These variants support different conditioning interfaces, allowing applications to choose between reproducing a reference voice and directing a new voice through text.

Figure 3 · Source
RADKA-CSS retrieves dialogue examples related to the current conversation in both meaning and delivery. A graph-based aggregation mechanism combines their style information with current context, conditioning expressive conversational speech synthesis.
Paper · GitHub · Project · Details

Figure 2 · Source
Prosody-guided codec language modeling.

Figure 1 · Source
Raon-OpenTTS is a family of reference-conditioned diffusion synthesizers trained on a large, documented English speech collection. The release pairs its models with data processing and evaluation resources, allowing robustness across varied acoustic conditions to be examined alongside clean-speech quality.

Figure 1 · Source
RapFlow-TTS regularizes the acoustic velocity field so longer generation steps remain consistent. Time-interval scheduling and adversarial objectives improve the resulting few-step synthesizer, reducing the iterations needed to render reference-conditioned speech.

Figure 1 · Source
ReGenVoice applies the ReGen representation-and-waveform modeling approach to text-to-speech. Multiple levels of generated conditioning help reconstruct detailed waveforms from compressed latents, linking efficient acoustic representation with reference-conditioned speech synthesis.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
ReStyle-TTS changes vocal attributes relative to a reference recording. Independent text and reference guidance, composable style adapters and timbre-consistency optimization allow continuous expressive edits while limiting changes to the speaker's identity.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
RTFree-F5 replaces the transcript normally associated with an F5-TTS reference recording with projected speech features. A lightweight adapter reuses the pretrained generator, enabling transcript-free voice conditioning, including references whose pronunciation makes transcription unreliable.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Revival with Voice learns voice identity from face images and delivery attributes from descriptions. Audio-only training data and stylized portrait augmentation broaden its input coverage, enabling controlled speech from real faces or artistic portraits.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
RWKVTTS uses the recurrent RWKV-7 architecture for speech synthesis in place of a conventional Transformer backbone. Its token-generation path targets efficient streaming and reduced state-management cost while conditioning audio on the supplied text.

Figure 2 · Source
S5-TTS adapts T5-TTS for word-by-word synthesis using limited future text. Lookahead-aware masks, convolutional auxiliary attention and distillation let the model start speaking before the complete sentence is available while retaining voice conditioning and alignment.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Sarashina2.2-TTS emphasizes reliable Japanese pronunciation, including characters with several possible readings. Its semantic language model and acoustic flow decoder use reference speech for voice conditioning, with balanced multilingual training to reduce dependence on the reference language.

Figure 1 · Source
SASLM derives expressive intent from its own evolving semantic states through an information bottleneck. Acoustic feedback aligns generated speech with that intent, reducing the need for externally supplied emotion labels in context-sensitive speech rendering.
Paper · GitHub · Project · Details

Figure 3 · Source
Autoregressive speech foundation model.

Figure 1 · Source
This zero-shot synthesizer learns linguistic content and reference-speaker attributes through separate representations. Two-stage self-distillation creates aligned examples that strengthen their separation, targeting stable voice cloning with a small inference footprint.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
SelfTTS learns separate representations of a speaker's identity and emotional delivery using contrastive and adversarial objectives. It then improves synthesis through self-generated training examples, allowing emotion transfer to speakers originally recorded with neutral expression.
Paper · GitHub · Project · Details

Figure 1 · Source
SemaVoice organizes its audio VAE latents using guidance from speech foundation-model representations. A continuous autoregressive backbone and patch-level diffusion head then synthesize reference-conditioned speech with greater emphasis on linguistic coherence.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
SemBridge uses discrete semantic targets during training to organize both acoustic latents and language-model hidden states. The resulting continuous generator supports zero-shot speech synthesis with stronger content alignment, without needing to generate the auxiliary semantic tokens during inference.

Figure 1 · Source
Shallow Flow Matching adds a lightweight head that predicts an intermediate acoustic state for a flow-based synthesizer. Starting refinement closer to the target reduces the remaining generation path, allowing coarse-to-fine speech synthesis with less iterative work.

Figure 2 · Source
SLED learns the conditional distribution of acoustic latents using an energy-distance objective rather than discrete token classification. Its autoregressive generator samples continuous speech representations, simplifying synthesis while retaining acoustic detail.

Figure 2 · Source
SlimSpeech reduces the parameter count of a rectified-flow TTS model and transfers knowledge into a lightweight generator. It targets efficient reference-conditioned speech synthesis while retaining the acoustic quality of a larger teacher.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
SMLLE uses a transducer to align incoming text with semantic speech tokens and duration information. A separate autoregressive stage generates acoustic frames, with controlled access to future text stabilizing incremental synthesis.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
SoulX-Podcast synthesizes conversational scripts with reference voices, dialect choices and nonverbal expression. Its long-form training targets consistent speaker identity and natural transitions across turns, while also supporting ordinary single-speaker TTS.
Paper · GitHub · Project · Details

Figure 3 · Source
Spark-TTS uses BiCodec to separate changing linguistic content from global speaker attributes, then predicts these tokens with a Qwen2.5 backbone. That separation supports both reference-based cloning and direct control of attributes such as speaking rate and pitch. It is useful for controllable speech generation where a reference recording alone is insufficient to specify the desired delivery.

Figure 3 · Source
SpeakStream trains on text interleaved with corresponding speech and generates audio as new text becomes available. The synthesis module remains compatible with an upstream text-streaming language model, supporting responsive conversational playback.
Paper · Project · GitHub: no author-linked repository found · Details

Paper figure · Source
Text-to-semantic and semantic-to-acoustic LMs.

Figure 1 · Source
SpeechAccentLLM uses a content tokenizer trained with transcription alignment and jointly learns accent conversion and synthesis. A reconstruction refinement stage improves generated speech while separating accent-related changes from the target speaker's identity.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
SpeechEdit combines text, instruction tokens and reference audio in a shared codec-language-model sequence. Paired examples that differ in selected attributes teach localized changes while retaining other aspects of the reference voice and delivery.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Shared encoder–decoder with modality interfaces.

Figure 2 · Source
Prompted neural codec language model.

Figure 1 · Source
Residual convolutional acoustic model with duration expansion.

Figure 3 · Source
Spotlight-TTS extracts expressive reference information primarily from voiced speech and adjusts the resulting style direction before acoustic generation. The method targets more faithful expression transfer; the catalog combines its related paper records under one model.
Paper 1 · Paper 2 · GitHub: no author-linked repository found · Details

Figure 1 · Source
StellarTTS encodes phoneme timing sparsely and uses a lightweight masked Transformer to generate speech tokens in parallel. A semantic-aware codec supports waveform reconstruction, while explicit timing representations provide control over pronunciation, duration and prosody.
Paper · Project · GitHub: no author-linked repository found · Details

Paper figure · Source
Step-Audio-EditX supports reference-based synthesis and repeated edits to emotion, speaking style or nonverbal delivery. Training on deliberately contrasting synthetic examples teaches the model to follow expressive changes without a separate attribute-embedding module.

Figure 2 · Source
Step-Audio-TTS is the compact synthesis component produced through the broader Step-Audio speech-data and distillation pipeline. Text and voice conditioning drive speech-token generation, while the released workflow exposes delivery controls. The TTS checkpoint is used to render supplied content; reasoning, tool use and dialogue management belong to other components of the system.

Figure 2 · Source
The TTS mode of StepAudio 2.5 uses a shared speech-language foundation with synthesis-specific decoding and preference training. Rich contextual supervision and feedback target controllable expression; this card describes its speech-rendering path within the broader system.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This synthesizer predicts continuous speech latents using a Gaussian-mixture conditional distribution. A stochastic monotonic alignment mechanism keeps the acoustic sequence ordered against the text, offering an alternative to autoregressive discrete-codec modeling.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
StreamMel alternates text tokens with continuous acoustic frames in one streaming synthesis model. This organization lets incoming text guide speech immediately while retaining reference-voice information across the generated audio stream.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Style-conditioned parallel synthesis with a transferable aligner.

Figure 1 · Source
Style diffusion and adversarial training with speech-model discriminators.

Figure 1, PDF p. 4 · Source
The SupertonicTTS research system compresses speech into continuous latents and predicts them from character-level text with flow matching. ConvNeXt blocks, temporal compression and a separate duration predictor keep synthesis compact. The later Supertonic ONNX release exposes preset voice-style assets; its packaged configurations should not be equated with the paper's 44M-parameter research model.
Model card · Paper · GitHub · Project · Details

Figure 1 · Source
Supertonic 2 extends the local ONNX synthesis line to five languages while retaining voice-style conditioning and a compact model. It provides a practical path to multilingual narration on devices that can run the supplied inference stack. Creating a new voice-style asset is a separate workflow from generating speech with an existing asset.
Editorial input/output diagram · Source
Supertonic 3 expands language coverage and adds expression tags while keeping local ONNX inference and preset voice styles. The release targets more reliable reading across short and long text, with controls for events such as breaths or laughter. Custom voice-style creation is offered through a separate service; downloaded styles can then condition local synthesis.
Editorial input/output diagram · Source
SwanVoice generates monologues or dialogues with up to four speakers using raw text, voice references and speaker-turn conditions. Pause markers and optional pronunciation substitutions provide textual control, while staged dialogue training supports longer expressive speech.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
SyncSpeech uses a temporal masking scheme to coordinate sequential speech structure with parallel token decoding. This hybrid organization targets faster first audio and higher throughput while retaining reference-conditioned synthesis quality.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Attention-based recurrent spectrogram synthesis.

Figure 1 · Source
Recurrent attention model and WaveNet vocoder.

Figure 1 · Source
TADA aligns text tokens one-to-one with continuous acoustic units. A language model with a flow-matching head predicts these synchronized representations, reducing the ambiguity of text-speech alignment during reference-conditioned synthesis.

Figure 2 · Source
TED-TTS modifies conditioning and decoding in a pretrained zero-shot synthesizer to control different parts of an utterance. Segment-specific emotion masks and duration steering support local changes while coordinating transitions and the overall stopping point.

Figure 1 · Source
Tibetan-TTS adapts a large speech generator through language-specific text representation, tokenizer changes and cross-language training. Its data preparation and quality enhancement target scarce Tibetan recordings and the differences between written forms and spoken pronunciation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
<a
Truncated — view the full README on GitHub.
6 commits
Python
100.0%
A visual catalog of text-to-speech architectures, with model diagrams, concise descriptions, and primary sources.
Python
18
6 commits
updated Sep 13, 2026
A visual catalog of text-to-speech models, from Tacotron to speech language models. Diagrams, primary sources and short notes for every entry.
380 models and families · Reviewed 2026-09-13
Model list · All diagrams · 2025+ TTS-arxiv-daily collection · Descriptions · Timeline · Methodology · Contribute
The TTS-arxiv-daily collection covers the source list's TTS systems with first paper submissions from January 1, 2025 onward. Every included family has an image, a description, paper links and an explicit GitHub availability status. The complete screening record is available as JSON.
T: text · S: speech or voice reference · A: other audio · I: image · V: video. Scope and labels.
| Model | Group | Input → output |
|---|---|---|
| A2TTS | Diffusion / flow | T, S → S |
| Affectron | Token LM | T, S → S |
| AgentSteerTTS | Token LM | T, S → S |
| AlignDiT | Diffusion / flow | T, S, V → S |
| AMNet | Parallel | T → S |
| ARCHI-TTS | Diffusion / flow | T, S → S |
| ATRIE | Token LM | T → S |
| Audiobook-CC | Token LM | T, S → S |
| AuEmoChat | Token LM | T, S, V → S |
| AuK | Diffusion / flow | T, S → S, A |
| Authentic-Dubber | Diffusion / flow | T, S, V → S |
| AutoSIFT | Diffusion / flow | T, S → S |
| AutoStyle-TTS | Token LM | T, S → S |
| AVLM (expressive speech) | Token LM | T, S, V → S |
| Bagpiper-TTS | Token LM | T → S |
| BareWave | Diffusion / flow | T, S → S |
| Bark | Token LM | T → S, A |
| BASE TTS | Token LM | T, S → S |
| BatonTTS (BatonVoice) | Token LM | T, S → S |
| BELLE | Continuous LM | T, S → S |
| BitTTS | Compact | T → S |
| Block-wise Mimi TTS | Token LM | T → S |
| BnTTS | Token LM | T, S → S |
| Bolbosh | Diffusion / flow | T → S |
| Borderless Long Speech Synthesis | Continuous LM | T, S → S |
| BreezyVoice | Token LM | T, S → S |
| BridgeTTS | Token LM | T, S → S |
| BVS | Token LM | T, V → S, A |
| CAM-TTS | Token LM | T, S → S |
| CapTalk | Token LM | T, S → S |
| CAST-TTS | Diffusion / flow | T, S → S |
| CaT-TTS | Token LM | T, S → S |
| Causal-prosody FastSpeech 2 | Parallel | T → S |
| CDE-StyleTTS | Diffusion / flow | T, S → S |
| Chain-of-Details TTS | Token LM | T, S → S |
| Chain-Talker | Token LM | T, S → S |
| Chatterbox | Token LM | T, S → S |
| Chatterbox-Flash | Token LM | T, S → S |
| ChatTTS | Token LM | T → S |
| CLaM-TTS | Token LM | T, S → S |
| CLEAR | Continuous LM | T, S → S |
| Clip-TTS | Parallel | T → S |
| Compact neural accessibility TTS | Compact | T → S |
| Compressed-to-fine speech LM | Token LM | T, S → S |
| Confucius4-TTS | Token LM | T, S → S |
| Continuous-token diffusion TTS | Continuous LM | T, S → S |
| Controllable masked-speech TTS | Token LM | T, S, A → S |
| CookVoice | Diffusion / flow | T, S → S |
| CosyEdit2 | Token LM | T, S → S |
| CoSyncDiT | Diffusion / flow | T, S, V → S |
| CosyVoice | Token LM | T, S → S |
| CosyVoice 2 | Token LM | T, S → S |
| CosyVoice 3 | Token LM | T, S → S |
| CosyWhisper (WhispSynth) | Token LM | T, S → S |
| CoVoMix2 | Diffusion / flow | T, S → S |
| Cross-Lingual F5-TTS | Diffusion / flow | T, S → S |
| CrossAccent-TTS | Token LM | T, S → S |
| CSM | Token LM | T, S → S |
| CTC-TTS | Token LM | T, S → S |
| CtrlSpeech | Continuous LM | T, S → S |
| CuteTTS | Continuous LM | T, S → S |
| DAIEN-TTS | Diffusion / flow | T, S, A → S |
| DARS | Diffusion / flow | T → S |
| DCAR | Token LM | T, S → S |
| Deep Voice | Autoregressive | T → S |
| Deep Voice 2 | Autoregressive | T → S |
| Deep Voice 3 | Autoregressive | T → S |
| DeepASMR | Token LM | T, S → S |
| DeepDubber-V1 | Diffusion / flow | T, V → S |
| DeepDubbing | Token LM | T, S → S |
| DelightfulTTS | Parallel | T → S |
| DELTA-TTS | Token LM | T, S → S |
| DepFlow | Diffusion / flow | T, S → S |
| Dia | Token LM | T, S → S |
| Dia2 | Token LM | T, S → S |
| DialoSpeech | Token LM | T, S → S |
| DiEmo-TTS | Parallel | T, S → S |
| Diff-TTS | Diffusion / flow | T → S |
| DiffCSS | Token LM | T, S → S |
| DiFlow-TTS | Token LM | T, S → S |
| DiFlowDubber | Token LM | T, S, V → S |
| DisCo-Speech | Token LM | T, S → S |
| DisSpeech | Token LM | T → S |
| DiSTAR | Token LM | T, S → S |
| DiTAR | Continuous LM | T, S → S |
| DiTTo-TTS | Diffusion / flow | T, S → S |
| DMOSpeech 2 | Diffusion / flow | T, S → S |
| DMP-TTS | Diffusion / flow | T, S → S |
| dots.tts | Continuous LM | T, S → S |
| Dragon-FM | Token LM | T, S → S |
| DrawSpeech | Diffusion / flow | T, I → S |
| DS-TTS | Diffusion / flow | T, S → S |
| DualDub | Token LM | T, V → S |
| DualSpeechLM | Token LM | T, S → S |
| E2 TTS | Diffusion / flow | T, S → S |
| ECTSpeech | Diffusion / flow | T, S → S |
| ELLA-V | Token LM | T, S → S |
| EME-TTS | Parallel | T → S |
| EMM-TTS | Token LM | T, S → S |
| EmojiVoice | Diffusion / flow | T → S |
| EmoShift | Token LM | T → S |
| EmoSSLSphere | Token LM | T → S |
| EmoSteer-TTS | Diffusion / flow | T, S → S |
| Emotion-timbre disentangled TTS | Parallel | T, S → S |
| EmotiVoice | Parallel | T → S |
| EmoTra-TTS | Token LM | T, S → S |
| EmoVoice | Token LM | T → S |
| End-to-end discrete-token TTS | Token LM | T, S → S |
| F5-TTS | Diffusion / flow | T, S → S |
| F5R-TTS | Diffusion / flow | T, S → S |
| Face-adapted StyleTTS 2 | Diffusion / flow | T, I → S |
| FaceSpeak | Diffusion / flow | T, I → S |
| FacialTalker | Token LM | T, S, V → S |
| FastPitch | Parallel | T → S |
| FastSpeech | Parallel | T → S |
| FastSpeech 2 | Parallel | T → S |
| FC-TTS | Token LM | T, S → S |
| FELLE | Continuous LM | T, S → S |
| FineCombo-TTS | Diffusion / flow | T, S → S |
| FireRedAudio | Continuous LM | T, S → S, A |
| FireRedTTS | Token LM | T, S → S |
| FireRedTTS-1S | Token LM | T, S → S |
| FireRedTTS-2 | Token LM | T, S → S |
| FireRedTTS3 | Continuous LM | T, S → S |
| Fish Audio S1 / OpenAudio S1 | Token LM | T, S → S |
| Fish Audio S2 | Token LM | T, S → S |
| Fish Speech | Token LM | T, S → S |
| Flamed-TTS | Diffusion / flow | T, S → S |
| FlashTTS | Token LM | T, S → S |
| FleSpeech | Token LM | T, S, I → S |
| FlexiVoice | Token LM | T, S → S |
| FlexSpeech | Diffusion / flow | T, S → S |
| Flowtron | Flow / VAE | T, S → S |
| FNH-TTS | Flow / VAE | T, S → S |
| Frame-stacked local Transformer TTS | Token LM | T, S → S |
| FreyaTTS | Diffusion / flow | T → S |
| Gemini 2.5 TTS | API | T → S |
| Gemini 3.1 Flash TTS | API | T → S |
| GibbsTTS | Token LM | T, S → S |
| GLM-TTS | Token LM | T, S → S |
| Glow-TTS | Flow / VAE | T → S |
| GOAT-TTS | Token LM | T, S → S |
| GPA | Token LM | T, S → S |
| GPT-4o Mini TTS | API | T → S |
| GPT-SoVITS | Token LM | T, S → S |
| Grad-TTS | Diffusion / flow | T → S |
| GRAFT | Token LM | T, S → S |
| GSA-TTS | Parallel | T, S → S |
| GST-Tacotron | Autoregressive | T, S → S |
| Habibi | Diffusion / flow | T, S → S |
| HD-PPT | Token LM | T, S → S |
| Higgs Audio v2 | Token LM | T, S → S |
| Higgs Audio v2.5 | Token LM | T, S → S |
| Higgs Audio v3 TTS | Token LM | T, S → S |
| HiStyle | Diffusion / flow | T → S |
| HoliDubber | Continuous LM | T, V → S |
| HoliTok (TTS) | Continuous LM | T, S → S |
| Hume Octave TTS | API | T, S → S |
| ImmersiveTTS | Diffusion / flow | T, S, A → S, A |
| IndexTTS | Token LM | T, S → S |
| IndexTTS 2.5 | Token LM | T, S → S |
| IndexTTS2 | Token LM | T, S → S |
| InstructAudio | Diffusion / flow | T → S |
| IntMeanFlow | Diffusion / flow | T, S → S |
| Inworld TTS-1 | Token LM | T → S |
| JaiTTS | Continuous LM | T, S → S |
| JAM-Flow | Diffusion / flow | T, S, V → S |
| JELLY | Token LM | T, S → S |
| JETS | Parallel | T → S |
| Joint non-autoregressive STT-TTS | Parallel | T → S |
| Joycent | Diffusion / flow | T, S → S |
| JoyVoice | Token LM | T, S → S |
| KABURI-TTS | Diffusion / flow | T → S |
| KittenTTS | Compact | T → S |
| Koel-TTS | Token LM | T, S → S |
| Kokoro | Compact | T → S |
| Kyutai TTS (DSM) | Token LM | T, S → S |
| LanStyleTTS | Parallel | T, S → S |
| LatinX | Token LM | T, S → S |
| LE2E-TTS | Compact | T → S |
| LightSpeech | Parallel | T → S |
| LLaDA-TTS | Token LM | T, S → S |
| Llasa | Token LM | T, S → S |
| Llasa+ | Token LM | T, S → S |
| LLMVoX | Token LM | T → S |
| Lombard Matcha-TTS | Diffusion / flow | T → S |
| LongCat-AudioDiT | Diffusion / flow | T, S → S |
| LoRP-TTS | Diffusion / flow | T, S → S |
| Luna-TTS | Token LM | T, S → S |
| M3-TTS | Diffusion / flow | T, S → S |
| MAGIC-TTS | Token LM | T, S → S |
| MagpieTTS-LF | Token LM | T, S → S |
| MambaVoiceCloning | Diffusion / flow | T, S → S |
| MamTra | Diffusion / flow | T, S → S |
| ManchuTTS | Diffusion / flow | T → S |
| Marco-Voice | Token LM | T, S → S |
| MARS6 | Token LM | T, S → S |
| Masked-style TTS | Token LM | T, S → S |
| MaskGCT | Token LM | T, S → S |
| Matcha-TTS | Diffusion / flow | T → S |
| MAVE | Token LM | T, S → S |
| Mega-TTS | Token LM | T, S → S |
| Mega-TTS 2 | Token LM | T, S → S |
| MegaTTS 3 | Diffusion / flow | T, S → S |
| Meitei Mayek TTS | Autoregressive | T → S |
| Mel-LLM (TTS) | Continuous LM | T → S |
| MELA-TTS | Continuous LM | T, S → S |
| MELD | Token LM | T, S → S |
| Mellotron | Autoregressive | T, S → S |
| MeloTTS | Flow / VAE | T → S |
| Metis | Token LM | T, S → S |
| MFCIG-CSS | Token LM | T, S, V → S |
| MiDashengLM-Gen | Continuous LM | T → S, A |
| MiniMax-Speech | Token LM | T, S → S |
| MixedG2P-T5 | Token LM | T, S → S |
| MM-MovieDubber | Diffusion / flow | T, V → S |
| MoE-TTS | Token LM | T → S |
| MoonCast | Token LM | T, S → S |
| MOSS-TTS | Token LM | T, S → S |
| MOSS-TTS-Nano | Token LM | T, S → S |
| MOSS-TTS-Realtime | Token LM | T, S → S |
| MOSS-TTSD | Token LM | T, S → S |
| MOSS-VoiceGenerator | Token LM | T → S |
| MP-ELD | Continuous LM | T, S → S |
| MPE-TTS | Token LM | T, S → S |
| Multistage multimodal TTS | Diffusion / flow | T, I → S |
| Muyan-TTS | Token LM | T, S → S |
| NaturalSpeech | Flow / VAE | T → S |
| NaturalSpeech 2 | Diffusion / flow | T, S → S |
| NaturalSpeech 3 | Diffusion / flow | T, S → S |
| NeuTTS Air | Token LM | T, S → S |
| NeuTTS Nano | Token LM | T, S → S |
| NeuTTS-2E | Token LM | T → S |
| NR-LauraTTS | Token LM | T, S → S |
| NVSpeech TTS | Token LM | T, S → S |
| Nüshu-PitchVITS | Flow / VAE | T → S |
| Ojibwe-Mi'kmaq-Maliseet TTS | Diffusion / flow | T → S |
| OmniVoice | Token LM | T, S → S |
| OpusLM | Token LM | T, S → S |
| Orpheus TTS | Token LM | T, S → S |
| OscillaTTS | Diffusion / flow | T, S → S |
| OuteTTS | Token LM | T, S → S |
| OV-InstructTTS | Token LM | T → S |
| OZSpeech | Token LM | T, S → S |
| PALLE | Token LM | T, S → S |
| Parallel GPT | Token LM | T, S → S |
| Parallel Tacotron | Parallel | T → S |
| Parallel Tacotron 2 | Parallel | T → S |
| ParaStyleTTS | Parallel | T → S |
| Parler-TTS | Token LM | T → S |
| Parler-TTS Hinglish adaptation | Token LM | T → S |
| PFluxTTS | Diffusion / flow | T, S → S |
| Phoenix TTS | Token LM | T, S → S |
| Phoneme-tone adaptive Thai TTS | Parallel | T, S → S |
| PilotTTS | Token LM | T, S → S |
| Piper (VITS voices) | Flow / VAE | T → S |
| Pocket TTS | Continuous LM | T, S → S |
| PortaSpeech | Flow / VAE | T → S |
| PROEMO | Parallel | T → S |
| Progressive face-conditioned TTS | Flow / VAE | T, I → S |
| Prompt-Unseen-Emotion | Token LM | T → S |
| PromptTTS | Parallel | T → S |
| PromptTTS 2 | Diffusion / flow | T → S |
| ProtoDisent-TTS | Flow / VAE | T, S → S |
| PS-TTS | Token LM | T, S → S |
| QTTS | Token LM | T, S → S |
| Qwen-Audio-3.0-TTS | Token LM | T, S → S |
| Qwen3-TTS | Token LM | T, S → S |
| RADKA-CSS | Token LM | T, S → S |
| RALL-E | Token LM | T, S → S |
| Raon-OpenTTS | Diffusion / flow | T, S → S |
| RapFlow-TTS | Diffusion / flow | T, S → S |
| ReGenVoice | Diffusion / flow | T, S → S |
| ReStyle-TTS | Diffusion / flow | T, S → S |
| RTFree-F5 | Diffusion / flow | T, S → S |
| RV-TTS | Diffusion / flow | T, I → S |
| RWKVTTS | Token LM | T, S → S |
| S5-TTS | Token LM | T, S → S |
| Sarashina2.2-TTS | Token LM | T, S → S |
| SASLM | Continuous LM | T, S → S |
| Seed-TTS | Token LM | T, S → S |
| Self-distilled zero-shot TTS | Compact | T, S → S |
| SelfTTS | Flow / VAE | T, S → S |
| SemaVoice | Continuous LM | T, S → S |
| SemBridge | Continuous LM | T, S → S |
| Shallow Flow Matching TTS | Diffusion / flow | T, S → S |
| SLED | Continuous LM | T, S → S |
| SlimSpeech | Compact | T, S → S |
| SMLLE | Token LM | T, S → S |
| SoulX-Podcast | Token LM | T, S → S |
| Spark-TTS | Token LM | T, S → S |
| SpeakStream | Continuous LM | T, S → S |
| SPEAR-TTS | Token LM | T, S → S |
| SpeechAccentLLM | Token LM | T, S → S |
| SpeechEdit | Token LM | T, S → S |
| SpeechT5 | Autoregressive | T, S → S |
| SpeechX | Token LM | T, S → S |
| SpeedySpeech | Parallel | T → S |
| Spotlight-TTS | Diffusion / flow | T, S → S |
| StellarTTS | Compact | T, S → S |
| Step-Audio-EditX | Token LM | T, S → S |
| Step-Audio-TTS | Token LM | T, S → S |
| StepAudio 2.5 TTS | Token LM | T, S → S |
| Stochastic-alignment continuous TTS | Continuous LM | T, S → S |
| StreamMel | Continuous LM | T, S → S |
| StyleTTS | Parallel | T, S → S |
| StyleTTS 2 | Diffusion / flow | T, S → S |
| Supertonic | Compact | T, S → S |
| Supertonic 2 | Compact | T → S |
| Supertonic 3 | Compact | T → S |
| SwanVoice | Diffusion / flow | T, S → S |
| SyncSpeech | Token LM | T, S → S |
| Tacotron | Autoregressive | T → S |
| Tacotron 2 | Autoregressive | T → S |
| TADA | Continuous LM | T, S → S |
| TED-TTS | Token LM | T, S → S |
| Tibetan-TTS | Token LM | T, S → S |
| TinyWave | Token LM | T, S → S |
| TLDR (TTS) | Token LM | T, S → S |
| TMD-TTS (formerly FMSD-TTS) | Diffusion / flow | T → S |
| TontaubeV1 | Token LM | T, S → S |
| Tortoise TTS | Token LM | T, S → S |
| Transformer TTS | Autoregressive | T → S |
| TTS-CtrlNet | Diffusion / flow | T, S → S |
| TTS-Transducer | Token LM | T, S → S |
| TTSYoruba | Concatenative | T → S |
| UDDETTS | Token LM | T → S |
| UmbraTTS | Diffusion / flow | T, S, A → S |
| UniFlow-Audio | Diffusion / flow | T, S, A, I, V → S, A |
| UNISON | Diffusion / flow | T, S, A → S, A |
| UniSonate | Diffusion / flow | T → S |
| UniSpeaker | Diffusion / flow | T, S, I → S |
| UniTAF | Token LM | T → S |
| UniTalker | Token LM | T, S, V → S |
| UniTTS | Token LM | T, S → S |
| UniVocal | Token LM | T, S → S |
| UniVoice (ASR and TTS) | Diffusion / flow | T, S → S |
| UniVoice (speech and singing) | Diffusion / flow | T, S → S |
| UniWav (TTS) | Diffusion / flow | T, S → S, A |
| USCF-conditioned TTS | Diffusion / flow | T, S → S |
| V-CASS | Token LM | T, V → S |
| VALL-E | Token LM | T, S → S |
| VALL-E 2 | Token LM | T, S → S |
| VALL-E X | Token LM | T, S → S |
| VALL-T | Token LM | T, S → S |
| Vclip | Flow / VAE | T, I → S |
| Vevo | Token LM | T, S → S |
| VibeVoice | Continuous LM | T, S → S |
| VibeVoice-Realtime | Continuous LM | T → S |
| VisualSpeech | Parallel | T, V → S |
| VITS | Flow / VAE | T → S |
| VITS2 | Flow / VAE | T → S |
| VividVoice | Diffusion / flow | T, I, V → S |
| VocalNet-M2 | Token LM | T, S → S |
| Voicebox | Diffusion / flow | T, S → S |
| VoiceChat-TTS | Continuous LM | T → S |
| VoiceCraft | Token LM | T, S → S |
| VoiceCraft-Dub | Token LM | T, S, V → S |
| VoiceDesigner | Diffusion / flow | T, S → S |
| VoiceSculptor | Token LM | T, S → S |
| VoxCPM | Continuous LM | T, S → S |
| VoxCPM2 | Continuous LM | T, S → S |
| Voxtral TTS | Token LM | T, S → S |
| VoXtream | Token LM | T, S → S |
| VoXtream2 | Token LM | T, S → S |
| VSpeechLM | Token LM | T, V → S |
| Wave-Tacotron | Autoregressive | T → S |
| WavTTS | Diffusion / flow | T, S → S |
| WenetSpeech-Wu TTS | Token LM | T, S → S |
| WeSCon | Token LM | T, S → S |
| WhisperSpeech | Token LM | T, S → S |
| WordVoice | Token LM | T, S → S |
| X-Voice | Diffusion / flow | T, S → S |
| X2Streaming-TTS | Token LM | T, S → S |
| XEmoRAG | Token LM | T, S → S |
| XTTS | Token LM | T, S → S |
| YourTTS | Flow / VAE | T, S → S |
| ZipVoice | Diffusion / flow | T, S → S |
| ZipVoice-Dialog | Diffusion / flow | T, S → S |
| Zonos | Token LM | T, S → S |
Figures are credited to their sources. Editorial input/output diagrams are labeled. Credits.
A2TTS extracts a voice embedding from a short recording and conditions a diffusion acoustic decoder on it. Reference-aware duration prediction improves timing consistency for multilingual synthesis in low-resource Indian languages.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Affectron extends a verbal-speech backbone to place nonverbal vocalizations in emotionally and contextually appropriate positions. Augmented training examples and structural masking enable expressive utterances containing events such as laughter while preserving the spoken content.
Paper · GitHub · Project · Details

Figure 2 · Source
AgentSteerTTS separates identity and emotional-prosodic representations, then grounds compound instructions in retrieved acoustic examples. A controller combines these conditions and uses feedback to refine expressive speech while preserving the intended speaker.
Paper · GitHub: no author-linked repository found · Details

Figure 4 · Source
AlignDiT aligns text, visual information and acoustic conditions before diffusion-based speech generation. Modality-specific guidance balances these inputs, targeting synchronized, intelligible speech that follows the timing and expression of the supplied scene.

Figure 1 · Source
AMNet adds phrase-structure information and local convolutional modeling to a parallel Mandarin acoustic model. These changes help capture contextual pauses, emphasis and intonation before the accompanying vocoder reconstructs the waveform.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
ARCHI-TTS uses a dedicated semantic alignment module to coordinate text and reference acoustic features. Reusing encoder features across denoising steps reduces repeated computation while the flow model generates the target speech.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
ATRIE converts character descriptions into separate voice-identity and dynamic prosody conditions. A compact adapter learns from a larger language-model teacher and modulates a GPT-SoVITS-based synthesizer, supporting expressive persona-driven speech generation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Audiobook-CC models context beyond individual sentences and separates style instructions from voice prompts. Distillation strengthens emotional expression, supporting multi-character narration with more consistent voices and performance across longer passages.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
AuEmoChat learns a discrete emotion representation from speech and compresses dialogue history while retaining emotionally relevant information. Its language model predicts emotion and speech tokens, which a context-conditioned flow decoder renders into expressive conversational audio.

Figure 2 · Source
AuK combines language-model conditioning, an audio VAE and successive multimodal and single-stream diffusion blocks. One model handles reference-based speech synthesis and instruction-guided editing; its distilled AuK-Flash variant reduces the number of generation steps.

Figure 4 · Source
Authentic-Dubber retrieves emotionally relevant audiovisual examples and progressively incorporates them into speech generation. Its director-actor formulation connects a scene's visual context and reference delivery to the target transcript for expressive movie dubbing.

Figure 2 · Source
AutoSIFT divides a reference voice's style into attribute-specific components and a residual representation. Text instructions replace selected attributes while unmentioned characteristics remain conditioned on the reference, allowing partial style editing during speech synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
AutoStyle-TTS matches the target text against a collection of expressive speech examples using learned textual embeddings. The selected recording provides style conditioning for synthesis, automatically adapting delivery to the content.
Paper · GitHub · Project · Details

Paper figure · Source
This audio-visual language model adds full-face information to an expressive speech backbone. Training on emotion and dialogue tasks connects facial cues with spoken delivery, enabling speech generation that uses visual as well as acoustic conversational context.

Figure 2 · Source
Bagpiper-TTS converts a natural-language request into a detailed speech plan containing words and delivery information. The generator follows that plan for tasks ranging from ordinary narration to multi-speaker rendering, role-play and singing.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
BareWave generates speech in waveform space using a single inference path. Representation alignment, staged noise scheduling and perceptual objectives guide training, replacing the separate acoustic-feature and waveform-reconstruction stages common in other TTS systems.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
Hierarchical autoregressive audio tokens.
Editorial input/output diagram · Source
Autoregressive speechcodes and convolutional decoder.

Figure 1 · Source
BatonVoice interprets a user's expressive request and translates it into controls for its dedicated BatonTTS generator. Separating instruction interpretation from acoustic rendering enables more explicit feature control, including transfer to languages outside the control-training data.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
BELLE predicts both speech values and their uncertainty in a continuous autoregressive synthesizer. Multiple synthetic renditions of the same text provide training support for the variance estimate, enabling richer acoustic distributions without adding an iterative inference stage.
Paper · GitHub · Project · Details

Figure 1 · Source
BitTTS reduces storage and computation through extremely low-bit trained weights and indexed parameter sharing. It targets speech generation on constrained devices, preserving a full synthesis path while shrinking the model representation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This streaming system replaces continuous acoustic regression with direct prediction of Mimi codec layers. A modified FastSpeech 2 backbone supplies aligned features and a depth-wise decoder fills residual codebooks, producing successive speech blocks without temporal autoregression.
Paper · GitHub: no author-linked repository found · Details

Figure 1 (paper page 10) · Source
BnTTS extends an XTTS-based multilingual pipeline to Bangla using language-specific phonetic adaptations. A small amount of target-speaker audio supports personalization, with the model designed for limited-resource speech synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Bolbosh adapts Matcha-TTS to Kashmiri with language-aware text processing and cross-language training. Its design targets the pronunciation and script challenges of a low-resource language while retaining efficient non-autoregressive acoustic generation.

Figure 1 · Source
This system organizes speech instructions at global, sentence and token levels to guide extended recordings. A continuous-token backbone uses explicit planning and condition dropout to combine voice design, multi-speaker rendering and changing acoustic or emotional context.
Paper · GitHub: no author-linked repository found · Details
Editorial input/output diagram · Source
BreezyVoice combines supervised speech tokens, a language model and flow-based acoustics with a pronunciation frontend. Its Taiwanese Mandarin adaptation provides explicit phonetic control for characters with multiple readings while retaining reference-based voice synthesis.

Figure 1 · Source
BridgeTTS uses the BridgeCode dual representation to shorten the sequence predicted by its autoregressive language model. Bridging modules reconstruct more detailed continuous acoustic features from those sparse tokens, balancing generation speed with voice fidelity.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
Beyond Video-to-SFX predicts audio semantic tokens from visual information and phonetic cues, then refines them into acoustic tokens. The two-stage generator produces intelligible speech whose timing and environmental sound fit the supplied video.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
CAM-TTS retains global narrative information and retrieves local details through an updatable memory block. Prefix attention combines those memories with preceding context, guiding sentence-level synthesis across longer paragraphs.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
CapTalk designs voices from descriptions for individual utterances and multi-speaker dialogue. Hierarchical conditioning separates stable speaker identity from changing turn-level delivery, while explicit planning tokens control dynamic expressive attributes.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
CAST-TTS maps a voice description or a reference recording into a common timbre-conditioning interface. Cross-attention delivers that information to the synthesizer, enabling voice design and reference-based cloning within the same generation framework.
Paper · GitHub · Project · Details

Figure 1 · Source
CaT-TTS separates textual understanding from acoustic generation in a two-Transformer architecture. During decoding, a masked parallel inference procedure guides speech-token predictions to reduce local errors in zero-shot voice synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
This FastSpeech 2 extension explicitly models emotion alongside duration, pitch and energy. Counterfactual training separates emotional changes from linguistic content, allowing users to modify prosody while preserving the intended words.
Paper · GitHub: no author-linked repository found · Details

Figure 1 (paper page 3) · Source
CDE-StyleTTS lets acoustic states evolve continuously along a phoneme sequence whose timing comes from durations. Sampling this trajectory supplies the acoustic decoder with timing-sensitive representations, providing a way to transfer changing expressive style instead of merely repeating phoneme embeddings.

Paper figure · Source
Chain-of-Details TTS progressively predicts speech at increasing temporal resolutions using a shared decoder. The coarsest stage provides an implicit phonetic plan, while subsequent stages recover timing detail without a separate phoneme-duration predictor.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Chain-Talker first derives an emotional description from dialogue history, then predicts semantic speech codes. A final rendering stage combines these plans to synthesize expressive responses whose delivery fits the conversational context.

Figure 2 · Source
Chatterbox synthesizes speech from text and a voice reference, with controls for expressive delivery. Its multilingual models focus on cross-language voice consistency, while Turbo and Nano use a smaller backbone and a single-step acoustic decoder. These variants serve different narration, conversational playback and local-device requirements; their capabilities are not interchangeable.
Editorial input/output diagram · Source
Chatterbox-Flash generates speech-token blocks in parallel while keeping block-by-block streaming. Calibration against common-token priors and confidence-based stopping improve its discrete diffusion decoding after adaptation from a pretrained autoregressive synthesizer.

Figure 3 · Source
Autoregressive speech-token generation.
Editorial input/output diagram · Source
Probabilistic residual quantization and multi-token LM.

Figure 1 · Source
CLEAR models speech directly in a continuous latent space, avoiding discrete codec-token prediction. Its zero-shot generator combines reference-voice conditioning with incremental audio production, targeting a balance between naturalness and response latency.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Clip-TTS trains its textual representation against corresponding mel-spectrogram information through a contrastive objective. The acoustic Transformer uses the resulting context-aware features to improve prosodic interpretation during speech generation.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
This compact synthesis system combines a shared-parameter text frontend with an efficient recurrent waveform generator. It targets responsive accessibility voices on low-power devices, where small storage requirements and immediate playback matter alongside naturalness.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This speech-language-model design keeps recent acoustic tokens and voice prompts at full detail while compressing distant context. The asymmetric representation reduces redundant long-sequence processing without discarding the local cues needed for pronunciation and vocal consistency.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Confucius4-TTS extracts voice characteristics from self-supervised speech features without requiring a transcript of the reference recording. A language model predicts semantic tokens and a flow decoder generates mel-spectrograms, enabling voice cloning within and across fourteen languages.

Figure 1 · Source
This model combines a language head that predicts boundaries with a diffusion head that generates continuous acoustic frames. Masked and staged training stabilize speaker-reference conditioning, providing a text-to-speech path within a multimodal language-model architecture.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
This synthesizer separates reference voice information from acoustic background conditions. An explicit task control selects whether background sound is retained or removed, allowing personalized speech generation under different environmental requirements.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
CookVoice aligns textual content, style and prosodic controls to acoustic frames before speech generation. The same compact model supports spoken and sung voices, reference imitation and editing, allowing individual voice attributes to be controlled within a shared synthesis pipeline.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
CosyEdit2 adapts a text-speech language model and acoustic decoder for consistent speech editing, then refines them with editing-specific rewards. The paper also evaluates the resulting improvement in zero-shot text-to-speech, connecting local editing consistency with reference-based synthesis.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
CoSyncDiT guides flow-based speech synthesis through acoustic-style adaptation, visual calibration and timed context alignment. These stages connect the supplied transcript and scene information to expressive, synchronized movie dubbing.

Figure 2 · Source
Supervised semantic tokens and flow decoder.

Figure 1 · Source
Text/speech LM and chunk-aware flow matching.

Figure 1 · Source
CosyVoice 3 extends streaming, reference-conditioned synthesis with a tokenizer trained on several speech-understanding tasks and a reward model for post-training. It targets reliable pronunciation, speaker identity and prosody across languages, dialects and less controlled text. The research scaling experiments and the downloadable Fun-CosyVoice3 checkpoint represent different model configurations.
Paper · GitHub · Project · Details

Figure 2 · Source
The WhispSynth generation pipeline combines a CosyVoice synthesizer with pitch-free digital signal processing to produce whispered speech. It supports multilingual whisper generation while avoiding the voiced pitch patterns of ordinary speech synthesis.

Figure 2 · Source
CoVoMix2 generates scripted dialogue directly with a flow-matching model, using reference voices without their transcripts. Speaker-disentangled text, sentence alignment and prompt masking support controlled timing and overlapping turns.

Figure 1 · Source
Cross-Lingual F5-TTS changes reference preparation and training so the generated text need not be paired with a reference transcript. Word-aligned acoustic prompts support cross-language voice cloning while reusing the flow-matching synthesis backbone.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
CrossAccent-TTS separates speaker identity from accent-related information in a speech synthesis model. Weighted language embeddings control the accent subspace, allowing gradual accent changes and cross-language synthesis while retaining the reference speaker's timbre.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
CSM uses the text and audio of preceding speaker turns to shape the delivery of the next utterance. A Llama backbone predicts speech representations, and a smaller decoder completes Mimi audio codes. It is a contextual speech renderer: an application supplies the words to say, including any responses written by a separate language model.
Editorial input/output diagram · Source
CTC-TTS uses automatically derived alignment and two-word interleaving to train an incremental speech language model. Its length-concatenated and feature-stacked variants make different trade-offs between speech quality and generation latency.

Figure 2 · Source
CtrlSpeech adds local pitch, loudness and duration conditioning to a patch-autoregressive diffusion synthesizer. A separate global speaker condition preserves the reference voice while users modify the delivery of individual words or phonemes.

Figure 2 · Source
CuteTTS combines a causal audio VAE, an autoregressive patch model and an explicitly speaker-conditioned flow head. Generating patches rather than individual frames reduces sequential work; a distilled variant targets faster streaming while preserving reference-voice synthesis.

Figure 1 · Source
DAIEN-TTS separates a reference recording into speech and environmental components, then conditions acoustic generation on them independently. Its extended formulation additionally models reverberation and uses separate guidance controls for speech, noise and room acoustics, enabling voice cloning into a chosen environment.
Paper 1 · Project 1 · Paper 2 · GitHub · Project 2 · Details

Paper figure · Source
DARS separately models pathological timing and acoustic style to synthesize dysarthric speech. A multistage rhythm predictor and conditioned flow model generate targeted examples for improving recognition under limited real-world speech data.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
DCAR changes the number of acoustic tokens predicted at each autoregressive step. Adapting the chunk size to the generation state shortens sequential processing while maintaining content alignment and reference-conditioned speech quality.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Neural TTS pipeline with autoregressive WaveNet synthesis.

Figure 1 · Source
Speaker-conditioned neural pipeline and WaveNet synthesis.

Figure 1 · Source
Convolutional attention encoder-decoder and converter.

Figure 1 · Source
DeepASMR separates ASMR delivery from the reference speaker's identity using discrete speech representations. A language model predicts content and style, while a flow-based acoustic decoder transfers the requested performance into the target voice.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
DeepDubber-V1 interprets visual scenes and dubbing requirements before generating the target speech. Its multimodal conditions distinguish narration, monologue and dialogue, guiding both expressive delivery and synchronization.

Figure 1 · Source
DeepDubbing assigns voices to characters and conditions speech rendering on the surrounding script. Its voice-design and instruction-synthesis stages support multi-participant audiobook production while maintaining character identity and context-sensitive expression.

Figure 1 · Source
Conformer acoustic model with explicit and implicit prosody.

Figure 1 · Source
DELTA-TTS converts a pretrained speech language model to parallel discrete diffusion using lightweight adaptation. Local convolution and confidence-based decoding help retain acoustic structure while the model fills speech-token positions in an order determined by prediction confidence.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
DepFlow separates a depression-related acoustic representation from speaker identity and spoken content. The representation conditions a flow-matching synthesizer, enabling controlled synthetic speech for investigating acoustic cues and data augmentation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Dia turns a speaker-tagged transcript into conversational audio, including supported nonverbal events such as laughter and coughing. Reference audio and its transcript can establish speaker identity and delivery. Its English checkpoint is useful for scripted exchanges and dialogue narration, with the conversation content supplied by the user rather than generated by the model.
Editorial input/output diagram · Source
Dia2 begins synthesizing before the complete script is available, allowing an application to feed words incrementally. Audio prefixes provide speaker and conversational context, while a streaming codec path reconstructs the output. The released 1B and 2B checkpoints focus on English dialogue; prefix conditioning helps keep voices consistent between generations.
Editorial input/output diagram · Source
DialoSpeech models two speakers on separate tracks and uses chunked flow matching for acoustic rendering. The design supports expressive scripted dialogue, including interactions whose timing is difficult to reproduce by joining independent utterances.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
DiEmo-TTS distills emotion information from speech while suppressing unrelated speaker characteristics. Cluster-based sampling and representation perturbation improve cross-speaker emotion transfer, including situations where extensive emotion labels are unavailable.

Figure 1 · Source
Text-conditioned denoising diffusion acoustic model.

Figure 3 · Source
DiffCSS samples prosody representations from multimodal conversational context using a diffusion model. A prosody-conditioned speech language model renders those samples, allowing several expressive deliveries that remain consistent with the same dialogue.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
DiFlow-TTS maps phonemes into linguistic content and generates separate prosody and acoustic token streams through discrete flow matching. Factorizing these responsibilities supports compact zero-shot synthesis with fewer sequential generation steps.

Figure 2 · Source
DiFlowDubber first learns linguistic content and separate prosodic-acoustic tokens through a discrete-flow TTS model. A subsequent video-dubbing stage aligns those representations to visual timing, connecting voice generation with synchronized lip movements.
Paper · GitHub · Project · Details

Figure 2 · Source
DisCo-Speech learns a codec that separates content, delivery and speaker identity. A language model predicts combined content-prosody tokens while the decoder receives a separate timbre representation, enabling reference cloning with independently controllable vocal attributes.
Paper · GitHub · Project · Details

Figure 1 · Source
DisSpeech maps Mandarin text and marked stuttering events to semantic speech tokens without temporal autoregression. Pitch and energy modeling guide acoustic reconstruction, enabling controlled repetitions and other disfluencies for speech synthesis and recognition-data augmentation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
DiSTAR first drafts blocks of residual-quantized speech tokens with a language model. A masked diffusion decoder then fills acoustic detail within each block, combining temporal planning and parallel refinement entirely in discrete codec space.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
DiTAR predicts a sequence of compressed acoustic patches using a language model, then generates each patch's detail through diffusion. Separating global temporal planning from local reconstruction supports zero-shot speech synthesis with controllable sampling diversity.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Latent diffusion Transformer with speech-length prediction.

Figure 1 · Source
DMOSpeech 2 adds reinforcement learning to duration prediction in an already metric-optimized synthesizer. Rewards derived from speaker similarity and transcription accuracy guide timing choices, linking prosodic planning with reference-voice and content objectives.
Paper · GitHub · Project · Details

Figure 1 · Source
DMP-TTS maps style descriptions and reference recordings into a shared conditioning space. Chained guidance controls content, timbre and style separately, allowing detailed synthesis adjustments within a latent diffusion Transformer.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
dots.tts predicts continuous acoustic representations from multilingual text and voice references. Its flow-matching output head supports both audio streaming and streaming text input; guidance-aware distillation reduces the work needed to generate each audio packet.
Paper · GitHub · Project · Details

Figure 1 · Source
Dragon-FM predicts successive speech chunks autoregressively while refining the tokens inside each chunk with bidirectional flow matching. Compact acoustic codes and cross-chunk caching reduce generation overhead and support longer content such as podcasts.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
DrawSpeech turns user-drawn prosodic curves into detailed pitch and energy conditions for synthesis. A diffusion model fills in the acoustic detail, enabling localized expressive control beyond an overall style description.

Paper figure · Source
DS-TTS extracts complementary voice characteristics through two style encoders. Dynamic modulation conditions the acoustic generator on these representations, supporting unseen speakers and adapting synthesis across different sentence lengths.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
DualDub generates spoken dialogue and background sound together from video and textual conditions. A cross-modal aligner coordinates the two decoding heads, targeting temporally synchronized soundtracks rather than speech rendered in isolation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
DualSpeechLM uses understanding-oriented speech tokens as input and acoustic codec tokens for generation. Semantic supervision and staged conditioning coordinate the two representations, supporting speech synthesis within a unified understanding-and-generation architecture.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
Flow matching with filler-token text conditioning.

Figure 1 · Source
ECTSpeech gradually tightens consistency constraints on a pretrained diffusion synthesizer. The resulting generator maps a noise state to speech in one step, reducing repeated denoising while retaining the conditioning of the original TTS model.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
Alignment-guided token reordering.

Figure 1 · Source
EME-TTS jointly models emotional delivery and local emphasis instead of treating them as independent effects. Automatically derived emphasis labels and variance-related features help users stress selected material while retaining a recognizable target emotion.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
EMM-TTS separates emotional content modeling from speaker-specific acoustic generation. Speaker-consistency objectives and adaptive normalization help transfer emotion across languages while retaining the reference speaker's timbre.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
EmojiVoice adds interpretable emoji prompts to the text encoder and flow predictor of Matcha-TTS. Changing prompts across phrases varies expression during longer robot utterances, providing a lightweight control interface for expressive synthesis.

Paper figure · Source
EmoShift adds a lightweight layer that learns emotion-dependent changes to a TTS model's hidden representation. The controls adjust expressive delivery and emotion intensity while leaving most of the underlying synthesis backbone intact.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
EmoSSLSphere combines an emotion representation constrained to a sphere with discrete units derived from self-supervised speech features. The model synthesizes emotional speech across languages while organizing expressive controls independently of the textual content.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
EmoSteer-TTS extracts emotion-related directions from a pretrained synthesizer's internal activations. Applying those directions during inference changes, mixes or removes emotional expression without retraining; the paper evaluates the method across several different TTS backbones.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
This emotional synthesizer learns separate reference encoders for timbre and emotion. A mutual-information objective reduces their overlap, while phoneme-level emotion prediction carries changing expression into the generated acoustic sequence.
Paper · GitHub · Project · Details

Figure 1 · Source
PromptTTS-derived style and content conditioning.
Editorial input/output diagram · Source
EmoTra-TTS introduces frame-level valence, arousal and dominance controls into both prosodic planning and acoustic decoding. Synthetic transition examples teach the model to move between emotions within an utterance while keeping its wording and speaker identity consistent.
Paper · GitHub · Project · Details

Figure 2 · Source
EmoVoice interprets free-form textual descriptions of emotional delivery. Its phoneme-boost variant predicts phonetic and acoustic information together to improve content consistency while retaining expressive style control.

Figure 1 · Source
This system trains the discrete speech representation together with the language and acoustic models that consume it. Reconstruction and recognition feedback also update token prediction, reducing mismatches between separately trained components in reference-conditioned speech synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
F5-TTS learns text-guided speech infilling with flow matching, refining character representations before a Transformer predicts the speech trajectory. A reference clip supplies voice context, and Sway Sampling controls how inference steps are distributed. The 2025 v1 Base release refines training and inference within the same general architecture for zero-shot speech synthesis.

Figure 1 · Source
F5R-TTS adapts a flow-based synthesizer to reinforcement learning through a probabilistic formulation of generation. Transcription accuracy and speaker-similarity rewards refine the model's ability to preserve both target content and the reference voice.

Figure 2 · Source
This model maps facial features into the style space of StyleTTS 2 through a lightweight learned adapter. It synthesizes text in a plausible face-conditioned voice without an audio reference; the paper evaluates transfer to unseen identities and another synthesis language.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FaceSpeak extracts speaker-related and expressive information from real or stylized portraits. These visual conditions guide text-to-speech generation, allowing the requested words to be rendered in a plausible voice and emotion associated with the image.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
FacialTalker quantizes facial action information and combines it with text and speech history in a conversational synthesizer. Joint preference training over visual and speech tokens helps the generated delivery reflect the interlocutor's facial expression and dialogue context.

Figure 2 · Source
Parallel Transformer with explicit pitch prediction.

Figure 1 · Source
Feed-forward Transformer and duration-based length regulator.

Figure 1, PDF p. 4 · Source
Parallel Transformer with duration, pitch and energy prediction.

Figure 1, PDF p. 3 · Source
FC-TTS conditions generation on two recordings, one supplying delivery style and the other speaker identity. Specialized representation processing and auxiliary training objectives aim to prevent either condition from leaking unwanted attributes into the synthesized voice.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FELLE generates continuous acoustic frames sequentially, using the preceding frame to shape the next flow-matching prior. A coarse-to-fine acoustic head refines each prediction, supporting reference-conditioned speech without discrete speech-token classification.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FineCombo-TTS interprets style descriptions relative to a supplied speech reference. A flow-based variance predictor models how acoustic attributes should change, enabling precise relative edits without requiring an explicit independent embedding for every voice attribute.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
FireRedAudio uses different acoustic encoders for understanding audio and conditioning speech generation. Its shared language model drives a flow-matching decoder over continuous RedAE latents, supporting voice cloning, instruction-controlled synthesis and speech editing within the broader audio model.

Figure 1 · Source
Text-to-semantic LM and speech decoder.

Figure 3 · Source
FireRedTTS-1S extends the FireRed synthesis line with incremental acoustic decoding. Its chunked flow-matching and frame-autoregressive multi-stream decoder options provide different trade-offs between initial latency and sustained generation speed.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FireRedTTS-2 models chronological sequences of speaker-labeled text and speech using a large Transformer plus a smaller codebook decoder. A low-rate streaming tokenizer reduces the number of audio steps. It targets long conversations and podcasts where speaker changes, turn-specific delivery and continuity across utterances matter.

Figure 1 · Source
FireRedTTS3 uses a semantically supervised audio autoencoder to make continuous speech representations easier to predict. The Base variant provides multilingual reference cloning; the Instruct variant adds natural-language voice design and editing of spoken content or acoustic attributes.

Figure 1 · Source
The S1 family combines multilingual voice conditioning with explicit markers for emotion, tone and nonverbal sounds. Its full model and distilled S1-mini offer different deployment sizes, with reinforcement learning used to refine generation. It is suited to expressive narration and character dialogue, although access and capabilities depend on the selected release.
Editorial input/output diagram · Source
Fish Audio S2 extends the Fish speech-model line with natural-language delivery instructions and multi-speaker, multi-turn synthesis. Its training pipeline uses speech descriptions, quality assessment and reward modeling to improve controllability. The released inference stack supports streaming, making the model relevant to both scripted audio production and incremental spoken responses.

Figure 2 · Source
Dual-autoregressive slow/fast transformers.

Figure 2 · Source
Flamed-TTS combines representations of different speech attributes in an attention-free generator. Its reformulated flow-matching process targets efficient zero-shot synthesis with flexible pacing and reduced sequential computation.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
FlashTTS processes incoming text and speech context on staggered tracks so synthesis can begin before sentence completion. Multi-token prediction accelerates the language model, while a distilled acoustic decoder reduces waveform-generation delay.
Paper · GitHub · Project · Details

Figure 1 · Source
FleSpeech unifies text, voice recordings and visual prompts into a common conditioning representation. Its multistage generator uses those controls to manipulate voice and delivery attributes flexibly rather than requiring one fixed prompt modality.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
FlexiVoice accepts an optional style instruction and an optional reference voice alongside the target text. Progressive preference training teaches the model to follow both conditions while reducing unwanted coupling between wording, speaker identity and delivery.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
FlexSpeech separates timing control from the acoustic synthesis component to balance stable pronunciation and natural expression. A small set of style examples can adapt the duration module without retraining the full generator, supporting efficient delivery customization.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Autoregressive normalizing flows over mel spectrograms.

Figure 1 · Source
FNH-TTS routes linguistic and speaker information through several duration experts to model varied timing patterns. It combines this predictor with changes to waveform generation, aiming for robust end-to-end speech synthesis across voices and prosodic conditions.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This architecture lets a global language model predict several speech frames at a time and delegates their codec entries to a smaller local Transformer. The paper compares sequential local decoding with iterative masked prediction, showing how frame stacking changes the trade-off between synthesis throughput and acoustic fidelity.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FreyaTTS is a Turkish-focused non-autoregressive synthesizer that maps character sequences to continuous audio latents. A frozen waveform autoencoder reconstructs the speech, while duration prediction and voice-focused post-training support efficient conversational playback without phoneme or discrete speech tokenization.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Gemini 2.5 TTS converts supplied text into speech with prompt-based control of accent, pace, style and emotion. The Flash and Pro interfaces support single-speaker narration and two-speaker scripts with separately assigned voices. These are dedicated speech-generation endpoints; their internal acoustic architecture is not fully disclosed in the public documentation.
Docs 1 · Docs 2 · Docs 3 · Details
Editorial input/output diagram · Source
Gemini 3.1 Flash TTS adds expressive audio tags to prompt-steered speech generation, giving authors more local control over narration and delivery. It targets natural, responsive multilingual synthesis through a managed API. The public preview documentation describes the interface and controls, without enough architectural detail to reconstruct the underlying speech generator.
Editorial input/output diagram · Source
GibbsTTS generates discrete speech tokens through a continuous-time jump process. Metric-aware transition scheduling and a finite-step correction improve how token states evolve, supporting zero-shot voice synthesis with discrete flow matching.
Paper · GitHub · Project · Details

Figure 1 · Source
GLM-TTS first predicts speech tokens autoregressively, then converts them into audio with a diffusion decoder. Pitch-aware tokenization and multi-reward reinforcement learning target pronunciation, speaker similarity and expression. Hybrid phoneme/text input and LoRA voice adaptation provide controls for applications that need repeatable pronunciation and customized voices.

Figure 1 · Source
Normalizing-flow acoustic model and monotonic alignment search.

Figure 1, PDF p. 3 · Source
GOAT-TTS encodes continuous voice information in one branch and predicts speech tokens in another. Partial language-model adaptation preserves textual knowledge, while multi-token prediction supports streaming synthesis with reference-based paralinguistic conditioning.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
General-Purpose Audio uses one autoregressive backbone to predict discrete speech tokens across synthesis, recognition and conversion tasks. Its TTS path combines target text with voice context, with shared multitask training and scalable inference.

Figure 1 · Source
GPT-4o Mini TTS combines the text to be spoken with instructions that steer accent, speed, tone and emotional delivery. The Speech API can stream audio before the full result is complete and supports several output formats. It provides a managed synthesis component for narration and voice applications; public documentation does not disclose the complete acoustic architecture.
Editorial input/output diagram · Source
GPT-SoVITS couples text-to-semantic token prediction with a reference-conditioned speech decoder. Its 2025 V3, V4 and V2 Pro releases extend a workflow that supports both zero-shot synthesis and voice adaptation from a small training set. The surrounding WebUI helps prepare data, while recognition and source-separation utilities remain separate components.
Editorial input/output diagram · Source
Score-based diffusion decoder and monotonic alignment search.

Figure 2 · Source
GRAFT attaches codec tokens from a spoken word example to that word's location in the text prompt. Separate target-speaker conditioning allows the pronunciation hint to come from another voice while the synthesized sentence retains the desired speaker.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
GSA-TTS extracts local style information at successive levels and combines it through attention into a global reference condition. This richer style representation guides the acoustic model when synthesizing an unseen speaker's voice.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Tacotron with a reference encoder and global style tokens.

Figure 1 · Source
Habibi trains an Arabic synthesizer progressively from standard language to regional dialects using curated public speech. It targets zero-shot voice cloning across dialects and reading without mandatory diacritic marks.
Paper · GitHub · Project · Details

Figure 1 · Source
HD-PPT learns speech codes that distinguish spoken content from instruction-related preferences. A language model predicts semantic information, expressive style and acoustic detail in sequence, improving the mapping from natural-language requests to controllable speech.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Higgs Audio v2 combines interleaved text/audio modeling with a unified speech tokenizer and DualFFN layers for acoustic prediction. Reference clips and scene context influence voices, while text context shapes prosody across narration and multi-speaker scripts. The generation checkpoint covers synthesis; the separate understanding branch in the family diagram is not another output mode of this checkpoint.

Official architecture diagram · Source
Higgs Audio v2.5, now documented as Higgs TTS 2.5, reduces the autoregressive audio Transformer to 1B parameters. GRPO-based alignment and a curated voice dataset refine pronunciation, cloning and expressive control tags. It targets production narration and conversational speech with lower computational requirements than the preceding 3B generation model.
Editorial input/output diagram · Source
Higgs TTS 3 uses an autoregressive decoder over interleaved text and eight speech codebooks, with a delay pattern and fused input/output projections. Reference audio establishes a voice, while inline tokens control emotion, style, pauses and sound effects. The 4B release targets multilingual conversational speech and expressive response rendering.

Official architecture diagram · Source
HiStyle predicts a voice's timbre first and finer delivery attributes afterward from textual descriptions. Contrastive text-audio alignment organizes these style representations before they condition a speech synthesizer.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
HoliDubber conditions audio generation on video and a text prompt describing speech and sound effects. A causal model plans successive latent patches and a local diffusion Transformer generates their detail, supporting synchronized dubbing within complex acoustic scenes.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
HoliTok combines linguistic and acoustic information in a continuous representation designed for both understanding and generation. Its downstream autoregressive model and diffusion decoder demonstrate a text-to-speech path using the same latents employed for recognition.

Figure 1 · Source
Octave uses text context and acting instructions to adjust pronunciation, emphasis, tempo and emotional delivery. Its API supports voice creation from descriptions, voice cloning and continuation across longer passages. Octave 1 and the Octave 2 preview have different feature coverage; the public interface is documented more fully than the internal speech-model architecture.
Editorial input/output diagram · Source
ImmersiveTTS jointly models spoken content and its surrounding acoustic scene in a multimodal diffusion Transformer. Speech and general-audio representations provide complementary training signals, helping generated wording remain intelligible within the requested environmental context.
Paper · GitHub · Project · Details

Figure 1 · Source
IndexTTS adapts the XTTS/Tortoise approach with a Conformer reference encoder and a BigVGAN2 speech decoder. Hybrid character/pinyin input gives explicit control over difficult Chinese pronunciations. Its central use case is zero-shot voice cloning with predictable text rendering, including content that benefits from pronunciation correction.
Paper · GitHub · Project · Details

Figure 1 · Source
IndexTTS 2.5 shortens semantic sequences with a lower-rate codec and replaces the acoustic module's backbone with a Zipformer design. Multilingual training strategies and reinforcement learning extend pronunciation and emotion transfer across languages. It retains reference-based voice conditioning while reducing the cost of semantic and acoustic generation.

Figure 1 · Source
IndexTTS2 separates speaker identity from emotional style so that different references can control timbre and delivery. Its autoregressive formulation also supports explicit output-token budgeting for duration control, alongside unconstrained generation. These mechanisms are intended for expressive speech and timing-sensitive work such as dubbing; availability of controls should be checked in the chosen implementation.

Figure 1 · Source
InstructAudio combines natural-language instructions with phonemes or lyrics in a common generation format. Joint and single-stream diffusion layers synthesize speech or music while controlling attributes such as voice, emotion, accent or musical character.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
IntMeanFlow distills a flow-based speech generator to predict integrated acoustic updates over larger intervals. A search for effective sampling steps further reduces decoding work, enabling reference-conditioned synthesis with fewer iterative refinements.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Inworld TTS-1 and TTS-1-Max are multilingual autoregressive synthesizers designed for low-latency speech output. Textual audio markup controls emotions and nonverbal vocalizations, with model variants offering different capacity and inference-cost trade-offs.

Figure 1 · Source
JaiTTS continually trains a VoxCPM-derived synthesizer on Thai-centered speech data. Its semantic planning, residual acoustic modeling and local diffusion decoding retain reference-based voice cloning while targeting fluent Thai pronunciation and delivery.

Figure 1 · Source
JAM-Flow couples audio and motion diffusion modules within a shared model. Text, voice references and optional motion conditions support synchronized speech and facial movement, with an infilling objective allowing several conditioning combinations.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
JELLY combines an emotion-aware Q-Former with several partially adapted language-model modules. Joint emotion recognition and contextual reasoning guide a speech synthesizer toward responses whose delivery matches the conversation.
Paper · GitHub · Project · Details

Paper figure · Source
Joint FastSpeech 2 and HiFi-GAN with learned alignment.

Figure 1, PDF p. 2 · Source
This model handles text and speech within one non-autoregressive architecture. Its TTS path predicts acoustic output from text, and feeding partial predictions back into the model improves generation through iterative refinement while also supporting recognition training.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Joycent separates accent information from speaker identity using an adversarially trained accent encoder. It injects accent and speaker features at different text-encoder layers, supporting accent-conditioned synthesis without requiring a separate accented phoneme prediction stage.
Paper · GitHub · Project · Details

Figure 1 · Source
JoyVoice conditions long-form speech on speaker-labeled text and shared conversational context. Autoregressive hidden states feed an acoustic diffusion decoder, allowing multiple speakers, changing expression and more flexible turn boundaries within one synthesis system.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
KABURI-TTS renders each participant on a separate audio channel from time-aligned phonemes and speaker activity. Supplying the timing layout explicitly lets the system synthesize overlapping speech, backchannels and interruptions for two-speaker conversations.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
KittenTTS provides small ONNX speech models with built-in voices and adjustable playback speed. The Mini, Micro and Nano releases offer different size and inference tradeoffs for CPU-oriented applications. The official README documents usage more fully than internal acoustic design, so the catalog presents it as a compact synthesis family without asserting an undisclosed architecture.
Editorial input/output diagram · Source
Koel-TTS explores several ways to condition a Transformer synthesizer on text and reference audio. Automatic speech-recognition and speaker-verification feedback, together with classifier-free guidance, improve adherence to the requested words and voice.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Kokoro's 2025 v1.0 release uses a compact StyleTTS 2-derived decoder with an iSTFTNet waveform generator. The released model relies on preset voice representations and omits style diffusion and a reference encoder. It is suited to lightweight narration and application speech where a small synthesis model and ready-made voices are useful.
Editorial input/output diagram · Source
Kyutai TTS treats text and speech as aligned streams separated by a controlled delay. A decoder-only language model can therefore emit audio as text arrives, instead of waiting for a complete utterance. This formulation supports incremental synthesis for voice interfaces and long streams, with speaker conditioning supplied through the TTS implementation.

Figure 1 · Source
LanStyleTTS standardizes phonetic inputs and introduces local style conditioning across languages. The framework can augment several parallel acoustic backbones, allowing one multilingual model to vary delivery at phoneme level.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
LatinX uses staged text-to-audio training, voice-cloning adaptation and automatic preference alignment. The resulting multilingual Transformer renders text in the source speaker's voice, supporting the synthesis stage of cross-language speech translation.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
LE2E-TTS trains a compact text-to-waveform pipeline end to end rather than separately optimizing acoustic and waveform stages. It targets local devices where model size, response time and compute cost constrain deployment.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
FastSpeech-derived architecture found by neural architecture search.
Editorial input/output diagram · Source
LLaDA-TTS adapts a speech language model to fill masked token sequences in parallel. Its bidirectional generation also supports inserting, replacing or deleting spoken words, combining reference-based TTS and speech editing through the same model.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Llasa maps text and an optional speech prompt to a single stream of codec tokens using a Llama-style Transformer. The simple representation makes standard language-model scaling and sampling techniques applicable to synthesis. Its research also explores speech-model verifiers that select samples for content accuracy, voice consistency or emotional expression.
Paper · GitHub · Project · Details

Figure 2 · Source
Llasa+ adds multi-token prediction modules to a frozen Llasa backbone and checks their proposals with that backbone. A causal codec decoder turns accepted tokens into streaming audio. The resulting design addresses autoregressive latency while retaining the original speech model, making it relevant to systems that need incremental playback without retraining an entire backbone.

Figure 1 · Source
LLMVoX connects to an upstream language model through a streaming queue interface. Its small speech generator renders incoming text incrementally, allowing long conversations without tightly coupling the synthesizer to one particular language-model backbone.
Paper · GitHub · Project · Details

Figure 2 · Source
This Matcha-TTS extension learns vocal effort and articulation from automatically derived labels. It provides continuous controls over speech clarity and loudness-related effort, together with word-level emphasis, to synthesize the clearer delivery used in noisy listening conditions.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
LongCat-AudioDiT maps text and reference speech into continuous waveform latents using a diffusion model. A jointly considered waveform autoencoder and adapted inference guidance address the interaction between acoustic reconstruction and zero-shot synthesis quality.

Figure 1 · Source
LoRP-TTS adapts a pretrained zero-shot synthesizer using small low-rank parameter updates. It focuses on preserving a target speaker from limited, potentially noisy or spontaneous recordings whose acoustic conditions differ from the original training data.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
Luna-TTS adapts an autoregressive text backbone into a speech diffusion language model. Its parallel and Realtime variants share a tokenizer and training lineage; Realtime predicts successive codec blocks while denoising each block in parallel for incremental audio delivery.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 (paper page 4) · Source
M3-TTS uses joint text-audio diffusion layers to learn alignment without first stretching text into a guessed acoustic timeline. Additional single-stream layers refine acoustic details, producing reference-conditioned speech through a compressed mel representation.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MAGIC-TTS exposes timing controls for selected speech tokens and pauses. Training includes incomplete control signals so the model can follow local edits where provided and infer natural timing elsewhere, supporting precise pacing without requiring every segment to be specified.

Figure 1 · Source
MagpieTTS-LF extends MagpieTTS at inference time by retaining acoustic and textual context across sentence boundaries. Soft alignment priors and history-aware encoding support coherent longer narration without retraining the synthesizer on long recordings.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MambaVoiceCloning uses state-space modules to encode phonemes, learn their timing and condition expressive synthesis. A training-only alignment teacher supplies timing supervision, while the generation path targets efficient long sequences and limited-lookahead speech streaming.

Figure 1 · Source
MamTra mixes state-space and attention layers to retain global context while reducing the cost of long sequences. Knowledge transfer from a pretrained Transformer initializes the hybrid synthesizer, combining efficient local processing with expressive acoustic modeling.
Paper · GitHub · Project · Details

Figure 1 · Source
ManchuTTS builds multilevel text representations suited to Manchu and feeds them into a convolutional diffusion Transformer. Its non-autoregressive generator and augmented training data target speech synthesis where naturally recorded material is scarce.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Marco-Voice learns separate speaker and emotion representations using contrastive training. Rotating the emotional representation provides smooth expressive control while preserving the reference voice across different delivery styles.

Figure 1 · Source
MARS6 encodes text and a speaker representation before generating hierarchical acoustic codes. Its compact encoder-decoder design targets expressive speech and reference-voice cloning with reduced inference cost.
Paper · GitHub · Project · Details

Paper figure · Source
This controllable system first predicts a masked-autoencoder-derived speech style representation from text and controls. A second model generates codec tokens, allowing speaker characteristics and expressive attributes to be specified separately during synthesis.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Masked generative codec transformers.

Figure 1 · Source
Conditional flow matching with an encoder-decoder acoustic model.

Figure 1 · Source
MAVE combines a state-space backbone with cross-attention to generate speech conditioned on text and acoustic context. It supports zero-shot voice synthesis and editing while reducing the attention-memory requirements of a comparable Transformer-based codec model.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Disentangled speech factors and prosody LM.

Figure 1 · Source
Prosody language model and multi-sentence prompting.

Figure 1 · Source
MegaTTS 3 guides a latent diffusion Transformer with sparse text-speech alignment boundaries, leaving the model room to learn finer timing. Classifier-free guidance controls accent strength, while piecewise rectified flow reduces sampling work. The design targets robust zero-shot voice synthesis with more flexible alignment than a fully fixed duration sequence.

Figure 1 · Source
This Manipuri speech synthesizer maps Meitei Mayek writing to an ARPAbet-based phoneme representation before acoustic generation with Tacotron 2. A HiFi-GAN vocoder reconstructs the waveform. The paper develops a single-speaker system for a language with limited training resources and tonal pronunciation requirements.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
The synthesis experiment in Mel-LLM extends a language model to predict mel-based acoustic information directly. Its next-token VAE decoder demonstrates a text-to-speech path within an otherwise understanding-focused model; the paper presents this as a proof of concept with quality limitations.
Paper · GitHub: no author-linked repository found · Details

Fig. 1 (paper page 2) · Source
MELA-TTS predicts continuous mel-spectrogram frames from text and speaker conditions. A training-time alignment module connects the decoder to recognition-derived semantic features, helping the joint Transformer-diffusion model retain linguistic structure without discrete speech tokenization.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MELD learns discrete latent variables from mel-spectrograms jointly with its speech language model. This shared optimization supports zero-shot synthesis and recognition while addressing omissions and excessive silence associated with less coordinated acoustic representations.
Paper · GitHub: no author-linked repository found · Details

Figures 1–2 (paper page 2; panels assembled) · Source
Tacotron 2 with global style tokens, pitch and rhythm conditioning.
Editorial input/output diagram · Source
VITS-family multilingual speech synthesis.
Editorial input/output diagram · Source
Metis pretrains on unlabeled speech before adapting to task-specific conditions such as text. Self-supervised semantic tokens and acoustic codes support a shared foundation for reference-based TTS, conversion and other speech-generation tasks.
Paper · GitHub · Project · Details

Paper figure · Source
MFCIG-CSS represents dialogue history through separate graphs of meaning and vocal expression. Fine-grained multimodal interactions condition the speech synthesizer, helping each scripted response fit the surrounding conversation.

Figure 1 · Source
MiDashengLM-Gen trains a language model together with a conditional flow-matching output head to generate variable-length audio. Its text conditioning supports scenes containing intelligible speech alongside music or other sounds, with generation performed over continuous audio representations.
Paper · GitHub · Project · Details

Paper figure · Source
MiniMax-Speech extracts speaker characteristics directly from reference audio without requiring its transcript, then generates speech with an autoregressive Transformer and Flow-VAE. The research emphasizes multilingual zero-shot cloning and expressive delivery. Additional adaptation mechanisms support emotion control, description-based voice creation and more specialized voice cloning without replacing the base model.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MixedG2P-T5 learns acoustic units from speech and uses a language-model synthesis path for text containing mixed scripts. It reduces dependence on manually designed grapheme-to-phoneme rules while retaining accent and intonation information in the speech representation.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
MM-MovieDubber interprets scene information to distinguish dialogue, narration and monologue delivery. A speech generator then uses the resulting multimodal conditions with the target content to render expressive movie dubbing.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
MoE-TTS augments a frozen text language model with speech-specific expert parameters. Retaining the original language knowledge helps the synthesizer interpret unfamiliar style descriptions while learning the acoustic generation task.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
MoonCast combines podcast script preparation with a synthesizer trained for longer, spontaneous-sounding delivery. Voice references allow unseen speakers to render the resulting conversation, while discourse-level context supports more natural transitions than isolated sentence synthesis.
Paper · GitHub · Project · Details

Figure 1 · Source
MOSS-TTS offers two generators over a shared discrete audio representation: a delay-pattern model and a model with a frame-local Transformer. They balance long-context control against efficient codebook prediction and speaker preservation. The family supports reference-conditioned synthesis, pronunciation and duration controls, with later checkpoints extending language handling and explicit pauses.

Figure 2, PDF p. 8 · Source
MOSS-TTS-Nano packages multilingual voice cloning into a roughly 100M-parameter speech generator with a compact audio tokenizer. Streaming output and an ONNX inference path make it relevant to CPU-based readers and local applications. Its published performance depends on the runtime and hardware, and the tokenizer is a separate part of the deployment footprint.

Official architecture diagram · Source
MOSS-TTS-Realtime uses a Qwen3-derived backbone for linguistic context and a smaller local Transformer to predict audio codebooks. Text and speech tokens are handled at different levels so the system can accept text and emit audio incrementally. It targets low-latency spoken responses while preserving context across the generated utterance.
Editorial input/output diagram · Source
MOSS-TTSD turns a dialogue script with explicit speaker tags into a continuous multi-party recording. Long-context modeling helps maintain speaker identity, turn assignment and acoustic continuity, while short references can define voices. It is designed for podcasts, commentary and other scripted conversations; the source paper evaluates dialogue-specific consistency as well as intelligibility.

Figure 2, PDF p. 4 · Source
MOSS-VoiceGenerator creates a speaking voice from a natural-language description rather than requiring an example speaker recording. Training on expressive cinematic speech exposes it to varied delivery and acoustic conditions. The model is intended for character design, storytelling and role-based narration where the desired voice must be specified in words.

Figure 1 · Source
MP-ELD predicts low-rate continuous speech tokens through several information paths with separate local encoders. A flow decoder combines their predictions, while the accompanying Locodec representation is designed to limit accumulated errors during long speech generation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
MPE-TTS combines reference speech and textual prompts to specify an unseen speaker and the desired emotion. A prosody predictor and emotion-consistency objective carry those controls into the synthesized acoustic performance.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This framework learns face and text conditioning in separate stages before using them for voice synthesis. Visual knowledge distillation and training across text-face and text-speech pairs reduce reliance on fully matched multimodal recordings.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Muyan-TTS trains a speech language model on a large podcast collection for expressive reference-based synthesis. The release documents data preparation, training and optimized inference, with an emphasis on reproducible podcast-style voice generation.

Figure 1 · Source
Text-to-waveform VAE with enhanced prior and duration modeling.

Figure 1 · Source
Latent diffusion over neural-codec representations.

Figure 1 · Source
Factorized speech codec and attribute-wise diffusion.

Figure 3 · Source
NeuTTS Air pairs a phoneme-conditioned language model with NeuCodec to synthesize a reference voice locally. Quantized GGUF backbones support incremental generation through the documented streaming backend. It targets embedded and desktop voice applications, with the codec's compute and memory requirements considered alongside those of the language-model backbone.
Editorial input/output diagram · Source
NeuTTS Nano reduces the speech-model backbone while retaining phoneme conditioning, reference-based cloning and NeuCodec reconstruction. Its English, German, French and Spanish models share a design but use separate language-specific checkpoints. Quantized variants support compact local deployments, and streaming depends on the selected inference backend.
Editorial input/output diagram · Source
NeuTTS-2E accepts text directly and adds explicit emotional delivery to the NeuTTS language-model-and-codec pipeline. The released configuration supplies four fixed speaker presets instead of arbitrary reference-based cloning. It is intended for compact expressive speech applications, with streaming available through the documented GGUF inference path.
Editorial input/output diagram · Source
NR-LauraTTS cleans the discrete representation of a noisy voice prompt before passing it to LauraTTS. Token prediction and embedding refinement reduce background contamination, supporting reference cloning when the available recording is acoustically imperfect.
Paper · GitHub · Project · Details

Figure 1 · Source
The NVSpeech pipeline includes a TTS model that renders text with explicitly marked nonverbal events. Word-level annotations connect ordinary speech with vocalizations such as laughter, providing a shared representation for recognition and controllable audio generation.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
Nüshu-PitchVITS uses pitch annotations from Nüshu's writing system to guide acoustic generation under very limited data. A frame-level pitch predictor conditions the VITS waveform path, allowing syllable recordings and linguistic tone knowledge to support sentence synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
This model family shares speech-synthesis training across three related Indigenous languages. The paper compares attention-based and attention-free flow architectures, demonstrating how joint linguistic coverage can help languages with limited recordings.

Figure 1 · Source
OmniVoice predicts multiple acoustic codebooks directly from text using a masked, non-autoregressive diffusion language model. Random masking across codebooks and initialization from a pretrained language model support multilingual generation without a separate text-to-semantic stage. It focuses on broad-language zero-shot synthesis and voice conditioning, rather than visual or general-purpose omni interaction.

Figure 1 · Source
OpusLM extends text language models through speech-text pretraining on public data. Its interleaved representation supports speech recognition, text-conditioned synthesis and textual continuation within a transparent family of shared backbones.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Orpheus TTS repurposes a Llama-family language model to generate speech codec tokens from text. Emotion tags and speaker conditioning guide expressive delivery, while the project supplies inference and adaptation workflows. English releases and multilingual previews have different coverage, so a checkpoint's documented capabilities matter when selecting it for narration or a voice application.
Editorial input/output diagram · Source
OscillaTTS changes the periodic nonlinearities used in a style-diffusion synthesis backbone. Adjustable oscillatory modulation is designed to capture rapid pitch and amplitude changes while a linear bypass stabilizes the acoustic signal, targeting sharper expressive prosody.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
OuteTTS represents speech in a form that can be generated by a decoder-only language model and reconstructed by an audio decoder. A speaker reference guides vocal identity, style and accent. Its 1.0 line supports standard LLM serving backends, but generation settings and repetition handling need to follow the matching model implementation.
Editorial input/output diagram · Source
OV-InstructTTS interprets voice and delivery descriptions beyond a fixed inventory of style labels. Its reasoning-based conditioning connects broader textual requests with expressive speech generation, supported by a dedicated instruction-speech dataset.

Figure 2 · Source
OZSpeech generates disentangled speech components with a flow model conditioned on a learned prior. The design targets single-step zero-shot synthesis while separately modeling content, prosody and speaker-related information from the voice prompt.
Paper · GitHub · Project · Details

Figure 1 · Source
PALLE generates variable-length speech spans at fixed decoding steps, combining temporal planning with parallel token prediction. A second non-autoregressive stage refines the initial sequence, supporting efficient zero-shot synthesis.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
Parallel GPT divides speech generation between a general autoregressive predictor and a non-autoregressive detail model. The parallel refinement stage conditions on the initial tokens, balancing independence and interaction between semantic and acoustic information.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Parallel acoustic model with a variational residual encoder.

Figure 1 · Source
Parallel synthesis with differentiable duration modeling.

Figure 1 · Source
ParaStyleTTS converts textual style prompts into separate controls for prosody and broader paralinguistic characteristics. The lightweight adaptation design targets expressive speech from descriptions while making the roles of the two conditioning levels explicit.
Paper · GitHub · Project · Details

Figure 1 · Source
Description-conditioned codec language model.

Figure 1 · Source
This Parler-TTS extension introduces language-specific phonetic alignment and emotion embeddings for Hindi and Indian English. Its conditioning targets code-switched utterances whose accent and emotional delivery change coherently across language boundaries.

Figure 2 · Source
PFluxTTS combines two acoustic-generation paths by fusing their predicted vector fields at inference. Sequential reference embeddings support transcript-free cross-language voice cloning, and a super-resolution vocoder reconstructs high-rate output audio.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Phoenix TTS aligns its speech tokenizer with the downstream flow-matching acoustic decoder during training. An autoregressive language model predicts the resulting discrete representation, supporting text-to-speech and voice conversion without separating token design from acoustic reconstruction.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
This Thai speech synthesizer encodes phonemes and tones with a language-specific BERT model, then predicts duration, pitch and energy for a GAN-trained waveform decoder. A reference-derived style vector supports voice cloning. Multilingual pretraining of acoustic feature extractors and Thai adaptation address limited language-specific data.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
PilotTTS uses paired recordings and Q-Former conditioning to separate a speaker's identity from delivery style. The model supports reference cloning, emotional and nonverbal expression, and Chinese dialect synthesis within a shared autoregressive pipeline.

Figure 3 · Source
VITS voice models exported for local inference.
Editorial input/output diagram · Source
Pocket TTS uses continuous autoregressive speech modeling with a flow-based output mechanism, avoiding long sequences of discrete acoustic codebooks. It combines a small language-model backbone with streaming audio reconstruction and reusable voice conditioning. The project targets CPU-based speech synthesis; language-specific models and runtime choices affect its speed and voice behavior.

Figure 1 · Source
Variational acoustic model and flow-based post-net.

Figure 1, PDF p. 4 · Source
PROEMO combines emotional prompts with an explicit intensity control in a multi-speaker synthesizer. The conditioning adjusts delivery strength and prosodic variation, allowing the same spoken text to be rendered with different emotional performances.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This face-conditioned synthesizer combines local facial regions into progressively broader visual representations. Joint visual and acoustic attribute learning and multiple photographs of each training speaker align the face representation with voice characteristics, conditioning speech generation on text and a face image.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Prompt-Unseen-Emotion learns the relationship between emotion descriptions and speech using a language-model synthesis backbone. Weighted combinations of known emotions and contextual language knowledge allow expressive delivery outside the original categorical training labels.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Style and content text encoders with a speech decoder.

Figure 1 · Source
Prompt-conditioned TTS with a diffusion variation network.

Figure 1 · Source
ProtoDisent-TTS learns a codebook of healthy and dysarthric articulation patterns separately from speaker identity. Adversarial constraints reduce pathological information in the speaker representation, enabling controlled synthesis of articulation characteristics in a target voice.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
PS-TTS uses vowel-based alignment to coordinate the timing and phonetic structure of dubbed speech. Its PS-Comet variant also considers semantic preservation when choosing translated text, connecting translation choices with a TTS rendering stage.
Paper · GitHub: no author-linked repository found · Details

Fig. 1 (paper page 3) · Source
QTTS predicts residual speech codes produced by its QDAC tokenizer. Hierarchical parallel and delayed multihead variants organize codebook dependencies differently, offering alternative balances between acoustic detail and sequential decoding cost.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
Qwen-Audio-3.0-TTS combines compact semantic speech tokens with progressively trained language and acoustic models. Natural-language instructions and inline tags control delivery, while multilingual reference conditioning supports voice cloning and longer speech generation under varied recording conditions.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 3 · Source
Qwen3-TTS combines a dual-track speech language model with tokenizers designed for compact streaming audio. The released 12Hz line separates Base voice cloning, CustomVoice preset-speaker control and VoiceDesign creation from descriptions. These variants support different conditioning interfaces, allowing applications to choose between reproducing a reference voice and directing a new voice through text.

Figure 3 · Source
RADKA-CSS retrieves dialogue examples related to the current conversation in both meaning and delivery. A graph-based aggregation mechanism combines their style information with current context, conditioning expressive conversational speech synthesis.
Paper · GitHub · Project · Details

Figure 2 · Source
Prosody-guided codec language modeling.

Figure 1 · Source
Raon-OpenTTS is a family of reference-conditioned diffusion synthesizers trained on a large, documented English speech collection. The release pairs its models with data processing and evaluation resources, allowing robustness across varied acoustic conditions to be examined alongside clean-speech quality.

Figure 1 · Source
RapFlow-TTS regularizes the acoustic velocity field so longer generation steps remain consistent. Time-interval scheduling and adversarial objectives improve the resulting few-step synthesizer, reducing the iterations needed to render reference-conditioned speech.

Figure 1 · Source
ReGenVoice applies the ReGen representation-and-waveform modeling approach to text-to-speech. Multiple levels of generated conditioning help reconstruct detailed waveforms from compressed latents, linking efficient acoustic representation with reference-conditioned speech synthesis.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
ReStyle-TTS changes vocal attributes relative to a reference recording. Independent text and reference guidance, composable style adapters and timbre-consistency optimization allow continuous expressive edits while limiting changes to the speaker's identity.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
RTFree-F5 replaces the transcript normally associated with an F5-TTS reference recording with projected speech features. A lightweight adapter reuses the pretrained generator, enabling transcript-free voice conditioning, including references whose pronunciation makes transcription unreliable.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Revival with Voice learns voice identity from face images and delivery attributes from descriptions. Audio-only training data and stylized portrait augmentation broaden its input coverage, enabling controlled speech from real faces or artistic portraits.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
RWKVTTS uses the recurrent RWKV-7 architecture for speech synthesis in place of a conventional Transformer backbone. Its token-generation path targets efficient streaming and reduced state-management cost while conditioning audio on the supplied text.

Figure 2 · Source
S5-TTS adapts T5-TTS for word-by-word synthesis using limited future text. Lookahead-aware masks, convolutional auxiliary attention and distillation let the model start speaking before the complete sentence is available while retaining voice conditioning and alignment.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
Sarashina2.2-TTS emphasizes reliable Japanese pronunciation, including characters with several possible readings. Its semantic language model and acoustic flow decoder use reference speech for voice conditioning, with balanced multilingual training to reduce dependence on the reference language.

Figure 1 · Source
SASLM derives expressive intent from its own evolving semantic states through an information bottleneck. Acoustic feedback aligns generated speech with that intent, reducing the need for externally supplied emotion labels in context-sensitive speech rendering.
Paper · GitHub · Project · Details

Figure 3 · Source
Autoregressive speech foundation model.

Figure 1 · Source
This zero-shot synthesizer learns linguistic content and reference-speaker attributes through separate representations. Two-stage self-distillation creates aligned examples that strengthen their separation, targeting stable voice cloning with a small inference footprint.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
SelfTTS learns separate representations of a speaker's identity and emotional delivery using contrastive and adversarial objectives. It then improves synthesis through self-generated training examples, allowing emotion transfer to speakers originally recorded with neutral expression.
Paper · GitHub · Project · Details

Figure 1 · Source
SemaVoice organizes its audio VAE latents using guidance from speech foundation-model representations. A continuous autoregressive backbone and patch-level diffusion head then synthesize reference-conditioned speech with greater emphasis on linguistic coherence.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
SemBridge uses discrete semantic targets during training to organize both acoustic latents and language-model hidden states. The resulting continuous generator supports zero-shot speech synthesis with stronger content alignment, without needing to generate the auxiliary semantic tokens during inference.

Figure 1 · Source
Shallow Flow Matching adds a lightweight head that predicts an intermediate acoustic state for a flow-based synthesizer. Starting refinement closer to the target reduces the remaining generation path, allowing coarse-to-fine speech synthesis with less iterative work.

Figure 2 · Source
SLED learns the conditional distribution of acoustic latents using an energy-distance objective rather than discrete token classification. Its autoregressive generator samples continuous speech representations, simplifying synthesis while retaining acoustic detail.

Figure 2 · Source
SlimSpeech reduces the parameter count of a rectified-flow TTS model and transfers knowledge into a lightweight generator. It targets efficient reference-conditioned speech synthesis while retaining the acoustic quality of a larger teacher.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
SMLLE uses a transducer to align incoming text with semantic speech tokens and duration information. A separate autoregressive stage generates acoustic frames, with controlled access to future text stabilizing incremental synthesis.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
SoulX-Podcast synthesizes conversational scripts with reference voices, dialect choices and nonverbal expression. Its long-form training targets consistent speaker identity and natural transitions across turns, while also supporting ordinary single-speaker TTS.
Paper · GitHub · Project · Details

Figure 3 · Source
Spark-TTS uses BiCodec to separate changing linguistic content from global speaker attributes, then predicts these tokens with a Qwen2.5 backbone. That separation supports both reference-based cloning and direct control of attributes such as speaking rate and pitch. It is useful for controllable speech generation where a reference recording alone is insufficient to specify the desired delivery.

Figure 3 · Source
SpeakStream trains on text interleaved with corresponding speech and generates audio as new text becomes available. The synthesis module remains compatible with an upstream text-streaming language model, supporting responsive conversational playback.
Paper · Project · GitHub: no author-linked repository found · Details

Paper figure · Source
Text-to-semantic and semantic-to-acoustic LMs.

Figure 1 · Source
SpeechAccentLLM uses a content tokenizer trained with transcription alignment and jointly learns accent conversion and synthesis. A reconstruction refinement stage improves generated speech while separating accent-related changes from the target speaker's identity.
Paper · GitHub: no author-linked repository found · Details

Figure 3 · Source
SpeechEdit combines text, instruction tokens and reference audio in a shared codec-language-model sequence. Paired examples that differ in selected attributes teach localized changes while retaining other aspects of the reference voice and delivery.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Shared encoder–decoder with modality interfaces.

Figure 2 · Source
Prompted neural codec language model.

Figure 1 · Source
Residual convolutional acoustic model with duration expansion.

Figure 3 · Source
Spotlight-TTS extracts expressive reference information primarily from voiced speech and adjusts the resulting style direction before acoustic generation. The method targets more faithful expression transfer; the catalog combines its related paper records under one model.
Paper 1 · Paper 2 · GitHub: no author-linked repository found · Details

Figure 1 · Source
StellarTTS encodes phoneme timing sparsely and uses a lightweight masked Transformer to generate speech tokens in parallel. A semantic-aware codec supports waveform reconstruction, while explicit timing representations provide control over pronunciation, duration and prosody.
Paper · Project · GitHub: no author-linked repository found · Details

Paper figure · Source
Step-Audio-EditX supports reference-based synthesis and repeated edits to emotion, speaking style or nonverbal delivery. Training on deliberately contrasting synthetic examples teaches the model to follow expressive changes without a separate attribute-embedding module.

Figure 2 · Source
Step-Audio-TTS is the compact synthesis component produced through the broader Step-Audio speech-data and distillation pipeline. Text and voice conditioning drive speech-token generation, while the released workflow exposes delivery controls. The TTS checkpoint is used to render supplied content; reasoning, tool use and dialogue management belong to other components of the system.

Figure 2 · Source
The TTS mode of StepAudio 2.5 uses a shared speech-language foundation with synthesis-specific decoding and preference training. Rich contextual supervision and feedback target controllable expression; this card describes its speech-rendering path within the broader system.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
This synthesizer predicts continuous speech latents using a Gaussian-mixture conditional distribution. A stochastic monotonic alignment mechanism keeps the acoustic sequence ordered against the text, offering an alternative to autoregressive discrete-codec modeling.
Paper · GitHub: no author-linked repository found · Details

Figure 1 · Source
StreamMel alternates text tokens with continuous acoustic frames in one streaming synthesis model. This organization lets incoming text guide speech immediately while retaining reference-voice information across the generated audio stream.
Paper · GitHub: no author-linked repository found · Details

Paper figure · Source
Style-conditioned parallel synthesis with a transferable aligner.

Figure 1 · Source
Style diffusion and adversarial training with speech-model discriminators.

Figure 1, PDF p. 4 · Source
The SupertonicTTS research system compresses speech into continuous latents and predicts them from character-level text with flow matching. ConvNeXt blocks, temporal compression and a separate duration predictor keep synthesis compact. The later Supertonic ONNX release exposes preset voice-style assets; its packaged configurations should not be equated with the paper's 44M-parameter research model.
Model card · Paper · GitHub · Project · Details

Figure 1 · Source
Supertonic 2 extends the local ONNX synthesis line to five languages while retaining voice-style conditioning and a compact model. It provides a practical path to multilingual narration on devices that can run the supplied inference stack. Creating a new voice-style asset is a separate workflow from generating speech with an existing asset.
Editorial input/output diagram · Source
Supertonic 3 expands language coverage and adds expression tags while keeping local ONNX inference and preset voice styles. The release targets more reliable reading across short and long text, with controls for events such as breaths or laughter. Custom voice-style creation is offered through a separate service; downloaded styles can then condition local synthesis.
Editorial input/output diagram · Source
SwanVoice generates monologues or dialogues with up to four speakers using raw text, voice references and speaker-turn conditions. Pause markers and optional pronunciation substitutions provide textual control, while staged dialogue training supports longer expressive speech.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 2 · Source
SyncSpeech uses a temporal masking scheme to coordinate sequential speech structure with parallel token decoding. This hybrid organization targets faster first audio and higher throughput while retaining reference-conditioned synthesis quality.
Paper · Project · GitHub: no author-linked repository found · Details

Figure 1 · Source
Attention-based recurrent spectrogram synthesis.

Figure 1 · Source
Recurrent attention model and WaveNet vocoder.

Figure 1 · Source
TADA aligns text tokens one-to-one with continuous acoustic units. A language model with a flow-matching head predicts these synchronized representations, reducing the ambiguity of text-speech alignment during reference-conditioned synthesis.

Figure 2 · Source
TED-TTS modifies conditioning and decoding in a pretrained zero-shot synthesizer to control different parts of an utterance. Segment-specific emotion masks and duration steering support local changes while coordinating transitions and the overall stopping point.

Figure 1 · Source
Tibetan-TTS adapts a large speech generator through language-specific text representation, tokenizer changes and cross-language training. Its data preparation and quality enhancement target scarce Tibetan recordings and the differences between written forms and spoken pronunciation.
Paper · GitHub: no author-linked repository found · Details

Figure 2 · Source
<a
Truncated — view the full README on GitHub.
6 commits
Python
100.0%