A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
1,196
stars
1,658
commits
Python
primary language
Sep 5, 2026
updated
Universal multi-engine TTS extension for ComfyUI - evolved from the original ChatterBox Voice project.
A comprehensive ComfyUI extension providing unified Text-to-Speech, Voice Conversion, Audio Editing, and integrated RVC model training through multiple engines including ChatterboxTTS, DramaBox, F5-TTS, Higgs Audio 2, Higgs Audio v3, Step Audio EditX, MOSS-TTS, Echo-TTS, and RVC (Real-time Voice Conversion), with modular architecture designed for extensibility, runtime isolation for fragile legacy stacks, and a modern Transformers 5 main environment.
Subtitle workflows are still a core focus: the suite can transcribe to SRT, rebuild subtitles from edited transcripts, or estimate fresh SRT timing from plain text using the same advanced readability rules, while preserving project control tags for downstream TTS.
| Engine | Languages | Size | Key Features |
|---|---|---|---|
| F5-TTS | 🇺🇸🇩🇪🇪🇸🇫🇷🇮🇹🇯🇵 +4 | ~1.2GB each | Targeted Word/Speech Editing, Speed control |
| ChatterBox | 🇺🇸🇩🇪🇫🇷🇮🇹🇯🇵🇰🇷 +4 | ~4.3GB | Expressiveness slider |
| ChatterBox 23L | 🌐 24 languages | ~4.3GB | V1, V2, and V3 official checkpoints |
| VibeVoice | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +21 | 5.4GB / 18GB | 90-min long-form, Native 4-speaker (Base models) |
| Higgs Audio 2 | 🇺🇸🇨🇳🇩🇪🇪🇸🇰🇷 | ~9GB | 3 multi-speaker, CUDA graphs (55+ tokens/sec) |
| Higgs Audio v3 | 🌐 100+ languages | ~8GB | Native inline emotion/style/prosody/SFX tags |
| IndexTTS 2 / 2.5 | 🇺🇸🇨🇳🇪🇸🇯🇵🇸🇦 | ~4.7GB / ~5.49GB | Emotion Control: 8 vectors, Text as reference |
| CosyVoice3 | 🇺🇸🇨🇳🇯🇵🇰🇷 | ~5.4GB | Paralinguistic tags |
| Qwen3-TTS | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +4 | ~3-6GB | Voice design, ASR (Automatic Speech Recognition) |
| Granite ASR | 🇺🇸🇩🇪🇪🇸🇫🇷🇯🇵🇵🇹 | ~4.6GB | Native speaker attribution / diarization (plus model variant), Native word-level timestamps (plus model variant) |
| Step Audio EditX | 🇺🇸🇨🇳🇯🇵🇰🇷 | ~7GB | Second Pass Speech Editing Node: 14 emotions, 32 speaking styles |
| Echo-TTS | 🇺🇸 | ~5.3GB + ~1.8GB | Diffusion-based (~30s best), Force Speaker KV (speaker drift control) |
| Fish Audio S2 Pro | 🌐 80+ languages | ~10.3GB / ~8.0GB | Free-form sub-word emotion/prosody tags, Native multi-speaker and multi-turn dialogue with dynamic speaker references |
| Dots TTS | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +13 | ~6GB | Official auto language detect / language control, SOAR and MeanFlow distilled variants |
| DramaBox | 🇺🇸 | ~16.4GB | Expressive scene prompting and stage directions, Native and SRT-aware duration targeting |
| OmniVoice | 🌐 600+ languages | ~3.7GB | Inline non-verbal tags and pronunciation overrides, Reference-free voice design |
| MOSS-TTS | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +18 | ~8.5GB tokenizer + ~6.1GB/17GB/18GB model | Reference-free voice design with MOSS-VoiceGenerator, Native 1-5 speaker TTSD dialogue |
| MOSS-SoundEffect v2 | 🇺🇸🇨🇳 | ~11.2GB | Durations up to 30 seconds, Native negative prompting, CFG, flow shift, and diffusion-step controls |
| RVC | 🌐 Any | 100-300MB | Real-time VC, Integrated training workflow |
📊 Full comparison tables → | Language matrix → | Feature matrix → | Model download sources → | Model folder layouts →
Note: These tables are generated automatically from source: tts_audio_suite_engines.yaml
🎭 ChatterBox Voice Era 🌟 Multi-Engine Era
| |
v1.0 ───────────► v1.1 ────────► v2.0 ──────────► v3.0 ─────────┐
Jun 25 Jun 25 Jun 25 Jul 25 │
│ │ │ │ │
Foundation SRT Modular F5-TTS + │
ChatterBox Subtitles Structure Audio │
Voice Cloning Timing Node Refactor Analyzer │
▼
v3.4 ◄──────────────── v3.2 ◄──────────────── v3.1 ◄────────────┘
Jul 25 Jul 25 Jul 25
│ │ │
Language Pause Character
Switching Tags Switching
[German:Bob] [pause:1s] [Alice]
│
│ ⚙️ TTS Audio Suite Era
▼ |
v4.0 ──────────► v4.3 ──────────► v4.4 ────────► v4.5 ──────────┐
Aug 25 Aug 25 Aug 25 Aug 25 │
│ │ │ │ │
BREAKING! RVC + Silent Higgs Audio 2 │
Project Voice Speech New TTS Engine │
Renamed Conversion Analyzer Voice Cloning │
TTS Audio Suite + Streaming │
▼
v4.9 ◄─────────────── v4.8 ◄────────────────── v4.6 ◄───────────┘
Sep 25 Sep 25 Aug 25
│ │ │
IndexTTS-2 Chatterbox VibeVoice
Emotion Multilingual New TTS Engine
Control Official (23-lang) 90min Generation
│
│ 🎨 Inline Editor Tags Era
▼ |
v4.12 ──────────────► v4.15 ────────────► v4.16 ───────────────┐
Oct 25 Dez 25 Dez 25 │
│ │ │ │
Per-Seg Parameter Step Audio EditX CosyVoice3 │
Switching [seed:24] Inline Edit tags TTS + VC │
<laughter:2> │
▼
v4.24◄─────────────── v4.22 ◄─────────────── v4.19 ◄───────────┘
Mar 26 Mar 26 Jan 26
│ │ │
Text to SRT Echo-TTS Qwen3-TTS
Builder English TTS TTS + ASR + VoiceDesign
|
|─── 🎓 Training Support Era 🧱 Runtime Isolation T5 Era
▼ |
v4.25 ──────────────► v4.26 ────────────► v5.00 ───────────────┐
Apr 26 May 26 Jun 26 │
│ │ │ │
RVC MOSS-TTS Transformers 5 │
Model Training Higgs Audio v3 TTS │
│
▼
v5.3 ◄─────────────── v5.2 ◄─────────────── v5.1 ◄─────────────┘
Jun 26 Mar 26 Jan 26
│ │ │
Native SRT Duration OmniVoice TTS Dots TTS
Granite ASR
Visual Tag Builder
│
▼
v5.4 ───────────────────────────────► v5.5
Jul 26 Jul 26
│ │
Fish Audio S2 Pro MOSS-TTS v1.5
IndexTTS-2 Emotion Blending Sound Effects
Faster Tag Editor Voice Designer
Character Alias Manager
Want to add support for a new TTS engine, Voice Changer, ASR, or special audio model?
Start with the New Engine Guide Hub. It is written for users guiding an LLM through the process: first research the official model, then check existing ComfyUI implementations, decide scope, implement in the suite architecture, and run the parity checklist before PR review. For TTS engines, Unified TTS Text and Unified SRT TTS are a required pair.
The "ChatterBox SRT Voice TTS" node allows TTS generation by processing SRT content (SubRip Subtitle) files, ensuring precise timing and synchronization with your audio.
Key SRT Features:
smart_natural Timing Mode: Intelligent shifting logic that prevents overlaps and ensures natural speech flowAdjusted_SRT Output: Provides actual timings for generated audio for accurate post-processingFor comprehensive technical information, refer to the SRT_IMPLEMENTATION.md file.
This is the new architectural baseline for the suite.
This matters because the suite now has a clearer split:
NEW: DramaBox is integrated as an English expressive TTS engine for both Unified TTS Text and Unified SRT TTS.
torch.compileImportant limitations:
See the DramaBox Prompting Guide for prompt syntax, controls, memory modes, duration behavior, and examples. See the DramaBox LoRA Training Guide for dataset formats, training workflow, adapter loading, and CPU-safe preflight.
NEW in v4.4.0: Video analysis and mouth movement detection for silent video processing!
Perfect for:
Important Notes:
NEW in v4.5.0: State-of-the-art voice cloning technology with advanced neural voice replication!
Key Capabilities:
[CharacterName] syntax for multi-speaker dialoguesTechnical Features:
ComfyUI/models/TTS/HiggsAudio/ structureQuick Start:
Higgs Audio Engine node to configure voice cloning parametersTTS Text or TTS SRT node for generationPerfect for:
NEW: Higgs Audio v3 is now integrated as a native main-environment engine on the modern Transformers 5 stack.
<|emotion:amusement|>, <|style:whispering|>, <|prosody:pause|>, and <|sfx:laughter|><emotion:amusement>-style input and normalizes it internally to the official Higgs formatHiggs Audio v3 inline tags mode instead of pretending all inline systems are Step Audio EditXImportant behavior note:
Good fit for:
[Alice], [Bob] character tags with voice files from the voices folder - supports unlimited characters with pause tags and per-character control[Character] tag auto-conversion and manual "Speaker 1: Hello" format for up to 4 speakersTechnical Features:
Quick Start:
⚙️ VibeVoice Engine node to configure model and multi-speaker modeTTS Text or TTS SRT node for generationVibeVoice and KugelAudio run directly in the suite's Transformers 5 main environment.
Perfect for:
NEW in v3.1.0: Seamless character switching for both F5TTS and ChatterBox engines!
[CharacterName] tags to switch between different voices[Alice] instead of [female_01] with #character_alias_map.txt📖 Complete Character Switching Guide
Example usage:
Hello! This is the narrator speaking.
[Alice] Hi there! I'm Alice, nice to meet you.
[Bob] And I'm Bob! Great to meet you both.
Back to the narrator for the conclusion.
NEW in v3.4.0: Seamless language switching using simple bracket notation!
[language:character] tags to switch languages and models automatically[German:Alice], [Brazil:Bob], [USA:], [Portugal:] - no need to remember language codes![fr:Alice], [de:Bob], or [es:] (language only) patternsSupported Languages:
Example usage:
Hello! This is English text with the default model.
[de:Alice] Hallo! Ich spreche Deutsch mit Alice's Stimme.
[fr:] Bonjour! Je parle français avec la voix du narrateur.
[es:Bob] ¡Hola! Soy Bob hablando en español.
Back to English with the original model.
Advanced SRT Integration:
1
00:00:01,000 --> 00:00:04,000
Hello! Welcome to our multilingual show.
2
00:00:04,500 --> 00:00:08,000
[de:female_01] Willkommen zu unserer mehrsprachigen Show!
3
00:00:08,500 --> 00:00:12,000
[fr:] Bienvenue à notre émission multilingue!
NEW: Progressive voice refinement with intelligent caching for instant experimentation!
How it works:
Intelligent caching examples:
Practical tip: Start with 1 pass, then test 2-5 passes to find the sweet spot for your audio. More passes can improve voice similarity, but there is no universal best value.
NEW in v4.1.0: Professional-grade Real-time Voice Conversion with .pth character models!
📖 See Model folder layouts for detailed setup paths
How it works:
NEW: Integrated RVC training inside the suite, using the same unified node style as the rest of the project instead of a detached external workflow.
🎓 Model Training accepts TTS_ENGINE and routes by engine type📦 RVC Dataset Prep handles dataset path/zip/folder upload or direct audio input, slicing, HuBERT features, F0 extraction, and reusable prep caches🎛️ RVC Training Config exposes practical training controls with tooltip guidance instead of raw upstream garbagecontinue_from for training further from a finished RVC model/artifactpretrained_v2 RVC training init checkpoints are auto-managedCurrent scope:
Typical flow:
⚙️ RVC Engine📦 RVC Dataset Prep🎛️ RVC Training Config🎓 Model Training🎭 Load RVC Character ModelImportant notes:
ComfyUI/output/tts_audio_suite_training/rvc/.pth models and .index files go under ComfyUI/models/TTS/RVC/save_best_model is only a low-loss inference candidate, not a magical quality oracle. You still need to listen.NEW: Intelligent pause insertion for natural speech timing control!
[pause:1.5], [wait:2s], [stop:3][pause:500ms], [wait:1200ms], [stop:800ms]pause, wait, stop (all work identically)Example usage:
Welcome to our show! [pause:1s] Today we'll discuss exciting topics.
[Alice] I'm really excited! [wait:500ms] This will be great.
[stop:2] Let's get started with the main content.
NEW in v4.6.29: ChatterBox TTS now supports 11 languages with community-finetuned models and automatic model management!
Supported Languages:
<haha>, <wow> tags[it] prefix for Italian textKey Features:
Usage: Select language from dropdown → First generation downloads model → Subsequent generations use cached model
NEW in v4.8.0: Official ResembleAI Chatterbox Multilingual TTS model with native support for 23 languages!
The Chatterbox Multilingual TTS (referred to internally as "ChatterBox Official 23-Lang" to distinguish from community models) is ResembleAI's first production-grade open-source TTS model supporting 23 languages out of the box. This is the official successor to the original ChatterBox model with enhanced multilingual capabilities.
🎯 Key Advantages over Community Models:
🌍 Supported Languages (23 total + Vietnamese community finetune):
Arabic (ar), Danish (da), German (de), Greek (el), English (en), Spanish (es), Finnish (fi), French (fr), Hebrew (he), Hindi (hi), Italian (it), Japanese (ja), Korean (ko), Malay (ms), Dutch (nl), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Swedish (sv), Swahili (sw), Turkish (tr), Chinese (zh)
🇻🇳 Vietnamese (Viterbox): Community finetune by Dolly AI 23 with expanded Vietnamese tokenization (dolly-vn/viterbox) - select from model version dropdown
🔧 Fully Integrated Features:
[CharacterName] support with per-character voice references[language:character] syntax with intelligent parameter switching[pause:Ns] support with character voice inheritance🆚 vs Community Models:
| Feature | Chatterbox Multilingual TTS | Community Models |
|---|---|---|
| Languages | 23 native languages | 11 finetuned variants |
| Model Loading | Single model, parameter switching | Separate model per language |
| Voice Cloning | Zero-shot across all languages | Per-model training |
| Official Support | ✅ ResembleAI official | Community maintained |
| Character Integration | ✅ Full integration | ✅ Full integration |
| SRT Support | ✅ Advanced timing modes | ✅ Advanced timing modes |
| Performance | Optimized single-model | Multiple model overhead |
🎭 Character Example:
[En:Alice] Hello everyone! [De:Hans] Guten Tag! [Es:Maria] ¡Hola! [pause:2s] [En:Alice] That was amazing multilingual switching!
This creates seamless multilingual character switching with proper voice inheritance and pause support - all within a single model.
🎭 NEW: v2 Special Emotion & Sound Tokens 🚧 Experimental
ChatterBox v2 vocabulary includes 30+ special tokens for emotions, sounds, and vocal effects. Note: These are experimental - tokens may produce minimal or no audible effects. ResembleAI has not officially documented their usage (see issue #186).
Try angle brackets <emotion> to experiment:
[Alice] Hello! <laughter> hahaha. [pause:0.5] <whisper> This might work slightly.
Available v2 Tokens:
<giggle>, <laughter>, <sigh>, <cry>, <gasp>, <groan><whisper>, <mumble>, <singing>, <humming><cough>, <sneeze>, <sniff>, <inhale>, <exhale>Model Version Selection:
Both versions fully support character switching, language switching, and pause tags. The v2 special tokens are experimental with limited effectiveness - our implementation is ready for when/if ResembleAI improves this feature. The angle bracket syntax <emotion> avoids conflicts with character tags [Name] and pause tags [pause:1s].
NEW in v4.3.0: Complete architectural overhaul implementing universal streaming system with parallel processing capabilities!
Key Features:
batch_size parameterPerformance Notes:
batch_size=0 for optimal performance (sequential processing)batch_size > 1 enables parallel workers but typically slower due to GPU inference characteristicsNEW in v4.9.0: Revolutionary IndexTTS-2 engine with advanced emotion control and dual-source emotion blending!
emotion_control and audio references to emotion_audio; both can be used together{seg} template processing for contextual per-segment emotionsemotion_audio as an emotion reference for natural expressionopt_narrator on emotion_audio, including per-character [Character:emotion_ref] references[Character:emotion_ref] syntax, blendable with vector/text emotionduration_factor scales the internal semantic feature sequence (0.5 shorter/faster, 1.0 unchanged, 2.0 longer/slower). It is not natural prosody or exact-duration planning, does not apply to 2.0, and is not used by SRT native-duration targeting<word|pronunciation> annotations through suite text processing2.0 versus 2.5: Treat 2.5 as a multilingual/efficiency alternative, not an automatic voice-cloning quality upgrade. In our manual listening, legacy 2.0 preserved speaker resemblance better when transferring a strong emotion from a different reference voice; 2.5 may still be preferable for Japanese, Spanish, Arabic, or cross-lingual generation. Strong external emotion settings can reduce perceived speaker identity, so compare both models for the target voice.
Key Features:
{seg} placeholder for contextual emotion analysis (e.g., "Worried parent speaking: {seg}")Example Usage:
Welcome to our show! [Alice:happy_sarah] I'm so excited to be here!
[Bob:angry_narrator] That's completely unacceptable behavior.
Perfect for:
NEW in v4.15: Revolutionary LLM-based audio post-processing with emotion, style, and paralinguistic control!
<Laughter:2>, <emotion:happy>, <style:whisper> tagsKey Features:
<Laughter:2|emotion:happy|style:whisper>Example Usage:
[Alice] Hello there <Laughter:2> my friend! <emotion:happy>
[Bob] Listen carefully <style:whisper|speed:slower>, this is important.
[Alice] I'm laughing so hard <Laughter:3> <restore>
Perfect for:
⚠️ Important Notes:
<restore> tag to recover original voice characterNEW in v4.16: Alibaba's fast multilingual voice cloning with native paralinguistic tags, instruct mode, and zero-shot voice conversion!
<breath>, <laughter>, <cough>, <sigh>, <laughing>text</laughing> - processed during generation → 📖 Tags GuideKey Features:
<breath>, <laughter>, <cough>, <sigh>, <gasp>, <laughing>text</laughing>, <strong>text</strong>[CharacterName] support with per-character voice references[en:], [zh:], [ja:], [ko:] bracket syntax or native <|en|> tags[pause:Ns] support for natural speech timingExample Usage:
[Alice] Hello everyone! This is zero-shot voice cloning.
[Bob] 你好!我说普通话。[pause:1s] 还可以说方言。
Instruct Mode Examples:
# Speak with Cantonese dialect
Instruct: 请用广东话表达。
# Speak with excitement
Instruct: 用兴奋的语气说话。
Perfect for:
NEW in v4.19: Alibaba's Qwen3-TTS with 3 distinct TTS model types - CustomVoice presets, dedicated text-to-voice design, and zero-shot voice cloning. The engine's model dropdown exposes every checkpoint and marks installed checkpoints with a local: prefix. Model-specific controls appear only when they apply.
NEW: ✏️ Unified ASR Transcribe support now includes Qwen3-ASR and Granite ASR, giving the suite a second ASR engine option with optional custom timestamps/SRT for Granite via the reused Qwen forced aligner. Granite 4.1 plus also adds native speaker diarization and native word timestamps.
Model Types:
🎭 CustomVoice Model (0.6B / 1.7B): 9 preset multilingual speakers (Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee)
✍️ VoiceDesign Model (1.7B only): Dedicated Qwen voice creation from text descriptions
🎤 Base Model (0.6B / 1.7B): Zero-shot voice cloning from 3-30s reference audio
🔤 ASR: Qwen3-TTS Engine can be connected to the ✏️ Unified ASR Transcribe node for transcription
⚠️ Style instructions only work with CustomVoice and VoiceDesign. Voice cloning (Base) ignores the instruction field entirely.
Technical Specs:
Unified Features Support:
Voice Designer Node:
The shared designer accepts Qwen3-TTS, MOSS-TTS, or OmniVoice engine configurations and outputs the same NARRATOR_VOICE format. The voice-design instruction lives on 🎨 Voice Designer; the engine keeps model, language, and generation settings. Select Qwen VoiceDesign or MOSS VoiceGenerator in the engine's model dropdown, or set OmniVoice to Voice Design mode. The corresponding engine instruction stays visible but is disabled because it would be ignored. Incompatible modes stop with a direct correction message. OmniVoice's controlled tag vocabulary can still be assembled with 📐 Visual Tag Builder. Connect the resulting opt_narrator to 💾 Save Character Voice when persistence is wanted.
💾 Save Character Voice accepts only opt_narrator, keeping persistence separate from voice construction. For existing audio, use 🎭 Character Voices with the audio and its exact transcription, then connect its opt_narrator output to Save Character Voice. The save node writes the established three-file format—name.wav, name.reference.txt, and metadata in name.txt—under models/voices/.
Description: "A deep, authoritative male voice with clear articulation"
→ Voice generated and cached → Use in TTS Text/SRT nodes
Perfect for:
NEW in v5.x: OmniVoice is now integrated into the unified suite with native duration-aware SRT generation and a generalized visual tag workflow.
<> form like <laughter>, then converted internally for generation → 📖 OmniVoice Tags GuidePractical note:
Use the built-in OmniVoice preset in 📐 Visual Tag Builder for the canonical voice-design workflow. If you need a different tag schema, the same node now supports reusable custom presets with saved column order.
NEW in v4.26: OpenMOSS engine family integration with unified support for single-speaker TTS and native multi-speaker dialogue.
Model Variants:
MOSS-TTS-Local-TransformerMOSS-TTS-v1.5 — 31 languages and more stable cloninglaion/moss-tts-v1.5-8b-voice-actingMOSS-TTSMOSS-TTSD-v1.0MOSS-VoiceGenerator — select it in the MOSS engine for Voice DesignerMOSS-Audio-TokenizerCompatible community full checkpoints can also be placed in models/TTS/moss_tts/<model-name>/.
They are listed as local:<model-name> and classified from config.json; unsupported layouts fail explicitly.
Supported Native Input Forms (TTSD):
[Character] tags[1] / [S1] numeric speaker tagsSpeaker 1: ... formatAll native forms are normalized internally to canonical [S1]...[S5] dialogue.
Speaker Modes:
Important Native Compatibility Rule:
Native TTSD mode now hard-fails (explicit error popup) instead of silently switching models when these are detected:
[] parameter changesIf you need those controls, switch to Custom Character Switching and use MOSS-TTS-Local-Transformer, MOSS-TTS-v1.5, or MOSS-TTS.
Official Prompt Fields Exposed:
instruction, quality, sound_event, ambient_sound, language, duration_tokens
Per-segment overrides are supported with [] parameter syntax for whole-segment conditioning.
Documentation:
Training status:
🎓 Model Training flow.models/TTS/moss_tts/loras/.NEW in v4.23: ASR subtitle generation is now modular instead of being buried inside the transcriber.
The flow is now:
✏️ ASR Transcribe → optional 📝 ASR Punctuation / Truecase → 📺 Text to SRT Builder
This matters because the suite can now:
🔧 SRT Advanced Options now belongs to the builder stage instead of being mixed into ASR4.1 plus can emit suite-native speaker tags like [Speaker 1] for downstream TTS/alias workflows, and if you need both diarization and word timings the node automatically falls back to the reused Qwen forced alignerasr_timing_data disconnected and the builder estimates subtitle timings from plain text using the same SRT options that later shape the final cuesCurrent intended use:
✏️ ASR Transcribe for timed transcription with Qwen3-ASR or Granite ASR📝 ASR Punctuation / Truecase mainly for low-punctuation ASR outputs like Granite📺 Text to SRT Builder to turn cleaned text + ASR timing data into final SRTGranite note:
granite-speech-4.1-2b keeps Japanese supportgranite-speech-4.1-2b-plus adds native diarization and native word timestamps, but drops JapaneseWorkflow example:
Use the new Unified ✏️ ASR Transcribe + SRT Builder workflow for both Granite ASR and Qwen3-ASR examples.
NEW in v4.22: Echo-TTS DiT-based voice cloning with reference audio support.
Usage:
⚙️ Echo-TTS Engine node🎤 TTS Text or 📺 TTS SRTComfyUI/models/TTS/echo-tts-base/📖 See Model folder layouts for detailed setup paths
NEW in v4.10.0: Universal multilingual text preprocessing node for improved TTS pronunciation quality across languages!
Perfect for:
⚠️ Experimental Feature: This is a new experimental feature - user feedback needed to validate pronunciation improvements across different languages. Please test with your target languages and report results!
📖 Try the workflow: F5 TTS integration + 📝 Phoneme Text Normalizer
NEW in v4.12.0: Fine-grained per-segment control over TTS generation parameters across all engines with an interactive tag editor!
Beyond character switching and language control, you can now override generation parameters (seed, temperature, CFG, speed, etc.) on a per-segment basis using inline tags. The new 🏷️ Multiline TTS Tag Editor node makes building complex tags easier and more visual with:
{seg} dynamic emotion controls directly from the Inline Tags panelThis enables dynamic control over individual audio segments without modifying node defaults.
🎯 Key Features:
[Alice|seed:42|temp:0.5] or [seed:42|Alice|temp:0.5]temp (temperature), cfg_weight (cfg), exag (exaggeration)[SEED:42], [Seed:42], [seed:42] all work identically[de:Alice|seed:42|temp:0.7]💡 Real-World Examples:
[Alice|seed:42] This is the first segment of Alice.
[Alice|seed:42] This is the second segment, sounding identical.
[Bob|temperature:0.3] Bob speaks carefully and precisely.
[Bob|temperature:0.8] Bob speaks more creatively and varied.
[Narrator|temp:0.4] Important introduction - precise delivery.
[Narrator|temp:0.7] Creative narrative section - more varied!
🔧 Supported Parameters by Engine:
📖 Guides: Per-Segment Parameter Switching | Multiline TTS Tag Editor | OmniVoice Tags Guide
Perfect for:
[!TIP] You might want to create a clean portable installation just for TTS-Audio-Suite, instead of adding it to your regular ComfyUI setup. This can make the installation process smoother and may reduce initialization time.
One-click installation with intelligent dependency management:
Python 3.13 Support:
Same intelligent installer, manual setup:
Clone the repository
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
cd TTS-Audio-Suite
Run the intelligent installer:
ComfyUI Portable:
# Windows:
..\..\..\python_embeded\python.exe install.py
# Linux/Mac:
../../../python_embeded/python.exe install.py
ComfyUI with venv/conda:
# First activate your ComfyUI environment, then:
python install.py
The installer automatically handles all dependency conflicts and Python version compatibility.
Run your first workflow
Try a Workflow
Restart ComfyUI and look for 🎤 TTS Audio Suite nodes
🧪 Python 3.13 Users: Installation is fully supported! The system automatically uses OpenSeeFace for mouth movement analysis when MediaPipe is unavailable.
Need offline/manual setup? Use docs/MODEL_DOWNLOAD_SOURCES.md and docs/MODEL_LAYOUTS.md
This section provides a detailed guide for installing TTS Audio Suite, covering different ComfyUI installation methods.
ComfyUI installation (Portable, Direct with venv, or through Manager)
Python 3.12 or higher
Optional system libraries (Linux only):
# Ubuntu/Debian - Optional audio features
sudo apt-get install portaudio19-dev libsamplerate0-dev
# Fedora/RHEL
sudo dnf install portaudio-devel libsamplerate-devel
📋 Optional:
libsamplerate0-devprovides additional audio-resampling support.portaudio19-devenables voice recording. Missing either package no longer blocks installation of the TTS engines.
Optional macOS dependencies:
brew install portaudio
Windows: No additional system dependencies needed (libraries come pre-compiled)
For portable installations, follow these steps:
Clone the repository into the ComfyUI/custom_nodes folder:
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
Navigate to the cloned directory:
cd TTS-Audio-Suite
Run the install script to automatically handle all dependencies. Important: Use the python.exe executable located in your ComfyUI portable installation.
../../../python_embeded/python.exe install.py
The script will automatically install all required Python packages and detect any missing system dependencies.
If you have a direct installation with a virtual environment (venv), follow these steps:
Clone the repository into the ComfyUI/custom_nodes folder:
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
Activate your ComfyUI virtual environment. This is crucial to ensure dependencies are installed in the correct environment. The method to activate the venv may vary depending on your setup. Here's a common example:
cd ComfyUI
. ./venv/bin/activate
or on Windows:
ComfyUI\venv\Scripts\activate
Navigate to the cloned directory:
cd custom_nodes/TTS-Audio-Suite
Run the install script to automatically handle all dependencies:
python install.py
The script will automatically install all required Python packages and detect any missing system dependencies.
Install the ComfyUI Manager if you haven't already.
Use the Manager to install the "TTS Audio Suite" node.
The manager might handle dependencies automatically, but it's still recommended to verify the installation. Navigate to the node's directory:
cd ComfyUI/custom_nodes/TTS-Audio-Suite
Activate your ComfyUI virtual environment (see instructions in "Direct Installation with venv").
If you encounter issues, run the install script to manually install dependencies:
python install.py
Our install script automatically detects missing optional system libraries and will display feature warnings like:
[!] Optional system dependencies are missing
============================================================
OPTIONAL LINUX SYSTEM DEPENDENCIES
============================================================
• libsamplerate0-dev (optional additional audio-resampling support)
• portaudio19-dev (optional voice recording)
Please install with:
# Ubuntu/Debian:
sudo apt-get install libsamplerate0-dev portaudio19-dev
# Fedora/RHEL:
sudo dnf install libsamplerate-devel portaudio-devel
============================================================
Core TTS installation will continue; only the listed features may be unavailable.
A common problem is installing dependencies in the wrong Python environment. Always ensure you are installing dependencies within your ComfyUI's Python environment.
If the engine comparison table shows Shared or Dedicated in the Isolation column, that engine has its own secondary-environment path for dependency conflicts. Configure that on the engine node with ⚠️ Runtime Isolation instead of trying to downgrade your main ComfyUI environment.
Verify your Python environment: After activating your venv or navigating to your portable ComfyUI installation, check the Python executable being used:
which python
This should point to the Python executable within your ComfyUI installation (e.g., ComfyUI/python_embeded/python.exe or ComfyUI/venv/bin/python).
If s3tokenizer fails to install: This dependency can be problematic. Try upgrading your pip and setuptools:
python -m pip install --upgrade pip setuptools wheel
Then, try installing the requirements again.
If you cloned the node manually (without the Manager): Make sure you install the requirements.txt file.
To update the node to the latest version:
Navigate to the node's directory:
cd ComfyUI/custom_nodes/TTS-Audio-Suite
Pull the latest changes from the repository:
git pull
Reinstall the dependencies (in case they have been updated):
python install.py
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
Some dependencies, particularly s3tokenizer, can occasionally cause installation issues on certain Python setups (e.g., Python 3.10, sometimes used by tools like Stability Matrix).
To minimize potential problems, it's highly recommended to first ensure your core packaging tools are up-to-date in your ComfyUI's virtual environment:
python -m pip install --upgrade pip setuptools wheel
After running the command above, install the node's specific requirements:
python install.py
ChatterBox Voice now supports FFmpeg for high-quality audio stretching. While not required, it's recommended for the best audio quality:
Windows:
winget install FFmpeg
# or with Chocolatey
choco install ffmpeg
macOS:
brew install ffmpeg
Linux:
# Ubuntu/Debian
sudo apt-get install ffmpeg
# Fedora
sudo dnf install ffmpeg
If FFmpeg is not available, ChatterBox will automatically fall back to using the built-in phase vocoder method for audio stretching - your workflows will continue to work without interruption.
No manual downloads needed. All engines download required models automatically on first use.
For offline/manual setup:
| Engine | Primary model path | Auto-download | Notes |
|---|---|---|---|
| ChatterBox | ComfyUI/models/TTS/chatterbox/ | ✅ | Legacy ComfyUI/models/chatterbox/ still works |
| ChatterBox 23-Lang | ComfyUI/models/TTS/chatterbox_official_23lang/ | ✅ | v1/v2/v3 coexist in same folder |
| F5-TTS | ComfyUI/models/TTS/F5-TTS/ | ✅ | Optional Vocos and voice refs |
| Higgs Audio 2 | ComfyUI/models/TTS/HiggsAudio/ | ✅ | Generation + tokenizer |
| Higgs Audio v3 | ComfyUI/models/TTS/higgs_audio_v3/ | ✅ | Official 4B multilingual TTS model |
| VibeVoice | ComfyUI/models/TTS/VibeVoice/ | ✅ | 1.5B and 7B variants |
| RVC | ComfyUI/models/TTS/RVC/ | ✅* | Base models auto; character .pth can be user-provided |
| IndexTTS-2 | ComfyUI/models/TTS/IndexTTS/ | ✅ | Emotion components included |
| Step Audio EditX | ComfyUI/models/TTS/step_audio_editx/ | ✅ | Main model + tokenizer stack |
| CosyVoice3 | ComfyUI/models/TTS/CosyVoice/ | ✅ | Variant-specific lazy downloads |
| Qwen3-TTS / ASR | ComfyUI/models/TTS/qwen3_tts/ | ✅ | Per-variant download + shared tokenizer |
| MOSS-TTS | ComfyUI/models/TTS/moss_tts/ | ✅ | Local/Delay/VoiceGenerator/SoundEffect v1/TTSD models plus shared MOSS-Audio-Tokenizer codec |
| MOSS-SoundEffect v2 | ComfyUI/models/TTS/moss_soundeffect_v2/ | ✅ | Official v2 diffusion pipeline; configured ComfyUI environment |
| Granite ASR | ComfyUI/models/TTS/granite_asr/ | ✅ | Granite ASR models; plus adds native diarization/timestamps, optional Qwen forced aligner reused lazily for timestamps/SRT fallback |
| Echo-TTS | ComfyUI/models/TTS/echo-tts-base/ | ✅ | ~7.1GB total (base + dac); CC-BY-NC-SA |
| Dots TTS | ComfyUI/models/TTS/dots_tts/ | ✅ | Official base / soar / mf checkpoints with tokenizer, vocoder, speaker encoder |
| DramaBox | ComfyUI/models/TTS/dramabox/DramaBox/ | ✅ | ~16.4GB download; fast mode roughly 24GB VRAM; experimental FP8, staged, and sequential options can reduce VRAM, but no minimum GPU size is guaranteed; conditional LTX-2 Community License |
| Fish Audio S2 Pro | ComfyUI/models/TTS/fish_audio_s2_pro/ | ✅ | Official BF16 or optional community FP8 checkpoint; the official checkpoint can be quantized on load with BNB INT8/NF4; main T5 environment with process teardown for Clear VRAM; Fish Audio Research License |
| OmniVoice | ComfyUI/models/TTS/omnivoice/ | ✅ | Official OmniVoice model. Voice cloning in this suite requires explicit reference text. |
Generated from tts_audio_suite_engines.yaml.
If TTS Audio Suite has been helpful for your projects, consider supporting its development:
Your support helps maintain and improve this project for the entire community!
Ready-to-use ComfyUI workflows - Download and drag into ComfyUI:
| Workflow | Description | Features | Status | Files |
|---|---|---|---|---|
| Unified 📺 TTS SRT | Universal SRT processing with all TTS engines | • ChatterBox/F5-TTS/Higgs Audio 2 • Multiple timing modes • Multi-character switching • Overlap SRT support | ✅ New in v4.5 | 📁 JSON |
| Unified 🔄 Voice Changer | Modern voice conversion with multiple engines | • RVC + ChatterBox VC • Iterative refinement • Real-time conversion | ✅ Updated for v4.3 | 📁 JSON |
| Unified ✏️ ASR Transcribe + SRT Builder | Modular ASR + subtitle workflow | • Granite ASR + Qwen3 ASR examples • Separate transcription and SRT building • Works with the new Text to SRT Builder flow | ✅ New in v4.23 | 📁 JSON |
| Unified 🌩️ Sound Effects | Text-to-sound generation with compatible engines | • MOSS-SoundEffect v1 and v2 • Per-segment parameters and pauses • Long-duration chunking and audio cache | ✅ New | 📁 JSON |
| Unified 🎨 Voice Designer | Reference-free character voice creation | • Qwen3-TTS, MOSS-TTS, and OmniVoice • Free-form descriptions or Visual Tag Builder • Preview and save reusable character voices | ✅ New | 📁 JSON · 🖼️ Cover |
| Workflow | Description | Status | Files |
|---|---|---|---|
| 🤐 Voice Cleaning | Audio restoration & cleanup with dual tool pipeline | ✅ New in v4.13 | 📁 JSON |
| DramaBox LoRA 🎓 Model Training | DramaBox IC-LoRA training workflow from staged speech clips | ✅ New | 📁 JSON |
| MOSS LoRA 🎓 Model Training | Initial MOSS LoRA training workflow from clipped speech dataset | ✅ New in v4.27 | 📁 JSON |
| RVC 🎓 Model Training | RVC voice model training workflow | ✅ New in v4.25 | 📁 JSON |
| 🎨 Step Audio EditX - Audio Editor | Step Audio EditX audio editing with inline edit tags | ✅ New in v4.14 | 📁 JSON |
| ⚙️ Step Audio EditX Integration | Step Audio EditX TTS engine with zero-shot voice cloning | ✅ New in v4.14 | 📁 JSON |
| ⚙️ Higgs Audio v3 Integration | Higgs Audio v3 TTS with zero-shot voice cloning and native inline tags | ✅ New in v4.27 | 📁 JSON |
| ⚙️ OmniVoice Engine Integration | OmniVoice multilingual TTS with cloning, voice design, and native duration control | ✅ New in v4.28 | 📁 JSON |
| ⚙️ Fish Audio S2 Pro Integration | Fish S2 Pro multilingual cloning with native multi-speaker dialogue, inline control, and long-form generation | ✅ New in v5.3 | 📁 JSON |
| ⚙️ DramaBox Integration | DramaBox expressive scene prompting with native SRT duration targeting | ✅ New in v5.6 | 📁 JSON |
| 🌈 IndexTTS-2 Integration | IndexTTS-2 engine with advanced emotion control | ✅ New in v4.9 | 📁 JSON |
| 📝 F5 TTS + Text Normalizer | F5-TTS with multilingual text processing and phonemization | ✅ New in v4.10.0 | 📁 JSON |
| Qwen3 integration + ASR | Qwen3-TTS voice generation with ASR transcription | ✅ New in v4.21 | 📁 JSON |
| VibeVoice Integration | VibeVoice long-form TTS with multi-speaker support | ✅ Compatible | 📁 JSON |
| ChatterBox Integration | General ChatterBox TTS and Voice Conversion | ✅ Compatible | 📁 JSON |
| F5-TTS Speech Editor | Interactive waveform analysis for F5-TTS editing | ✅ Updated for v4 | 📁 JSON |
💡 Recommended: Use the new Unified 📺 TTS SRT workflow which showcases the unified TTS flow in one comprehensive workflow. It demonstrates SRT processing, timing modes, multi-character switching, and modern engine integration across the suite.
📥 Usage: Download the
.jsonfiles and drag them directly into your ComfyUI interface. The workflows will automatically load with proper node connections.
The TTS Audio Suite project code is licensed under MIT.
Python
93.1%
JavaScript
6.2%
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
1,196
stars
1,658
commits
Python
primary language
Sep 5, 2026
updated
Universal multi-engine TTS extension for ComfyUI - evolved from the original ChatterBox Voice project.
A comprehensive ComfyUI extension providing unified Text-to-Speech, Voice Conversion, Audio Editing, and integrated RVC model training through multiple engines including ChatterboxTTS, DramaBox, F5-TTS, Higgs Audio 2, Higgs Audio v3, Step Audio EditX, MOSS-TTS, Echo-TTS, and RVC (Real-time Voice Conversion), with modular architecture designed for extensibility, runtime isolation for fragile legacy stacks, and a modern Transformers 5 main environment.
Subtitle workflows are still a core focus: the suite can transcribe to SRT, rebuild subtitles from edited transcripts, or estimate fresh SRT timing from plain text using the same advanced readability rules, while preserving project control tags for downstream TTS.
| Engine | Languages | Size | Key Features |
|---|---|---|---|
| F5-TTS | 🇺🇸🇩🇪🇪🇸🇫🇷🇮🇹🇯🇵 +4 | ~1.2GB each | Targeted Word/Speech Editing, Speed control |
| ChatterBox | 🇺🇸🇩🇪🇫🇷🇮🇹🇯🇵🇰🇷 +4 | ~4.3GB | Expressiveness slider |
| ChatterBox 23L | 🌐 24 languages | ~4.3GB | V1, V2, and V3 official checkpoints |
| VibeVoice | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +21 | 5.4GB / 18GB | 90-min long-form, Native 4-speaker (Base models) |
| Higgs Audio 2 | 🇺🇸🇨🇳🇩🇪🇪🇸🇰🇷 | ~9GB | 3 multi-speaker, CUDA graphs (55+ tokens/sec) |
| Higgs Audio v3 | 🌐 100+ languages | ~8GB | Native inline emotion/style/prosody/SFX tags |
| IndexTTS 2 / 2.5 | 🇺🇸🇨🇳🇪🇸🇯🇵🇸🇦 | ~4.7GB / ~5.49GB | Emotion Control: 8 vectors, Text as reference |
| CosyVoice3 | 🇺🇸🇨🇳🇯🇵🇰🇷 | ~5.4GB | Paralinguistic tags |
| Qwen3-TTS | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +4 | ~3-6GB | Voice design, ASR (Automatic Speech Recognition) |
| Granite ASR | 🇺🇸🇩🇪🇪🇸🇫🇷🇯🇵🇵🇹 | ~4.6GB | Native speaker attribution / diarization (plus model variant), Native word-level timestamps (plus model variant) |
| Step Audio EditX | 🇺🇸🇨🇳🇯🇵🇰🇷 | ~7GB | Second Pass Speech Editing Node: 14 emotions, 32 speaking styles |
| Echo-TTS | 🇺🇸 | ~5.3GB + ~1.8GB | Diffusion-based (~30s best), Force Speaker KV (speaker drift control) |
| Fish Audio S2 Pro | 🌐 80+ languages | ~10.3GB / ~8.0GB | Free-form sub-word emotion/prosody tags, Native multi-speaker and multi-turn dialogue with dynamic speaker references |
| Dots TTS | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +13 | ~6GB | Official auto language detect / language control, SOAR and MeanFlow distilled variants |
| DramaBox | 🇺🇸 | ~16.4GB | Expressive scene prompting and stage directions, Native and SRT-aware duration targeting |
| OmniVoice | 🌐 600+ languages | ~3.7GB | Inline non-verbal tags and pronunciation overrides, Reference-free voice design |
| MOSS-TTS | 🇺🇸🇨🇳🇩🇪🇪🇸🇫🇷🇮🇹 +18 | ~8.5GB tokenizer + ~6.1GB/17GB/18GB model | Reference-free voice design with MOSS-VoiceGenerator, Native 1-5 speaker TTSD dialogue |
| MOSS-SoundEffect v2 | 🇺🇸🇨🇳 | ~11.2GB | Durations up to 30 seconds, Native negative prompting, CFG, flow shift, and diffusion-step controls |
| RVC | 🌐 Any | 100-300MB | Real-time VC, Integrated training workflow |
📊 Full comparison tables → | Language matrix → | Feature matrix → | Model download sources → | Model folder layouts →
Note: These tables are generated automatically from source: tts_audio_suite_engines.yaml
🎭 ChatterBox Voice Era 🌟 Multi-Engine Era
| |
v1.0 ───────────► v1.1 ────────► v2.0 ──────────► v3.0 ─────────┐
Jun 25 Jun 25 Jun 25 Jul 25 │
│ │ │ │ │
Foundation SRT Modular F5-TTS + │
ChatterBox Subtitles Structure Audio │
Voice Cloning Timing Node Refactor Analyzer │
▼
v3.4 ◄──────────────── v3.2 ◄──────────────── v3.1 ◄────────────┘
Jul 25 Jul 25 Jul 25
│ │ │
Language Pause Character
Switching Tags Switching
[German:Bob] [pause:1s] [Alice]
│
│ ⚙️ TTS Audio Suite Era
▼ |
v4.0 ──────────► v4.3 ──────────► v4.4 ────────► v4.5 ──────────┐
Aug 25 Aug 25 Aug 25 Aug 25 │
│ │ │ │ │
BREAKING! RVC + Silent Higgs Audio 2 │
Project Voice Speech New TTS Engine │
Renamed Conversion Analyzer Voice Cloning │
TTS Audio Suite + Streaming │
▼
v4.9 ◄─────────────── v4.8 ◄────────────────── v4.6 ◄───────────┘
Sep 25 Sep 25 Aug 25
│ │ │
IndexTTS-2 Chatterbox VibeVoice
Emotion Multilingual New TTS Engine
Control Official (23-lang) 90min Generation
│
│ 🎨 Inline Editor Tags Era
▼ |
v4.12 ──────────────► v4.15 ────────────► v4.16 ───────────────┐
Oct 25 Dez 25 Dez 25 │
│ │ │ │
Per-Seg Parameter Step Audio EditX CosyVoice3 │
Switching [seed:24] Inline Edit tags TTS + VC │
<laughter:2> │
▼
v4.24◄─────────────── v4.22 ◄─────────────── v4.19 ◄───────────┘
Mar 26 Mar 26 Jan 26
│ │ │
Text to SRT Echo-TTS Qwen3-TTS
Builder English TTS TTS + ASR + VoiceDesign
|
|─── 🎓 Training Support Era 🧱 Runtime Isolation T5 Era
▼ |
v4.25 ──────────────► v4.26 ────────────► v5.00 ───────────────┐
Apr 26 May 26 Jun 26 │
│ │ │ │
RVC MOSS-TTS Transformers 5 │
Model Training Higgs Audio v3 TTS │
│
▼
v5.3 ◄─────────────── v5.2 ◄─────────────── v5.1 ◄─────────────┘
Jun 26 Mar 26 Jan 26
│ │ │
Native SRT Duration OmniVoice TTS Dots TTS
Granite ASR
Visual Tag Builder
│
▼
v5.4 ───────────────────────────────► v5.5
Jul 26 Jul 26
│ │
Fish Audio S2 Pro MOSS-TTS v1.5
IndexTTS-2 Emotion Blending Sound Effects
Faster Tag Editor Voice Designer
Character Alias Manager
Want to add support for a new TTS engine, Voice Changer, ASR, or special audio model?
Start with the New Engine Guide Hub. It is written for users guiding an LLM through the process: first research the official model, then check existing ComfyUI implementations, decide scope, implement in the suite architecture, and run the parity checklist before PR review. For TTS engines, Unified TTS Text and Unified SRT TTS are a required pair.
The "ChatterBox SRT Voice TTS" node allows TTS generation by processing SRT content (SubRip Subtitle) files, ensuring precise timing and synchronization with your audio.
Key SRT Features:
smart_natural Timing Mode: Intelligent shifting logic that prevents overlaps and ensures natural speech flowAdjusted_SRT Output: Provides actual timings for generated audio for accurate post-processingFor comprehensive technical information, refer to the SRT_IMPLEMENTATION.md file.
This is the new architectural baseline for the suite.
This matters because the suite now has a clearer split:
NEW: DramaBox is integrated as an English expressive TTS engine for both Unified TTS Text and Unified SRT TTS.
torch.compileImportant limitations:
See the DramaBox Prompting Guide for prompt syntax, controls, memory modes, duration behavior, and examples. See the DramaBox LoRA Training Guide for dataset formats, training workflow, adapter loading, and CPU-safe preflight.
NEW in v4.4.0: Video analysis and mouth movement detection for silent video processing!
Perfect for:
Important Notes:
NEW in v4.5.0: State-of-the-art voice cloning technology with advanced neural voice replication!
Key Capabilities:
[CharacterName] syntax for multi-speaker dialoguesTechnical Features:
ComfyUI/models/TTS/HiggsAudio/ structureQuick Start:
Higgs Audio Engine node to configure voice cloning parametersTTS Text or TTS SRT node for generationPerfect for:
NEW: Higgs Audio v3 is now integrated as a native main-environment engine on the modern Transformers 5 stack.
<|emotion:amusement|>, <|style:whispering|>, <|prosody:pause|>, and <|sfx:laughter|><emotion:amusement>-style input and normalizes it internally to the official Higgs formatHiggs Audio v3 inline tags mode instead of pretending all inline systems are Step Audio EditXImportant behavior note:
Good fit for:
[Alice], [Bob] character tags with voice files from the voices folder - supports unlimited characters with pause tags and per-character control[Character] tag auto-conversion and manual "Speaker 1: Hello" format for up to 4 speakersTechnical Features:
Quick Start:
⚙️ VibeVoice Engine node to configure model and multi-speaker modeTTS Text or TTS SRT node for generationVibeVoice and KugelAudio run directly in the suite's Transformers 5 main environment.
Perfect for:
NEW in v3.1.0: Seamless character switching for both F5TTS and ChatterBox engines!
[CharacterName] tags to switch between different voices[Alice] instead of [female_01] with #character_alias_map.txt📖 Complete Character Switching Guide
Example usage:
Hello! This is the narrator speaking.
[Alice] Hi there! I'm Alice, nice to meet you.
[Bob] And I'm Bob! Great to meet you both.
Back to the narrator for the conclusion.
NEW in v3.4.0: Seamless language switching using simple bracket notation!
[language:character] tags to switch languages and models automatically[German:Alice], [Brazil:Bob], [USA:], [Portugal:] - no need to remember language codes![fr:Alice], [de:Bob], or [es:] (language only) patternsSupported Languages:
Example usage:
Hello! This is English text with the default model.
[de:Alice] Hallo! Ich spreche Deutsch mit Alice's Stimme.
[fr:] Bonjour! Je parle français avec la voix du narrateur.
[es:Bob] ¡Hola! Soy Bob hablando en español.
Back to English with the original model.
Advanced SRT Integration:
1
00:00:01,000 --> 00:00:04,000
Hello! Welcome to our multilingual show.
2
00:00:04,500 --> 00:00:08,000
[de:female_01] Willkommen zu unserer mehrsprachigen Show!
3
00:00:08,500 --> 00:00:12,000
[fr:] Bienvenue à notre émission multilingue!
NEW: Progressive voice refinement with intelligent caching for instant experimentation!
How it works:
Intelligent caching examples:
Practical tip: Start with 1 pass, then test 2-5 passes to find the sweet spot for your audio. More passes can improve voice similarity, but there is no universal best value.
NEW in v4.1.0: Professional-grade Real-time Voice Conversion with .pth character models!
📖 See Model folder layouts for detailed setup paths
How it works:
NEW: Integrated RVC training inside the suite, using the same unified node style as the rest of the project instead of a detached external workflow.
🎓 Model Training accepts TTS_ENGINE and routes by engine type📦 RVC Dataset Prep handles dataset path/zip/folder upload or direct audio input, slicing, HuBERT features, F0 extraction, and reusable prep caches🎛️ RVC Training Config exposes practical training controls with tooltip guidance instead of raw upstream garbagecontinue_from for training further from a finished RVC model/artifactpretrained_v2 RVC training init checkpoints are auto-managedCurrent scope:
Typical flow:
⚙️ RVC Engine📦 RVC Dataset Prep🎛️ RVC Training Config🎓 Model Training🎭 Load RVC Character ModelImportant notes:
ComfyUI/output/tts_audio_suite_training/rvc/.pth models and .index files go under ComfyUI/models/TTS/RVC/save_best_model is only a low-loss inference candidate, not a magical quality oracle. You still need to listen.NEW: Intelligent pause insertion for natural speech timing control!
[pause:1.5], [wait:2s], [stop:3][pause:500ms], [wait:1200ms], [stop:800ms]pause, wait, stop (all work identically)Example usage:
Welcome to our show! [pause:1s] Today we'll discuss exciting topics.
[Alice] I'm really excited! [wait:500ms] This will be great.
[stop:2] Let's get started with the main content.
NEW in v4.6.29: ChatterBox TTS now supports 11 languages with community-finetuned models and automatic model management!
Supported Languages:
<haha>, <wow> tags[it] prefix for Italian textKey Features:
Usage: Select language from dropdown → First generation downloads model → Subsequent generations use cached model
NEW in v4.8.0: Official ResembleAI Chatterbox Multilingual TTS model with native support for 23 languages!
The Chatterbox Multilingual TTS (referred to internally as "ChatterBox Official 23-Lang" to distinguish from community models) is ResembleAI's first production-grade open-source TTS model supporting 23 languages out of the box. This is the official successor to the original ChatterBox model with enhanced multilingual capabilities.
🎯 Key Advantages over Community Models:
🌍 Supported Languages (23 total + Vietnamese community finetune):
Arabic (ar), Danish (da), German (de), Greek (el), English (en), Spanish (es), Finnish (fi), French (fr), Hebrew (he), Hindi (hi), Italian (it), Japanese (ja), Korean (ko), Malay (ms), Dutch (nl), Norwegian (no), Polish (pl), Portuguese (pt), Russian (ru), Swedish (sv), Swahili (sw), Turkish (tr), Chinese (zh)
🇻🇳 Vietnamese (Viterbox): Community finetune by Dolly AI 23 with expanded Vietnamese tokenization (dolly-vn/viterbox) - select from model version dropdown
🔧 Fully Integrated Features:
[CharacterName] support with per-character voice references[language:character] syntax with intelligent parameter switching[pause:Ns] support with character voice inheritance🆚 vs Community Models:
| Feature | Chatterbox Multilingual TTS | Community Models |
|---|---|---|
| Languages | 23 native languages | 11 finetuned variants |
| Model Loading | Single model, parameter switching | Separate model per language |
| Voice Cloning | Zero-shot across all languages | Per-model training |
| Official Support | ✅ ResembleAI official | Community maintained |
| Character Integration | ✅ Full integration | ✅ Full integration |
| SRT Support | ✅ Advanced timing modes | ✅ Advanced timing modes |
| Performance | Optimized single-model | Multiple model overhead |
🎭 Character Example:
[En:Alice] Hello everyone! [De:Hans] Guten Tag! [Es:Maria] ¡Hola! [pause:2s] [En:Alice] That was amazing multilingual switching!
This creates seamless multilingual character switching with proper voice inheritance and pause support - all within a single model.
🎭 NEW: v2 Special Emotion & Sound Tokens 🚧 Experimental
ChatterBox v2 vocabulary includes 30+ special tokens for emotions, sounds, and vocal effects. Note: These are experimental - tokens may produce minimal or no audible effects. ResembleAI has not officially documented their usage (see issue #186).
Try angle brackets <emotion> to experiment:
[Alice] Hello! <laughter> hahaha. [pause:0.5] <whisper> This might work slightly.
Available v2 Tokens:
<giggle>, <laughter>, <sigh>, <cry>, <gasp>, <groan><whisper>, <mumble>, <singing>, <humming><cough>, <sneeze>, <sniff>, <inhale>, <exhale>Model Version Selection:
Both versions fully support character switching, language switching, and pause tags. The v2 special tokens are experimental with limited effectiveness - our implementation is ready for when/if ResembleAI improves this feature. The angle bracket syntax <emotion> avoids conflicts with character tags [Name] and pause tags [pause:1s].
NEW in v4.3.0: Complete architectural overhaul implementing universal streaming system with parallel processing capabilities!
Key Features:
batch_size parameterPerformance Notes:
batch_size=0 for optimal performance (sequential processing)batch_size > 1 enables parallel workers but typically slower due to GPU inference characteristicsNEW in v4.9.0: Revolutionary IndexTTS-2 engine with advanced emotion control and dual-source emotion blending!
emotion_control and audio references to emotion_audio; both can be used together{seg} template processing for contextual per-segment emotionsemotion_audio as an emotion reference for natural expressionopt_narrator on emotion_audio, including per-character [Character:emotion_ref] references[Character:emotion_ref] syntax, blendable with vector/text emotionduration_factor scales the internal semantic feature sequence (0.5 shorter/faster, 1.0 unchanged, 2.0 longer/slower). It is not natural prosody or exact-duration planning, does not apply to 2.0, and is not used by SRT native-duration targeting<word|pronunciation> annotations through suite text processing2.0 versus 2.5: Treat 2.5 as a multilingual/efficiency alternative, not an automatic voice-cloning quality upgrade. In our manual listening, legacy 2.0 preserved speaker resemblance better when transferring a strong emotion from a different reference voice; 2.5 may still be preferable for Japanese, Spanish, Arabic, or cross-lingual generation. Strong external emotion settings can reduce perceived speaker identity, so compare both models for the target voice.
Key Features:
{seg} placeholder for contextual emotion analysis (e.g., "Worried parent speaking: {seg}")Example Usage:
Welcome to our show! [Alice:happy_sarah] I'm so excited to be here!
[Bob:angry_narrator] That's completely unacceptable behavior.
Perfect for:
NEW in v4.15: Revolutionary LLM-based audio post-processing with emotion, style, and paralinguistic control!
<Laughter:2>, <emotion:happy>, <style:whisper> tagsKey Features:
<Laughter:2|emotion:happy|style:whisper>Example Usage:
[Alice] Hello there <Laughter:2> my friend! <emotion:happy>
[Bob] Listen carefully <style:whisper|speed:slower>, this is important.
[Alice] I'm laughing so hard <Laughter:3> <restore>
Perfect for:
⚠️ Important Notes:
<restore> tag to recover original voice characterNEW in v4.16: Alibaba's fast multilingual voice cloning with native paralinguistic tags, instruct mode, and zero-shot voice conversion!
<breath>, <laughter>, <cough>, <sigh>, <laughing>text</laughing> - processed during generation → 📖 Tags GuideKey Features:
<breath>, <laughter>, <cough>, <sigh>, <gasp>, <laughing>text</laughing>, <strong>text</strong>[CharacterName] support with per-character voice references[en:], [zh:], [ja:], [ko:] bracket syntax or native <|en|> tags[pause:Ns] support for natural speech timingExample Usage:
[Alice] Hello everyone! This is zero-shot voice cloning.
[Bob] 你好!我说普通话。[pause:1s] 还可以说方言。
Instruct Mode Examples:
# Speak with Cantonese dialect
Instruct: 请用广东话表达。
# Speak with excitement
Instruct: 用兴奋的语气说话。
Perfect for:
NEW in v4.19: Alibaba's Qwen3-TTS with 3 distinct TTS model types - CustomVoice presets, dedicated text-to-voice design, and zero-shot voice cloning. The engine's model dropdown exposes every checkpoint and marks installed checkpoints with a local: prefix. Model-specific controls appear only when they apply.
NEW: ✏️ Unified ASR Transcribe support now includes Qwen3-ASR and Granite ASR, giving the suite a second ASR engine option with optional custom timestamps/SRT for Granite via the reused Qwen forced aligner. Granite 4.1 plus also adds native speaker diarization and native word timestamps.
Model Types:
🎭 CustomVoice Model (0.6B / 1.7B): 9 preset multilingual speakers (Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee)
✍️ VoiceDesign Model (1.7B only): Dedicated Qwen voice creation from text descriptions
🎤 Base Model (0.6B / 1.7B): Zero-shot voice cloning from 3-30s reference audio
🔤 ASR: Qwen3-TTS Engine can be connected to the ✏️ Unified ASR Transcribe node for transcription
⚠️ Style instructions only work with CustomVoice and VoiceDesign. Voice cloning (Base) ignores the instruction field entirely.
Technical Specs:
Unified Features Support:
Voice Designer Node:
The shared designer accepts Qwen3-TTS, MOSS-TTS, or OmniVoice engine configurations and outputs the same NARRATOR_VOICE format. The voice-design instruction lives on 🎨 Voice Designer; the engine keeps model, language, and generation settings. Select Qwen VoiceDesign or MOSS VoiceGenerator in the engine's model dropdown, or set OmniVoice to Voice Design mode. The corresponding engine instruction stays visible but is disabled because it would be ignored. Incompatible modes stop with a direct correction message. OmniVoice's controlled tag vocabulary can still be assembled with 📐 Visual Tag Builder. Connect the resulting opt_narrator to 💾 Save Character Voice when persistence is wanted.
💾 Save Character Voice accepts only opt_narrator, keeping persistence separate from voice construction. For existing audio, use 🎭 Character Voices with the audio and its exact transcription, then connect its opt_narrator output to Save Character Voice. The save node writes the established three-file format—name.wav, name.reference.txt, and metadata in name.txt—under models/voices/.
Description: "A deep, authoritative male voice with clear articulation"
→ Voice generated and cached → Use in TTS Text/SRT nodes
Perfect for:
NEW in v5.x: OmniVoice is now integrated into the unified suite with native duration-aware SRT generation and a generalized visual tag workflow.
<> form like <laughter>, then converted internally for generation → 📖 OmniVoice Tags GuidePractical note:
Use the built-in OmniVoice preset in 📐 Visual Tag Builder for the canonical voice-design workflow. If you need a different tag schema, the same node now supports reusable custom presets with saved column order.
NEW in v4.26: OpenMOSS engine family integration with unified support for single-speaker TTS and native multi-speaker dialogue.
Model Variants:
MOSS-TTS-Local-TransformerMOSS-TTS-v1.5 — 31 languages and more stable cloninglaion/moss-tts-v1.5-8b-voice-actingMOSS-TTSMOSS-TTSD-v1.0MOSS-VoiceGenerator — select it in the MOSS engine for Voice DesignerMOSS-Audio-TokenizerCompatible community full checkpoints can also be placed in models/TTS/moss_tts/<model-name>/.
They are listed as local:<model-name> and classified from config.json; unsupported layouts fail explicitly.
Supported Native Input Forms (TTSD):
[Character] tags[1] / [S1] numeric speaker tagsSpeaker 1: ... formatAll native forms are normalized internally to canonical [S1]...[S5] dialogue.
Speaker Modes:
Important Native Compatibility Rule:
Native TTSD mode now hard-fails (explicit error popup) instead of silently switching models when these are detected:
[] parameter changesIf you need those controls, switch to Custom Character Switching and use MOSS-TTS-Local-Transformer, MOSS-TTS-v1.5, or MOSS-TTS.
Official Prompt Fields Exposed:
instruction, quality, sound_event, ambient_sound, language, duration_tokens
Per-segment overrides are supported with [] parameter syntax for whole-segment conditioning.
Documentation:
Training status:
🎓 Model Training flow.models/TTS/moss_tts/loras/.NEW in v4.23: ASR subtitle generation is now modular instead of being buried inside the transcriber.
The flow is now:
✏️ ASR Transcribe → optional 📝 ASR Punctuation / Truecase → 📺 Text to SRT Builder
This matters because the suite can now:
🔧 SRT Advanced Options now belongs to the builder stage instead of being mixed into ASR4.1 plus can emit suite-native speaker tags like [Speaker 1] for downstream TTS/alias workflows, and if you need both diarization and word timings the node automatically falls back to the reused Qwen forced alignerasr_timing_data disconnected and the builder estimates subtitle timings from plain text using the same SRT options that later shape the final cuesCurrent intended use:
✏️ ASR Transcribe for timed transcription with Qwen3-ASR or Granite ASR📝 ASR Punctuation / Truecase mainly for low-punctuation ASR outputs like Granite📺 Text to SRT Builder to turn cleaned text + ASR timing data into final SRTGranite note:
granite-speech-4.1-2b keeps Japanese supportgranite-speech-4.1-2b-plus adds native diarization and native word timestamps, but drops JapaneseWorkflow example:
Use the new Unified ✏️ ASR Transcribe + SRT Builder workflow for both Granite ASR and Qwen3-ASR examples.
NEW in v4.22: Echo-TTS DiT-based voice cloning with reference audio support.
Usage:
⚙️ Echo-TTS Engine node🎤 TTS Text or 📺 TTS SRTComfyUI/models/TTS/echo-tts-base/📖 See Model folder layouts for detailed setup paths
NEW in v4.10.0: Universal multilingual text preprocessing node for improved TTS pronunciation quality across languages!
Perfect for:
⚠️ Experimental Feature: This is a new experimental feature - user feedback needed to validate pronunciation improvements across different languages. Please test with your target languages and report results!
📖 Try the workflow: F5 TTS integration + 📝 Phoneme Text Normalizer
NEW in v4.12.0: Fine-grained per-segment control over TTS generation parameters across all engines with an interactive tag editor!
Beyond character switching and language control, you can now override generation parameters (seed, temperature, CFG, speed, etc.) on a per-segment basis using inline tags. The new 🏷️ Multiline TTS Tag Editor node makes building complex tags easier and more visual with:
{seg} dynamic emotion controls directly from the Inline Tags panelThis enables dynamic control over individual audio segments without modifying node defaults.
🎯 Key Features:
[Alice|seed:42|temp:0.5] or [seed:42|Alice|temp:0.5]temp (temperature), cfg_weight (cfg), exag (exaggeration)[SEED:42], [Seed:42], [seed:42] all work identically[de:Alice|seed:42|temp:0.7]💡 Real-World Examples:
[Alice|seed:42] This is the first segment of Alice.
[Alice|seed:42] This is the second segment, sounding identical.
[Bob|temperature:0.3] Bob speaks carefully and precisely.
[Bob|temperature:0.8] Bob speaks more creatively and varied.
[Narrator|temp:0.4] Important introduction - precise delivery.
[Narrator|temp:0.7] Creative narrative section - more varied!
🔧 Supported Parameters by Engine:
📖 Guides: Per-Segment Parameter Switching | Multiline TTS Tag Editor | OmniVoice Tags Guide
Perfect for:
[!TIP] You might want to create a clean portable installation just for TTS-Audio-Suite, instead of adding it to your regular ComfyUI setup. This can make the installation process smoother and may reduce initialization time.
One-click installation with intelligent dependency management:
Python 3.13 Support:
Same intelligent installer, manual setup:
Clone the repository
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
cd TTS-Audio-Suite
Run the intelligent installer:
ComfyUI Portable:
# Windows:
..\..\..\python_embeded\python.exe install.py
# Linux/Mac:
../../../python_embeded/python.exe install.py
ComfyUI with venv/conda:
# First activate your ComfyUI environment, then:
python install.py
The installer automatically handles all dependency conflicts and Python version compatibility.
Run your first workflow
Try a Workflow
Restart ComfyUI and look for 🎤 TTS Audio Suite nodes
🧪 Python 3.13 Users: Installation is fully supported! The system automatically uses OpenSeeFace for mouth movement analysis when MediaPipe is unavailable.
Need offline/manual setup? Use docs/MODEL_DOWNLOAD_SOURCES.md and docs/MODEL_LAYOUTS.md
This section provides a detailed guide for installing TTS Audio Suite, covering different ComfyUI installation methods.
ComfyUI installation (Portable, Direct with venv, or through Manager)
Python 3.12 or higher
Optional system libraries (Linux only):
# Ubuntu/Debian - Optional audio features
sudo apt-get install portaudio19-dev libsamplerate0-dev
# Fedora/RHEL
sudo dnf install portaudio-devel libsamplerate-devel
📋 Optional:
libsamplerate0-devprovides additional audio-resampling support.portaudio19-devenables voice recording. Missing either package no longer blocks installation of the TTS engines.
Optional macOS dependencies:
brew install portaudio
Windows: No additional system dependencies needed (libraries come pre-compiled)
For portable installations, follow these steps:
Clone the repository into the ComfyUI/custom_nodes folder:
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
Navigate to the cloned directory:
cd TTS-Audio-Suite
Run the install script to automatically handle all dependencies. Important: Use the python.exe executable located in your ComfyUI portable installation.
../../../python_embeded/python.exe install.py
The script will automatically install all required Python packages and detect any missing system dependencies.
If you have a direct installation with a virtual environment (venv), follow these steps:
Clone the repository into the ComfyUI/custom_nodes folder:
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
Activate your ComfyUI virtual environment. This is crucial to ensure dependencies are installed in the correct environment. The method to activate the venv may vary depending on your setup. Here's a common example:
cd ComfyUI
. ./venv/bin/activate
or on Windows:
ComfyUI\venv\Scripts\activate
Navigate to the cloned directory:
cd custom_nodes/TTS-Audio-Suite
Run the install script to automatically handle all dependencies:
python install.py
The script will automatically install all required Python packages and detect any missing system dependencies.
Install the ComfyUI Manager if you haven't already.
Use the Manager to install the "TTS Audio Suite" node.
The manager might handle dependencies automatically, but it's still recommended to verify the installation. Navigate to the node's directory:
cd ComfyUI/custom_nodes/TTS-Audio-Suite
Activate your ComfyUI virtual environment (see instructions in "Direct Installation with venv").
If you encounter issues, run the install script to manually install dependencies:
python install.py
Our install script automatically detects missing optional system libraries and will display feature warnings like:
[!] Optional system dependencies are missing
============================================================
OPTIONAL LINUX SYSTEM DEPENDENCIES
============================================================
• libsamplerate0-dev (optional additional audio-resampling support)
• portaudio19-dev (optional voice recording)
Please install with:
# Ubuntu/Debian:
sudo apt-get install libsamplerate0-dev portaudio19-dev
# Fedora/RHEL:
sudo dnf install libsamplerate-devel portaudio-devel
============================================================
Core TTS installation will continue; only the listed features may be unavailable.
A common problem is installing dependencies in the wrong Python environment. Always ensure you are installing dependencies within your ComfyUI's Python environment.
If the engine comparison table shows Shared or Dedicated in the Isolation column, that engine has its own secondary-environment path for dependency conflicts. Configure that on the engine node with ⚠️ Runtime Isolation instead of trying to downgrade your main ComfyUI environment.
Verify your Python environment: After activating your venv or navigating to your portable ComfyUI installation, check the Python executable being used:
which python
This should point to the Python executable within your ComfyUI installation (e.g., ComfyUI/python_embeded/python.exe or ComfyUI/venv/bin/python).
If s3tokenizer fails to install: This dependency can be problematic. Try upgrading your pip and setuptools:
python -m pip install --upgrade pip setuptools wheel
Then, try installing the requirements again.
If you cloned the node manually (without the Manager): Make sure you install the requirements.txt file.
To update the node to the latest version:
Navigate to the node's directory:
cd ComfyUI/custom_nodes/TTS-Audio-Suite
Pull the latest changes from the repository:
git pull
Reinstall the dependencies (in case they have been updated):
python install.py
cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
Some dependencies, particularly s3tokenizer, can occasionally cause installation issues on certain Python setups (e.g., Python 3.10, sometimes used by tools like Stability Matrix).
To minimize potential problems, it's highly recommended to first ensure your core packaging tools are up-to-date in your ComfyUI's virtual environment:
python -m pip install --upgrade pip setuptools wheel
After running the command above, install the node's specific requirements:
python install.py
ChatterBox Voice now supports FFmpeg for high-quality audio stretching. While not required, it's recommended for the best audio quality:
Windows:
winget install FFmpeg
# or with Chocolatey
choco install ffmpeg
macOS:
brew install ffmpeg
Linux:
# Ubuntu/Debian
sudo apt-get install ffmpeg
# Fedora
sudo dnf install ffmpeg
If FFmpeg is not available, ChatterBox will automatically fall back to using the built-in phase vocoder method for audio stretching - your workflows will continue to work without interruption.
No manual downloads needed. All engines download required models automatically on first use.
For offline/manual setup:
| Engine | Primary model path | Auto-download | Notes |
|---|---|---|---|
| ChatterBox | ComfyUI/models/TTS/chatterbox/ | ✅ | Legacy ComfyUI/models/chatterbox/ still works |
| ChatterBox 23-Lang | ComfyUI/models/TTS/chatterbox_official_23lang/ | ✅ | v1/v2/v3 coexist in same folder |
| F5-TTS | ComfyUI/models/TTS/F5-TTS/ | ✅ | Optional Vocos and voice refs |
| Higgs Audio 2 | ComfyUI/models/TTS/HiggsAudio/ | ✅ | Generation + tokenizer |
| Higgs Audio v3 | ComfyUI/models/TTS/higgs_audio_v3/ | ✅ | Official 4B multilingual TTS model |
| VibeVoice | ComfyUI/models/TTS/VibeVoice/ | ✅ | 1.5B and 7B variants |
| RVC | ComfyUI/models/TTS/RVC/ | ✅* | Base models auto; character .pth can be user-provided |
| IndexTTS-2 | ComfyUI/models/TTS/IndexTTS/ | ✅ | Emotion components included |
| Step Audio EditX | ComfyUI/models/TTS/step_audio_editx/ | ✅ | Main model + tokenizer stack |
| CosyVoice3 | ComfyUI/models/TTS/CosyVoice/ | ✅ | Variant-specific lazy downloads |
| Qwen3-TTS / ASR | ComfyUI/models/TTS/qwen3_tts/ | ✅ | Per-variant download + shared tokenizer |
| MOSS-TTS | ComfyUI/models/TTS/moss_tts/ | ✅ | Local/Delay/VoiceGenerator/SoundEffect v1/TTSD models plus shared MOSS-Audio-Tokenizer codec |
| MOSS-SoundEffect v2 | ComfyUI/models/TTS/moss_soundeffect_v2/ | ✅ | Official v2 diffusion pipeline; configured ComfyUI environment |
| Granite ASR | ComfyUI/models/TTS/granite_asr/ | ✅ | Granite ASR models; plus adds native diarization/timestamps, optional Qwen forced aligner reused lazily for timestamps/SRT fallback |
| Echo-TTS | ComfyUI/models/TTS/echo-tts-base/ | ✅ | ~7.1GB total (base + dac); CC-BY-NC-SA |
| Dots TTS | ComfyUI/models/TTS/dots_tts/ | ✅ | Official base / soar / mf checkpoints with tokenizer, vocoder, speaker encoder |
| DramaBox | ComfyUI/models/TTS/dramabox/DramaBox/ | ✅ | ~16.4GB download; fast mode roughly 24GB VRAM; experimental FP8, staged, and sequential options can reduce VRAM, but no minimum GPU size is guaranteed; conditional LTX-2 Community License |
| Fish Audio S2 Pro | ComfyUI/models/TTS/fish_audio_s2_pro/ | ✅ | Official BF16 or optional community FP8 checkpoint; the official checkpoint can be quantized on load with BNB INT8/NF4; main T5 environment with process teardown for Clear VRAM; Fish Audio Research License |
| OmniVoice | ComfyUI/models/TTS/omnivoice/ | ✅ | Official OmniVoice model. Voice cloning in this suite requires explicit reference text. |
Generated from tts_audio_suite_engines.yaml.
If TTS Audio Suite has been helpful for your projects, consider supporting its development:
Your support helps maintain and improve this project for the entire community!
Ready-to-use ComfyUI workflows - Download and drag into ComfyUI:
| Workflow | Description | Features | Status | Files |
|---|---|---|---|---|
| Unified 📺 TTS SRT | Universal SRT processing with all TTS engines | • ChatterBox/F5-TTS/Higgs Audio 2 • Multiple timing modes • Multi-character switching • Overlap SRT support | ✅ New in v4.5 | 📁 JSON |
| Unified 🔄 Voice Changer | Modern voice conversion with multiple engines | • RVC + ChatterBox VC • Iterative refinement • Real-time conversion | ✅ Updated for v4.3 | 📁 JSON |
| Unified ✏️ ASR Transcribe + SRT Builder | Modular ASR + subtitle workflow | • Granite ASR + Qwen3 ASR examples • Separate transcription and SRT building • Works with the new Text to SRT Builder flow | ✅ New in v4.23 | 📁 JSON |
| Unified 🌩️ Sound Effects | Text-to-sound generation with compatible engines | • MOSS-SoundEffect v1 and v2 • Per-segment parameters and pauses • Long-duration chunking and audio cache | ✅ New | 📁 JSON |
| Unified 🎨 Voice Designer | Reference-free character voice creation | • Qwen3-TTS, MOSS-TTS, and OmniVoice • Free-form descriptions or Visual Tag Builder • Preview and save reusable character voices | ✅ New | 📁 JSON · 🖼️ Cover |
| Workflow | Description | Status | Files |
|---|---|---|---|
| 🤐 Voice Cleaning | Audio restoration & cleanup with dual tool pipeline | ✅ New in v4.13 | 📁 JSON |
| DramaBox LoRA 🎓 Model Training | DramaBox IC-LoRA training workflow from staged speech clips | ✅ New | 📁 JSON |
| MOSS LoRA 🎓 Model Training | Initial MOSS LoRA training workflow from clipped speech dataset | ✅ New in v4.27 | 📁 JSON |
| RVC 🎓 Model Training | RVC voice model training workflow | ✅ New in v4.25 | 📁 JSON |
| 🎨 Step Audio EditX - Audio Editor | Step Audio EditX audio editing with inline edit tags | ✅ New in v4.14 | 📁 JSON |
| ⚙️ Step Audio EditX Integration | Step Audio EditX TTS engine with zero-shot voice cloning | ✅ New in v4.14 | 📁 JSON |
| ⚙️ Higgs Audio v3 Integration | Higgs Audio v3 TTS with zero-shot voice cloning and native inline tags | ✅ New in v4.27 | 📁 JSON |
| ⚙️ OmniVoice Engine Integration | OmniVoice multilingual TTS with cloning, voice design, and native duration control | ✅ New in v4.28 | 📁 JSON |
| ⚙️ Fish Audio S2 Pro Integration | Fish S2 Pro multilingual cloning with native multi-speaker dialogue, inline control, and long-form generation | ✅ New in v5.3 | 📁 JSON |
| ⚙️ DramaBox Integration | DramaBox expressive scene prompting with native SRT duration targeting | ✅ New in v5.6 | 📁 JSON |
| 🌈 IndexTTS-2 Integration | IndexTTS-2 engine with advanced emotion control | ✅ New in v4.9 | 📁 JSON |
| 📝 F5 TTS + Text Normalizer | F5-TTS with multilingual text processing and phonemization | ✅ New in v4.10.0 | 📁 JSON |
| Qwen3 integration + ASR | Qwen3-TTS voice generation with ASR transcription | ✅ New in v4.21 | 📁 JSON |
| VibeVoice Integration | VibeVoice long-form TTS with multi-speaker support | ✅ Compatible | 📁 JSON |
| ChatterBox Integration | General ChatterBox TTS and Voice Conversion | ✅ Compatible | 📁 JSON |
| F5-TTS Speech Editor | Interactive waveform analysis for F5-TTS editing | ✅ Updated for v4 | 📁 JSON |
💡 Recommended: Use the new Unified 📺 TTS SRT workflow which showcases the unified TTS flow in one comprehensive workflow. It demonstrates SRT processing, timing modes, multi-character switching, and modern engine integration across the suite.
📥 Usage: Download the
.jsonfiles and drag them directly into your ComfyUI interface. The workflows will automatically load with proper node connections.
The TTS Audio Suite project code is licensed under MIT.
Python
93.1%
JavaScript
6.2%