Open-weights voice acting data pipeline combining structured taxonomy sampling, DramaBox TTS synthesis, Sidon speech restoration, and ChatterboxVC augmentation with best-of-N ranking across 46 scoring methods.
Live Demo: Sidon+VC Sample Groups β 20 groups x 25 candidates with LLM-guided CUT TO: splitting, Whisper turbo ASR, and Gemma 4 re-annotation.
Benchmarks (landing page):
- π¬ Vanilla DramaBox TTS β generate + reward-rank β raw two-scene
CUT TO:DramaBox prompts fed directly to the 8B MOSS voice-acting TTS, 4 seeds, scored + reward-ranked; listenable takes, best-of-k quality/compute trade-off, and the full k=1..32 seed-scaling walltime table.- β‘ Local-LLM DramaBox prompt-generation throughput β per-pathway/language token + throughput stats with a 1M-prompt estimate, plus the DramaBox TTS seed-scaling table for best-of-k planning.
- π How the DramaBox prompt dataset is sampled & generated β a plain-English, reproducible walkthrough of the
laion/dramabox-cutscene-promptsdataset: the 5 sampling pathways, every taxonomy it draws from (VoiceNet, 40 EmoNet emotions, archetypes, situations, 180 vocal bursts), the exact prompts the model receives, and real generated examples with their sampled metadata.
Dataset Plan: See the full technical white paper β Towards an Emotionally Expressive Audio Omni-Model β for the complete LAION Voice and LAION Voice Acting dataset construction plan, model inventory, and annotation strategy.
End-to-end voice prompt generation and audio synthesis using the DramaBox TTS model (22B DiT transformer) and structured voice taxonomy sampling. Based on the voice taxonomy research from Schuhmann et al., 2025 and EmoNet-Voice (Schuhmann et al., 2025).
This pipeline generates richly annotated voice performance prompts in the DramaBox format β single-speaker scenes with stage directions (English) and spoken dialogue (target language) β then synthesizes them into audio. Each prompt is procedurally constructed by sampling from structured taxonomies, then expanded by an LLM (Gemma 4 E4B-it) into a full performance script.
The full audio processing chain:
Taxonomy Sampling LLM Prompt Gen DramaBox TTS (22B)
(VoiceNet/Archetype/ (Gemma 4 E4B-it) CFG=2.5, STG=1.5
Situation/ActingChall) 25 candidates/prompt
| | |
+------------------------+ |
v
+---------------------------+
| Per-Sample Augment |
| |
| Path A: Sidon only |
| (16kHz -> 48kHz) |
| Path B: ChatterboxVC |
| + Sidon (VC -> restore)|
| |
| Pick best by |
| DNS-MOS OVR score |
+-------------+-------------+
|
+-------------v-------------+
| Whisper Turbo ASR |
| (word-level timestamps) |
+-------------+-------------+
|
+-------------v-------------+
| LLM-Guided CUT TO: |
| Split (Gemma 4 E4B-it) |
| + quiet-spot detection |
| + LLM fade strategy |
+-------------+-------------+
|
+-------------v-------------+
| Best-of-25 Ranking |
| (WER, VoiceCLAP, |
| EmoNet, Content) |
+-------------+-------------+
|
+-------------v-------------+
| Gemma 4 Re-annotation |
| (ASR -> refined prompt) |
+---------------------------+
For each raw TTS candidate, two enhancement paths run and the best is selected:
Raw TTS Audio βββ¬βββΊ Sidon Speech Restoration βββΊ DNS-MOS βββ
(from DramaBox) β (w2v-BERT LoRA + DAC) ββββΊ Pick higher OVR
β (16kHz β 48kHz) β
ββββΊ ChatterboxVC βββΊ Sidon βββΊ DNS-MOS βββββ
(S3Gen flow-matching VC,
self-VC or ref-VC)
For two-scene audio, a self-VC of the full audio provides the VC target for speaker consistency:
Full Audio βββΊ Self-VC βββΊ Sidon βββΊ full_enhanced (VC target)
β
βββββββββββββββββββββββββββββββββ
β
Part 1 βββ¬βββΊ Sidon only βββΊ DNS-MOS βββ
β ββββΊ Pick best
ββββΊ VC(βfull_enhanced) + Sidon βββΊ DNS-MOS βββ
Part 2 βββ¬βββΊ Sidon only βββΊ DNS-MOS βββ
β ββββΊ Pick best
ββββΊ VC(βfull_enhanced) + Sidon βββΊ DNS-MOS βββ
Scoring methods (46 total):
Listen to the latest Sidon+VC experiment β 20 groups x 25 candidates = 500 audio clips with LLM-guided CUT TO: splitting and Gemma 4 re-annotation.
| Demo | Description | Link |
|---|---|---|
| Sidon+VC Sample Groups | 20 groups, best-of-25, LLM-guided splits | sidon_vc_sample_groups.html |
50 groups x 25 candidates = 1,250 audio clips across 10 pages. Each page has an interactive ranking dropdown with 46 methods.
| Page | Groups | Link |
|---|---|---|
| Page 1 | Groups 0-4 | accc_lavasr_p1.html |
| Page 2 | Groups 5-9 | accc_lavasr_p2.html |
| Page 3 | Groups 10-14 | accc_lavasr_p3.html |
| Page 4 | Groups 15-19 | accc_lavasr_p4.html |
| Page 5 | Groups 20-24 | accc_lavasr_p5.html |
| Page 6 | Groups 25-29 | accc_lavasr_p6.html |
| Page 7 | Groups 30-34 | accc_lavasr_p7.html |
| Page 8 | Groups 35-39 | accc_lavasr_p8.html |
| Page 9 | Groups 40-44 | accc_lavasr_p9.html |
| Page 10 | Groups 45-49 | accc_lavasr_p10.html |
Also available:
The pipeline supports 12 generation paths organized into three families. Each path uses a different sampling strategy to produce diverse voice acting data.
| Path | Sampling | Description | Details |
|---|---|---|---|
| A (VoiceNet) | 57 VoiceNet dims + EmoNet + Vocal Bursts | Full taxonomy sampling: 3 mandatory dims (Tempo, Gender, Age) + 5 random, 1-3 emotions, flow style, mandatory words | Path A Details |
| B (Archetype) | 920 archetypes x 92 genres | Genre/character archetype-based: random archetype + emotions + Tempo/Arousal | Path B Details |
| C (Archetype Named) | Same as B + explicit naming | Archetype with explicit role naming in the DramaBox script (e.g. "a battle-hardened noble knight") | Path C Details |
| D (Reference Audio) | Timbre whisper + VoiceNet + Chatterbox VC | Reference audio pipeline: timbre caption guides prompt, DramaBox TTS + voice conversion to match reference speaker | Path D Details |
| AC (Acting Challenge) | 19,247 acting challenges + VoiceNet gender/age | Audition-style method acting from challenge scenarios β naturalistic, genuine, dynamic emotional arc | AC Details |
| SIT (Situation) | 289 situations x EmoNet emotions | Situation-driven acting: actor is physically/socially IN a specific situation from the Situation Taxonomy (body posture, activity, social context, environment, health, climate, fatigue, pain) with sampled emotions | SIT Details |
All CC paths generate two scenes with the same speaker in contrasting emotional states, separated by a "CUT TO:" marker. The speaker's fundamental voice (age, gender, timbre) stays identical β only the emotional delivery changes. Audio is split into Scene 1 / Scene 2 using LLM-guided splitting (Gemma 4 E4B-it + Whisper turbo word-level timestamps + quiet-spot detection).
| Path | Sampling | Key Improvement | Details |
|---|---|---|---|
| CC-A (VoiceNet) | VoiceNet + contrasting emotions | Original two-scene format | CC Details |
| CC-B (Archetype) | Archetype + contrasting emotions | Original two-scene format | CC Details |
| CC-C (Archetype Named) | Archetype named + contrasting emotions | Original two-scene format | CC Details |
| CC2-A (VoiceNet v2) | VoiceNet + contrasting emotions | Enhanced: explicit emotional scene setup + dramatic transition descriptions | CC2 Details |
| CC2-B (Archetype v2) | Archetype + contrasting emotions | Enhanced: genuine/spontaneous/authentic delivery emphasis | CC2 Details |
| CC2-C (Archetype Named v2) | Archetype named + contrasting emotions | Enhanced: visceral emotional contrast, human-sounding | CC2 Details |
| ACCC (Acting Challenge CC) | Acting challenge + VoiceNet gender/age | Challenge-driven two-scene format β same actor, same challenge, contrasting emotional moments | ACCC Details |
| SIT-CC (Situation CC) | Situation + EmoNet + contrasting emotions | Two-scene situation-driven format β same actor IN the same situation, two contrasting emotional moments (5,749 pre-generated prompts in en/fr/es/de) | SIT-CC Details |
All paths use Sidon (w2v-BERT LoRA encoder + DAC decoder) for speech restoration:
Optionally, Chatterbox VC (S3Gen flow-matching VC) is applied before Sidon:
Each enhanced candidate is scored using a native PyTorch DNS-MOS model:
Two-scene audio is split using a three-phase LLM-guided pipeline:
For each group of candidates, 46 ranking methods are available:
Quality/CLAP methods (6):
| Method | Formula |
|---|---|
| v_snr_L (default) | (1 - WER) x (san_L - neg_san_L + 2) |
| v_snr_S | (1 - WER) x (san_S - neg_san_S + 2) |
| v_san_L | (1 - WER) x (san_L + 1) |
| v_san_S | (1 - WER) x (san_S + 1) |
| Content Enjoyment | Raw Empathic Insight Plus score |
| Standard | (1 - WER) x Content Enjoyment |
Where:
EmoNet emotion methods (40):
Each of the 40 EmoNet emotion dimensions from Empathic Insight Plus is a separate ranking method. Audio is encoded with BUD-E-Whisper (768-dim), pooled (mean+min+max+std = 3072-dim), then scored by 40 specialized MLP expert heads.
Emotion rankings use a WER < 10% hard cutoff β only candidates that said the right words qualify. Within qualifying candidates, they are ranked by descending emotion score.
The 40 emotions: Affection, Amusement, Anger, Astonishment/Surprise, Awe, Bitterness, Concentration, Confusion, Contemplation, Contempt, Contentment, Disappointment, Disgust, Distress, Doubt, Elation, Embarrassment, Emotional Numbness, Fatigue/Exhaustion, Fear, Helplessness, Hope/Enthusiasm/Optimism, Impatience/Irritability, Infatuation, Interest, Intoxication/Altered States, Jealousy/Envy, Longing, Malevolence/Malice, Pain, Pleasure/Ecstasy, Pride, Relief, Sadness, Sexual Lust, Shame, Sourness, Teasing, Thankfulness/Gratitude, Triumph.
The pipeline samples from several structured taxonomies to create diverse, controlled voice performances:
| Taxonomy | Size | Format | Documentation |
|---|---|---|---|
| VoiceNet | 57 dimensions x 7 levels | HTML | Taxonomy docs / Interactive viewer |
| VoiceNet Extension | Situation-dependent dims | HTML | Interactive viewer |
| EmoNet | 40 emotions x 4 intensity levels | JSON | Taxonomy docs |
| Vocal Bursts | 120 non-linguistic sounds | JSON | Taxonomy docs |
| Character Archetypes | 920 archetypes x 92 genres | JSON | Taxonomy docs |
| Acting Challenges | 19,247 challenge scenarios | JSON | Preview (100 samples) |
| Extreme Physical | 6 categories, 60 subcategories, 600 challenges | JSON | Tension (100), Breathlessness (100), Pain (100), Temperature (100), Taste (100), Surprise (100) |
| Situation Taxonomy | 11 dimensions, 289 situations | JSON | Data file β Body posture (32), physical activity (69), speaking target (25), social context (56), environment (22), health (18), face/head gear (14), climate (10), substances (12), fatigue (19), pain (12) |
Paper references:
79,087 ready-to-use DramaBox two-scene CUT TO: prompts across all pathways and languages:
| File | Pathway | Count | Language | Examples |
|---|---|---|---|---|
dramabox_cca_voicenet.json | CC-A (VoiceNet) | 19,332 | English | Examples |
dramabox_cc2c_archetype.json | CC2-C (Archetype) | 9,999 | English | Examples |
dramabox_accc_acting_challenge.json | ACCC (Acting Challenge) | 12,893 | English | Examples |
dramabox_sit_situation.json | SIT (Situation) | 5,749 | en/fr/es/de | Examples |
dramabox_extreme_physical.json | Extreme Physical | 600 | English | Examples |
dramabox_cca_voicenet_de.json | CC-A (VoiceNet) | 9,983 | German | Examples |
dramabox_cc2c_archetype_de.json | CC2-C (Archetype) | 9,983 | German | Examples |
dramabox_accc_acting_challenge_de.json | ACCC (Acting Challenge) | 9,948 | German | Examples |
dramabox_extreme_physical_de.json | Extreme Physical | 600 | German | Examples |
German prompts use oe/ae/ue instead of umlauts (ΓΆ/Γ€/ΓΌ). Directions and speaker descriptions are in English; only spoken dialogue (in "double quotes") is in the target language.
Full 57-dimension voice attribute sampling. The most granular control over voice performance.
See docs/path_a_voicenet.md for full details.
Genre/character archetype-based sampling. Focuses on character identity over individual vocal dimensions.
See docs/path_b_archetype.md for full details.
Same as Path B but with explicit instruction to name the archetype role in the DramaBox script output (e.g. "a battle-hardened noble knight" in the speaker description and stage directions). This gives DramaBox TTS a stronger character signal.
See docs/path_c_archetype_named.md for full details.
The most promising path for voice cloning. Uses reference audio's timbre whisper caption to guide prompt generation, then voice-converts the DramaBox TTS output to match the reference speaker.
laion/timbre-whisper)voice_ref directly to DramaBox leads to unstable/garbled generationsWhy text-only TTS + VC? The timbre whisper caption gives Gemma 4 a rich description of the target speaker's vocal qualities, which guides the LLM to produce a speaker-consistent DramaBox script. Chatterbox VC then handles the actual voice transfer. This two-stage approach is far more stable than passing
voice_refdirectly to DramaBox, which causes garbled or incoherent audio output.
See docs/path_d_reference.md for full details.
Audition-style method acting performances driven by acting challenge scenarios. Samples from 19,247 structured challenges covering diverse emotional and situational contexts.
Key characteristics:
See docs/path_ac_acting_challenge.md for full details.
Situation-driven acting challenges where the actor is physically and socially embedded in a specific situation from the Situation Taxonomy. The taxonomy covers 11 dimensions with 289 total situations describing how body posture, physical activity, speaking target, social context, environment, health conditions, face/head gear, climate, substances, fatigue, and pain affect the voice.
The situation taxonomy is based on the Extended VoiceNet Taxonomy (interactive viewer) which extends the core 57 VoiceNet voice attribute dimensions with situation-dependent dimensions that describe how the speaker's physical state and environment affect their vocal output.
See docs/path_sit_situation.md for full details.
All CC paths produce two scenes with the same speaker in contrasting emotional states, separated by a "CUT TO:" marker.
The original two-scene format. Three sampling variants matching standalone Paths A, B, C:
Emotion contrast logic: If Scene 1 has positive emotions -> Scene 2 samples from negative emotions (and vice versa). Word count: 50-80 total (~25-40 per scene).
See docs/path_cc_character_consistent.md for full details.
Improved version of CC with enhanced LLM prompting:
See docs/path_cc2_character_consistent_v2.md for full details.
Challenge-driven two-scene format: same actor performing the same acting challenge at two different emotional moments with dramatically shifted delivery. This is the primary path used in the current experiment.
See docs/path_ac_acting_challenge.md#accc-character-consistent for full details.
git clone https://github.com/LAION-AI/Voice-Acting-Pipeline.git
cd Voice-Acting-Pipeline
pip install -e .
For TTS synthesis (requires GPU with ~24GB VRAM):
pip install -e ".[tts]"
For audio refinement and scoring:
pip install -e ".[refinement,scoring]"
# Generate 1000 DramaBox prompts using GPUs 0 and 1
dramabox generate-prompts --config config.json --total 1000 --gpus 0,1
# Synthesize audio from an existing CSV
dramabox synthesize --csv output/dramabox_chunk_000.csv --gpus 0,1,2,3
# Generate prompts and immediately synthesize audio
dramabox run --config config.json --total 1000 --gpus 0,1,2,3
dramabox reference --config config.json --ref-dir /path/to/references --total 10 --gpus 6,7
# Full 4-path demo: A + B + C + D, 10 prompts each, best-of-3 scoring
dramabox demo --config config.json --full --n-prompts 10 --best-of-n 3 --gpus 6,7
dramabox score --audio output/audio/sample_000000_raw.wav --prompt "prompt text" --gpu 0
All parameters are in config.json. See config_schema.md for full documentation of every field.
| Section | Parameter | Default | Description |
|---|---|---|---|
prompt_generation | llm_model | google/gemma-4-E4B-it | LLM for prompt generation |
prompt_generation | total_prompts | 100000 | Number of prompts to generate |
sampling | archetype_ratio | 0.20 | Fraction using archetype path |
sampling | word_count_min/max | 10 / 60 | Target dialogue word count range |
tts | cfg_scale | 2.0 | Classifier-free guidance scale |
tts | steps | 30 | Euler flow matching steps |
best_of_n | n_candidates | 3 | Candidates per Best-of-N ranking |
Languages are configured in config.json. Currently active: English, German, French, Spanish. Ready to enable: Italian, Dutch, Russian, Portuguese, Chinese, Japanese, Korean, Arabic, Hindi, Turkish, Polish, Swedish.
| Model | Purpose | VRAM |
|---|---|---|
google/gemma-4-E4B-it | DramaBox prompt generation | ~16GB |
ResembleAI/Dramabox | TTS synthesis (22B DiT) | ~24GB |
sarulab-speech/sidon_raw_weight | Speech restoration (w2v-BERT LoRA + DAC, 16kHz->48kHz) | ~4GB |
| Chatterbox VC | Voice conversion (S3Gen flow-matching VC) | ~4GB |
| DNS-MOS (PyTorch) | Quality scoring (SIG/BAK/OVR, 1-5 scale) | ~0.1GB |
laion/VoiceCLAP | Audio-text similarity scoring (Large 3584-dim + Small 768-dim) | ~2GB |
laion/Empathic-Insight-Voice-Plus | 40 EmoNet emotion scoring + content enjoyment (BUD-E-Whisper + MLP) | ~2GB |
laion/BUD-E-Whisper | Audio encoder for emotion scoring (768-dim embeddings) | ~1GB |
| Whisper turbo | Word-level timestamps for ASR + audio splitting | ~3GB |
nvidia/parakeet-tdt-0.6b-v3 | ASR for WER scoring | ~2GB |
laion/timbre-whisper | On-the-fly timbre captioning (Path D) | ~2GB |
| MOSS-Audio-8B-Thinking | Audio-guided prompt re-annotation (3 passes per sample) | ~8GB |
Voice-Acting-Pipeline/
βββ README.md # This file
βββ LAION-Voice-Whitepaper.md # Dataset plan: LAION Voice + Voice Acting corpus
βββ config.json # All configurable parameters
βββ config_schema.md # Documentation for config fields
βββ pyproject.toml # Python packaging
βββ run_sample_groups.py # Full pipeline: TTS + Sidon/VC + ASR + LLM splitting + scoring + HTML
βββ run_reannotate.py # Standalone Gemma 4 re-annotation + HTML report rebuild
βββ run_resplit.py # Standalone LLM-guided CUT TO: re-splitting
βββ data/
β βββ voicenet_ext_taxonomy.html # VoiceNet (57 dims x 7 levels)
β βββ all_acting_challenges.json # 19,247 acting challenge scenarios
β βββ acting_challenges_situation_inspired.json # 5,749 situation-inspired challenges
β βββ acting_challenges_eric_morris_inspired.json # 4,030 Eric Morris-inspired challenges
β βββ acting_challenges_existing_inspired.json # 7,390 existing challenge variants
β βββ dramabox_cca_voicenet.json # 19,332 pre-generated CC-A DramaBox prompts
β βββ dramabox_cc2c_archetype.json # 9,999 pre-generated CC2-C DramaBox prompts
β βββ dramabox_accc_acting_challenge.json # 12,893 pre-generated ACCC DramaBox prompts
β βββ dramabox_sit_situation.json # 5,749 pre-generated SIT DramaBox prompts (en/fr/es/de)
β βββ dramabox_cca_voicenet_de.json # 9,983 German CC-A DramaBox prompts (no umlauts)
β βββ dramabox_cc2c_archetype_de.json # 9,983 German CC2-C DramaBox prompts (no umlauts)
β βββ dramabox_accc_acting_challenge_de.json # 9,948 German ACCC DramaBox prompts (no umlauts)
β βββ dramabox_extreme_physical.json # 600 extreme physical DramaBox prompts
β βββ dramabox_extreme_physical_de.json # 600 German extreme physical DramaBox prompts
β βββ acting_challenges_extreme_physical.json # 600 extreme physical challenges
β βββ extreme_physical_taxonomy.json # 6 categories x 10 subcategories taxonomy
β βββ situation_taxonomy.json # Situation taxonomy (poses, activities, contexts)
β βββ emonet_taxonomy.json # EmoNet (40 emotions x 4 intensity levels)
β βββ vocal_bursts_taxonomy.json # Vocal bursts (120 types)
β βββ archetypes.json # Archetypes (920 x 92 genres)
β βββ wordlists/ # Per-language word lists
βββ dramabox/
β βββ cli.py # CLI entry point
β βββ config_loader.py # Config loading and validation
β βββ taxonomy.py # Taxonomy parsers and loaders
β βββ sampling.py # Path A + Path B sampling
β βββ reference_sampling.py # Path D: reference audio sampling
β βββ prompts.py # LLM prompt construction
β βββ prompt_generator.py # Multi-GPU LLM batch generation
β βββ tts_synthesizer.py # Multi-GPU DramaBox TTS
β βββ sidon_enhance.py # Sidon + ChatterboxVC augmentation (replaces RE-USE + LavaSR)
β βββ reuse_enhance.py # RE-USE speech enhancement (legacy)
β βββ moss_refine.py # MOSS-Audio re-annotation (audio-guided prompt rewriting)
β βββ moss_pipeline.py # MOSS orchestrator (multi-GPU job distribution)
β βββ scoring.py # ASR WER + content enjoyment + EmoNet scoring
β βββ demo_grid.py # HTML demo grid generator
β βββ pipeline.py # Mode 1-6 orchestrator
βββ scripts/
β βββ _accc_lavasr_pipeline.py # ACCC LavaSR experiment pipeline (current)
β βββ _emonet_worker.py # EmoNet 40-emotion GPU worker
β βββ _score_worker.py # WER + content enjoyment GPU worker
β βββ _lavasr_worker.py # LavaSR BWE GPU worker
β βββ _lavasr_clap_worker.py # VoiceCLAP scoring GPU worker
βββ docs/
β βββ voicenet_taxonomy.md # VoiceNet 57-dim taxonomy
β βββ voicenet_extension_taxonomy.html # Interactive VoiceNet viewer
β βββ emonet_taxonomy.md # EmoNet 40 emotions
β βββ vocal_bursts_taxonomy.md # 120 vocal bursts
β βββ archetypes.md # 920 archetypes
β βββ acting_challenges_preview.html # Acting challenge preview (100 samples)
β βββ paper_reference.md # Citation and BibTeX
β βββ path_a_voicenet.md # Path A detailed docs
β βββ path_b_archetype.md # Path B detailed docs
β βββ path_c_archetype_named.md # Path C detailed docs
β βββ path_d_reference.md # Path D detailed docs
β βββ path_ac_acting_challenge.md # AC + ACCC detailed docs
β βββ path_sit_situation.md # SIT + SIT-CC situation pathway docs
β βββ path_cc_character_consistent.md # CC v1 detailed docs
β βββ path_cc2_character_consistent_v2.md # CC2 v2 detailed docs
β βββ dramabox_cca_voicenet_examples.md # CC-A English examples
β βββ dramabox_cc2c_archetype_examples.md # CC2-C English examples
β βββ dramabox_accc_acting_challenge_examples.md # ACCC English examples
β βββ dramabox_sit_situation_examples.md # SIT multilingual examples
β βββ dramabox_cca_voicenet_de_examples.md # CC-A German examples
β βββ dramabox_cc2c_archetype_de_examples.md # CC2-C German examples
β βββ dramabox_accc_acting_challenge_de_examples.md # ACCC German examples
β βββ dramabox_extreme_physical_examples.md # Extreme Physical English examples
β βββ dramabox_extreme_physical_de_examples.md # Extreme Physical German examples
β βββ demo/ # HTML demo grids with embedded audio
β βββ sidon_vc_sample_groups.html # Sidon+VC experiment (20 groups, LLM splits)
β βββ accc_lavasr.html # ACCC LavaSR index (redirects to page 1)
β βββ accc_lavasr_p1.html # ACCC LavaSR grid pages 1-10
β βββ ...
β βββ pitch_analysis.html # Pitch analysis
βββ examples/
βββ example_prompt.txt # Sample DramaBox prompt
| Component | Minimum | Recommended |
|---|---|---|
| Prompt generation | 1 GPU, 16GB VRAM | 4+ GPUs, 16GB+ each |
| TTS synthesis | 1 GPU, 24GB VRAM | 4+ GPUs, 24GB+ each |
| Sidon + ChatterboxVC augmentation | 1 GPU, 12GB VRAM | 8 GPUs (parallel workers) |
| MOSS re-annotation | 1 GPU, 8GB VRAM (4-bit) | 8 GPUs (parallel workers) |
| VoiceCLAP + EmoNet scoring | 1 GPU, 4GB VRAM | 8 GPUs (parallel workers) |
| RAM | 32GB | 64GB+ |
After post-processing, the pipeline runs a MOSS re-annotation pass (dramabox/moss_refine.py) that closes the loop between intended and actual performance.
Original DramaBox prompts are directions β the TTS model interprets them, and the actual audio may differ from what was requested. MOSS-Audio-8B-Thinking listens to each generated audio clip together with the original prompt and ASR transcript, then rewrites the prompt to match what was actually performed.
Original DramaBox Prompt
+
ASR Transcript (from Whisper)
+
Generated Audio (MP3)
|
v
+---------------------------+
| MOSS-Audio-8B-Thinking |
| (4-bit, per-GPU worker) |
| |
| Listens to audio + |
| reads text context |
| |
| Rewrites prompt to |
| match actual performance |
+---------------------------+
|
v
moss_refined_prompt_full
moss_refined_prompt_part1
moss_refined_prompt_part2
# Run MOSS re-annotation on all post-processed samples
python dramabox/moss_refine.py # All GPUs
python dramabox/moss_refine.py --num-gpus 4 # 4 GPUs
python dramabox/moss_refine.py --test # First 10 samples, 1 GPU
Requires /tmp/moss_venv with transformers==4.57.1 (MOSS is incompatible with transformers >= 5.x).
We trained one LoRA adapter per EmoNet emotion for the MOSS-TTS-Local-Transformer-4.55B voice-acting model, so an emotion can be applied on top of ordinary voice cloning. Each adapter is the best of a rank {16, 32, 64} sweep: for scarce data a lower rank curbs overfitting, and the winner is picked by a combined score over emotion intensity, vocal-burst-blend, genuineness and speaker similarity.
Data (per emotion, ~1,000β1,400 clips): natural in-the-wild speech (EmoLia) is preferred and filled
with the most intense LAION's Got Talent clips; ~25% of rows carry a same-voice, contrasting-emotion
reference; instructions are procedural emotion captions from
procedural-voice-captions.
TTS-AGI/moss-emotion-loras-40The demo plays five neutral sentences per emotion in three languages, comparing the LoRA with and without an emotion reference; the LoRA raises the automatic EmoNet emotion score while keeping the voice and naturalness.
Python
62.2%
HTML
37.8%
Open-weights voice acting data pipeline combining structured taxonomy sampling, DramaBox TTS synthesis, Sidon speech restoration, and ChatterboxVC augmentation with best-of-N ranking across 46 scoring methods.
Live Demo: Sidon+VC Sample Groups β 20 groups x 25 candidates with LLM-guided CUT TO: splitting, Whisper turbo ASR, and Gemma 4 re-annotation.
Benchmarks (landing page):
- π¬ Vanilla DramaBox TTS β generate + reward-rank β raw two-scene
CUT TO:DramaBox prompts fed directly to the 8B MOSS voice-acting TTS, 4 seeds, scored + reward-ranked; listenable takes, best-of-k quality/compute trade-off, and the full k=1..32 seed-scaling walltime table.- β‘ Local-LLM DramaBox prompt-generation throughput β per-pathway/language token + throughput stats with a 1M-prompt estimate, plus the DramaBox TTS seed-scaling table for best-of-k planning.
- π How the DramaBox prompt dataset is sampled & generated β a plain-English, reproducible walkthrough of the
laion/dramabox-cutscene-promptsdataset: the 5 sampling pathways, every taxonomy it draws from (VoiceNet, 40 EmoNet emotions, archetypes, situations, 180 vocal bursts), the exact prompts the model receives, and real generated examples with their sampled metadata.
Dataset Plan: See the full technical white paper β Towards an Emotionally Expressive Audio Omni-Model β for the complete LAION Voice and LAION Voice Acting dataset construction plan, model inventory, and annotation strategy.
End-to-end voice prompt generation and audio synthesis using the DramaBox TTS model (22B DiT transformer) and structured voice taxonomy sampling. Based on the voice taxonomy research from Schuhmann et al., 2025 and EmoNet-Voice (Schuhmann et al., 2025).
This pipeline generates richly annotated voice performance prompts in the DramaBox format β single-speaker scenes with stage directions (English) and spoken dialogue (target language) β then synthesizes them into audio. Each prompt is procedurally constructed by sampling from structured taxonomies, then expanded by an LLM (Gemma 4 E4B-it) into a full performance script.
The full audio processing chain:
Taxonomy Sampling LLM Prompt Gen DramaBox TTS (22B)
(VoiceNet/Archetype/ (Gemma 4 E4B-it) CFG=2.5, STG=1.5
Situation/ActingChall) 25 candidates/prompt
| | |
+------------------------+ |
v
+---------------------------+
| Per-Sample Augment |
| |
| Path A: Sidon only |
| (16kHz -> 48kHz) |
| Path B: ChatterboxVC |
| + Sidon (VC -> restore)|
| |
| Pick best by |
| DNS-MOS OVR score |
+-------------+-------------+
|
+-------------v-------------+
| Whisper Turbo ASR |
| (word-level timestamps) |
+-------------+-------------+
|
+-------------v-------------+
| LLM-Guided CUT TO: |
| Split (Gemma 4 E4B-it) |
| + quiet-spot detection |
| + LLM fade strategy |
+-------------+-------------+
|
+-------------v-------------+
| Best-of-25 Ranking |
| (WER, VoiceCLAP, |
| EmoNet, Content) |
+-------------+-------------+
|
+-------------v-------------+
| Gemma 4 Re-annotation |
| (ASR -> refined prompt) |
+---------------------------+
For each raw TTS candidate, two enhancement paths run and the best is selected:
Raw TTS Audio βββ¬βββΊ Sidon Speech Restoration βββΊ DNS-MOS βββ
(from DramaBox) β (w2v-BERT LoRA + DAC) ββββΊ Pick higher OVR
β (16kHz β 48kHz) β
ββββΊ ChatterboxVC βββΊ Sidon βββΊ DNS-MOS βββββ
(S3Gen flow-matching VC,
self-VC or ref-VC)
For two-scene audio, a self-VC of the full audio provides the VC target for speaker consistency:
Full Audio βββΊ Self-VC βββΊ Sidon βββΊ full_enhanced (VC target)
β
βββββββββββββββββββββββββββββββββ
β
Part 1 βββ¬βββΊ Sidon only βββΊ DNS-MOS βββ
β ββββΊ Pick best
ββββΊ VC(βfull_enhanced) + Sidon βββΊ DNS-MOS βββ
Part 2 βββ¬βββΊ Sidon only βββΊ DNS-MOS βββ
β ββββΊ Pick best
ββββΊ VC(βfull_enhanced) + Sidon βββΊ DNS-MOS βββ
Scoring methods (46 total):
Listen to the latest Sidon+VC experiment β 20 groups x 25 candidates = 500 audio clips with LLM-guided CUT TO: splitting and Gemma 4 re-annotation.
| Demo | Description | Link |
|---|---|---|
| Sidon+VC Sample Groups | 20 groups, best-of-25, LLM-guided splits | sidon_vc_sample_groups.html |
50 groups x 25 candidates = 1,250 audio clips across 10 pages. Each page has an interactive ranking dropdown with 46 methods.
| Page | Groups | Link |
|---|---|---|
| Page 1 | Groups 0-4 | accc_lavasr_p1.html |
| Page 2 | Groups 5-9 | accc_lavasr_p2.html |
| Page 3 | Groups 10-14 | accc_lavasr_p3.html |
| Page 4 | Groups 15-19 | accc_lavasr_p4.html |
| Page 5 | Groups 20-24 | accc_lavasr_p5.html |
| Page 6 | Groups 25-29 | accc_lavasr_p6.html |
| Page 7 | Groups 30-34 | accc_lavasr_p7.html |
| Page 8 | Groups 35-39 | accc_lavasr_p8.html |
| Page 9 | Groups 40-44 | accc_lavasr_p9.html |
| Page 10 | Groups 45-49 | accc_lavasr_p10.html |
Also available:
The pipeline supports 12 generation paths organized into three families. Each path uses a different sampling strategy to produce diverse voice acting data.
| Path | Sampling | Description | Details |
|---|---|---|---|
| A (VoiceNet) | 57 VoiceNet dims + EmoNet + Vocal Bursts | Full taxonomy sampling: 3 mandatory dims (Tempo, Gender, Age) + 5 random, 1-3 emotions, flow style, mandatory words | Path A Details |
| B (Archetype) | 920 archetypes x 92 genres | Genre/character archetype-based: random archetype + emotions + Tempo/Arousal | Path B Details |
| C (Archetype Named) | Same as B + explicit naming | Archetype with explicit role naming in the DramaBox script (e.g. "a battle-hardened noble knight") | Path C Details |
| D (Reference Audio) | Timbre whisper + VoiceNet + Chatterbox VC | Reference audio pipeline: timbre caption guides prompt, DramaBox TTS + voice conversion to match reference speaker | Path D Details |
| AC (Acting Challenge) | 19,247 acting challenges + VoiceNet gender/age | Audition-style method acting from challenge scenarios β naturalistic, genuine, dynamic emotional arc | AC Details |
| SIT (Situation) | 289 situations x EmoNet emotions | Situation-driven acting: actor is physically/socially IN a specific situation from the Situation Taxonomy (body posture, activity, social context, environment, health, climate, fatigue, pain) with sampled emotions | SIT Details |
All CC paths generate two scenes with the same speaker in contrasting emotional states, separated by a "CUT TO:" marker. The speaker's fundamental voice (age, gender, timbre) stays identical β only the emotional delivery changes. Audio is split into Scene 1 / Scene 2 using LLM-guided splitting (Gemma 4 E4B-it + Whisper turbo word-level timestamps + quiet-spot detection).
| Path | Sampling | Key Improvement | Details |
|---|---|---|---|
| CC-A (VoiceNet) | VoiceNet + contrasting emotions | Original two-scene format | CC Details |
| CC-B (Archetype) | Archetype + contrasting emotions | Original two-scene format | CC Details |
| CC-C (Archetype Named) | Archetype named + contrasting emotions | Original two-scene format | CC Details |
| CC2-A (VoiceNet v2) | VoiceNet + contrasting emotions | Enhanced: explicit emotional scene setup + dramatic transition descriptions | CC2 Details |
| CC2-B (Archetype v2) | Archetype + contrasting emotions | Enhanced: genuine/spontaneous/authentic delivery emphasis | CC2 Details |
| CC2-C (Archetype Named v2) | Archetype named + contrasting emotions | Enhanced: visceral emotional contrast, human-sounding | CC2 Details |
| ACCC (Acting Challenge CC) | Acting challenge + VoiceNet gender/age | Challenge-driven two-scene format β same actor, same challenge, contrasting emotional moments | ACCC Details |
| SIT-CC (Situation CC) | Situation + EmoNet + contrasting emotions | Two-scene situation-driven format β same actor IN the same situation, two contrasting emotional moments (5,749 pre-generated prompts in en/fr/es/de) | SIT-CC Details |
All paths use Sidon (w2v-BERT LoRA encoder + DAC decoder) for speech restoration:
Optionally, Chatterbox VC (S3Gen flow-matching VC) is applied before Sidon:
Each enhanced candidate is scored using a native PyTorch DNS-MOS model:
Two-scene audio is split using a three-phase LLM-guided pipeline:
For each group of candidates, 46 ranking methods are available:
Quality/CLAP methods (6):
| Method | Formula |
|---|---|
| v_snr_L (default) | (1 - WER) x (san_L - neg_san_L + 2) |
| v_snr_S | (1 - WER) x (san_S - neg_san_S + 2) |
| v_san_L | (1 - WER) x (san_L + 1) |
| v_san_S | (1 - WER) x (san_S + 1) |
| Content Enjoyment | Raw Empathic Insight Plus score |
| Standard | (1 - WER) x Content Enjoyment |
Where:
EmoNet emotion methods (40):
Each of the 40 EmoNet emotion dimensions from Empathic Insight Plus is a separate ranking method. Audio is encoded with BUD-E-Whisper (768-dim), pooled (mean+min+max+std = 3072-dim), then scored by 40 specialized MLP expert heads.
Emotion rankings use a WER < 10% hard cutoff β only candidates that said the right words qualify. Within qualifying candidates, they are ranked by descending emotion score.
The 40 emotions: Affection, Amusement, Anger, Astonishment/Surprise, Awe, Bitterness, Concentration, Confusion, Contemplation, Contempt, Contentment, Disappointment, Disgust, Distress, Doubt, Elation, Embarrassment, Emotional Numbness, Fatigue/Exhaustion, Fear, Helplessness, Hope/Enthusiasm/Optimism, Impatience/Irritability, Infatuation, Interest, Intoxication/Altered States, Jealousy/Envy, Longing, Malevolence/Malice, Pain, Pleasure/Ecstasy, Pride, Relief, Sadness, Sexual Lust, Shame, Sourness, Teasing, Thankfulness/Gratitude, Triumph.
The pipeline samples from several structured taxonomies to create diverse, controlled voice performances:
| Taxonomy | Size | Format | Documentation |
|---|---|---|---|
| VoiceNet | 57 dimensions x 7 levels | HTML | Taxonomy docs / Interactive viewer |
| VoiceNet Extension | Situation-dependent dims | HTML | Interactive viewer |
| EmoNet | 40 emotions x 4 intensity levels | JSON | Taxonomy docs |
| Vocal Bursts | 120 non-linguistic sounds | JSON | Taxonomy docs |
| Character Archetypes | 920 archetypes x 92 genres | JSON | Taxonomy docs |
| Acting Challenges | 19,247 challenge scenarios | JSON | Preview (100 samples) |
| Extreme Physical | 6 categories, 60 subcategories, 600 challenges | JSON | Tension (100), Breathlessness (100), Pain (100), Temperature (100), Taste (100), Surprise (100) |
| Situation Taxonomy | 11 dimensions, 289 situations | JSON | Data file β Body posture (32), physical activity (69), speaking target (25), social context (56), environment (22), health (18), face/head gear (14), climate (10), substances (12), fatigue (19), pain (12) |
Paper references:
79,087 ready-to-use DramaBox two-scene CUT TO: prompts across all pathways and languages:
| File | Pathway | Count | Language | Examples |
|---|---|---|---|---|
dramabox_cca_voicenet.json | CC-A (VoiceNet) | 19,332 | English | Examples |
dramabox_cc2c_archetype.json | CC2-C (Archetype) | 9,999 | English | Examples |
dramabox_accc_acting_challenge.json | ACCC (Acting Challenge) | 12,893 | English | Examples |
dramabox_sit_situation.json | SIT (Situation) | 5,749 | en/fr/es/de | Examples |
dramabox_extreme_physical.json | Extreme Physical | 600 | English | Examples |
dramabox_cca_voicenet_de.json | CC-A (VoiceNet) | 9,983 | German | Examples |
dramabox_cc2c_archetype_de.json | CC2-C (Archetype) | 9,983 | German | Examples |
dramabox_accc_acting_challenge_de.json | ACCC (Acting Challenge) | 9,948 | German | Examples |
dramabox_extreme_physical_de.json | Extreme Physical | 600 | German | Examples |
German prompts use oe/ae/ue instead of umlauts (ΓΆ/Γ€/ΓΌ). Directions and speaker descriptions are in English; only spoken dialogue (in "double quotes") is in the target language.
Full 57-dimension voice attribute sampling. The most granular control over voice performance.
See docs/path_a_voicenet.md for full details.
Genre/character archetype-based sampling. Focuses on character identity over individual vocal dimensions.
See docs/path_b_archetype.md for full details.
Same as Path B but with explicit instruction to name the archetype role in the DramaBox script output (e.g. "a battle-hardened noble knight" in the speaker description and stage directions). This gives DramaBox TTS a stronger character signal.
See docs/path_c_archetype_named.md for full details.
The most promising path for voice cloning. Uses reference audio's timbre whisper caption to guide prompt generation, then voice-converts the DramaBox TTS output to match the reference speaker.
laion/timbre-whisper)voice_ref directly to DramaBox leads to unstable/garbled generationsWhy text-only TTS + VC? The timbre whisper caption gives Gemma 4 a rich description of the target speaker's vocal qualities, which guides the LLM to produce a speaker-consistent DramaBox script. Chatterbox VC then handles the actual voice transfer. This two-stage approach is far more stable than passing
voice_refdirectly to DramaBox, which causes garbled or incoherent audio output.
See docs/path_d_reference.md for full details.
Audition-style method acting performances driven by acting challenge scenarios. Samples from 19,247 structured challenges covering diverse emotional and situational contexts.
Key characteristics:
See docs/path_ac_acting_challenge.md for full details.
Situation-driven acting challenges where the actor is physically and socially embedded in a specific situation from the Situation Taxonomy. The taxonomy covers 11 dimensions with 289 total situations describing how body posture, physical activity, speaking target, social context, environment, health conditions, face/head gear, climate, substances, fatigue, and pain affect the voice.
The situation taxonomy is based on the Extended VoiceNet Taxonomy (interactive viewer) which extends the core 57 VoiceNet voice attribute dimensions with situation-dependent dimensions that describe how the speaker's physical state and environment affect their vocal output.
See docs/path_sit_situation.md for full details.
All CC paths produce two scenes with the same speaker in contrasting emotional states, separated by a "CUT TO:" marker.
The original two-scene format. Three sampling variants matching standalone Paths A, B, C:
Emotion contrast logic: If Scene 1 has positive emotions -> Scene 2 samples from negative emotions (and vice versa). Word count: 50-80 total (~25-40 per scene).
See docs/path_cc_character_consistent.md for full details.
Improved version of CC with enhanced LLM prompting:
See docs/path_cc2_character_consistent_v2.md for full details.
Challenge-driven two-scene format: same actor performing the same acting challenge at two different emotional moments with dramatically shifted delivery. This is the primary path used in the current experiment.
See docs/path_ac_acting_challenge.md#accc-character-consistent for full details.
git clone https://github.com/LAION-AI/Voice-Acting-Pipeline.git
cd Voice-Acting-Pipeline
pip install -e .
For TTS synthesis (requires GPU with ~24GB VRAM):
pip install -e ".[tts]"
For audio refinement and scoring:
pip install -e ".[refinement,scoring]"
# Generate 1000 DramaBox prompts using GPUs 0 and 1
dramabox generate-prompts --config config.json --total 1000 --gpus 0,1
# Synthesize audio from an existing CSV
dramabox synthesize --csv output/dramabox_chunk_000.csv --gpus 0,1,2,3
# Generate prompts and immediately synthesize audio
dramabox run --config config.json --total 1000 --gpus 0,1,2,3
dramabox reference --config config.json --ref-dir /path/to/references --total 10 --gpus 6,7
# Full 4-path demo: A + B + C + D, 10 prompts each, best-of-3 scoring
dramabox demo --config config.json --full --n-prompts 10 --best-of-n 3 --gpus 6,7
dramabox score --audio output/audio/sample_000000_raw.wav --prompt "prompt text" --gpu 0
All parameters are in config.json. See config_schema.md for full documentation of every field.
| Section | Parameter | Default | Description |
|---|---|---|---|
prompt_generation | llm_model | google/gemma-4-E4B-it | LLM for prompt generation |
prompt_generation | total_prompts | 100000 | Number of prompts to generate |
sampling | archetype_ratio | 0.20 | Fraction using archetype path |
sampling | word_count_min/max | 10 / 60 | Target dialogue word count range |
tts | cfg_scale | 2.0 | Classifier-free guidance scale |
tts | steps | 30 | Euler flow matching steps |
best_of_n | n_candidates | 3 | Candidates per Best-of-N ranking |
Languages are configured in config.json. Currently active: English, German, French, Spanish. Ready to enable: Italian, Dutch, Russian, Portuguese, Chinese, Japanese, Korean, Arabic, Hindi, Turkish, Polish, Swedish.
| Model | Purpose | VRAM |
|---|---|---|
google/gemma-4-E4B-it | DramaBox prompt generation | ~16GB |
ResembleAI/Dramabox | TTS synthesis (22B DiT) | ~24GB |
sarulab-speech/sidon_raw_weight | Speech restoration (w2v-BERT LoRA + DAC, 16kHz->48kHz) | ~4GB |
| Chatterbox VC | Voice conversion (S3Gen flow-matching VC) | ~4GB |
| DNS-MOS (PyTorch) | Quality scoring (SIG/BAK/OVR, 1-5 scale) | ~0.1GB |
laion/VoiceCLAP | Audio-text similarity scoring (Large 3584-dim + Small 768-dim) | ~2GB |
laion/Empathic-Insight-Voice-Plus | 40 EmoNet emotion scoring + content enjoyment (BUD-E-Whisper + MLP) | ~2GB |
laion/BUD-E-Whisper | Audio encoder for emotion scoring (768-dim embeddings) | ~1GB |
| Whisper turbo | Word-level timestamps for ASR + audio splitting | ~3GB |
nvidia/parakeet-tdt-0.6b-v3 | ASR for WER scoring | ~2GB |
laion/timbre-whisper | On-the-fly timbre captioning (Path D) | ~2GB |
| MOSS-Audio-8B-Thinking | Audio-guided prompt re-annotation (3 passes per sample) | ~8GB |
Voice-Acting-Pipeline/
βββ README.md # This file
βββ LAION-Voice-Whitepaper.md # Dataset plan: LAION Voice + Voice Acting corpus
βββ config.json # All configurable parameters
βββ config_schema.md # Documentation for config fields
βββ pyproject.toml # Python packaging
βββ run_sample_groups.py # Full pipeline: TTS + Sidon/VC + ASR + LLM splitting + scoring + HTML
βββ run_reannotate.py # Standalone Gemma 4 re-annotation + HTML report rebuild
βββ run_resplit.py # Standalone LLM-guided CUT TO: re-splitting
βββ data/
β βββ voicenet_ext_taxonomy.html # VoiceNet (57 dims x 7 levels)
β βββ all_acting_challenges.json # 19,247 acting challenge scenarios
β βββ acting_challenges_situation_inspired.json # 5,749 situation-inspired challenges
β βββ acting_challenges_eric_morris_inspired.json # 4,030 Eric Morris-inspired challenges
β βββ acting_challenges_existing_inspired.json # 7,390 existing challenge variants
β βββ dramabox_cca_voicenet.json # 19,332 pre-generated CC-A DramaBox prompts
β βββ dramabox_cc2c_archetype.json # 9,999 pre-generated CC2-C DramaBox prompts
β βββ dramabox_accc_acting_challenge.json # 12,893 pre-generated ACCC DramaBox prompts
β βββ dramabox_sit_situation.json # 5,749 pre-generated SIT DramaBox prompts (en/fr/es/de)
β βββ dramabox_cca_voicenet_de.json # 9,983 German CC-A DramaBox prompts (no umlauts)
β βββ dramabox_cc2c_archetype_de.json # 9,983 German CC2-C DramaBox prompts (no umlauts)
β βββ dramabox_accc_acting_challenge_de.json # 9,948 German ACCC DramaBox prompts (no umlauts)
β βββ dramabox_extreme_physical.json # 600 extreme physical DramaBox prompts
β βββ dramabox_extreme_physical_de.json # 600 German extreme physical DramaBox prompts
β βββ acting_challenges_extreme_physical.json # 600 extreme physical challenges
β βββ extreme_physical_taxonomy.json # 6 categories x 10 subcategories taxonomy
β βββ situation_taxonomy.json # Situation taxonomy (poses, activities, contexts)
β βββ emonet_taxonomy.json # EmoNet (40 emotions x 4 intensity levels)
β βββ vocal_bursts_taxonomy.json # Vocal bursts (120 types)
β βββ archetypes.json # Archetypes (920 x 92 genres)
β βββ wordlists/ # Per-language word lists
βββ dramabox/
β βββ cli.py # CLI entry point
β βββ config_loader.py # Config loading and validation
β βββ taxonomy.py # Taxonomy parsers and loaders
β βββ sampling.py # Path A + Path B sampling
β βββ reference_sampling.py # Path D: reference audio sampling
β βββ prompts.py # LLM prompt construction
β βββ prompt_generator.py # Multi-GPU LLM batch generation
β βββ tts_synthesizer.py # Multi-GPU DramaBox TTS
β βββ sidon_enhance.py # Sidon + ChatterboxVC augmentation (replaces RE-USE + LavaSR)
β βββ reuse_enhance.py # RE-USE speech enhancement (legacy)
β βββ moss_refine.py # MOSS-Audio re-annotation (audio-guided prompt rewriting)
β βββ moss_pipeline.py # MOSS orchestrator (multi-GPU job distribution)
β βββ scoring.py # ASR WER + content enjoyment + EmoNet scoring
β βββ demo_grid.py # HTML demo grid generator
β βββ pipeline.py # Mode 1-6 orchestrator
βββ scripts/
β βββ _accc_lavasr_pipeline.py # ACCC LavaSR experiment pipeline (current)
β βββ _emonet_worker.py # EmoNet 40-emotion GPU worker
β βββ _score_worker.py # WER + content enjoyment GPU worker
β βββ _lavasr_worker.py # LavaSR BWE GPU worker
β βββ _lavasr_clap_worker.py # VoiceCLAP scoring GPU worker
βββ docs/
β βββ voicenet_taxonomy.md # VoiceNet 57-dim taxonomy
β βββ voicenet_extension_taxonomy.html # Interactive VoiceNet viewer
β βββ emonet_taxonomy.md # EmoNet 40 emotions
β βββ vocal_bursts_taxonomy.md # 120 vocal bursts
β βββ archetypes.md # 920 archetypes
β βββ acting_challenges_preview.html # Acting challenge preview (100 samples)
β βββ paper_reference.md # Citation and BibTeX
β βββ path_a_voicenet.md # Path A detailed docs
β βββ path_b_archetype.md # Path B detailed docs
β βββ path_c_archetype_named.md # Path C detailed docs
β βββ path_d_reference.md # Path D detailed docs
β βββ path_ac_acting_challenge.md # AC + ACCC detailed docs
β βββ path_sit_situation.md # SIT + SIT-CC situation pathway docs
β βββ path_cc_character_consistent.md # CC v1 detailed docs
β βββ path_cc2_character_consistent_v2.md # CC2 v2 detailed docs
β βββ dramabox_cca_voicenet_examples.md # CC-A English examples
β βββ dramabox_cc2c_archetype_examples.md # CC2-C English examples
β βββ dramabox_accc_acting_challenge_examples.md # ACCC English examples
β βββ dramabox_sit_situation_examples.md # SIT multilingual examples
β βββ dramabox_cca_voicenet_de_examples.md # CC-A German examples
β βββ dramabox_cc2c_archetype_de_examples.md # CC2-C German examples
β βββ dramabox_accc_acting_challenge_de_examples.md # ACCC German examples
β βββ dramabox_extreme_physical_examples.md # Extreme Physical English examples
β βββ dramabox_extreme_physical_de_examples.md # Extreme Physical German examples
β βββ demo/ # HTML demo grids with embedded audio
β βββ sidon_vc_sample_groups.html # Sidon+VC experiment (20 groups, LLM splits)
β βββ accc_lavasr.html # ACCC LavaSR index (redirects to page 1)
β βββ accc_lavasr_p1.html # ACCC LavaSR grid pages 1-10
β βββ ...
β βββ pitch_analysis.html # Pitch analysis
βββ examples/
βββ example_prompt.txt # Sample DramaBox prompt
| Component | Minimum | Recommended |
|---|---|---|
| Prompt generation | 1 GPU, 16GB VRAM | 4+ GPUs, 16GB+ each |
| TTS synthesis | 1 GPU, 24GB VRAM | 4+ GPUs, 24GB+ each |
| Sidon + ChatterboxVC augmentation | 1 GPU, 12GB VRAM | 8 GPUs (parallel workers) |
| MOSS re-annotation | 1 GPU, 8GB VRAM (4-bit) | 8 GPUs (parallel workers) |
| VoiceCLAP + EmoNet scoring | 1 GPU, 4GB VRAM | 8 GPUs (parallel workers) |
| RAM | 32GB | 64GB+ |
After post-processing, the pipeline runs a MOSS re-annotation pass (dramabox/moss_refine.py) that closes the loop between intended and actual performance.
Original DramaBox prompts are directions β the TTS model interprets them, and the actual audio may differ from what was requested. MOSS-Audio-8B-Thinking listens to each generated audio clip together with the original prompt and ASR transcript, then rewrites the prompt to match what was actually performed.
Original DramaBox Prompt
+
ASR Transcript (from Whisper)
+
Generated Audio (MP3)
|
v
+---------------------------+
| MOSS-Audio-8B-Thinking |
| (4-bit, per-GPU worker) |
| |
| Listens to audio + |
| reads text context |
| |
| Rewrites prompt to |
| match actual performance |
+---------------------------+
|
v
moss_refined_prompt_full
moss_refined_prompt_part1
moss_refined_prompt_part2
# Run MOSS re-annotation on all post-processed samples
python dramabox/moss_refine.py # All GPUs
python dramabox/moss_refine.py --num-gpus 4 # 4 GPUs
python dramabox/moss_refine.py --test # First 10 samples, 1 GPU
Requires /tmp/moss_venv with transformers==4.57.1 (MOSS is incompatible with transformers >= 5.x).
We trained one LoRA adapter per EmoNet emotion for the MOSS-TTS-Local-Transformer-4.55B voice-acting model, so an emotion can be applied on top of ordinary voice cloning. Each adapter is the best of a rank {16, 32, 64} sweep: for scarce data a lower rank curbs overfitting, and the winner is picked by a combined score over emotion intensity, vocal-burst-blend, genuineness and speaker similarity.
Data (per emotion, ~1,000β1,400 clips): natural in-the-wild speech (EmoLia) is preferred and filled
with the most intense LAION's Got Talent clips; ~25% of rows carry a same-voice, contrasting-emotion
reference; instructions are procedural emotion captions from
procedural-voice-captions.
TTS-AGI/moss-emotion-loras-40The demo plays five neutral sentences per emotion in three languages, comparing the LoRA with and without an emotion reference; the LoRA raises the automatic EmoNet emotion score while keeping the voice and naturalness.
Python
62.2%
HTML
37.8%