Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in DramaBox stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
unsloth/gemma-3n-E4B-it (bf16, temperature=0.9, max_tokens=1024)data/Prompts are sampled from the Voice-Acting-Pipeline taxonomy across
5 CUT-TO pathways, each producing a character-consistent two-scene prompt with a strong
emotional contrast across the CUT TO::
Every system prompt carries an emotional-emphasis instruction (make it raw, push the
contrast to its limit) and a compact vocal-burst menu (SFW taxonomy, 180 bursts)
so the model can place natural non-speech bursts as director notes and use full dynamic
range (shout/roar when angry, scream/gasp/sob when afraid or in pain). For ~50% of
rows a single vocal burst is randomly injected as a required element (recorded in
injected_burst). Dialogue length follows the pipeline default (~50 spoken words, ~25 per
scene). Each generation independently samples pathway, language, taxonomy attributes and
word-seeds with strong (os.urandom) randomness, so the input state is effectively unique.
sfw-v1)| language | count |
|---|---|
| German | 2,029,071 |
| English | 2,027,929 |
| pathway | count |
|---|---|
| CC2-C (Archetype) | 812,289 |
| ACCC (Acting Challenge) | 811,863 |
| SIT (Situations) | 811,033 |
| Extreme Physical | 810,939 |
| CCA (VoiceNet) | 810,876 |
| gender | count |
|---|---|
| female | 2,363,360 |
| male | 1,304,464 |
| other | 389,176 |
| age group | count |
|---|---|
| young adult | 1,287,024 |
| middle-aged | 850,806 |
| elderly | 761,296 |
| unspecified | 620,629 |
| youth | 530,227 |
| adult | 7,018 |
| emotion | count |
|---|---|
| Contentment | 127,291 |
| Pride | 126,820 |
| Relief | 126,558 |
| Sadness | 125,508 |
| Triumph | 124,426 |
| Hope/Optimism | 124,140 |
| Intoxication/Altered States | 124,137 |
| Confusion | 123,958 |
| Embarrassment | 123,859 |
| Fatigue/Exhaustion | 123,632 |
| Disgust | 123,553 |
| Bitterness | 123,523 |
| Shame | 123,157 |
| Disappointment | 122,785 |
| Doubt | 122,755 |
| burst | count |
|---|---|
| Conversational 'Mhm' (Yes) | 11,522 |
| Breathy 'Oh no' | 11,495 |
| Aggressive Snarl | 11,485 |
| Warrior Battle Cry | 11,476 |
| Whistling a Victory Tune | 11,475 |
| Polite Covered Yawn | 11,467 |
| Tension-Releasing Whoosh | 11,459 |
| Relaxing Exhale | 11,456 |
| Post-Thirst Gulp | 11,450 |
| Humming to Drown Out Noise | 11,449 |
| Absent-Minded Humming | 11,436 |
| Startle Grunt | 11,428 |
| Busking / Street Singing | 11,427 |
| Cackle | 11,426 |
| Stress Gulp | 11,426 |
Each row (one generated prompt) has these columns:
uid, global_index, gen_seed, worker_gpu, pathway, pathway_label, lang, dramabox_prompt, gender, age_group, sampled_gender, sampled_age, arousal, emotions, voicenet_attributes, flow_style, emotion_alignment, direction_style, archetype, genre, tempo, situation, situation_dim, challenge_title, challenge_id, category, subcategory, word_seeds, n_word_seeds, injected_burst, burst_taxonomy_source, burst_taxonomy_version, model, precision, temperature, top_p, max_tokens, input_tokens, output_tokens, gen_timestamp
Key fields: dramabox_prompt (the generated text), pathway/lang, sampled attributes
(gender, age_group, arousal, emotions, voicenet_attributes, archetype/genre,
situation, challenge_title, category), word_seeds, injected_burst, generation
settings (model, temperature, max_tokens), input_tokens/output_tokens, a uid
and a global global_index.
Released under CC-BY-4.0. Synthetic data generated by a local Gemma model; intended for research on expressive voice acting / TTS. No audio is included.
500 commits
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in DramaBox stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
unsloth/gemma-3n-E4B-it (bf16, temperature=0.9, max_tokens=1024)data/Prompts are sampled from the Voice-Acting-Pipeline taxonomy across
5 CUT-TO pathways, each producing a character-consistent two-scene prompt with a strong
emotional contrast across the CUT TO::
Every system prompt carries an emotional-emphasis instruction (make it raw, push the
contrast to its limit) and a compact vocal-burst menu (SFW taxonomy, 180 bursts)
so the model can place natural non-speech bursts as director notes and use full dynamic
range (shout/roar when angry, scream/gasp/sob when afraid or in pain). For ~50% of
rows a single vocal burst is randomly injected as a required element (recorded in
injected_burst). Dialogue length follows the pipeline default (~50 spoken words, ~25 per
scene). Each generation independently samples pathway, language, taxonomy attributes and
word-seeds with strong (os.urandom) randomness, so the input state is effectively unique.
sfw-v1)| language | count |
|---|---|
| German | 2,029,071 |
| English | 2,027,929 |
| pathway | count |
|---|---|
| CC2-C (Archetype) | 812,289 |
| ACCC (Acting Challenge) | 811,863 |
| SIT (Situations) | 811,033 |
| Extreme Physical | 810,939 |
| CCA (VoiceNet) | 810,876 |
| gender | count |
|---|---|
| female | 2,363,360 |
| male | 1,304,464 |
| other | 389,176 |
| age group | count |
|---|---|
| young adult | 1,287,024 |
| middle-aged | 850,806 |
| elderly | 761,296 |
| unspecified | 620,629 |
| youth | 530,227 |
| adult | 7,018 |
| emotion | count |
|---|---|
| Contentment | 127,291 |
| Pride | 126,820 |
| Relief | 126,558 |
| Sadness | 125,508 |
| Triumph | 124,426 |
| Hope/Optimism | 124,140 |
| Intoxication/Altered States | 124,137 |
| Confusion | 123,958 |
| Embarrassment | 123,859 |
| Fatigue/Exhaustion | 123,632 |
| Disgust | 123,553 |
| Bitterness | 123,523 |
| Shame | 123,157 |
| Disappointment | 122,785 |
| Doubt | 122,755 |
| burst | count |
|---|---|
| Conversational 'Mhm' (Yes) | 11,522 |
| Breathy 'Oh no' | 11,495 |
| Aggressive Snarl | 11,485 |
| Warrior Battle Cry | 11,476 |
| Whistling a Victory Tune | 11,475 |
| Polite Covered Yawn | 11,467 |
| Tension-Releasing Whoosh | 11,459 |
| Relaxing Exhale | 11,456 |
| Post-Thirst Gulp | 11,450 |
| Humming to Drown Out Noise | 11,449 |
| Absent-Minded Humming | 11,436 |
| Startle Grunt | 11,428 |
| Busking / Street Singing | 11,427 |
| Cackle | 11,426 |
| Stress Gulp | 11,426 |
Each row (one generated prompt) has these columns:
uid, global_index, gen_seed, worker_gpu, pathway, pathway_label, lang, dramabox_prompt, gender, age_group, sampled_gender, sampled_age, arousal, emotions, voicenet_attributes, flow_style, emotion_alignment, direction_style, archetype, genre, tempo, situation, situation_dim, challenge_title, challenge_id, category, subcategory, word_seeds, n_word_seeds, injected_burst, burst_taxonomy_source, burst_taxonomy_version, model, precision, temperature, top_p, max_tokens, input_tokens, output_tokens, gen_timestamp
Key fields: dramabox_prompt (the generated text), pathway/lang, sampled attributes
(gender, age_group, arousal, emotions, voicenet_attributes, archetype/genre,
situation, challenge_title, category), word_seeds, injected_burst, generation
settings (model, temperature, max_tokens), input_tokens/output_tokens, a uid
and a global global_index.
Released under CC-BY-4.0. Synthetic data generated by a local Gemma model; intended for research on expressive voice acting / TTS. No audio is included.
500 commits