🏆 Leaderboard | 🛠️ Evaluation Suite
TTS Voice Design is a high-quality benchmark of 1,000 character voice-design tasks spanning a broad range of media genres and real-world creative use cases. It evaluates whether a text-to-speech model can turn an open-ended character profile into a distinctive, appropriate, and usable voice.
Unlike benchmarks built around a fixed set of speakers or isolated acoustic attributes, this dataset covers complete character concepts: vocal identity, age and gender presentation, accent, timbre, pitch, speaking behavior, personality, and performance style.
The source pool is CharacterCodex, a collection of approximately 15,000 popular characters from games, film and television, animation, literature, mythology, web media, and factual or historical sources.
We select 1,000 tasks using a diversity-aware curation process that:
The anonymization step reduces IP-related leakage and makes the task about realizing the described role rather than recognizing or reproducing a known character.
| Coverage category | Examples | Media types |
|---|---|---|
| Games and Interactive Play | 133 | Video Games, Board Games, Card Games, Tabletop Role-Playing Games |
| Film and Television | 110 | Movies, Television Shows |
| Illustrated and Animated Media | 140 | Manga, Graphic Novels, Anime, Comic Books, Webcomics |
| Literature, Stage and Lyric Works | 178 | Novels, Plays, Short Stories, Songs, Poetry |
| Myth, Folklore and Historical Lore | 142 | Mythology, Urban Legends, Historical Texts, Folklore, Fairy Tales |
| Periodical and Web Media | 131 | Magazines, Blogs, Online Articles |
| Factual, News and Research Sources | 166 | Biographies, Documentaries, Newspapers, Scientific Papers, Research Journals |
The benchmark contains a broad range of human, non-human, and synthetic voice
concepts. The profile metadata is intended for analysis and stratified
evaluation; the voice_design_prompt remains the complete model input
specification.
| Attribute | Distribution |
|---|---|
| Gender presentation | masculine 558, feminine 393, neutral 21, unspecified 27, not applicable 1 |
| Age | child 3, teen 81, young adult 231, adult 380, older 98, unspecified 107, not applicable 100 |
| Speech rate | slow 334, medium 257, fast 254, dynamic 155 |
| Pitch | low 283, mid 426, high 176, dynamic 115 |
| Energy | low 290, mid 280, high 256, dynamic 174 |
| Humanity | human 900, non-human 77, synthetic 12, ambiguous 11 |
Accent coverage includes North American, British and Irish, Continental European, Latin American and Caribbean, East Asian, Oceanian, African, South Asian, Middle Eastern, Southeast Asian, and fictional accents, as well as cases where no accent is specified.
Frequently represented timbres include clear, breathy, rough, resonant, smooth, warm, gravelly, raspy, bright, and airy voices. Style labels span calm and reflective, dramatic and theatrical, authoritative, warm and nurturing, energetic, intimate, melancholic, scholarly, comedic, heroic, formal, eccentric, and sinister performances.
Each line in the JSONL file is one task:
{
"id": "tts-design-0001",
"language": "en",
"voice_design_prompt": "Male, middle-aged, European Portuguese. Warm, resonant journalist. Rhythmic, melodic cadence, moderate pace, precise articulation. Inquisitive, respectful, conveys curiosity and community empathy.",
"transcript": "Tell me, when the music starts and the square fills with families, what do you feel this celebration says about your neighborhood, and about the pride people carry here every day?",
"coverage_category": "Factual, News and Research Sources",
"media_type": "Newspapers",
"voice_profile": {
"gender": "masculine",
"age": "adult",
"accent_region": "continental_european",
"accent_raw": "European Portuguese",
"speech_rate": "medium",
"pitch": "mid",
"energy": "mid",
"timbre_tags": ["warm", "resonant", "clear"],
"style_tags": ["formal_professional", "calm_reflective"],
"humanity": "human"
}
}
| Field | Description |
|---|---|
id | Stable public task ID in the form tts-design-XXXX |
language | Transcript language; currently en |
voice_design_prompt | Complete natural-language specification of the target voice and role |
transcript | Preview text that the model should synthesize |
coverage_category | One of the seven broad media-coverage categories |
media_type | Fine-grained source type; 29 values in total |
voice_profile | Structured annotations for demographic, acoustic, and stylistic analysis |
For every record:
voice_design_prompt as the voice-design instruction.transcript.tts-design-0001.wav.Do not use coverage_category, media_type, or voice_profile as additional
model inputs. They are provided for analysis and visualization.
The benchmark reports three complementary metrics:
The self-contained judge prompt, evaluation script, WavLM embedding code, and metric documentation are available in the evaluation suite.
4 commits
🏆 Leaderboard | 🛠️ Evaluation Suite
TTS Voice Design is a high-quality benchmark of 1,000 character voice-design tasks spanning a broad range of media genres and real-world creative use cases. It evaluates whether a text-to-speech model can turn an open-ended character profile into a distinctive, appropriate, and usable voice.
Unlike benchmarks built around a fixed set of speakers or isolated acoustic attributes, this dataset covers complete character concepts: vocal identity, age and gender presentation, accent, timbre, pitch, speaking behavior, personality, and performance style.
The source pool is CharacterCodex, a collection of approximately 15,000 popular characters from games, film and television, animation, literature, mythology, web media, and factual or historical sources.
We select 1,000 tasks using a diversity-aware curation process that:
The anonymization step reduces IP-related leakage and makes the task about realizing the described role rather than recognizing or reproducing a known character.
| Coverage category | Examples | Media types |
|---|---|---|
| Games and Interactive Play | 133 | Video Games, Board Games, Card Games, Tabletop Role-Playing Games |
| Film and Television | 110 | Movies, Television Shows |
| Illustrated and Animated Media | 140 | Manga, Graphic Novels, Anime, Comic Books, Webcomics |
| Literature, Stage and Lyric Works | 178 | Novels, Plays, Short Stories, Songs, Poetry |
| Myth, Folklore and Historical Lore | 142 | Mythology, Urban Legends, Historical Texts, Folklore, Fairy Tales |
| Periodical and Web Media | 131 | Magazines, Blogs, Online Articles |
| Factual, News and Research Sources | 166 | Biographies, Documentaries, Newspapers, Scientific Papers, Research Journals |
The benchmark contains a broad range of human, non-human, and synthetic voice
concepts. The profile metadata is intended for analysis and stratified
evaluation; the voice_design_prompt remains the complete model input
specification.
| Attribute | Distribution |
|---|---|
| Gender presentation | masculine 558, feminine 393, neutral 21, unspecified 27, not applicable 1 |
| Age | child 3, teen 81, young adult 231, adult 380, older 98, unspecified 107, not applicable 100 |
| Speech rate | slow 334, medium 257, fast 254, dynamic 155 |
| Pitch | low 283, mid 426, high 176, dynamic 115 |
| Energy | low 290, mid 280, high 256, dynamic 174 |
| Humanity | human 900, non-human 77, synthetic 12, ambiguous 11 |
Accent coverage includes North American, British and Irish, Continental European, Latin American and Caribbean, East Asian, Oceanian, African, South Asian, Middle Eastern, Southeast Asian, and fictional accents, as well as cases where no accent is specified.
Frequently represented timbres include clear, breathy, rough, resonant, smooth, warm, gravelly, raspy, bright, and airy voices. Style labels span calm and reflective, dramatic and theatrical, authoritative, warm and nurturing, energetic, intimate, melancholic, scholarly, comedic, heroic, formal, eccentric, and sinister performances.
Each line in the JSONL file is one task:
{
"id": "tts-design-0001",
"language": "en",
"voice_design_prompt": "Male, middle-aged, European Portuguese. Warm, resonant journalist. Rhythmic, melodic cadence, moderate pace, precise articulation. Inquisitive, respectful, conveys curiosity and community empathy.",
"transcript": "Tell me, when the music starts and the square fills with families, what do you feel this celebration says about your neighborhood, and about the pride people carry here every day?",
"coverage_category": "Factual, News and Research Sources",
"media_type": "Newspapers",
"voice_profile": {
"gender": "masculine",
"age": "adult",
"accent_region": "continental_european",
"accent_raw": "European Portuguese",
"speech_rate": "medium",
"pitch": "mid",
"energy": "mid",
"timbre_tags": ["warm", "resonant", "clear"],
"style_tags": ["formal_professional", "calm_reflective"],
"humanity": "human"
}
}
| Field | Description |
|---|---|
id | Stable public task ID in the form tts-design-XXXX |
language | Transcript language; currently en |
voice_design_prompt | Complete natural-language specification of the target voice and role |
transcript | Preview text that the model should synthesize |
coverage_category | One of the seven broad media-coverage categories |
media_type | Fine-grained source type; 29 values in total |
voice_profile | Structured annotations for demographic, acoustic, and stylistic analysis |
For every record:
voice_design_prompt as the voice-design instruction.transcript.tts-design-0001.wav.Do not use coverage_category, media_type, or voice_profile as additional
model inputs. They are provided for analysis and visualization.
The benchmark reports three complementary metrics:
The self-contained judge prompt, evaluation script, WavLM embedding code, and metric documentation are available in the evaluation suite.
4 commits