This dataset accompanies the paper EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge. It contains 1645 diverse test cases designed to evaluate Text-to-Speech (TTS) models on six challenging scenarios: emotions, paralinguistics, foreign words, syntactic complexity, complex pronunciation (e.g., URLs, formulas), and questions.
The dataset is structured as follows: Each sample contains a category, the text to synthesize, the evolution depth, the language, and the corresponding baseline audio generated by gpt-4o-mini-tts alloy voice, against which we compute win-rate. Details on the data structure can be found in the dataset's metadata. See the linked Github repository for more details on usage and evaluation.
5 commits
1 commits
This dataset accompanies the paper EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge. It contains 1645 diverse test cases designed to evaluate Text-to-Speech (TTS) models on six challenging scenarios: emotions, paralinguistics, foreign words, syntactic complexity, complex pronunciation (e.g., URLs, formulas), and questions.
The dataset is structured as follows: Each sample contains a category, the text to synthesize, the evolution depth, the language, and the corresponding baseline audio generated by gpt-4o-mini-tts alloy voice, against which we compute win-rate. Details on the data structure can be found in the dataset's metadata. See the linked Github repository for more details on usage and evaluation.
5 commits
1 commits