nick-mccormick/tts-voices-sampler

Dataset

0

stars

10

commits

2

linked in READMEs

Sep 3, 2025

updated

not-for-all-audiences

README

TTS Voices Sampler

A small sample set of audio voices, with transcriptions, from other datasets.

Sources

EmoV-DB

Link. EmoV-DB is a dataset with a forced emotional reading of an arbitrary text. Multiple emotions are supplied with the same speaker. There are only four speakers in the dataset. All have a U.S. accent.

Unique non-commercial license https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md

Emilia

Link. Dataset which seems very clearly to be taken from real youtube videos. This dataset is absolutely gigantic (200,000+ hours, 2.4TB!). Attribution AND non-commercial are both required. And to be honest this could still be dicey as it is probably scraped indiscriminately.

CC-BY-NC license

Emilia YODAS

Link. Extension of the Emilia dataset with seemingly lower average quality data but only attribution is required (commercial usage OK). This is also 2.1TB and extremely huge.

CC-BY license

VCTK

Link. This dataset is from the Voice Cloning Toolkit but our sample here is taken from https://huggingface.co/kyutai/tts-voices (another good resource for voice clips).

CC BY 4.0 International license

"Testnew" dataset

Link. NSFW voice clips dataset. There are more explicit ones in the original dataset. There are only three actors provided here but dozens can be found in the larger dataset. Additionally we have created a smaller subset of this dataset which is only a few gigabytes instead of 155gb here.

Apache 2.0 license.

Contributors

nick-mccormick

10 commits

nick-mccormick/tts-voices-sampler

Dataset

0

stars

10

commits

2

linked in READMEs

Sep 3, 2025

updated

not-for-all-audiences

README

TTS Voices Sampler

A small sample set of audio voices, with transcriptions, from other datasets.

Sources

EmoV-DB

Link. EmoV-DB is a dataset with a forced emotional reading of an arbitrary text. Multiple emotions are supplied with the same speaker. There are only four speakers in the dataset. All have a U.S. accent.

Unique non-commercial license https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md

Emilia

Link. Dataset which seems very clearly to be taken from real youtube videos. This dataset is absolutely gigantic (200,000+ hours, 2.4TB!). Attribution AND non-commercial are both required. And to be honest this could still be dicey as it is probably scraped indiscriminately.

CC-BY-NC license

Emilia YODAS

Link. Extension of the Emilia dataset with seemingly lower average quality data but only attribution is required (commercial usage OK). This is also 2.1TB and extremely huge.

CC-BY license

VCTK

Link. This dataset is from the Voice Cloning Toolkit but our sample here is taken from https://huggingface.co/kyutai/tts-voices (another good resource for voice clips).

CC BY 4.0 International license

"Testnew" dataset

Link. NSFW voice clips dataset. There are more explicit ones in the original dataset. There are only three actors provided here but dozens can be found in the larger dataset. Additionally we have created a smaller subset of this dataset which is only a few gigabytes instead of 155gb here.

Apache 2.0 license.

Contributors

nick-mccormick

10 commits