A small sample set of audio voices, with transcriptions, from other datasets.
Link. EmoV-DB is a dataset with a forced emotional reading of an arbitrary text. Multiple emotions are supplied with the same speaker. There are only four speakers in the dataset. All have a U.S. accent.
Unique non-commercial license https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md
Link. Dataset which seems very clearly to be taken from real youtube videos. This dataset is absolutely gigantic (200,000+ hours, 2.4TB!). Attribution AND non-commercial are both required. And to be honest this could still be dicey as it is probably scraped indiscriminately.
CC-BY-NC license
Link. Extension of the Emilia dataset with seemingly lower average quality data but only attribution is required (commercial usage OK). This is also 2.1TB and extremely huge.
CC-BY license
Link. This dataset is from the Voice Cloning Toolkit but our sample here is taken from https://huggingface.co/kyutai/tts-voices (another good resource for voice clips).
CC BY 4.0 International license
Link. NSFW voice clips dataset. There are more explicit ones in the original dataset. There are only three actors provided here but dozens can be found in the larger dataset. Additionally we have created a smaller subset of this dataset which is only a few gigabytes instead of 155gb here.
Apache 2.0 license.
10 commits
A small sample set of audio voices, with transcriptions, from other datasets.
Link. EmoV-DB is a dataset with a forced emotional reading of an arbitrary text. Multiple emotions are supplied with the same speaker. There are only four speakers in the dataset. All have a U.S. accent.
Unique non-commercial license https://github.com/numediart/EmoV-DB/blob/master/LICENSE.md
Link. Dataset which seems very clearly to be taken from real youtube videos. This dataset is absolutely gigantic (200,000+ hours, 2.4TB!). Attribution AND non-commercial are both required. And to be honest this could still be dicey as it is probably scraped indiscriminately.
CC-BY-NC license
Link. Extension of the Emilia dataset with seemingly lower average quality data but only attribution is required (commercial usage OK). This is also 2.1TB and extremely huge.
CC-BY license
Link. This dataset is from the Voice Cloning Toolkit but our sample here is taken from https://huggingface.co/kyutai/tts-voices (another good resource for voice clips).
CC BY 4.0 International license
Link. NSFW voice clips dataset. There are more explicit ones in the original dataset. There are only three actors provided here but dozens can be found in the larger dataset. Additionally we have created a smaller subset of this dataset which is only a few gigabytes instead of 155gb here.
Apache 2.0 license.
10 commits