This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset includes an annotation of English MLS. Refers to this dataset card for the other languages.
The text_description column provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech repository.
This dataset was used alongside its original version and LibriTTS-R to train Parler-TTS Mini v1 and Large v1. A training recipe is available in the Parler-TTS library.
This dataset is a reproduction of work from the paper Natural language guidance of high-fidelity text-to-speech with synthetic annotations by Dan Lyth and Simon King, from Stability AI and Edinburgh University respectively. It was designed to train the Parler-TTS Mini v1 and Large v1 models
Contrarily to other TTS models, Parler-TTS is a fully open-source release. All of the datasets, pre-processing, training code and weights are released publicly under permissive license, enabling the community to build on our work and develop their own powerful TTS models. Parler-TTS was released alongside:
Here is an example on how to load the only the train split.
from dataset import load_dataset
load_dataset("parler-tts/mls-eng-speaker-descriptions", split="train")
Streaming is also supported.
from dataset import load_dataset
load_dataset("parler-tts/mls-eng-speaker-descriptions", streaming=True)
Note: This dataset doesn't actually keep track of the audio column of the original version. You can merge it back to the original dataset using this script from Parler-TTS or, even better, get inspiration from the training script of Parler-TTS, that efficiently process multiple annotated datasets.
Public Domain, Creative Commons Attribution 4.0 International Public License (CC-BY-4.0)
@article{Pratap2020MLSAL,
title={MLS: A Large-Scale Multilingual Dataset for Speech Research},
author={Vineel Pratap and Qiantong Xu and Anuroop Sriram and Gabriel Synnaeve and Ronan Collobert},
journal={ArXiv},
year={2020},
volume={abs/2012.03411}
}
@misc{lacombe-etal-2024-dataspeech,
author = {Yoach Lacombe and Vaibhav Srivastav and Sanchit Gandhi},
title = {Data-Speech},
year = {2024},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/ylacombe/dataspeech}}
}
@misc{lyth2024natural,
title={Natural language guidance of high-fidelity text-to-speech with synthetic annotations},
author={Dan Lyth and Simon King},
year={2024},
eprint={2402.01912},
archivePrefix={arXiv},
primaryClass={cs.SD}
}
6 commits
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset includes an annotation of English MLS. Refers to this dataset card for the other languages.
The text_description column provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech repository.
This dataset was used alongside its original version and LibriTTS-R to train Parler-TTS Mini v1 and Large v1. A training recipe is available in the Parler-TTS library.
This dataset is a reproduction of work from the paper Natural language guidance of high-fidelity text-to-speech with synthetic annotations by Dan Lyth and Simon King, from Stability AI and Edinburgh University respectively. It was designed to train the Parler-TTS Mini v1 and Large v1 models
Contrarily to other TTS models, Parler-TTS is a fully open-source release. All of the datasets, pre-processing, training code and weights are released publicly under permissive license, enabling the community to build on our work and develop their own powerful TTS models. Parler-TTS was released alongside:
Here is an example on how to load the only the train split.
from dataset import load_dataset
load_dataset("parler-tts/mls-eng-speaker-descriptions", split="train")
Streaming is also supported.
from dataset import load_dataset
load_dataset("parler-tts/mls-eng-speaker-descriptions", streaming=True)
Note: This dataset doesn't actually keep track of the audio column of the original version. You can merge it back to the original dataset using this script from Parler-TTS or, even better, get inspiration from the training script of Parler-TTS, that efficiently process multiple annotated datasets.
Public Domain, Creative Commons Attribution 4.0 International Public License (CC-BY-4.0)
@article{Pratap2020MLSAL,
title={MLS: A Large-Scale Multilingual Dataset for Speech Research},
author={Vineel Pratap and Qiantong Xu and Anuroop Sriram and Gabriel Synnaeve and Ronan Collobert},
journal={ArXiv},
year={2020},
volume={abs/2012.03411}
}
@misc{lacombe-etal-2024-dataspeech,
author = {Yoach Lacombe and Vaibhav Srivastav and Sanchit Gandhi},
title = {Data-Speech},
year = {2024},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/ylacombe/dataspeech}}
}
@misc{lyth2024natural,
title={Natural language guidance of high-fidelity text-to-speech with synthetic annotations},
author={Dan Lyth and Simon King},
year={2024},
eprint={2402.01912},
archivePrefix={arXiv},
primaryClass={cs.SD}
}
6 commits