parler-tts/mls-eng-speaker-descriptions

Dataset

13

stars

6

commits

2

linked in READMEs

Aug 8, 2024

updated

README

Dataset Card for Annotations of English MLS

This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.

MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.

This dataset includes an annotation of English MLS. Refers to this dataset card for the other languages.

The text_description column provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech repository.

This dataset was used alongside its original version and LibriTTS-R to train Parler-TTS Mini v1 and Large v1. A training recipe is available in the Parler-TTS library.

Motivation

This dataset is a reproduction of work from the paper Natural language guidance of high-fidelity text-to-speech with synthetic annotations by Dan Lyth and Simon King, from Stability AI and Edinburgh University respectively. It was designed to train the Parler-TTS Mini v1 and Large v1 models

Contrarily to other TTS models, Parler-TTS is a fully open-source release. All of the datasets, pre-processing, training code and weights are released publicly under permissive license, enabling the community to build on our work and develop their own powerful TTS models. Parler-TTS was released alongside:

Usage

Here is an example on how to load the only the train split.

from dataset import load_dataset

load_dataset("parler-tts/mls-eng-speaker-descriptions", split="train")

Streaming is also supported.

from dataset import load_dataset

load_dataset("parler-tts/mls-eng-speaker-descriptions", streaming=True)

Note: This dataset doesn't actually keep track of the audio column of the original version. You can merge it back to the original dataset using this script from Parler-TTS or, even better, get inspiration from the training script of Parler-TTS, that efficiently process multiple annotated datasets.

License

Public Domain, Creative Commons Attribution 4.0 International Public License (CC-BY-4.0)

Citation

@article{Pratap2020MLSAL,
  title={MLS: A Large-Scale Multilingual Dataset for Speech Research},
  author={Vineel Pratap and Qiantong Xu and Anuroop Sriram and Gabriel Synnaeve and Ronan Collobert},
  journal={ArXiv},
  year={2020},
  volume={abs/2012.03411}
}
@misc{lacombe-etal-2024-dataspeech,
  author = {Yoach Lacombe and Vaibhav Srivastav and Sanchit Gandhi},
  title = {Data-Speech},
  year = {2024},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/ylacombe/dataspeech}}
}
@misc{lyth2024natural,
      title={Natural language guidance of high-fidelity text-to-speech with synthetic annotations},
      author={Dan Lyth and Simon King},
      year={2024},
      eprint={2402.01912},
      archivePrefix={arXiv},
      primaryClass={cs.SD}
}

Contributors

ylacombe

6 commits

parler-tts/mls-eng-speaker-descriptions

Dataset

13

stars

6

commits

2

linked in READMEs

Aug 8, 2024

updated

README

Dataset Card for Annotations of English MLS

This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.

MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.

This dataset includes an annotation of English MLS. Refers to this dataset card for the other languages.

The text_description column provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech repository.

This dataset was used alongside its original version and LibriTTS-R to train Parler-TTS Mini v1 and Large v1. A training recipe is available in the Parler-TTS library.

Motivation

This dataset is a reproduction of work from the paper Natural language guidance of high-fidelity text-to-speech with synthetic annotations by Dan Lyth and Simon King, from Stability AI and Edinburgh University respectively. It was designed to train the Parler-TTS Mini v1 and Large v1 models

Contrarily to other TTS models, Parler-TTS is a fully open-source release. All of the datasets, pre-processing, training code and weights are released publicly under permissive license, enabling the community to build on our work and develop their own powerful TTS models. Parler-TTS was released alongside:

Usage

Here is an example on how to load the only the train split.

from dataset import load_dataset

load_dataset("parler-tts/mls-eng-speaker-descriptions", split="train")

Streaming is also supported.

from dataset import load_dataset

load_dataset("parler-tts/mls-eng-speaker-descriptions", streaming=True)

Note: This dataset doesn't actually keep track of the audio column of the original version. You can merge it back to the original dataset using this script from Parler-TTS or, even better, get inspiration from the training script of Parler-TTS, that efficiently process multiple annotated datasets.

License

Public Domain, Creative Commons Attribution 4.0 International Public License (CC-BY-4.0)

Citation

@article{Pratap2020MLSAL,
  title={MLS: A Large-Scale Multilingual Dataset for Speech Research},
  author={Vineel Pratap and Qiantong Xu and Anuroop Sriram and Gabriel Synnaeve and Ronan Collobert},
  journal={ArXiv},
  year={2020},
  volume={abs/2012.03411}
}
@misc{lacombe-etal-2024-dataspeech,
  author = {Yoach Lacombe and Vaibhav Srivastav and Sanchit Gandhi},
  title = {Data-Speech},
  year = {2024},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/ylacombe/dataspeech}}
}
@misc{lyth2024natural,
      title={Natural language guidance of high-fidelity text-to-speech with synthetic annotations},
      author={Dan Lyth and Simon King},
      year={2024},
      eprint={2402.01912},
      archivePrefix={arXiv},
      primaryClass={cs.SD}
}

Contributors

ylacombe

6 commits