This dataset contains data to train SPEAR TTS-like text-to-speech models that utilized semantic tokens derived from the OpenAI Whisper speech recognition model.
We currently provide semantic and acoustic tokens for the LibriLight and LibriTTS datasets (English only).
Acoustic tokens:
Semantic tokens:
Available LibriLight subsets:
small/medium/large (following the original dataset division but with large excluding the speaker 6454)6454 speaker from the large subset for training single-speaker TTS modelsWe plan to add more acoustic tokens from other codecs in the future.
11 commits
This dataset contains data to train SPEAR TTS-like text-to-speech models that utilized semantic tokens derived from the OpenAI Whisper speech recognition model.
We currently provide semantic and acoustic tokens for the LibriLight and LibriTTS datasets (English only).
Acoustic tokens:
Semantic tokens:
Available LibriLight subsets:
small/medium/large (following the original dataset division but with large excluding the speaker 6454)6454 speaker from the large subset for training single-speaker TTS modelsWe plan to add more acoustic tokens from other codecs in the future.
11 commits