ajd12342/paraspeechcaps

Dataset

23

stars

21

commits

1

linked in READMEs

Nov 22, 2025

updated

README

ParaSpeechCaps

We release ParaSpeechCaps (Paralinguistic Speech Captions), a large-scale dataset that annotates speech utterances with rich style captions ('A male speaker with a husky, raspy voice delivers happy and admiring remarks at a slow speed in a very noisy American environment. His speech is enthusiastic and confident, with occasional high-pitched inflections.'). It supports 59 style tags covering styles like pitch, rhythm, emotion, and more, spanning speaker-level intrinsic style tags and utterance-level situational style tags.

We also release Parler-TTS models finetuned on ParaSpeechCaps at ajd12342/parler-tts-mini-v1-paraspeechcaps and ajd12342/parler-tts-mini-v1-paraspeechcaps-only-base.

Please take a look at our paper, our codebase and our demo website for more information.

NOTE: We release style captions and a host of other useful style-related metadata, but not the source audio files. Please refer to our codebase for setup instructions on how to download them from their respective datasets (VoxCeleb, Expresso, EARS, Emilia).

License: CC BY-NC SA 4.0

Overview

ParaSpeechCaps is a large-scale dataset that annotates speech utterances with rich style captions. It consists of a human-annotated subset ParaSpeechCaps-Base and a large automatically-annotated subset ParaSpeechCaps-Scaled. Our novel pipeline combining off-the-shelf text and speech embedders, classifiers and an audio language model allows us to automatically scale rich tag annotations for such a wide variety of style tags for the first time.

Usage

This repository has been tested with Python 3.11 (conda create -n paraspeechcaps python=3.11), but most other versions should probably work. Install using

pip install datasets

You can use the dataset as follows:

from datasets import load_dataset

# Load the entire dataset
dataset = load_dataset("ajd12342/paraspeechcaps")

# Load specific splits of the dataset
train_scaled = load_dataset("ajd12342/paraspeechcaps", split="train_scaled")
train_base = load_dataset("ajd12342/paraspeechcaps", split="train_base")
dev = load_dataset("ajd12342/paraspeechcaps", split="dev")
holdout = load_dataset("ajd12342/paraspeechcaps", split="holdout")

# View a single example
example = train_base[0]
print(example)

Dataset Structure

The dataset contains the following columns:

ColumnTypeDescription
sourcestringSource dataset (e.g., Expresso, EARS, VoxCeleb, Emilia)
relative_audio_pathstringRelative path to identify the specific audio file being annotated
text_descriptionlist of strings1-2 Style Descriptions for the utterance
transcriptionstringTranscript of the speech
intrinsic_tagslist of stringsTags tied to a speaker's identity (e.g., shrill, guttural) (null if non-existent)
situational_tagslist of stringsTags that characterize individual utterances (e.g., happy, whispered) (null if non-existent)
basic_tagslist of stringsBasic tags (pitch, speed, gender, noise conditions)
all_tagslist of stringsCombination of all tag types
speakeridstringUnique identifier for the speaker
namestringName of the speaker
durationfloatDuration of the audio in seconds
genderstringSpeaker's gender
accentstringSpeaker's accent (null if non-existent)
pitchstringDescription of the pitch level
speaking_ratestringDescription of the speaking rate
noisestringDescription of background noise
utterance_pitch_meanfloatMean pitch value of the utterance
snrfloatSignal-to-noise ratio
phonemesstringPhonetic transcription
tag_of_intereststringThe rich tag of interest (only applicable for the 'test' split for evaluation, null for other splits)

The text_description field is a list because each example may have 1 or 2 text descriptions:

  • For Expresso and Emilia examples, all have 2 descriptions:
    • One with just situational tags
    • One with both intrinsic and situational tags
  • For Emilia examples that were found by both our intrinsic and situational automatic annotation pipelines, there are 2 descriptions:
    • One with just intrinsic tags
    • One with both intrinsic and situational tags

The relative_audio_path field contains relative paths, functioning as a unique identifier for the specific audio file being annotated. The repository contains setup instructions that can properly link the annotations to the source audio files.

Dataset Statistics

The dataset covers a total of 59 style tags, including both speaker-level intrinsic tags (33) and utterance-level situational tags (26). It consists of 282 train hours of human-labeled data and 2427 train hours of automatically annotated data (PSC-Scaled). It contains 2518 train hours with intrinsic tag annotations and 298 train hours with situational tag annotations, with 106 hours of overlap.

SplitNumber of ExamplesNumber of Unique SpeakersDuration (hours)
train_scaled924,65139,0022,427.16
train_base116,516641282.54
dev11,96762426.29
holdout14,75616733.04

Citation

If you use this dataset, the models or the repository, please cite our work as follows:

@misc{diwan2025scalingrichstylepromptedtexttospeech,
      title={Scaling Rich Style-Prompted Text-to-Speech Datasets}, 
      author={Anuj Diwan and Zhisheng Zheng and David Harwath and Eunsol Choi},
      year={2025},
      eprint={2503.04713},
      archivePrefix={arXiv},
      primaryClass={eess.AS},
      url={https://arxiv.org/abs/2503.04713}, 
}

Contributors

ajd12342

21 commits

ajd12342/paraspeechcaps

Dataset

23

stars

21

commits

1

linked in READMEs

Nov 22, 2025

updated

README

ParaSpeechCaps

We release ParaSpeechCaps (Paralinguistic Speech Captions), a large-scale dataset that annotates speech utterances with rich style captions ('A male speaker with a husky, raspy voice delivers happy and admiring remarks at a slow speed in a very noisy American environment. His speech is enthusiastic and confident, with occasional high-pitched inflections.'). It supports 59 style tags covering styles like pitch, rhythm, emotion, and more, spanning speaker-level intrinsic style tags and utterance-level situational style tags.

We also release Parler-TTS models finetuned on ParaSpeechCaps at ajd12342/parler-tts-mini-v1-paraspeechcaps and ajd12342/parler-tts-mini-v1-paraspeechcaps-only-base.

Please take a look at our paper, our codebase and our demo website for more information.

NOTE: We release style captions and a host of other useful style-related metadata, but not the source audio files. Please refer to our codebase for setup instructions on how to download them from their respective datasets (VoxCeleb, Expresso, EARS, Emilia).

License: CC BY-NC SA 4.0

Overview

ParaSpeechCaps is a large-scale dataset that annotates speech utterances with rich style captions. It consists of a human-annotated subset ParaSpeechCaps-Base and a large automatically-annotated subset ParaSpeechCaps-Scaled. Our novel pipeline combining off-the-shelf text and speech embedders, classifiers and an audio language model allows us to automatically scale rich tag annotations for such a wide variety of style tags for the first time.

Usage

This repository has been tested with Python 3.11 (conda create -n paraspeechcaps python=3.11), but most other versions should probably work. Install using

pip install datasets

You can use the dataset as follows:

from datasets import load_dataset

# Load the entire dataset
dataset = load_dataset("ajd12342/paraspeechcaps")

# Load specific splits of the dataset
train_scaled = load_dataset("ajd12342/paraspeechcaps", split="train_scaled")
train_base = load_dataset("ajd12342/paraspeechcaps", split="train_base")
dev = load_dataset("ajd12342/paraspeechcaps", split="dev")
holdout = load_dataset("ajd12342/paraspeechcaps", split="holdout")

# View a single example
example = train_base[0]
print(example)

Dataset Structure

The dataset contains the following columns:

ColumnTypeDescription
sourcestringSource dataset (e.g., Expresso, EARS, VoxCeleb, Emilia)
relative_audio_pathstringRelative path to identify the specific audio file being annotated
text_descriptionlist of strings1-2 Style Descriptions for the utterance
transcriptionstringTranscript of the speech
intrinsic_tagslist of stringsTags tied to a speaker's identity (e.g., shrill, guttural) (null if non-existent)
situational_tagslist of stringsTags that characterize individual utterances (e.g., happy, whispered) (null if non-existent)
basic_tagslist of stringsBasic tags (pitch, speed, gender, noise conditions)
all_tagslist of stringsCombination of all tag types
speakeridstringUnique identifier for the speaker
namestringName of the speaker
durationfloatDuration of the audio in seconds
genderstringSpeaker's gender
accentstringSpeaker's accent (null if non-existent)
pitchstringDescription of the pitch level
speaking_ratestringDescription of the speaking rate
noisestringDescription of background noise
utterance_pitch_meanfloatMean pitch value of the utterance
snrfloatSignal-to-noise ratio
phonemesstringPhonetic transcription
tag_of_intereststringThe rich tag of interest (only applicable for the 'test' split for evaluation, null for other splits)

The text_description field is a list because each example may have 1 or 2 text descriptions:

  • For Expresso and Emilia examples, all have 2 descriptions:
    • One with just situational tags
    • One with both intrinsic and situational tags
  • For Emilia examples that were found by both our intrinsic and situational automatic annotation pipelines, there are 2 descriptions:
    • One with just intrinsic tags
    • One with both intrinsic and situational tags

The relative_audio_path field contains relative paths, functioning as a unique identifier for the specific audio file being annotated. The repository contains setup instructions that can properly link the annotations to the source audio files.

Dataset Statistics

The dataset covers a total of 59 style tags, including both speaker-level intrinsic tags (33) and utterance-level situational tags (26). It consists of 282 train hours of human-labeled data and 2427 train hours of automatically annotated data (PSC-Scaled). It contains 2518 train hours with intrinsic tag annotations and 298 train hours with situational tag annotations, with 106 hours of overlap.

SplitNumber of ExamplesNumber of Unique SpeakersDuration (hours)
train_scaled924,65139,0022,427.16
train_base116,516641282.54
dev11,96762426.29
holdout14,75616733.04

Citation

If you use this dataset, the models or the repository, please cite our work as follows:

@misc{diwan2025scalingrichstylepromptedtexttospeech,
      title={Scaling Rich Style-Prompted Text-to-Speech Datasets}, 
      author={Anuj Diwan and Zhisheng Zheng and David Harwath and Eunsol Choi},
      year={2025},
      eprint={2503.04713},
      archivePrefix={arXiv},
      primaryClass={eess.AS},
      url={https://arxiv.org/abs/2503.04713}, 
}

Contributors

ajd12342

21 commits