ash56/ShiftySpeech

Dataset

This repository introduces: 🌀 *ShiftySpeech*: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

22

50 commits

2 linked in READMEs

updated Sep 2, 2026

See the code

README

This repository introduces: 🌀 ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

🔥 Key Features

  • 3000+ hours of synthetic speech
  • Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including:
    • 📖 Reading Style
    • 🎙️ Podcast
    • 🎥 YouTube
    • 🗣️ Languages (Three different languages)
    • 🌎 Demographics (including variations in age, accent, and gender)
  • Multiple Speech Generation Systems: Includes data synthesized from various TTS models and vocoders.

💡 Why We Built This Dataset

Driven by advances in self-supervised learning for speech, state-of-the-art synthetic speech detectors have achieved low error rates on popular benchmarks such as ASVspoof. However, prior benchmarks do not address the wide range of real-world variability in speech. Are reported error rates realistic in real-world conditions? To assess detector failure modes and robustness under controlled distribution shifts, we introduce ShiftySpeech, a benchmark with more than 3000 hours of synthetic speech from 7 domains, 6 TTS systems, 12 vocoders, and 3 languages.

⚙️ Usage

Ensure that you have soundfile or librosa installed for proper audio decoding:

pip install soundfile librosa
📌 Example: Loading the AISHELL Dataset Vocoded with APNet2
from datasets import load_dataset

dataset = load_dataset("ash56/ShiftySpeech", data_files={"data": f"Vocoders/apnet2/apnet2_aishell_flac.tar.gz"})["data"]

⚠️ Note: It is recommended to load data from a specific folder to avoid unnecessary memory usage.

The source datasets covered by different TTS and Vocoder systems are listed in tts.yaml and vocoders.yaml

📄 More Information

For detailed information on dataset sources and analysis, see our paper: ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

You can also find the full implementation on GitHub

Citation

If you find this dataset useful, please cite our work:

@article{garg2025shiftyspeech,
  title={ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts},
  author={Garg, Ashi and Cai, Zexin and Zhang, Lin and Xinyuan, Henry Li and Garc{\'\i}a-Perera, Leibny Paola and Duh, Kevin and Khudanpur, Sanjeev and Wiesner, Matthew and Andrews, Nicholas},
  journal={arXiv preprint arXiv:2502.05674},
  year={2025}
}

✉️ Contact

If you have any questions or comments about the resource, please feel free to reach out to us at: gargg.ashi@gmail.com or noa@jhu.edu

anti-spoofing
audio
deepfake
deepfake-audio
security
synthetic-speech-detection
voice-spoofing

ash56/ShiftySpeech

Dataset

This repository introduces: 🌀 *ShiftySpeech*: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

22

50 commits

2 linked in READMEs

updated Sep 2, 2026

See the code

README

This repository introduces: 🌀 ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

🔥 Key Features

  • 3000+ hours of synthetic speech
  • Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including:
    • 📖 Reading Style
    • 🎙️ Podcast
    • 🎥 YouTube
    • 🗣️ Languages (Three different languages)
    • 🌎 Demographics (including variations in age, accent, and gender)
  • Multiple Speech Generation Systems: Includes data synthesized from various TTS models and vocoders.

💡 Why We Built This Dataset

Driven by advances in self-supervised learning for speech, state-of-the-art synthetic speech detectors have achieved low error rates on popular benchmarks such as ASVspoof. However, prior benchmarks do not address the wide range of real-world variability in speech. Are reported error rates realistic in real-world conditions? To assess detector failure modes and robustness under controlled distribution shifts, we introduce ShiftySpeech, a benchmark with more than 3000 hours of synthetic speech from 7 domains, 6 TTS systems, 12 vocoders, and 3 languages.

⚙️ Usage

Ensure that you have soundfile or librosa installed for proper audio decoding:

pip install soundfile librosa
📌 Example: Loading the AISHELL Dataset Vocoded with APNet2
from datasets import load_dataset

dataset = load_dataset("ash56/ShiftySpeech", data_files={"data": f"Vocoders/apnet2/apnet2_aishell_flac.tar.gz"})["data"]

⚠️ Note: It is recommended to load data from a specific folder to avoid unnecessary memory usage.

The source datasets covered by different TTS and Vocoder systems are listed in tts.yaml and vocoders.yaml

📄 More Information

For detailed information on dataset sources and analysis, see our paper: ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

You can also find the full implementation on GitHub

Citation

If you find this dataset useful, please cite our work:

@article{garg2025shiftyspeech,
  title={ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts},
  author={Garg, Ashi and Cai, Zexin and Zhang, Lin and Xinyuan, Henry Li and Garc{\'\i}a-Perera, Leibny Paola and Duh, Kevin and Khudanpur, Sanjeev and Wiesner, Matthew and Andrews, Nicholas},
  journal={arXiv preprint arXiv:2502.05674},
  year={2025}
}

✉️ Contact

If you have any questions or comments about the resource, please feel free to reach out to us at: gargg.ashi@gmail.com or noa@jhu.edu

anti-spoofing
audio
deepfake
deepfake-audio
security
synthetic-speech-detection
voice-spoofing