MADIS-5 (Multi-domain Arabic Dialect Identification in Speech) is a manually curated dataset designed to facilitate evaluation of cross-domain robustness of Arabic Dialect Identification (ADI) systems. This dataset provides a comprehensive benchmark for testing out-of-domain generalization across different speech domains with diverse recording conditions and speaking styles.
Our dataset comprises speech samples from four different public sources, each offering varying degrees of similarity to the TV broadcast domain commonly used in ADI research:
This dataset is ideal for:
If you use this dataset in your research, please cite our paper:
@inproceedings{abdullah2025voice,
title={Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification},
author={Abdullah, Badr M. and Matthew Baas and Bernd Möbius and Dietrich Klakow},
year={2025},
publisher={Interspeech},
url={arxiv.org/abs/2505.24713}
}
Creative Commons Attribution-NonCommercial-NoDerivs 4.0 (CC BY-NC-ND 4.0)
We thank the contributors to the source datasets and platforms that made this compilation possible, including radio.garden, SARA archive, and the Multilingual TEDx dataset.
MADIS-5 (Multi-domain Arabic Dialect Identification in Speech) is a manually curated dataset designed to facilitate evaluation of cross-domain robustness of Arabic Dialect Identification (ADI) systems. This dataset provides a comprehensive benchmark for testing out-of-domain generalization across different speech domains with diverse recording conditions and speaking styles.
Our dataset comprises speech samples from four different public sources, each offering varying degrees of similarity to the TV broadcast domain commonly used in ADI research:
This dataset is ideal for:
If you use this dataset in your research, please cite our paper:
@inproceedings{abdullah2025voice,
title={Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification},
author={Abdullah, Badr M. and Matthew Baas and Bernd Möbius and Dietrich Klakow},
year={2025},
publisher={Interspeech},
url={arxiv.org/abs/2505.24713}
}
Creative Commons Attribution-NonCommercial-NoDerivs 4.0 (CC BY-NC-ND 4.0)
We thank the contributors to the source datasets and platforms that made this compilation possible, including radio.garden, SARA archive, and the Multilingual TEDx dataset.