Dataset Card for SADA - Saudi Audio Dataset for Arabic
3
15 commits
1 linked in READMEs
updated Sep 12, 2024
The SADA (Saudi Audio Dataset for Arabic) is a comprehensive dataset consisting of audio recordings from over 57 TV shows aired by the Saudi Broadcasting Authority (SBA). The dataset contains approximately 667 hours of audio data with transcripts, the majority of which are in various Saudi dialects (Najdi, Hijazi, Khaliji, etc.).
The dataset is intended for use in automatic speech recognition (ASR), text-to-speech (TTS) systems, and dialect identification tasks. It provides rich linguistic diversity across various dialects, making it suitable for training and fine-tuning speech models for the Arabic language.
The dataset should not be used for commercial purposes or any misuse involving speech or text processing in harmful or inappropriate contexts.
The dataset contains the following features:
.wav format.The dataset was created to provide a comprehensive resource for developing Arabic speech models, specifically targeting the variety of Saudi dialects and their applications in speech recognition and synthesis.
The data was collected from TV shows aired by the Saudi Broadcasting Authority. The audio files were segmented and transcribed manually, with a focus on ensuring transcription accuracy and uniformity in formatting.
The original content comes from the Saudi Broadcasting Authority's shows, curated by SDAIA.
The dataset does not contain sensitive personal information, as the focus is on broadcast TV shows with public speakers.
Given the focus on Saudi dialects, the dataset may not fully represent other Arabic dialects or languages. Users should be aware that models trained on this dataset may show biases toward Saudi dialects.
@inproceedings{SADA2023, title={SADA - SBA & SDAIA Audio Dataset for Arabic}, author={Areeb Alowisheq, Abdullah Alrajeh, Sadeen Alharbi, Abdulmajeed Alrowithi, Aljawharah Bin Tamran, Asma Ibrahim, Raghad Aloraini, Raneem Alnajim, Ranya Alkahtani, Renad Almuasaad, Sara Alrasheed, Shaykhah Alsubaie, Yaser Alonaizan}, booktitle={To be published}, affiliation={NCAI-SDAIA}, year={2023} }
For any inquiries, please contact the National Center for Artificial Intelligence at SDAIA.
Dataset Card for SADA - Saudi Audio Dataset for Arabic
3
15 commits
1 linked in READMEs
updated Sep 12, 2024
The SADA (Saudi Audio Dataset for Arabic) is a comprehensive dataset consisting of audio recordings from over 57 TV shows aired by the Saudi Broadcasting Authority (SBA). The dataset contains approximately 667 hours of audio data with transcripts, the majority of which are in various Saudi dialects (Najdi, Hijazi, Khaliji, etc.).
The dataset is intended for use in automatic speech recognition (ASR), text-to-speech (TTS) systems, and dialect identification tasks. It provides rich linguistic diversity across various dialects, making it suitable for training and fine-tuning speech models for the Arabic language.
The dataset should not be used for commercial purposes or any misuse involving speech or text processing in harmful or inappropriate contexts.
The dataset contains the following features:
.wav format.The dataset was created to provide a comprehensive resource for developing Arabic speech models, specifically targeting the variety of Saudi dialects and their applications in speech recognition and synthesis.
The data was collected from TV shows aired by the Saudi Broadcasting Authority. The audio files were segmented and transcribed manually, with a focus on ensuring transcription accuracy and uniformity in formatting.
The original content comes from the Saudi Broadcasting Authority's shows, curated by SDAIA.
The dataset does not contain sensitive personal information, as the focus is on broadcast TV shows with public speakers.
Given the focus on Saudi dialects, the dataset may not fully represent other Arabic dialects or languages. Users should be aware that models trained on this dataset may show biases toward Saudi dialects.
@inproceedings{SADA2023, title={SADA - SBA & SDAIA Audio Dataset for Arabic}, author={Areeb Alowisheq, Abdullah Alrajeh, Sadeen Alharbi, Abdulmajeed Alrowithi, Aljawharah Bin Tamran, Asma Ibrahim, Raghad Aloraini, Raneem Alnajim, Ranya Alkahtani, Renad Almuasaad, Sara Alrasheed, Shaykhah Alsubaie, Yaser Alonaizan}, booktitle={To be published}, affiliation={NCAI-SDAIA}, year={2023} }
For any inquiries, please contact the National Center for Artificial Intelligence at SDAIA.