UBC-NLP/Casablanca

Dataset

35

stars

3

commits

1

linked in READMEs

Nov 14, 2024

updated

algeria
arabic
asr
dialects
egypt
jordan
mauritania
morocco
palestine
speech
speech_processing
speech_recognition
uae
yemen

README

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

Casablanca

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic inclusion. This challenge is largely due to the absence of datasets that can empower diverse speech systems. In this paper, we seek to mitigate this obstacle for a number of Arabic dialects by presenting Casablanca, a large-scale community-driven effort to collect and transcribe a multi-dialectal Arabic dataset. The dataset covers eight dialects: Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni, and includes annotations for transcription, gender, dialect, and code-switching. We also develop a number of strong baselines exploiting Casablanca. The project page for Casablanca is accessible at: https://www.dlnlp.ai/speech/casablanca/

https://arxiv.org/abs/2410.04527

** Please note that in this version, we are releasing only the validation and test sets. **

Citation

If you use Casablanca work, please cite the paper where it was introduced:

BibTeX:

@article{talafha2024casablanca,
   title={Casablanca: Data and Models for Multidialectal Arabic Speech Recognition},
   author={Talafha, Bashar and Kadaoui, Karima and Magdy, Samar Mohamed and Habiboullah, Mariem
           and Chafei, Chafei Mohamed and El-Shangiti, Ahmed Oumar and Zayed,
           Hiba and Alhamouri, Rahaf and Assi, Rwaa and Alraeesi, Aisha and others},
   journal={arXiv preprint arXiv:2410.04527},
   year={2024}
   
}

Contributors

BT

UBC-NLP/Casablanca

Dataset

35

stars

3

commits

1

linked in READMEs

Nov 14, 2024

updated

algeria
arabic
asr
dialects
egypt
jordan
mauritania
morocco
palestine
speech
speech_processing
speech_recognition
uae
yemen

README

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

Casablanca

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic inclusion. This challenge is largely due to the absence of datasets that can empower diverse speech systems. In this paper, we seek to mitigate this obstacle for a number of Arabic dialects by presenting Casablanca, a large-scale community-driven effort to collect and transcribe a multi-dialectal Arabic dataset. The dataset covers eight dialects: Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni, and includes annotations for transcription, gender, dialect, and code-switching. We also develop a number of strong baselines exploiting Casablanca. The project page for Casablanca is accessible at: https://www.dlnlp.ai/speech/casablanca/

https://arxiv.org/abs/2410.04527

** Please note that in this version, we are releasing only the validation and test sets. **

Citation

If you use Casablanca work, please cite the paper where it was introduced:

BibTeX:

@article{talafha2024casablanca,
   title={Casablanca: Data and Models for Multidialectal Arabic Speech Recognition},
   author={Talafha, Bashar and Kadaoui, Karima and Magdy, Samar Mohamed and Habiboullah, Mariem
           and Chafei, Chafei Mohamed and El-Shangiti, Ahmed Oumar and Zayed,
           Hiba and Alhamouri, Rahaf and Assi, Rwaa and Alraeesi, Aisha and others},
   journal={arXiv preprint arXiv:2410.04527},
   year={2024}
   
}

Contributors

BT