CIDAR contains 10,000 instructions and their output. The dataset was created by selecting around 9,109 samples from Alpagasus dataset then translating it to Arabic using ChatGPT. In addition, we append that with around 891 Arabic grammar instructions from the webiste Ask the teacher. All the 10,000 samples were reviewed by around 12 reviewers.
| Name | Explanation |
|---|---|
| CIDAR | 10,000 instructions and responses in Arabic |
| CIDAR-EVAL-100 | 100 instructions to evaluate LLMs on cultural relevance |
| CIDAR-MCQ-100 | 100 Multiple choice questions and answers to evaluate LLMs on cultural relevance |
index(int): Unique number of each instruction-output pair.instruction(str): Required instruction.output(str): Response for the given instruction.You can download the dataset directly from HuggingFace or use the following code:
from datasets import load_dataset
cidar = load_dataset('arbml/CIDAR')
This sample shows an Arabic-specific instruction related to البلاغة(rhetoric) and الإعراب(grammatical status). The sentence is originally in Arabic.
index: 10030
instruction: لو سمحت اين صيغة المبالغة في " الجاحظ اديب سيارة اراؤه" و ما اعراب الجملة؟
output:
'صيغة المبالغة في هذه الجملة هي "سيّارة"، فاسم الفاعل من "السير" هو '
'"سائر/ سائرة" وإذا أردت الدلالة على المبالغة في سير الآراء (أي '
'شهرتها وكثرة تداولها) قلت "آراء سيّارة".\r\n'
'أما إعراب الجملة فهو: "الجاحظُ": مبتدأ مرفوع بالضمة. "أديبٌ": خبر '
'مرفوع بالضمة. "سيارةٌ": خبر متقدم، و"آراؤُه": مبتدأ متأخر، وهو '
'مضاف والهاء ضمير متصل مضاف إليه في محل جر. ويمكن اعتبار "سيارة" '
'مبتدأ وهو وصف يعمل عمل فعله، و"آراؤُه" فاعل سدّ مسدّ الخبر.\r\n'
'وفي الحالتين فجملة "سيارة آراؤه" جملة اسمية في محل رفع نعت '
'لـ"أديب".'
There were at least 12 contributors to the annotation of CIDAR. You can check the list here.
CIDAR is intended for research purposes only. The authors disclaim any responsibility for misuse and condemn any use contrary to Arabic culture or Islamic values. Even though subjected to human verification, there is no guarantee that responses are entirely aligned with Arabic culture and Islamic values. Users of the dataset are urged to exercise caution, employ critical thinking, and seek guidance from representative figures when necessary.
@inproceedings{alyafeai-etal-2024-cidar,
title = "{{CIDAR}: Culturally Relevant Instruction Dataset For {A}rabic}",
author = "Alyafeai, Zaid and
Almubarak, Khalid and
Ashraf, Ahmed and
Alnuhait, Deema and
Alshahrani, Saied and
Abdulrahman, Gubran and
Ahmed, Gamil and
Gawah, Qais and
Saleh, Zead and
Ghaleb, Mustafa and
Ali, Yousef and
Al-shaibani, Maged",
editor = "Ku, Lun-Wei and
Martins, Andre and
Srikumar, Vivek",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-acl.764/",
doi = "10.18653/v1/2024.findings-acl.764",
pages = "12878--12901",
}
@misc{alyafeai2024cidar,
title={{CIDAR: Culturally Relevant Instruction Dataset For Arabic}},
author={Zaid Alyafeai and Khalid Almubarak and Ahmed Ashraf and Deema Alnuhait and Saied Alshahrani and Gubran A. Q. Abdulrahman and Gamil Ahmed and Qais Gawah and Zead Saleh and Mustafa Ghaleb and Yousef Ali and Maged S. Al-Shaibani},
year={2024},
eprint={2402.03177},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
CIDAR contains 10,000 instructions and their output. The dataset was created by selecting around 9,109 samples from Alpagasus dataset then translating it to Arabic using ChatGPT. In addition, we append that with around 891 Arabic grammar instructions from the webiste Ask the teacher. All the 10,000 samples were reviewed by around 12 reviewers.
| Name | Explanation |
|---|---|
| CIDAR | 10,000 instructions and responses in Arabic |
| CIDAR-EVAL-100 | 100 instructions to evaluate LLMs on cultural relevance |
| CIDAR-MCQ-100 | 100 Multiple choice questions and answers to evaluate LLMs on cultural relevance |
index(int): Unique number of each instruction-output pair.instruction(str): Required instruction.output(str): Response for the given instruction.You can download the dataset directly from HuggingFace or use the following code:
from datasets import load_dataset
cidar = load_dataset('arbml/CIDAR')
This sample shows an Arabic-specific instruction related to البلاغة(rhetoric) and الإعراب(grammatical status). The sentence is originally in Arabic.
index: 10030
instruction: لو سمحت اين صيغة المبالغة في " الجاحظ اديب سيارة اراؤه" و ما اعراب الجملة؟
output:
'صيغة المبالغة في هذه الجملة هي "سيّارة"، فاسم الفاعل من "السير" هو '
'"سائر/ سائرة" وإذا أردت الدلالة على المبالغة في سير الآراء (أي '
'شهرتها وكثرة تداولها) قلت "آراء سيّارة".\r\n'
'أما إعراب الجملة فهو: "الجاحظُ": مبتدأ مرفوع بالضمة. "أديبٌ": خبر '
'مرفوع بالضمة. "سيارةٌ": خبر متقدم، و"آراؤُه": مبتدأ متأخر، وهو '
'مضاف والهاء ضمير متصل مضاف إليه في محل جر. ويمكن اعتبار "سيارة" '
'مبتدأ وهو وصف يعمل عمل فعله، و"آراؤُه" فاعل سدّ مسدّ الخبر.\r\n'
'وفي الحالتين فجملة "سيارة آراؤه" جملة اسمية في محل رفع نعت '
'لـ"أديب".'
There were at least 12 contributors to the annotation of CIDAR. You can check the list here.
CIDAR is intended for research purposes only. The authors disclaim any responsibility for misuse and condemn any use contrary to Arabic culture or Islamic values. Even though subjected to human verification, there is no guarantee that responses are entirely aligned with Arabic culture and Islamic values. Users of the dataset are urged to exercise caution, employ critical thinking, and seek guidance from representative figures when necessary.
@inproceedings{alyafeai-etal-2024-cidar,
title = "{{CIDAR}: Culturally Relevant Instruction Dataset For {A}rabic}",
author = "Alyafeai, Zaid and
Almubarak, Khalid and
Ashraf, Ahmed and
Alnuhait, Deema and
Alshahrani, Saied and
Abdulrahman, Gubran and
Ahmed, Gamil and
Gawah, Qais and
Saleh, Zead and
Ghaleb, Mustafa and
Ali, Yousef and
Al-shaibani, Maged",
editor = "Ku, Lun-Wei and
Martins, Andre and
Srikumar, Vivek",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-acl.764/",
doi = "10.18653/v1/2024.findings-acl.764",
pages = "12878--12901",
}
@misc{alyafeai2024cidar,
title={{CIDAR: Culturally Relevant Instruction Dataset For Arabic}},
author={Zaid Alyafeai and Khalid Almubarak and Ahmed Ashraf and Deema Alnuhait and Saied Alshahrani and Gubran A. Q. Abdulrahman and Gamil Ahmed and Qais Gawah and Zead Saleh and Mustafa Ghaleb and Yousef Ali and Maged S. Al-Shaibani},
year={2024},
eprint={2402.03177},
archivePrefix={arXiv},
primaryClass={cs.CL}
}