BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset for fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated at the sentence level across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset supports multi-class readability classification in the following formats:
{'ID': 10100010008, 'Sentence': 'عيد سعيد', 'Word_Count': 2, 'Word': 'عيد سعيد', 'Lex': 'عيد سعيد', 'D3Tok': 'عيد سعيد', 'D3Lex': 'عيد سعيد', 'Readability_Level': '2-ba', 'Readability_Level_19': 2, 'Readability_Level_7': 1, 'Readability_Level_5': 1, 'Readability_Level_3': 1, 'Annotator': 'A4', 'Document': 'BAREC_Majed_0229_1983_001.txt', 'Source': 'Majed', 'Book': 'Edition: 229', 'Author': '#', 'Domain': 'Arts & Humanities', 'Text_Class': 'Foundational'}
19-levels scheme, ranging from 1-alif to 19-qaf.19-levels scheme, ranging from 1 to 19.7-levels scheme, ranging from 1 to 7.5-levels scheme, ranging from 1 to 5.3-levels scheme, ranging from 1 to 3.A1-A5 or IAA).Arts & Humanities, STEM or Social Sciences).Foundational, Advanced or Specialized).We define the Readability Assessment task as an ordinal classification task. The following metrics are used for evaluation:
If you use BAREC in your work, please cite the following papers:
@inproceedings{elmadani-etal-2025-readability,
title = "A Large and Balanced Corpus for Fine-grained {A}rabic Readability Assessment",
author = "Elmadani, Khalid N. and
Habash, Nizar and
Taha-Thomure, Hanada",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-acl.842/"
}
@inproceedings{habash-etal-2025-guidelines,
title = "Guidelines for Fine-grained Sentence-level {A}rabic Readability Annotation",
author = "Habash, Nizar and
Taha-Thomure, Hanada and
Elmadani, Khalid N. and
Zeino, Zeina and
Abushmaes, Abdallah",
booktitle = "Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.law-1.30/"
}
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset for fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated at the sentence level across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset supports multi-class readability classification in the following formats:
{'ID': 10100010008, 'Sentence': 'عيد سعيد', 'Word_Count': 2, 'Word': 'عيد سعيد', 'Lex': 'عيد سعيد', 'D3Tok': 'عيد سعيد', 'D3Lex': 'عيد سعيد', 'Readability_Level': '2-ba', 'Readability_Level_19': 2, 'Readability_Level_7': 1, 'Readability_Level_5': 1, 'Readability_Level_3': 1, 'Annotator': 'A4', 'Document': 'BAREC_Majed_0229_1983_001.txt', 'Source': 'Majed', 'Book': 'Edition: 229', 'Author': '#', 'Domain': 'Arts & Humanities', 'Text_Class': 'Foundational'}
19-levels scheme, ranging from 1-alif to 19-qaf.19-levels scheme, ranging from 1 to 19.7-levels scheme, ranging from 1 to 7.5-levels scheme, ranging from 1 to 5.3-levels scheme, ranging from 1 to 3.A1-A5 or IAA).Arts & Humanities, STEM or Social Sciences).Foundational, Advanced or Specialized).We define the Readability Assessment task as an ordinal classification task. The following metrics are used for evaluation:
If you use BAREC in your work, please cite the following papers:
@inproceedings{elmadani-etal-2025-readability,
title = "A Large and Balanced Corpus for Fine-grained {A}rabic Readability Assessment",
author = "Elmadani, Khalid N. and
Habash, Nizar and
Taha-Thomure, Hanada",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-acl.842/"
}
@inproceedings{habash-etal-2025-guidelines,
title = "Guidelines for Fine-grained Sentence-level {A}rabic Readability Annotation",
author = "Habash, Nizar and
Taha-Thomure, Hanada and
Elmadani, Khalid N. and
Zeino, Zeina and
Abushmaes, Abdallah",
booktitle = "Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.law-1.30/"
}