tiiuae/Arabic-LLM-Benchmarks

List of Arabic Benchmarks for Arabic LLMs

20

27 commits

updated Feb 24, 2026

See the code

README

Arabic LLM Benchmarks

A comprehensive repository of Arabic LLMs benchmarks and evaluation benchmarks, curated from systematic research on evaluating Arabic Large Language Models.

📚 Survey paper: Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps


Table of Contents


📋 Overview

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing benchmarks into four categories: Knowledge and STEM, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. Our analysis reveals significant progress in benchmark diversity while identifying critical gaps: limited temporal evaluation, insufficient multi-turn dialogue assessment, and cultural misalignment in translated datasets. We examine three primary approaches: native collection, translation, and synthetic generation discussing their trade-offs regarding authenticity, scale, and cost. This work serves as a comprehensive reference for Arabic NLP researchers, providing insights into benchmark methodologies, reproducibility standards, and evaluation metrics while offering recommendations for future development.


🗂️ Taxonomy

taxonmoy The taxonomy formed based on existing work/benchmarks is depicted above. Each of the categories is defined below:

  • Knowledge includes benchmarks evaluating acquired knowledge and reasoning capabilities, along with domain-specific benchmarks in fields such as law and medicine.
  • Natural Language Processing (NLP) encompasses early task-specific benchmarks and comprehensive multi-task benchmarks, reflecting the evolution from narrow task evaluation to unified assessment across diverse dialects and domains.
  • Culture and Dialects groups benchmarks assessing cultural knowledge and dialect understanding, addressing the essential property of cultural awareness in Arabic LLMs.
  • Target-Specific covers benchmarks designed to assess particular LLM properties such as safety, hallucination detection, instruction-following, and vision capabilities.

🔬 Knowledge

⚙️ General Knowledge & STEM

NamePaper/BenchmarkLinksAccess
MMLUMeasuring Massive Multitask Language Understandingpaper • data • lightevalPublic
EXAMSEXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answeringpaper • data • lightevalPublic
ArabicMMLUArabicMMLU: Assessing Massive Multitask Language Understanding in Arabicpaper • data • repoPublic
AraSTEMAraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM SubjectspaperPrivate
GATA bilingual benchmark for evaluating large language modelspaperPrivate
QiyasThe Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in ArabicpaperPrivate
GATmath & GATLcGATmath and GATLc: Comprehensive benchmarks for evaluating Arabic large language modelspaper • dataPublic
3LM3LM: Bridging Arabic, STEM, and Code through Benchmarkingpaper • data • repoPublic

🏛️ Domain Knowledge

NamePaper/BenchmarkTopicLinksAccess
ArabLegalEvalArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language ModelsLawpaper • repoPublic
AraMedAraMed: Arabic Medical Question Answering using Pretrained Transformer Language ModelsMedicalpaper • repoPublic
MizanQAMizanQA: Benchmarking Large Language Models on Moroccan Legal Question AnsweringLawpaper • dataPublic
Fann or FlopFann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMsPoetrypaper • data • repoPublic
MedArabiQMedArabiQ: Benchmarking Large Language Models on Arabic Medical TasksMedicalpaper • repoPublic
Arabic-Gsm8kArabic-Gsm8kReasoningdataPublic
Hajj-FQAHajj-FQA: A benchmark Arabic dataset for developing question-answering systems on Hajj fatwasReligionpaperPrivate
MedAraBenchMedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMedicalpaper • repoPublic

💬 NLP Tasks

NamePaper/BenchmarkLinksAccess
SOQALNeural Arabic Question Answeringpaper • data • repoPublic
TyDi QATyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languagespaper • repoPublic
ORCAORCA: A Challenging Benchmark for Arabic Language Understandingpaper • dataPublic
DolphinDolphin: A Challenging and Diverse Benchmark for Arabic NLGpaperPrivate
LAraBenchLAraBench: Benchmarking Arabic AI with Large Language Modelspaper • repoPublic
BALSAMBALSAM: A Platform for Benchmarking Arabic Large Language Modelspaper • websitePublic
AlGhafaAlGhafa Evaluation Benchmark for Arabic Language Modelspaper • data • repoPublic

🌍 Culture & Dialects

NamePaper/BenchmarkLinksAccess
JawaherJawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarkingpaper • dataPublic
PalmPalm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMspaper • data • repoPublic
PalmXPalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culturepaper • data • data • websitePublic
ArabCultureCommonsense Reasoning in Arab Culturepaper • data • lm-evalPublic
AraDiCEAraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMspaper • data • lm-evalPublic
AL-QASIDAAL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabicpaper • repoPublic
AbsherAbsher: A Benchmark for Evaluating Large Language Models Understanding of Saudi DialectspaperPrivate
ADABADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational SociopragmaticspaperPrivate
DialectalArabicMMLUDialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Modelspaper • dataPublic

🎯 Specific-Targets

NamePaper/BenchmarkTopicLinksAccess
CamelEvalCamelEval: Advancing Culturally Aligned Arabic Language Models and BenchmarksInstruction-followingpaperPrivate
HalwasaHalwasa: Quantify and Analyze Hallucinations in Large Language Models: Arabic as a Case StudyHallucinationpaperPrivate
PeacockPeacock: A Family of Arabic Multimodal Large Language Models and BenchmarksMultimodalpaper • data • repoPublic
CAMEL-BenchCAMEL-Bench: A Comprehensive Arabic LMM BenchmarkMultimodalpaper • data • repoPublic
AraTrustAraTrust: An Evaluation of Trustworthiness for LLMs in ArabicSafetypaper • dataPublic
Arabic Safety DatasetArabic Dataset for LLM Safeguard EvaluationSafetypaper • repoPublic
ALRAGEALRAGEContext-Based (RAG)dataPublic
AraTableAraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular DataContext-Based (Tabular)paper • repoPublic
AraHalluEvalAraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMsHallucinationpaper • repoPublic
HalluVerse25HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM HallucinationsPoetrypaperPrivate
ARBARB: A Comprehensive Arabic Multimodal Reasoning BenchmarkMultimodalpaper • data • repoPublic
ASASRedteaming Frontier LLMs with AI Astrolabe Arabic Safety Index (ASAS - أساس)SafetyblogPublic

🤝 Contributing

We welcome contributions! If you'd like to add new benchmarks or improve existing information, please feel free to submit issues or pull requests.


📖 Citation

If you use this repository or find the survey helpful, please cite:

@misc{alzubaidi2025evaluatingarabiclargelanguage,
      title={Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps}, 
      author={Ahmed Alzubaidi and Shaikha Alsuwaidi and Basma El Amel Boussaha and Leen AlQadi and Omar Alkaabi and Mohammed Alyafeai and Hamza Alobeidli and Hakim Hacid},
      year={2025},
      eprint={2510.13430},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2510.13430}, 
}

tiiuae/Arabic-LLM-Benchmarks

List of Arabic Benchmarks for Arabic LLMs

20

27 commits

updated Feb 24, 2026

See the code

README

Arabic LLM Benchmarks

A comprehensive repository of Arabic LLMs benchmarks and evaluation benchmarks, curated from systematic research on evaluating Arabic Large Language Models.

📚 Survey paper: Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps


Table of Contents


📋 Overview

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing benchmarks into four categories: Knowledge and STEM, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. Our analysis reveals significant progress in benchmark diversity while identifying critical gaps: limited temporal evaluation, insufficient multi-turn dialogue assessment, and cultural misalignment in translated datasets. We examine three primary approaches: native collection, translation, and synthetic generation discussing their trade-offs regarding authenticity, scale, and cost. This work serves as a comprehensive reference for Arabic NLP researchers, providing insights into benchmark methodologies, reproducibility standards, and evaluation metrics while offering recommendations for future development.


🗂️ Taxonomy

taxonmoy The taxonomy formed based on existing work/benchmarks is depicted above. Each of the categories is defined below:

  • Knowledge includes benchmarks evaluating acquired knowledge and reasoning capabilities, along with domain-specific benchmarks in fields such as law and medicine.
  • Natural Language Processing (NLP) encompasses early task-specific benchmarks and comprehensive multi-task benchmarks, reflecting the evolution from narrow task evaluation to unified assessment across diverse dialects and domains.
  • Culture and Dialects groups benchmarks assessing cultural knowledge and dialect understanding, addressing the essential property of cultural awareness in Arabic LLMs.
  • Target-Specific covers benchmarks designed to assess particular LLM properties such as safety, hallucination detection, instruction-following, and vision capabilities.

🔬 Knowledge

⚙️ General Knowledge & STEM

NamePaper/BenchmarkLinksAccess
MMLUMeasuring Massive Multitask Language Understandingpaper • data • lightevalPublic
EXAMSEXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answeringpaper • data • lightevalPublic
ArabicMMLUArabicMMLU: Assessing Massive Multitask Language Understanding in Arabicpaper • data • repoPublic
AraSTEMAraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM SubjectspaperPrivate
GATA bilingual benchmark for evaluating large language modelspaperPrivate
QiyasThe Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in ArabicpaperPrivate
GATmath & GATLcGATmath and GATLc: Comprehensive benchmarks for evaluating Arabic large language modelspaper • dataPublic
3LM3LM: Bridging Arabic, STEM, and Code through Benchmarkingpaper • data • repoPublic

🏛️ Domain Knowledge

NamePaper/BenchmarkTopicLinksAccess
ArabLegalEvalArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language ModelsLawpaper • repoPublic
AraMedAraMed: Arabic Medical Question Answering using Pretrained Transformer Language ModelsMedicalpaper • repoPublic
MizanQAMizanQA: Benchmarking Large Language Models on Moroccan Legal Question AnsweringLawpaper • dataPublic
Fann or FlopFann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMsPoetrypaper • data • repoPublic
MedArabiQMedArabiQ: Benchmarking Large Language Models on Arabic Medical TasksMedicalpaper • repoPublic
Arabic-Gsm8kArabic-Gsm8kReasoningdataPublic
Hajj-FQAHajj-FQA: A benchmark Arabic dataset for developing question-answering systems on Hajj fatwasReligionpaperPrivate
MedAraBenchMedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMedicalpaper • repoPublic

💬 NLP Tasks

NamePaper/BenchmarkLinksAccess
SOQALNeural Arabic Question Answeringpaper • data • repoPublic
TyDi QATyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languagespaper • repoPublic
ORCAORCA: A Challenging Benchmark for Arabic Language Understandingpaper • dataPublic
DolphinDolphin: A Challenging and Diverse Benchmark for Arabic NLGpaperPrivate
LAraBenchLAraBench: Benchmarking Arabic AI with Large Language Modelspaper • repoPublic
BALSAMBALSAM: A Platform for Benchmarking Arabic Large Language Modelspaper • websitePublic
AlGhafaAlGhafa Evaluation Benchmark for Arabic Language Modelspaper • data • repoPublic

🌍 Culture & Dialects

NamePaper/BenchmarkLinksAccess
JawaherJawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarkingpaper • dataPublic
PalmPalm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMspaper • data • repoPublic
PalmXPalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culturepaper • data • data • websitePublic
ArabCultureCommonsense Reasoning in Arab Culturepaper • data • lm-evalPublic
AraDiCEAraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMspaper • data • lm-evalPublic
AL-QASIDAAL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabicpaper • repoPublic
AbsherAbsher: A Benchmark for Evaluating Large Language Models Understanding of Saudi DialectspaperPrivate
ADABADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational SociopragmaticspaperPrivate
DialectalArabicMMLUDialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Modelspaper • dataPublic

🎯 Specific-Targets

NamePaper/BenchmarkTopicLinksAccess
CamelEvalCamelEval: Advancing Culturally Aligned Arabic Language Models and BenchmarksInstruction-followingpaperPrivate
HalwasaHalwasa: Quantify and Analyze Hallucinations in Large Language Models: Arabic as a Case StudyHallucinationpaperPrivate
PeacockPeacock: A Family of Arabic Multimodal Large Language Models and BenchmarksMultimodalpaper • data • repoPublic
CAMEL-BenchCAMEL-Bench: A Comprehensive Arabic LMM BenchmarkMultimodalpaper • data • repoPublic
AraTrustAraTrust: An Evaluation of Trustworthiness for LLMs in ArabicSafetypaper • dataPublic
Arabic Safety DatasetArabic Dataset for LLM Safeguard EvaluationSafetypaper • repoPublic
ALRAGEALRAGEContext-Based (RAG)dataPublic
AraTableAraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular DataContext-Based (Tabular)paper • repoPublic
AraHalluEvalAraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMsHallucinationpaper • repoPublic
HalluVerse25HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM HallucinationsPoetrypaperPrivate
ARBARB: A Comprehensive Arabic Multimodal Reasoning BenchmarkMultimodalpaper • data • repoPublic
ASASRedteaming Frontier LLMs with AI Astrolabe Arabic Safety Index (ASAS - أساس)SafetyblogPublic

🤝 Contributing

We welcome contributions! If you'd like to add new benchmarks or improve existing information, please feel free to submit issues or pull requests.


📖 Citation

If you use this repository or find the survey helpful, please cite:

@misc{alzubaidi2025evaluatingarabiclargelanguage,
      title={Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps}, 
      author={Ahmed Alzubaidi and Shaikha Alsuwaidi and Basma El Amel Boussaha and Leen AlQadi and Omar Alkaabi and Mohammed Alyafeai and Hamza Alobeidli and Hakim Hacid},
      year={2025},
      eprint={2510.13430},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2510.13430}, 
}