Sara Ghaboura *
Ketan More *
Wafa Alghallabi
Omkar Thawakar
Jorma Laaksonen
Hisham Cholakkal
Salman Khan
Rao M. Anwer
*Equal Contribution
ARB is the first benchmark focused on step-by-step reasoning in Arabic cross both textual and visual modalities, covering 11 diverse domains spanning science, culture, OCR, and historical interpretation.
We evaluated 12 open- and closed-source LMMs using:
Stepwise Evaluation Using LLM-as-Judge for Closed-Source Models: | Metric ↓ / Model → | GPT-4o | GPT-4o-mini | GPT-4.1 | o4-mini | Gemini 1.5 Pro | Gemini 2.0 Flash | |----------------------------|--------|-------------|---------|---------|----------------|------------------| | Final Answer (%) | 60.22 | 52.22 | 59.43 | 58.93 | 56.70 | 57.80 | | Reasoning Steps (%) | 64.29 | 61.02 | 80.41 | 80.75| 64.34 | 64.09 |
Stepwise Evaluation Using LLM-as-Judge for Open-Source Models: | Metric ↓ / Model → | Qwen2.5-VL | LLaMA-3.2 | AIN | LLaMA-4 Scout | Aya-Vision | InternVL3 | |----------------------------|------------|-----------|-------|----------------|-------------|-----------| | Final Answer (%) | 37.02 | 25.58 | 27.35 | 48.52 | 28.81 | 31.04 | | Reasoning Steps (%) | 64.03 | 53.20 | 52.77 | 77.70 | 63.64 | 54.50 |
Each sample includes:
image_id: Visual inputquestion: Arabic question grounded in image reasoningchoices: The choices for the MCQsteps: Ordered reasoning chainanswer: Final solution (Arabic)category: One of 11 categories (e.g., OCR, Scientific, Visual, Math)If you use ARB dataset in your research, please consider citing:
@misc{ghaboura2025arbcomprehensivearabicmultimodal,
title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark},
author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},
year={2025},
eprint={2505.17021},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2505.17021},
}
Sara Ghaboura *
Ketan More *
Wafa Alghallabi
Omkar Thawakar
Jorma Laaksonen
Hisham Cholakkal
Salman Khan
Rao M. Anwer
*Equal Contribution
ARB is the first benchmark focused on step-by-step reasoning in Arabic cross both textual and visual modalities, covering 11 diverse domains spanning science, culture, OCR, and historical interpretation.
We evaluated 12 open- and closed-source LMMs using:
Stepwise Evaluation Using LLM-as-Judge for Closed-Source Models: | Metric ↓ / Model → | GPT-4o | GPT-4o-mini | GPT-4.1 | o4-mini | Gemini 1.5 Pro | Gemini 2.0 Flash | |----------------------------|--------|-------------|---------|---------|----------------|------------------| | Final Answer (%) | 60.22 | 52.22 | 59.43 | 58.93 | 56.70 | 57.80 | | Reasoning Steps (%) | 64.29 | 61.02 | 80.41 | 80.75| 64.34 | 64.09 |
Stepwise Evaluation Using LLM-as-Judge for Open-Source Models: | Metric ↓ / Model → | Qwen2.5-VL | LLaMA-3.2 | AIN | LLaMA-4 Scout | Aya-Vision | InternVL3 | |----------------------------|------------|-----------|-------|----------------|-------------|-----------| | Final Answer (%) | 37.02 | 25.58 | 27.35 | 48.52 | 28.81 | 31.04 | | Reasoning Steps (%) | 64.03 | 53.20 | 52.77 | 77.70 | 63.64 | 54.50 |
Each sample includes:
image_id: Visual inputquestion: Arabic question grounded in image reasoningchoices: The choices for the MCQsteps: Ordered reasoning chainanswer: Final solution (Arabic)category: One of 11 categories (e.g., OCR, Scientific, Visual, Math)If you use ARB dataset in your research, please consider citing:
@misc{ghaboura2025arbcomprehensivearabicmultimodal,
title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark},
author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},
year={2025},
eprint={2505.17021},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2505.17021},
}