MBZUAI/ARB

Dataset

Sara Ghaboura *

11

28 commits

1 linked in READMEs

updated Jun 21, 2025

See the code

README

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

[![arXiv](https://img.shields.io/badge/arXiv-2505.17021-C0DAD9)](https://arxiv.org/abs/2505.17021) [![Our Page](https://img.shields.io/badge/Visit-Our%20Page-D4EBDB?style=flat)](https://mbzuai-oryx.github.io/ARB/)

🪔✨ ARB Scope and Diversity

ARB is the first benchmark focused on step-by-step reasoning in Arabic cross both textual and visual modalities, covering 11 diverse domains spanning science, culture, OCR, and historical interpretation.

Figure: ARB Dataset Coverage

🌟 Key Features

  • Includes 1,356 multimodal samples with 5,119 curated reasoning steps.
  • Spans 11 diverse domains, from visual reasoning to historical and scientific analysis.
  • Emphasizes step-by-step reasoning, beyond just final answer prediction.
  • Each sample contains a chain of 2–6+ reasoning steps aligned to human logic.
  • Curated and verified by native Arabic speakers and domain experts for linguistic and cultural fidelity.
  • Built from hybrid sources: original Arabic data, high-quality translations, and synthetic samples.
  • Features a robust evaluation framework measuring both final answer accuracy and reasoning quality.
  • Fully open-source dataset and toolkit to support research in Arabic reasoning and multimodal AI.

🏗️ ARB Construction Pipeline

Figure: ARB Pipeline Overview

🗂️ ARB Collection

Figure: ARB Collection

🗂️ ARB Distribution

Figure: ARB dist

🧪 Evaluation Protocol

We evaluated 12 open- and closed-source LMMs using:

  • Lexical and Semantic Similarity Scoes: BLEU, ROUGE, BERTScore, LaBSE
  • Stepwise Evaluation Using LLM-as-Judge: Our curated metric includes 10 factors like faithfulness, interpretive depth, coherence, hallucination, and more.

🏆 Evaluation Results

  • Stepwise Evaluation Using LLM-as-Judge for Closed-Source Models: | Metric ↓ / Model → | GPT-4o | GPT-4o-mini | GPT-4.1 | o4-mini | Gemini 1.5 Pro | Gemini 2.0 Flash | |----------------------------|--------|-------------|---------|---------|----------------|------------------| | Final Answer (%) | 60.22 | 52.22 | 59.43 | 58.93 | 56.70 | 57.80 | | Reasoning Steps (%) | 64.29 | 61.02 | 80.41 | 80.75| 64.34 | 64.09 |

  • Stepwise Evaluation Using LLM-as-Judge for Open-Source Models: | Metric ↓ / Model → | Qwen2.5-VL | LLaMA-3.2 | AIN | LLaMA-4 Scout | Aya-Vision | InternVL3 | |----------------------------|------------|-----------|-------|----------------|-------------|-----------| | Final Answer (%) | 37.02 | 25.58 | 27.35 | 48.52 | 28.81 | 31.04 | | Reasoning Steps (%) | 64.03 | 53.20 | 52.77 | 77.70 | 63.64 | 54.50 |

📂 Dataset Structure

Each sample includes:

  • image_id: Visual input
  • question: Arabic question grounded in image reasoning
  • choices: The choices for the MCQ
  • steps: Ordered reasoning chain
  • answer: Final solution (Arabic)
  • category: One of 11 categories (e.g., OCR, Scientific, Visual, Math)

Example JSON: ```json { "image_id":"Chart_2.png", "question":"من خلال الرسم البياني لعدد القطع لكل عضو في الكشف عن السرطان، إذا جمعنا نسبة 'أخرى' مع نسبة 'الرئة'، فكيف يقاربان نسبة 'الكلى' تقريبًا؟", "answer":"ج", "choices":"['أ. مجموعهما أكبر بكثير من نسبة الكلى', 'ب. مجموعهما يساوي تقريبًا نسبة الكلى', 'ج. مجموعهما أقل بشكل ملحوظ من نسبة الكلى']", "steps":"الخطوة 1: تحديد النسب المئوية لكل من 'أخرى' و'الرئة' و'الكلى' من الرسم البياني.\nالإجراء 1: 'أخرى' = 0.7%، 'الرئة' = 1.8%، 'الكلى' = 4.3%.\n\nالخطوة 2: حساب مجموع النسب المئوية لـ 'أخرى' و'الرئة'.\nالإجراء 2: 0.7% + 1.8% = 2.5%.\n\nالخطوة 3: مقارنة مجموع النسب المئوية لـ 'أخرى' و'الرئة' مع نسبة 'الكلى'.\nالإجراء 3: 2.5% (مجموع 'أخرى' و'الرئة') أقل من 4.3% (نسبة 'الكلى').\n\nالخطوة 4: اختيار الإجابة الصحيحة بناءً على المقارنة.\nالإجراء 4: اختيار 'ج' لأن مجموعهما أقل بشكل ملحوظ من نسبة 'الكلى'.", "category ":"CDT", }, ```

📚 Citation

If you use ARB dataset in your research, please consider citing:

@misc{ghaboura2025arbcomprehensivearabicmultimodal,
      title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark}, 
      author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},
      year={2025},
      eprint={2505.17021},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2505.17021}, 
}

MBZUAI/ARB

Dataset

Sara Ghaboura *

11

28 commits

1 linked in READMEs

updated Jun 21, 2025

See the code

README

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

[![arXiv](https://img.shields.io/badge/arXiv-2505.17021-C0DAD9)](https://arxiv.org/abs/2505.17021) [![Our Page](https://img.shields.io/badge/Visit-Our%20Page-D4EBDB?style=flat)](https://mbzuai-oryx.github.io/ARB/)

🪔✨ ARB Scope and Diversity

ARB is the first benchmark focused on step-by-step reasoning in Arabic cross both textual and visual modalities, covering 11 diverse domains spanning science, culture, OCR, and historical interpretation.

Figure: ARB Dataset Coverage

🌟 Key Features

  • Includes 1,356 multimodal samples with 5,119 curated reasoning steps.
  • Spans 11 diverse domains, from visual reasoning to historical and scientific analysis.
  • Emphasizes step-by-step reasoning, beyond just final answer prediction.
  • Each sample contains a chain of 2–6+ reasoning steps aligned to human logic.
  • Curated and verified by native Arabic speakers and domain experts for linguistic and cultural fidelity.
  • Built from hybrid sources: original Arabic data, high-quality translations, and synthetic samples.
  • Features a robust evaluation framework measuring both final answer accuracy and reasoning quality.
  • Fully open-source dataset and toolkit to support research in Arabic reasoning and multimodal AI.

🏗️ ARB Construction Pipeline

Figure: ARB Pipeline Overview

🗂️ ARB Collection

Figure: ARB Collection

🗂️ ARB Distribution

Figure: ARB dist

🧪 Evaluation Protocol

We evaluated 12 open- and closed-source LMMs using:

  • Lexical and Semantic Similarity Scoes: BLEU, ROUGE, BERTScore, LaBSE
  • Stepwise Evaluation Using LLM-as-Judge: Our curated metric includes 10 factors like faithfulness, interpretive depth, coherence, hallucination, and more.

🏆 Evaluation Results

  • Stepwise Evaluation Using LLM-as-Judge for Closed-Source Models: | Metric ↓ / Model → | GPT-4o | GPT-4o-mini | GPT-4.1 | o4-mini | Gemini 1.5 Pro | Gemini 2.0 Flash | |----------------------------|--------|-------------|---------|---------|----------------|------------------| | Final Answer (%) | 60.22 | 52.22 | 59.43 | 58.93 | 56.70 | 57.80 | | Reasoning Steps (%) | 64.29 | 61.02 | 80.41 | 80.75| 64.34 | 64.09 |

  • Stepwise Evaluation Using LLM-as-Judge for Open-Source Models: | Metric ↓ / Model → | Qwen2.5-VL | LLaMA-3.2 | AIN | LLaMA-4 Scout | Aya-Vision | InternVL3 | |----------------------------|------------|-----------|-------|----------------|-------------|-----------| | Final Answer (%) | 37.02 | 25.58 | 27.35 | 48.52 | 28.81 | 31.04 | | Reasoning Steps (%) | 64.03 | 53.20 | 52.77 | 77.70 | 63.64 | 54.50 |

📂 Dataset Structure

Each sample includes:

  • image_id: Visual input
  • question: Arabic question grounded in image reasoning
  • choices: The choices for the MCQ
  • steps: Ordered reasoning chain
  • answer: Final solution (Arabic)
  • category: One of 11 categories (e.g., OCR, Scientific, Visual, Math)

Example JSON: ```json { "image_id":"Chart_2.png", "question":"من خلال الرسم البياني لعدد القطع لكل عضو في الكشف عن السرطان، إذا جمعنا نسبة 'أخرى' مع نسبة 'الرئة'، فكيف يقاربان نسبة 'الكلى' تقريبًا؟", "answer":"ج", "choices":"['أ. مجموعهما أكبر بكثير من نسبة الكلى', 'ب. مجموعهما يساوي تقريبًا نسبة الكلى', 'ج. مجموعهما أقل بشكل ملحوظ من نسبة الكلى']", "steps":"الخطوة 1: تحديد النسب المئوية لكل من 'أخرى' و'الرئة' و'الكلى' من الرسم البياني.\nالإجراء 1: 'أخرى' = 0.7%، 'الرئة' = 1.8%، 'الكلى' = 4.3%.\n\nالخطوة 2: حساب مجموع النسب المئوية لـ 'أخرى' و'الرئة'.\nالإجراء 2: 0.7% + 1.8% = 2.5%.\n\nالخطوة 3: مقارنة مجموع النسب المئوية لـ 'أخرى' و'الرئة' مع نسبة 'الكلى'.\nالإجراء 3: 2.5% (مجموع 'أخرى' و'الرئة') أقل من 4.3% (نسبة 'الكلى').\n\nالخطوة 4: اختيار الإجابة الصحيحة بناءً على المقارنة.\nالإجراء 4: اختيار 'ج' لأن مجموعهما أقل بشكل ملحوظ من نسبة 'الكلى'.", "category ":"CDT", }, ```

📚 Citation

If you use ARB dataset in your research, please consider citing:

@misc{ghaboura2025arbcomprehensivearabicmultimodal,
      title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark}, 
      author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},
      year={2025},
      eprint={2505.17021},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2505.17021}, 
}