HeshamHaroon/ArabicRAGB

Dataset

ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark

13

14 commits

1 linked in READMEs

updated Dec 15, 2025

See the code

README

ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark

Arabic 5 Dialects 13K+ Records

Dataset Description

ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content.

Key Features

  • Passage-Grounded Queries: Each query is generated from and answerable by its paired passage
  • Multi-Dialect Coverage: MSA, Egyptian, Gulf, Levantine, and Maghrebi Arabic
  • Complexity Levels: Simple, Moderate, Complex, and Multi-hop queries
  • Unified Format: Query and passage together in each record

Dataset Statistics

ComponentCount
Training Records10,530
Test Records2,633
Total Records13,163

Dialect Distribution

DialectCount
MSA4,012
Egyptian3,999
Gulf3,783
Levantine697
Maghrebi672

Complexity Distribution

LevelCount
Simple4,092
Moderate4,030
Complex4,215
Multi-hop826

Usage

Loading the Dataset

from datasets import load_dataset

# Load the dataset
dataset = load_dataset("HeshamHaroon/ArabicRAGB")

train_data = dataset["train"]
test_data = dataset["test"]

print(f"Train records: {len(train_data)}")
print(f"Test records: {len(test_data)}")

# Example record
example = train_data[0]
print(f"Query: {example['query']}")
print(f"Dialect: {example['query_dialect']}")
print(f"Passage: {example['passage_text'][:200]}...")

Example: RAG Evaluation

from datasets import load_dataset
from sentence_transformers import SentenceTransformer
import numpy as np

# Load dataset
dataset = load_dataset("HeshamHaroon/ArabicRAGB")
test_data = dataset["test"]

# Load multilingual model
model = SentenceTransformer('sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2')

# Build passage index from unique passages
passages = list({r["passage_id"]: r["passage_text"] for r in test_data}.values())
passage_embeddings = model.encode(passages, show_progress_bar=True)

# Evaluate retrieval
def retrieve(query, top_k=5):
    query_emb = model.encode(query)
    scores = np.dot(passage_embeddings, query_emb)
    top_indices = np.argsort(scores)[-top_k:][::-1]
    return [passages[i] for i in top_indices]

# Test with a query
results = retrieve(test_data[0]["query"])
print(f"Query: {test_data[0]['query']}")
print(f"Top result: {results[0][:200]}...")

Data Format

Each record contains a query paired with its source passage:

{
  "id": "r_abc123def456",
  "query": "ما هي شروط الحصول على تأشيرة دخول السعودية؟",
  "query_dialect": "msa",
  "query_complexity": "simple",
  "passage_id": "p_xyz789",
  "passage_title": "تأشيرة السعودية",
  "passage_text": "للحصول على تأشيرة دخول المملكة العربية السعودية...",
  "source_url": "https://example.com/visa",
  "source_category": "government"
}

Field Descriptions

FieldDescription
idUnique record identifier
queryArabic question answerable from the passage
query_dialectDialect: msa, egyptian, gulf, levantine, maghrebi
query_complexityComplexity: simple, moderate, complex, multi_hop
passage_idUnique passage identifier
passage_titleTitle of the source document
passage_textFull passage text (query answer is within)
source_urlOriginal source URL
source_categoryCategory: wikipedia, government, healthcare, etc.

Topics Covered

The corpus covers diverse topics relevant to Arabic speakers:

CategoryTopics
GeographyArab countries, major cities
HistoryIslamic history, regional history
CultureArabic literature, poetry
ScienceAI, computing, technology
HealthMedical conditions, healthcare
LawLegal systems, rights
EconomyFinance, investments

Evaluation Metrics

Recommended metrics for benchmarking:

Retrieval:

  • MRR@k (Mean Reciprocal Rank)
  • Recall@k
  • NDCG@k

Generation:

  • BLEU, ROUGE-L
  • BERTScore (multilingual)

Limitations

  • Queries are synthetically generated (grounded in passages)
  • No human-annotated relevance scores
  • Content reflects December 2025 snapshot

Citation

@dataset{arabicragb2025,
  title={ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark},
  author={Hesham Haroun},
  year={2025},
  publisher={Hugging Face},
  url={https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB}
}

License

This dataset is released under CC BY-SA 4.0.


Built for advancing Arabic NLP research
arabic
benchmark
egyptian-arabic
gulf-arabic
levantine-arabic
maghrebi-arabic
multi-dialect
retrieval

HeshamHaroon/ArabicRAGB

Dataset

ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark

13

14 commits

1 linked in READMEs

updated Dec 15, 2025

See the code

README

ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark

Arabic 5 Dialects 13K+ Records

Dataset Description

ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content.

Key Features

  • Passage-Grounded Queries: Each query is generated from and answerable by its paired passage
  • Multi-Dialect Coverage: MSA, Egyptian, Gulf, Levantine, and Maghrebi Arabic
  • Complexity Levels: Simple, Moderate, Complex, and Multi-hop queries
  • Unified Format: Query and passage together in each record

Dataset Statistics

ComponentCount
Training Records10,530
Test Records2,633
Total Records13,163

Dialect Distribution

DialectCount
MSA4,012
Egyptian3,999
Gulf3,783
Levantine697
Maghrebi672

Complexity Distribution

LevelCount
Simple4,092
Moderate4,030
Complex4,215
Multi-hop826

Usage

Loading the Dataset

from datasets import load_dataset

# Load the dataset
dataset = load_dataset("HeshamHaroon/ArabicRAGB")

train_data = dataset["train"]
test_data = dataset["test"]

print(f"Train records: {len(train_data)}")
print(f"Test records: {len(test_data)}")

# Example record
example = train_data[0]
print(f"Query: {example['query']}")
print(f"Dialect: {example['query_dialect']}")
print(f"Passage: {example['passage_text'][:200]}...")

Example: RAG Evaluation

from datasets import load_dataset
from sentence_transformers import SentenceTransformer
import numpy as np

# Load dataset
dataset = load_dataset("HeshamHaroon/ArabicRAGB")
test_data = dataset["test"]

# Load multilingual model
model = SentenceTransformer('sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2')

# Build passage index from unique passages
passages = list({r["passage_id"]: r["passage_text"] for r in test_data}.values())
passage_embeddings = model.encode(passages, show_progress_bar=True)

# Evaluate retrieval
def retrieve(query, top_k=5):
    query_emb = model.encode(query)
    scores = np.dot(passage_embeddings, query_emb)
    top_indices = np.argsort(scores)[-top_k:][::-1]
    return [passages[i] for i in top_indices]

# Test with a query
results = retrieve(test_data[0]["query"])
print(f"Query: {test_data[0]['query']}")
print(f"Top result: {results[0][:200]}...")

Data Format

Each record contains a query paired with its source passage:

{
  "id": "r_abc123def456",
  "query": "ما هي شروط الحصول على تأشيرة دخول السعودية؟",
  "query_dialect": "msa",
  "query_complexity": "simple",
  "passage_id": "p_xyz789",
  "passage_title": "تأشيرة السعودية",
  "passage_text": "للحصول على تأشيرة دخول المملكة العربية السعودية...",
  "source_url": "https://example.com/visa",
  "source_category": "government"
}

Field Descriptions

FieldDescription
idUnique record identifier
queryArabic question answerable from the passage
query_dialectDialect: msa, egyptian, gulf, levantine, maghrebi
query_complexityComplexity: simple, moderate, complex, multi_hop
passage_idUnique passage identifier
passage_titleTitle of the source document
passage_textFull passage text (query answer is within)
source_urlOriginal source URL
source_categoryCategory: wikipedia, government, healthcare, etc.

Topics Covered

The corpus covers diverse topics relevant to Arabic speakers:

CategoryTopics
GeographyArab countries, major cities
HistoryIslamic history, regional history
CultureArabic literature, poetry
ScienceAI, computing, technology
HealthMedical conditions, healthcare
LawLegal systems, rights
EconomyFinance, investments

Evaluation Metrics

Recommended metrics for benchmarking:

Retrieval:

  • MRR@k (Mean Reciprocal Rank)
  • Recall@k
  • NDCG@k

Generation:

  • BLEU, ROUGE-L
  • BERTScore (multilingual)

Limitations

  • Queries are synthetically generated (grounded in passages)
  • No human-annotated relevance scores
  • Content reflects December 2025 snapshot

Citation

@dataset{arabicragb2025,
  title={ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark},
  author={Hesham Haroun},
  year={2025},
  publisher={Hugging Face},
  url={https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB}
}

License

This dataset is released under CC BY-SA 4.0.


Built for advancing Arabic NLP research
arabic
benchmark
egyptian-arabic
gulf-arabic
levantine-arabic
maghrebi-arabic
multi-dialect
retrieval