SARD: Synthetic Arabic Recognition Dataset
14
500 commits
2 linked in READMEs
updated May 20, 2026
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Massive Scale: 2,621,075 document images containing 794.6 million words
Typographic Diversity: Five distinct Arabic fonts (Amiri, Sakkal Majalla, Arial, Calibri, Scheherazade New and Traditional Arabic)
Structured Formatting: Designed to mimic real-world book layouts with consistent typography
Clean Data: Synthetically generated with no scanning artifacts, blur, or distortions
Content Diversity: Text spans multiple domains including culture, literature, Shariah, social topics, and more
Storage Size: 1.54 TB
The dataset is divided into five splits based on font name:
π Sample Images
![]() | ![]() |
![]() | ![]() |
Each split contains data specific to a single font with the following attributes:
image_name: Unique identifier for each imagechunk: The text content associated with the imagefont_name: The font used in text renderingimage_base64: Base64-encoded image representationsample_id: Unique IDarticle_link: Link of the source article| Category | Number of Articles |
|---|---|
| Culture | 13,253 |
| Fatawa & Counsels | 8,096 |
| Literature & Language | 11,581 |
| Bibliography | 26,393 |
| Publications & Competitions | 1,123 |
| Shariah | 46,665 |
| Social | 8,827 |
| Translations | 443 |
| Muslim's News | 16,724 |
| Total Articles | 133,105 |
| Font | Words Per Page | Font Size |
|---|---|---|
| Sakkal Majalla | 50β300 | 14 pt |
| Arial | 50β500 | 12 pt |
| Calibri | 50β500 | 12 pt |
| Amiri | 50β300 | 12 pt |
| Scheherazade | 50β250 | 12 pt |
| Traditional Arabic | 50β350 | 14 pt |
| Specification | Measurement |
|---|---|
| Left Margin | 0.9 inches |
| Right Margin | 0.9 inches |
| Top Margin | 1.0 inch |
| Bottom Margin | 1.0 inch |
| Gutter Margin | 0.2 inches |
| Page Width | 8.27 inches (A4) |
| Page Height | 11.69 inches (A4) |
from datasets import load_dataset
import base64
from io import BytesIO
from PIL import Image
import matplotlib.pyplot as plt
# Load dataset with streaming enabled
ds = load_dataset("riotu-lab/SARD", streaming=True)
print(ds)
# Iterate over a specific font dataset (e.g., Amiri)
for sample in ds["Amiri"]:
image_name = sample["image_name"]
chunk = sample["chunk"] # Arabic text transcription
font_name = sample["font_name"]
# Decode Base64 image
image_data = base64.b64decode(sample["image_base64"])
image = Image.open(BytesIO(image_data))
# Display the image
plt.figure(figsize=(10, 10))
plt.imshow(image)
plt.axis('off')
plt.title(f"Font: {font_name}")
plt.show()
# Print the details
print(f"Image Name: {image_name}")
print(f"Font Name: {font_name}")
print(f"Text Chunk: {chunk}")
# Break after one sample for testing
break
SAND is designed to support various Arabic text recognition tasks:
The authors thank Prince Sultan University for their support in developing this dataset.
If you use SARD in your work, please cite the following paper:
APA:
Nacar, O., Al-Habashi, Y., Sibaee, S., Ammar, A., & Boulila, W. (2025). SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition. arXiv preprint arXiv:2505.24600.
https://arxiv.org/abs/2505.24600
BibTeX:
@misc{nacar2025sard,
title={SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition},
author={Omer Nacar and Yasser Al-Habashi and Serry Sibaee and Adel Ammar and Wadii Boulila},
year={2025},
eprint={2505.24600},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2505.24600},
}
Usage Note:
This dataset includes text sourced from Alukah. Use is limited to non-commercial purposes in accordance with the source terms. Redistribution, republication, or commercial use is not permitted without permission and proper attribution to the original source and author.
SARD: Synthetic Arabic Recognition Dataset
14
500 commits
2 linked in READMEs
updated May 20, 2026
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Massive Scale: 2,621,075 document images containing 794.6 million words
Typographic Diversity: Five distinct Arabic fonts (Amiri, Sakkal Majalla, Arial, Calibri, Scheherazade New and Traditional Arabic)
Structured Formatting: Designed to mimic real-world book layouts with consistent typography
Clean Data: Synthetically generated with no scanning artifacts, blur, or distortions
Content Diversity: Text spans multiple domains including culture, literature, Shariah, social topics, and more
Storage Size: 1.54 TB
The dataset is divided into five splits based on font name:
π Sample Images
![]() | ![]() |
![]() | ![]() |
Each split contains data specific to a single font with the following attributes:
image_name: Unique identifier for each imagechunk: The text content associated with the imagefont_name: The font used in text renderingimage_base64: Base64-encoded image representationsample_id: Unique IDarticle_link: Link of the source article| Category | Number of Articles |
|---|---|
| Culture | 13,253 |
| Fatawa & Counsels | 8,096 |
| Literature & Language | 11,581 |
| Bibliography | 26,393 |
| Publications & Competitions | 1,123 |
| Shariah | 46,665 |
| Social | 8,827 |
| Translations | 443 |
| Muslim's News | 16,724 |
| Total Articles | 133,105 |
| Font | Words Per Page | Font Size |
|---|---|---|
| Sakkal Majalla | 50β300 | 14 pt |
| Arial | 50β500 | 12 pt |
| Calibri | 50β500 | 12 pt |
| Amiri | 50β300 | 12 pt |
| Scheherazade | 50β250 | 12 pt |
| Traditional Arabic | 50β350 | 14 pt |
| Specification | Measurement |
|---|---|
| Left Margin | 0.9 inches |
| Right Margin | 0.9 inches |
| Top Margin | 1.0 inch |
| Bottom Margin | 1.0 inch |
| Gutter Margin | 0.2 inches |
| Page Width | 8.27 inches (A4) |
| Page Height | 11.69 inches (A4) |
from datasets import load_dataset
import base64
from io import BytesIO
from PIL import Image
import matplotlib.pyplot as plt
# Load dataset with streaming enabled
ds = load_dataset("riotu-lab/SARD", streaming=True)
print(ds)
# Iterate over a specific font dataset (e.g., Amiri)
for sample in ds["Amiri"]:
image_name = sample["image_name"]
chunk = sample["chunk"] # Arabic text transcription
font_name = sample["font_name"]
# Decode Base64 image
image_data = base64.b64decode(sample["image_base64"])
image = Image.open(BytesIO(image_data))
# Display the image
plt.figure(figsize=(10, 10))
plt.imshow(image)
plt.axis('off')
plt.title(f"Font: {font_name}")
plt.show()
# Print the details
print(f"Image Name: {image_name}")
print(f"Font Name: {font_name}")
print(f"Text Chunk: {chunk}")
# Break after one sample for testing
break
SAND is designed to support various Arabic text recognition tasks:
The authors thank Prince Sultan University for their support in developing this dataset.
If you use SARD in your work, please cite the following paper:
APA:
Nacar, O., Al-Habashi, Y., Sibaee, S., Ammar, A., & Boulila, W. (2025). SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition. arXiv preprint arXiv:2505.24600.
https://arxiv.org/abs/2505.24600
BibTeX:
@misc{nacar2025sard,
title={SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition},
author={Omer Nacar and Yasser Al-Habashi and Serry Sibaee and Adel Ammar and Wadii Boulila},
year={2025},
eprint={2505.24600},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2505.24600},
}
Usage Note:
This dataset includes text sourced from Alukah. Use is limited to non-commercial purposes in accordance with the source terms. Redistribution, republication, or commercial use is not permitted without permission and proper attribution to the original source and author.