This repository provides a pipeline for generating an Arabic OCR dataset. The pipeline processes textual data, converts it into various font styles, generates PDF and image representations, and stores the output in a structured format suitable for OCR training.
Install the required Python libraries using:
pip install datasets python-docx pdf2image PIL requests huggingface_hub
To ensure full reproducibility of the SARD dataset generation and proper handling of Arabic Complex Text Layout (CTL), the following software environment was utilized:
Ensure your dataset is available on Hugging Face and update the script with:
DATASET_NAME: Your dataset name.repo_id: Your Hugging Face dataset repository ID.YOUR_API_TOKEN: Your Hugging Face API token.Execute the script to process text and generate the dataset:
python text2image.py
If the script stops, it will resume from the last processed index using processing_state.json.
Each batch of processed data is stored as a CSV file with the following columns:
image_name: Unique identifier for each image
chunk: The text content associated with the image
font_name: The font used in text rendering
image_base64: Base64-encoded image representation
sample_id: Unique ID
article_link: Link of the source article
For books_links.pkl it has the links for the used books in the statistical study for choosing the fonts in SARD dataset
import pickle
with open("books_links.pkl", "rb") as f:
data = f.load()
print(data[0:])
Python
100.0%
This repository provides a pipeline for generating an Arabic OCR dataset. The pipeline processes textual data, converts it into various font styles, generates PDF and image representations, and stores the output in a structured format suitable for OCR training.
Install the required Python libraries using:
pip install datasets python-docx pdf2image PIL requests huggingface_hub
To ensure full reproducibility of the SARD dataset generation and proper handling of Arabic Complex Text Layout (CTL), the following software environment was utilized:
Ensure your dataset is available on Hugging Face and update the script with:
DATASET_NAME: Your dataset name.repo_id: Your Hugging Face dataset repository ID.YOUR_API_TOKEN: Your Hugging Face API token.Execute the script to process text and generate the dataset:
python text2image.py
If the script stops, it will resume from the last processed index using processing_state.json.
Each batch of processed data is stored as a CSV file with the following columns:
image_name: Unique identifier for each image
chunk: The text content associated with the image
font_name: The font used in text rendering
image_base64: Base64-encoded image representation
sample_id: Unique ID
article_link: Link of the source article
For books_links.pkl it has the links for the used books in the statistical study for choosing the fonts in SARD dataset
import pickle
with open("books_links.pkl", "rb") as f:
data = f.load()
print(data[0:])
Python
100.0%