The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:
This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text.
The dataset was used to train two models:
arabic-base-nougatarabic-large-nougatThese models are designed for OCR tasks and converting PDF content to Markdown in the Arabic language context.
The dataset supports the findings of the research paper: Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction.
This dataset is released under the GPL-3.0 License, ensuring its open-source availability.
If you use this dataset, please cite the corresponding research paper:
@misc{rashad2024arabicnougatfinetuningvisiontransformers,
title={Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction},
author={Mohamed Rashad},
year={2024},
eprint={2411.17835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2411.17835},
}
The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:
This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text.
The dataset was used to train two models:
arabic-base-nougatarabic-large-nougatThese models are designed for OCR tasks and converting PDF content to Markdown in the Arabic language context.
The dataset supports the findings of the research paper: Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction.
This dataset is released under the GPL-3.0 License, ensuring its open-source availability.
If you use this dataset, please cite the corresponding research paper:
@misc{rashad2024arabicnougatfinetuningvisiontransformers,
title={Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction},
author={Mohamed Rashad},
year={2024},
eprint={2411.17835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2411.17835},
}