MohamedRashad/arabic-img2md

Dataset

Arabic Img2MD

14

51 commits

2 linked in READMEs

updated Nov 28, 2024

See the code

README

Arabic Img2MD

Dataset Summary

The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:

  • Train: 13,700 examples
  • Test: 1,520 examples

This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text.

Usage

The dataset was used to train two models:

  • arabic-base-nougat
  • arabic-large-nougat

These models are designed for OCR tasks and converting PDF content to Markdown in the Arabic language context.

Research Context

The dataset supports the findings of the research paper: Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction.

Licensing

This dataset is released under the GPL-3.0 License, ensuring its open-source availability.

Citation

If you use this dataset, please cite the corresponding research paper:

@misc{rashad2024arabicnougatfinetuningvisiontransformers,
      title={Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction}, 
      author={Mohamed Rashad},
      year={2024},
      eprint={2411.17835},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2411.17835}, 
}

MohamedRashad/arabic-img2md

Dataset

Arabic Img2MD

14

51 commits

2 linked in READMEs

updated Nov 28, 2024

See the code

README

Arabic Img2MD

Dataset Summary

The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:

  • Train: 13,700 examples
  • Test: 1,520 examples

This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text.

Usage

The dataset was used to train two models:

  • arabic-base-nougat
  • arabic-large-nougat

These models are designed for OCR tasks and converting PDF content to Markdown in the Arabic language context.

Research Context

The dataset supports the findings of the research paper: Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction.

Licensing

This dataset is released under the GPL-3.0 License, ensuring its open-source availability.

Citation

If you use this dataset, please cite the corresponding research paper:

@misc{rashad2024arabicnougatfinetuningvisiontransformers,
      title={Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction}, 
      author={Mohamed Rashad},
      year={2024},
      eprint={2411.17835},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2411.17835}, 
}