End-to-End Structured OCR for Arabic books.
Github ๐ค Hugging Face ๐ Paper ๐๏ธ Data ๐ฝ๏ธ Demo
The arabic-base-nougat OCR is an end-to-end structured Optical Character Recognition (OCR) system designed specifically for the Arabic language.
The model is based on the facebook/nougat-base architecture and has been fine-tuned using MohamedRashad/arabic-img2md.
Demo: https://huggingface.co/spaces/MohamedRashad/Arabic-Nougat
Or, use the code below to get started with the model locally.
Don't forget to update transformers:
pip install -U transformers
from PIL import Image
import torch
from transformers import NougatProcessor, VisionEncoderDecoderModel
# Load the model and processor
processor = NougatProcessor.from_pretrained("MohamedRashad/arabic-base-nougat")
model = VisionEncoderDecoderModel.from_pretrained("MohamedRashad/arabic-base-nougat", torch_dtype=torch.bfloat16, attn_implementation={"decoder": "flash_attention_2", "encoder": "eager"})
# Get the max context length of the model & dtype of the weights
context_length = model.decoder.config.max_position_embeddings
torch_dtype = model.dtype
# Move the model to GPU if available
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
def predict(img_path):
# prepare PDF image for the model
image = Image.open(img_path)
pixel_values = processor(image, return_tensors="pt").pixel_values.to(torch_dtype).to(device)
# generate transcription
outputs = model.generate(
pixel_values.to(device),
repetition_penalty=1.5,
min_length=1,
max_new_tokens=context_length,
bad_words_ids=[[processor.tokenizer.unk_token_id]],
)
page_sequence = processor.batch_decode(outputs, skip_special_tokens=True)[0]
page_sequence = processor.post_process_generation(page_sequence, fix_markdown=False)
return page_sequence
print(predict("path/to/page_image.jpg"))
The arabic-base-nougat OCR is designed for tasks that involve converting images of Arabic book pages into structured text, especially when Markdown format is desired. It is suitable for applications in the field of digitizing Arabic literature and facilitating text extraction from printed materials.
It is crucial to be aware of the model's limitations, particularly in instances where accurate OCR results are critical. Users are advised to verify and review the output, especially in scenarios where precision is paramount.
If you use or build upon the arabic-base-nougat OCR, please acknowledge the model developer and the open-source community for their contributions. Additionally, be sure to include a copy of the GPL 3.0 license with any redistributed or modified versions of the model.
By selecting the GPL 3.0 license, you promote the principles of open source and ensure that the benefits of the model are shared with the broader community.
If you find this model useful, please cite the corresponding research paper:
@misc{rashad2024arabicnougatfinetuningvisiontransformers,
title={Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction},
author={Mohamed Rashad},
year={2024},
eprint={2411.17835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2411.17835},
}
The arabic-base-nougat OCR is a tool provided "as is," and the developers make no guarantees regarding its suitability for specific tasks. Users are encouraged to thoroughly evaluate the model's output for their particular use cases and requirements.
End-to-End Structured OCR for Arabic books.
Github ๐ค Hugging Face ๐ Paper ๐๏ธ Data ๐ฝ๏ธ Demo
The arabic-base-nougat OCR is an end-to-end structured Optical Character Recognition (OCR) system designed specifically for the Arabic language.
The model is based on the facebook/nougat-base architecture and has been fine-tuned using MohamedRashad/arabic-img2md.
Demo: https://huggingface.co/spaces/MohamedRashad/Arabic-Nougat
Or, use the code below to get started with the model locally.
Don't forget to update transformers:
pip install -U transformers
from PIL import Image
import torch
from transformers import NougatProcessor, VisionEncoderDecoderModel
# Load the model and processor
processor = NougatProcessor.from_pretrained("MohamedRashad/arabic-base-nougat")
model = VisionEncoderDecoderModel.from_pretrained("MohamedRashad/arabic-base-nougat", torch_dtype=torch.bfloat16, attn_implementation={"decoder": "flash_attention_2", "encoder": "eager"})
# Get the max context length of the model & dtype of the weights
context_length = model.decoder.config.max_position_embeddings
torch_dtype = model.dtype
# Move the model to GPU if available
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
def predict(img_path):
# prepare PDF image for the model
image = Image.open(img_path)
pixel_values = processor(image, return_tensors="pt").pixel_values.to(torch_dtype).to(device)
# generate transcription
outputs = model.generate(
pixel_values.to(device),
repetition_penalty=1.5,
min_length=1,
max_new_tokens=context_length,
bad_words_ids=[[processor.tokenizer.unk_token_id]],
)
page_sequence = processor.batch_decode(outputs, skip_special_tokens=True)[0]
page_sequence = processor.post_process_generation(page_sequence, fix_markdown=False)
return page_sequence
print(predict("path/to/page_image.jpg"))
The arabic-base-nougat OCR is designed for tasks that involve converting images of Arabic book pages into structured text, especially when Markdown format is desired. It is suitable for applications in the field of digitizing Arabic literature and facilitating text extraction from printed materials.
It is crucial to be aware of the model's limitations, particularly in instances where accurate OCR results are critical. Users are advised to verify and review the output, especially in scenarios where precision is paramount.
If you use or build upon the arabic-base-nougat OCR, please acknowledge the model developer and the open-source community for their contributions. Additionally, be sure to include a copy of the GPL 3.0 license with any redistributed or modified versions of the model.
By selecting the GPL 3.0 license, you promote the principles of open source and ensure that the benefits of the model are shared with the broader community.
If you find this model useful, please cite the corresponding research paper:
@misc{rashad2024arabicnougatfinetuningvisiontransformers,
title={Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction},
author={Mohamed Rashad},
year={2024},
eprint={2411.17835},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2411.17835},
}
The arabic-base-nougat OCR is a tool provided "as is," and the developers make no guarantees regarding its suitability for specific tasks. Users are encouraged to thoroughly evaluate the model's output for their particular use cases and requirements.