Document Understanding & OCR Models

14 repos

Deep learning models and frameworks for optical character recognition (OCR) and document understanding, particularly focused on vision transformer architectures for extracting text and structured data from scanned documents and images. The cluster centers on the Donut model family—a transformer-based approach to document understanding—along with supporting tools for multimodal document processing, layout analysis, and training on datasets like CORD (receipts) and RVL-CDIP (document classification).

image-to-text ·1,497
pytorch ·1,490
image-text-to-text ·1,431
vision-encoder-decoder ·1,431
endpoints_compatible ·1,431
transformers ·1,431
vision ·748
donut ·685
safetensors ·629
trocr ·517