14 repos
Deep learning models and frameworks for optical character recognition (OCR) and document understanding, particularly focused on vision transformer architectures for extracting text and structured data from scanned documents and images. The cluster centers on the Donut model family—a transformer-based approach to document understanding—along with supporting tools for multimodal document processing, layout analysis, and training on datasets like CORD (receipts) and RVL-CDIP (document classification).