42 repos across 2 sub-areas
Libraries and model implementations for vision-language tasks, particularly image captioning and image-to-text generation using transformer architectures. The cluster centers on Japanese language variants of LLaVA models and related multimodal AI systems that combine visual understanding with text generation. While the language composition appears sparse in the metadata, the consistent presence of image-captioning, transformers, and image-text-to-text topics across repositories indicates a focused area around building and deploying models that process both images and text.
Cluster 462344
28 repos
Document Understanding & OCR Models
14 repos
Deep learning models and frameworks for optical character recognition (OCR) and document understanding, particularly focused on vision transformer architectures for extracting text and structured data from scanned documents and images. The cluster centers on the Donut model family—a transformer-based approach to document understanding—along with supporting tools for multimodal document processing, layout analysis, and training on datasets like CORD (receipts) and RVL-CDIP (document classification).