2 repos
Optical character recognition (OCR) and document parsing systems for extracting text, structure, and layout information from images and PDF documents. The cluster includes both specialized OCR engines (like HunyuanOCR) and supporting infrastructure for document analysis, with particular emphasis on structured document understanding for downstream applications like retrieval-augmented generation (RAG) and document translation. Most repos are Python-based tools and libraries for building OCR pipelines and document extraction workflows.