15 repos
Python-based tools and libraries for extracting, parsing, and converting content from PDF documents using optical character recognition (OCR) and document layout analysis. This cluster focuses on the challenges of understanding document structure, handling complex layouts, and accurately extracting text and tabular data from PDFs—a common need in document automation, data pipeline, and knowledge extraction workflows. Central projects like Docling, MinerU, and related tools demonstrate both academic and production-focused approaches to this problem.