PDF Document Parsing and OCR

15 repos

Python-based tools and libraries for extracting, parsing, and converting content from PDF documents using optical character recognition (OCR) and document layout analysis. This cluster focuses on the challenges of understanding document structure, handling complex layouts, and accurately extracting text and tabular data from PDFs—a common need in document automation, data pipeline, and knowledge extraction workflows. Central projects like Docling, MinerU, and related tools demonstrate both academic and production-focused approaches to this problem.

Python · 12
Rust · 2
Java · 1
ocr ·284,158
pdf-parser ·232,282
pdf ·212,027
pdf-extractor-rag ·169,890
ai4science ·169,890
python ·125,055
document-parsing ·120,706
rag ·119,064
pdf-converter ·109,525
extract-data ·90,873

pymupdf/PyMuPDF4LLM

PyMuPDF4LLM

Python

2,184

290 commits