Optical Character Recognition and Document Processing

16 repos

A cluster focused on OCR (optical character recognition) technology, centered around the Tesseract engine and related tools for extracting text from images and PDFs. The repositories span multiple implementations—Python libraries like tesserocr and OCRmyPDF for practical document automation, JavaScript bindings for browser-based OCR, C++ core engines, and supporting datasets (tessdata_fast). The cluster also includes modern approaches like LLM-aided OCR that combine traditional vision with language models, reflecting the area's evolution from rule-based to hybrid AI-driven text extraction.

Python · 7
C++ · 2
JavaScript · 2
C · 1
ocr ·253,915
tesseract ·174,347
machine-learning ·151,402
lstm ·111,037
pdf ·90,372
tesseract-ocr ·81,052
optical-character-recognition ·77,131
ocr-engine ·76,443
hacktoberfest ·76,443
python ·75,383