An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
2,083
stars
144
commits
Python
primary language
Mar 17, 2026
updated
An on-premises document information extraction and benchmarking toolkit.

We're excited to announce the release of Nanonets-OCR-s, a compact 3B parameter model specifically trained for efficient image to markdown conversion with semantic understanding for images, signatures, watermarks, etc.!
📢 Read the full announcement | 🤗 Hugging Face model
docext is a comprehensive on-premises document intelligence toolkit powered by vision-language models (VLMs). It provides three core capabilities:
📄 PDF & Image to Markdown Conversion: Transform documents into structured markdown with intelligent content recognition, including LaTeX equations, signatures, watermarks, tables, and semantic tagging.
🔍 Document Information Extraction: OCR-free extraction of structured information (fields, tables, etc.) from documents such as invoices, passports, and other document types, with confidence scoring.
📊 Intelligent Document Processing Leaderboard: A comprehensive benchmarking platform that tracks and evaluates vision-language model performance across OCR, Key Information Extraction (KIE), document classification, table extraction, and other intelligent document processing tasks.
Convert both PDF and images to markdown with content recognition and semantic tagging.
<img></img> tags.<signature></signature> tags.<watermark></watermark> tags.<page_number></page_number> tags.🔍 For in-depth information, see the release blog.
For setup instructions and additional details, check out the full feature guide for the pdf to markdown.
This benchmark evaluates performance across seven key document intelligence challenges:
🔍 For in-depth information, see the release blog.
📊 Live leaderboard: https://idp-leaderboard.org
For setup instructions and additional details, check out the full feature guide for the Intelligent Document Processing Leaderboard.
For more details (Installation, Usage, and so on), please check out the feature guide.
gemini-2.5-pro-preview-06-05 evaluation metrics to the leaderboard.docext extraction.gemini-2.5-pro-preview-03-25, claude-sonnet-4 evaluation metrics to the leaderboard.InternVL3-38B-Instruct, qwen2.5-vl-32b-instruct evaluation metrics to the leaderboard.gemma-3-27b-it evaluation metrics to the leaderboard.Claude 3.7 sonnet, mistral-medium-3 evaluation metrics to the leaderboard.docext is developed by Nanonets, a leader in document AI and intelligent document processing solutions. Nanonets is committed to advancing the field of document understanding through open-source contributions and innovative AI technologies. If you are looking for information extraction solutions for your business, please visit our website to learn more.
We welcome contributions! Please see contribution.md for guidelines. If you have a feature request or need support for a new model, feel free to open an issue—we'd love to discuss it further!
If you encounter any issues while using docext, please refer to our Troubleshooting guide for common problems and solutions.
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
Python
96.7%
Jupyter Notebook
2.6%
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
2,083
stars
144
commits
Python
primary language
Mar 17, 2026
updated
An on-premises document information extraction and benchmarking toolkit.

We're excited to announce the release of Nanonets-OCR-s, a compact 3B parameter model specifically trained for efficient image to markdown conversion with semantic understanding for images, signatures, watermarks, etc.!
📢 Read the full announcement | 🤗 Hugging Face model
docext is a comprehensive on-premises document intelligence toolkit powered by vision-language models (VLMs). It provides three core capabilities:
📄 PDF & Image to Markdown Conversion: Transform documents into structured markdown with intelligent content recognition, including LaTeX equations, signatures, watermarks, tables, and semantic tagging.
🔍 Document Information Extraction: OCR-free extraction of structured information (fields, tables, etc.) from documents such as invoices, passports, and other document types, with confidence scoring.
📊 Intelligent Document Processing Leaderboard: A comprehensive benchmarking platform that tracks and evaluates vision-language model performance across OCR, Key Information Extraction (KIE), document classification, table extraction, and other intelligent document processing tasks.
Convert both PDF and images to markdown with content recognition and semantic tagging.
<img></img> tags.<signature></signature> tags.<watermark></watermark> tags.<page_number></page_number> tags.🔍 For in-depth information, see the release blog.
For setup instructions and additional details, check out the full feature guide for the pdf to markdown.
This benchmark evaluates performance across seven key document intelligence challenges:
🔍 For in-depth information, see the release blog.
📊 Live leaderboard: https://idp-leaderboard.org
For setup instructions and additional details, check out the full feature guide for the Intelligent Document Processing Leaderboard.
For more details (Installation, Usage, and so on), please check out the feature guide.
gemini-2.5-pro-preview-06-05 evaluation metrics to the leaderboard.docext extraction.gemini-2.5-pro-preview-03-25, claude-sonnet-4 evaluation metrics to the leaderboard.InternVL3-38B-Instruct, qwen2.5-vl-32b-instruct evaluation metrics to the leaderboard.gemma-3-27b-it evaluation metrics to the leaderboard.Claude 3.7 sonnet, mistral-medium-3 evaluation metrics to the leaderboard.docext is developed by Nanonets, a leader in document AI and intelligent document processing solutions. Nanonets is committed to advancing the field of document understanding through open-source contributions and innovative AI technologies. If you are looking for information extraction solutions for your business, please visit our website to learn more.
We welcome contributions! Please see contribution.md for guidelines. If you have a feature request or need support for a new model, feel free to open an issue—we'd love to discuss it further!
If you encounter any issues while using docext, please refer to our Troubleshooting guide for common problems and solutions.
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
Python
96.7%
Jupyter Notebook
2.6%