StarDoc-AI/navidc-ocr-demo

Space

NaviDC-OCR

42

10 commits

1 linked in READMEs

updated Aug 18, 2026

See the code
gradio
mcp-server

README

NaviDC-OCR

Demo of StarDoc-AI/NaviDC-OCR, a 1.2B document-parsing vision-language model that unifies digital and camera-captured documents (paper · code).

The app follows the authors' two-stage pipeline:

  1. Layout — the page is resized to 1036×1036 and the model emits reading-ordered regions. Detection mode returns axis-aligned boxes (digital pages, flat scans); Segmentation mode returns multi-point polygons, which is the paper's geometry-aware path for photographed, curved or crumpled pages.
  2. Recognition — each region is cropped (polygon-masked and de-rotated where needed) and recognized with the block-type-specific prompt and sampling parameters from the reference implementation. Tables come back as OTSL and are converted to HTML, equations to LaTeX, using the authors' post-processors (vendored under NaviOCR/).

A third Single region mode skips layout and runs the authors' one-block path (block_parse) over the whole image with the prompt for a chosen block type — table, formula, code, seal, or a chart / scientific figure, which the model converts into the table it implies.

Outputs: rendered document, Markdown source, layout overlay, and the raw block list as JSON.

Credits

Example pages are the official assets from the NaviDC-OCR model card (Apache-2.0). The NaviOCR/ package is a trimmed copy of the authors' reference implementation (Apache-2.0), limited to the modules needed for the transformers backend.

Contributors

multimodalart

10 commits

StarDoc-AI/navidc-ocr-demo

Space

NaviDC-OCR

42

10 commits

1 linked in READMEs

updated Aug 18, 2026

See the code
gradio
mcp-server

README

NaviDC-OCR

Demo of StarDoc-AI/NaviDC-OCR, a 1.2B document-parsing vision-language model that unifies digital and camera-captured documents (paper · code).

The app follows the authors' two-stage pipeline:

  1. Layout — the page is resized to 1036×1036 and the model emits reading-ordered regions. Detection mode returns axis-aligned boxes (digital pages, flat scans); Segmentation mode returns multi-point polygons, which is the paper's geometry-aware path for photographed, curved or crumpled pages.
  2. Recognition — each region is cropped (polygon-masked and de-rotated where needed) and recognized with the block-type-specific prompt and sampling parameters from the reference implementation. Tables come back as OTSL and are converted to HTML, equations to LaTeX, using the authors' post-processors (vendored under NaviOCR/).

A third Single region mode skips layout and runs the authors' one-block path (block_parse) over the whole image with the prompt for a chosen block type — table, formula, code, seal, or a chart / scientific figure, which the model converts into the table it implies.

Outputs: rendered document, Markdown source, layout overlay, and the raw block list as JSON.

Credits

Example pages are the official assets from the NaviDC-OCR model card (Apache-2.0). The NaviOCR/ package is a trimmed copy of the authors' reference implementation (Apache-2.0), limited to the modules needed for the transformers backend.

Contributors

multimodalart

10 commits