zhihengli-casia/puredocbench

PureDocBench: source-traceable benchmark for document parsing across clean, degraded, and real-world settings

Python

45

27 commits

updated Aug 29, 2026

See the code
benchmark
document-ai
document-parsing
ocr
synthetic-data

README

PureDocBench

How far is document parsing from solved?
A source-traceable benchmark for OCR and document parsing across clean, digitally degraded, and real-degraded document settings.

Hugging Face Dataset Data License Code License Paper

中文说明 | Dataset | Paper | GT Review & Corrections

PureDocBench uses HTML/CSS document sources as hidden anchors: each page is rendered into images and annotated from the same structured source. This gives a benchmark where text, tables, formulas, captions, and reading order can be scored with less post-hoc annotation noise.

PureDocBench 是一个源可追踪的 OCR / 文档解析 benchmark。数据由 HTML/CSS 源文件渲染而来,GT 标注从同源结构中抽取,覆盖 clean、digital-degraded、real-degraded 三条图像轨道。

Updates

  • Current GT: The stable alias is puredocbench-gt-latest; the exact revision and timestamp are recorded in Hugging Face gt/latest.json.
  • 2026-06-14: Updated GT annotations and opened a GT Review app for community corrections.
  • 2026-05-08: Initial public release of PureDocBench, including the paper PDF and full dataset on Hugging Face.

GT Annotation Examples

The examples below show colored coordinate boxes over clean rendered pages from an academic paper, a patent form, and a tuition invoice.

PureDocBench GT coordinate annotation examples

PureDocBench overview

At A Glance

ItemCount
Official pages1,475
Official images4,425
Top-level domains10
Fine-grained subcategories66
Image tracksclean, digital-degraded, real-degraded
Scored structurestext, formulas, tables, reading order

Main Leaderboard

The paper evaluates 40 systems across pipeline specialists, end-to-end document parsers, and general-purpose VLMs. The live leaderboard additionally tracks public post-publication results. Each track reports Overall, TextEdit, FormulaCDM, TableTEDS, and ROEdit; Avg3 averages the three track Overall scores.

Open the interactive sortable leaderboard ↗

Three-track leaderboard on PureDocBench
ModelParamsCleanDigital DegradedReal DegradedAvg3
Overall↑TextEditFormulaCDMTableTEDSROEditOverall↑TextEditFormulaCDMTableTEDSROEditOverall↑TextEditFormulaCDMTableTEDSROEdit
Pipeline / Multi-stage Specialists
NaviDC-OCR1.2B86.900.11181.0191.0977.470.20672.5980.4570.850.30265.1177.6678.41
DotsMOCR3B76.270.15166.2377.650.27373.160.19864.3274.950.30961.730.31254.3961.970.39370.39
Dolphin-v23B65.900.34259.8072.120.42960.240.39352.2067.860.46144.920.55339.9850.040.55857.02
MonkeyOCR-pro-3B3B62.230.34648.4672.830.49257.400.39745.5766.320.52646.490.51138.1852.430.60055.37
YouTu-Parsing2B75.020.23067.3480.740.35869.660.27061.4474.490.38860.290.36052.2064.690.43068.32
MinerU2.5-Pro1.2B75.870.22265.1484.680.34671.770.27261.7980.730.37862.560.37552.7072.470.44670.07
MinerU2.51.2B74.900.18462.0881.040.32768.920.24556.9974.240.37459.150.37049.0165.410.44667.66
MonkeyOCR-pro-1.2B1.2B61.090.35847.4371.600.49855.720.41643.9164.830.52943.820.55636.9450.070.60953.54
PaddleOCR-VL-1.50.9B73.010.26663.5382.120.42866.730.33958.0376.070.47860.500.39854.0067.330.51066.75
GLM-OCR0.9B68.650.31457.8979.440.47063.060.38353.2374.210.52058.310.43350.3467.830.54363.34
OpenOCR0.1B32.700.35433.500.000.50730.030.41031.090.000.54125.730.48625.810.000.59129.49
End-to-End Specialists
OvisOCR2*0.8B81.5577.0966.5675.06
olmOCR-2-7B7B69.360.28456.8979.590.35865.870.31854.5774.810.37856.100.41748.7961.250.43963.78
olmOCR-7B7B62.560.38858.6967.770.46657.840.43655.4461.660.49947.300.54246.2649.800.56855.90
FD-RL4B78.380.19368.2186.220.33476.330.21467.1683.220.35067.040.29858.8272.080.39173.92
Logics-Parsing-v24B76.350.21367.6782.670.34273.850.24867.3379.020.37567.640.30461.6571.640.41672.61
OCRVerse4B73.180.27363.7883.090.39371.360.30263.9580.360.41563.660.36357.0370.300.45269.40
Qianfan-OCR4B57.220.37049.7958.830.44350.850.43844.4151.960.48545.060.49439.0845.530.50951.04
Nanonets-OCR23B64.830.25444.9874.940.37761.230.30745.4068.970.40849.030.43535.5055.090.46858.36
DeepSeek-OCR-23B55.530.35446.0056.010.46649.410.41240.7848.670.49343.600.48637.3042.060.53349.51
OCRFlux-3B3B47.140.45438.3548.460.42441.820.48631.9042.170.43737.210.55932.6534.870.49142.06
DeepSeek-OCR3B53.500.41945.3957.060.51446.950.47839.9948.640.54840.480.53734.0441.120.57546.98
dots.ocr2.9B72.010.24861.3779.510.37965.950.30756.6771.860.41755.680.40347.7059.630.46764.55
FireRed-OCR2B70.810.28763.8677.230.39668.490.31962.6474.770.42257.420.41551.6062.160.47465.57
HunyuanOCR1B65.610.26955.7468.020.38261.490.30851.6263.680.40054.580.42148.3057.540.45960.56
UniRec-0.1B0.1B58.910.42251.3167.600.52652.420.50148.3759.040.57834.440.65830.9738.160.68548.59
OpenDoc-0.1B0.1B60.280.41153.0968.860.51952.460.50148.4159.040.57744.270.54738.4649.060.60352.00
General VLMs: Qwen3.5
Qwen3.5-397B-A17B397B/17B69.120.23365.2665.400.36668.340.24463.9165.530.37662.700.28760.7056.120.39966.72
Qwen3.5-122B-A10B122B/10B76.140.22667.9683.030.37576.340.22067.8283.210.36669.850.28162.1975.440.40174.11
Qwen3.5-35B-A3B35B/3B68.400.23264.9463.450.37468.040.24564.7863.860.37960.590.31059.6853.070.41965.68
Qwen3.5-27B27B72.070.22766.3672.510.36270.730.23664.6171.170.36765.920.28361.2364.820.39069.57
Qwen3.5-9B9B73.870.25467.6079.390.38873.340.26067.0079.010.39665.450.33260.9168.590.43770.89
Qwen3.5-4B4B73.450.27669.9678.020.41072.530.28168.8876.780.41263.470.38061.2767.170.47769.82
Qwen3.5-2B2B66.240.34862.8470.700.47365.220.35058.3072.360.47755.920.44050.9960.790.52162.46
Qwen3.5-0.8B0.8B60.770.37654.3965.540.50059.280.38654.2262.220.51047.930.49844.6048.980.55755.99
General VLMs: Qwen3-VL
Qwen3-VL-8B8B72.440.26165.1078.350.41172.030.26664.8877.820.40962.730.34255.5566.810.44869.07
Qwen3-VL-4B4B72.040.26265.1077.170.41870.840.27263.5476.130.42559.610.37855.1561.470.48067.50
Qwen3-VL-2B2B66.370.30059.0470.030.43965.810.31460.2568.520.44854.090.42851.0553.990.51162.09
General VLMs: Other
Kimi K2.61T/32B72.320.30366.9380.300.46669.950.32264.6977.310.47568.020.33562.4475.140.48170.10
Step3-VL10B53.650.49653.4157.160.50952.740.51653.6256.150.52945.060.57945.4247.660.57350.48
MiniCPM-V-4.58B51.810.43945.9753.360.48149.380.46142.7951.500.48937.590.58332.0139.060.55246.26
Gemini-3.1-Pro---70.040.30665.6375.080.40969.280.32265.8174.240.41771.980.30068.6277.260.38670.43

* OvisOCR2 scores are author-reported post-publication results from its official model card. Only per-track Overall and Avg3 were published; unreported component metrics are shown as —.

NaviDC-OCR scores are author-reported post-publication results from its official repository. The source reports TextEdit, FormulaCDM, and TableTEDS for each track, but not ROEdit; unreported ROEdit metrics are shown as —.

Bold marks the best reported score in each column; underlined marks the runner-up. GitHub README tables cannot run sorting scripts, so use the interactive leaderboard to sort any metric in either direction.

Diagnostics

The diagnostic panel shows where current systems still have headroom. Formula recognition is the largest single bottleneck, and real degradation changes rankings more sharply than digital degradation.

Diagnostic panels

Case Studies

The four case studies below are all taken from the paper. They show failures that aggregate scores can hide: notation loss, reading-order mistakes, annotation contamination, table-structure errors, character-level corruption, and missing visual authentication cues.

Case 1: Academic

Case study 1: academic structured lab report

Case 2: Business

Case study 2: business product specification table

Case 3: Finance

Case study 3: finance actuarial valuation report

Case 4: Certificate

Case study 4: Chinese product quality certificate

Appendix Highlights

The appendix documents the degradation design, per-category behavior, and source-validity checks used to make the benchmark reproducible.

Degradation operations

Degradation scenarios

Per-category overview

Source-validity dashboard

Download

The full image/GT/HTML release is hosted on Hugging Face:

# After downloading all files from Hugging Face:
shasum -a 256 -c SHA256SUMS.txt
cat pdb_full.tar.part-* | tar -xf -

Verify the split archive and reconstructed release:

python scripts/verify_split_archive.py /path/to/downloaded/files

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv

GT Review

Stable public identifiers: puredocbench-gt-latest for the GT archive and puredocbench-review-v1 for this review UI. Exact update timestamps remain in the Hugging Face metadata and changelog rather than public URLs or directory names. Use the review app to inspect annotations and export correction patches.

Local launch:

mkdir -p review/assets
ln -s /path/to/puredocbench/images/clean review/assets/images
python3 -m http.server 8767 --directory review

Open:

http://127.0.0.1:8767/index.html

Static app URL:

https://zhihengli-casia.github.io/PureDocBench/review/

The GitHub repository does not include the full image release. For visual review on GitHub Pages, click Load Images and select the downloaded images/clean folder. Local launch can also use the symlink above.

GT Coordinates

If you need spatial labels, regenerate clean-render coordinates from the HTML/CSS sources:

python scripts/add_gt_coordinates.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --in-place \
  --include-bbox \
  --include-coordinate-system \
  --report coordinate_report.json

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --require-coordinates \
  --require-bbox

The script follows the OmniDocBench GT convention and adds a rectangular poly field to each layout_dets item. poly is a flat list of clean-image pixel coordinates in top-left, top-right, bottom-right, bottom-left order: [x1, y1, x2, y1, x2, y2, x1, y2]. A derived bbox: [x1, y1, x2, y2] can also be written with --include-bbox, but poly is the primary coordinate field. Run playwright install chromium first if the Playwright browser is not installed, or pass --browser-channel chrome to use a local Chrome installation.

Inference And Scoring

PureDocBench includes a public CLI for model-agnostic inference, lightweight scoring, and OmniDocBench export:

pip install -e .

puredocbench infer \
  --images /path/to/puredocbench/images/clean \
  --output-dir predictions/my_model_clean \
  --command-template 'python my_model_infer.py --image {image} --out {output}'

puredocbench score \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --pred-dir predictions/my_model_clean \
  --track clean \
  --out-dir scores/my_model_clean

See docs/INFERENCE_SCORING.md for the full interface and OmniDocBench export path.

Repository Contents

manifests/                         Release and sample manifests
metadata/                          Dataset card and Croissant metadata
scripts/                           Rendering, degradation, validation, leaderboard tools
puredocbench/                      Public inference, scoring, and OmniDocBench export CLI
model_inference/                   Sanitized model inference configs and runners
supplemental_inference_scoring/    API/local inference and scoring utilities
assets/figures/                    Figures from the paper
paper/                             Paper PDF

Quick Start

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium

Render one HTML page:

python scripts/render_single_image.py \
  --html /path/to/page.html \
  --out /path/to/page.png \
  --dpi 300

Apply a deterministic degradation profile:

python scripts/apply_degradation_ablation.py \
  --input /path/to/clean_images \
  --output /path/to/degraded_images \
  --profile full_medium

License

  • Dataset assets are released under CC BY 4.0; see LICENSE_DATA.
  • Code in this repository is released under the license in LICENSE.
  • Model weights are not redistributed.

Citation

@misc{puredocbench,
  title        = {How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings},
  author       = {Li, Zhiheng and collaborators},
  year         = {2026},
  howpublished = {\url{https://github.com/zhihengli-casia/puredocbench}},
  note         = {Dataset and benchmark release}
}

Contributors

zhihengli-casia

27 commits

zhihengli-casia/puredocbench

PureDocBench: source-traceable benchmark for document parsing across clean, degraded, and real-world settings

Python

45

27 commits

updated Aug 29, 2026

See the code
benchmark
document-ai
document-parsing
ocr
synthetic-data

README

PureDocBench

How far is document parsing from solved?
A source-traceable benchmark for OCR and document parsing across clean, digitally degraded, and real-degraded document settings.

Hugging Face Dataset Data License Code License Paper

中文说明 | Dataset | Paper | GT Review & Corrections

PureDocBench uses HTML/CSS document sources as hidden anchors: each page is rendered into images and annotated from the same structured source. This gives a benchmark where text, tables, formulas, captions, and reading order can be scored with less post-hoc annotation noise.

PureDocBench 是一个源可追踪的 OCR / 文档解析 benchmark。数据由 HTML/CSS 源文件渲染而来,GT 标注从同源结构中抽取,覆盖 clean、digital-degraded、real-degraded 三条图像轨道。

Updates

  • Current GT: The stable alias is puredocbench-gt-latest; the exact revision and timestamp are recorded in Hugging Face gt/latest.json.
  • 2026-06-14: Updated GT annotations and opened a GT Review app for community corrections.
  • 2026-05-08: Initial public release of PureDocBench, including the paper PDF and full dataset on Hugging Face.

GT Annotation Examples

The examples below show colored coordinate boxes over clean rendered pages from an academic paper, a patent form, and a tuition invoice.

PureDocBench GT coordinate annotation examples

PureDocBench overview

At A Glance

ItemCount
Official pages1,475
Official images4,425
Top-level domains10
Fine-grained subcategories66
Image tracksclean, digital-degraded, real-degraded
Scored structurestext, formulas, tables, reading order

Main Leaderboard

The paper evaluates 40 systems across pipeline specialists, end-to-end document parsers, and general-purpose VLMs. The live leaderboard additionally tracks public post-publication results. Each track reports Overall, TextEdit, FormulaCDM, TableTEDS, and ROEdit; Avg3 averages the three track Overall scores.

Open the interactive sortable leaderboard ↗

Three-track leaderboard on PureDocBench
ModelParamsCleanDigital DegradedReal DegradedAvg3
Overall↑TextEditFormulaCDMTableTEDSROEditOverall↑TextEditFormulaCDMTableTEDSROEditOverall↑TextEditFormulaCDMTableTEDSROEdit
Pipeline / Multi-stage Specialists
NaviDC-OCR1.2B86.900.11181.0191.0977.470.20672.5980.4570.850.30265.1177.6678.41
DotsMOCR3B76.270.15166.2377.650.27373.160.19864.3274.950.30961.730.31254.3961.970.39370.39
Dolphin-v23B65.900.34259.8072.120.42960.240.39352.2067.860.46144.920.55339.9850.040.55857.02
MonkeyOCR-pro-3B3B62.230.34648.4672.830.49257.400.39745.5766.320.52646.490.51138.1852.430.60055.37
YouTu-Parsing2B75.020.23067.3480.740.35869.660.27061.4474.490.38860.290.36052.2064.690.43068.32
MinerU2.5-Pro1.2B75.870.22265.1484.680.34671.770.27261.7980.730.37862.560.37552.7072.470.44670.07
MinerU2.51.2B74.900.18462.0881.040.32768.920.24556.9974.240.37459.150.37049.0165.410.44667.66
MonkeyOCR-pro-1.2B1.2B61.090.35847.4371.600.49855.720.41643.9164.830.52943.820.55636.9450.070.60953.54
PaddleOCR-VL-1.50.9B73.010.26663.5382.120.42866.730.33958.0376.070.47860.500.39854.0067.330.51066.75
GLM-OCR0.9B68.650.31457.8979.440.47063.060.38353.2374.210.52058.310.43350.3467.830.54363.34
OpenOCR0.1B32.700.35433.500.000.50730.030.41031.090.000.54125.730.48625.810.000.59129.49
End-to-End Specialists
OvisOCR2*0.8B81.5577.0966.5675.06
olmOCR-2-7B7B69.360.28456.8979.590.35865.870.31854.5774.810.37856.100.41748.7961.250.43963.78
olmOCR-7B7B62.560.38858.6967.770.46657.840.43655.4461.660.49947.300.54246.2649.800.56855.90
FD-RL4B78.380.19368.2186.220.33476.330.21467.1683.220.35067.040.29858.8272.080.39173.92
Logics-Parsing-v24B76.350.21367.6782.670.34273.850.24867.3379.020.37567.640.30461.6571.640.41672.61
OCRVerse4B73.180.27363.7883.090.39371.360.30263.9580.360.41563.660.36357.0370.300.45269.40
Qianfan-OCR4B57.220.37049.7958.830.44350.850.43844.4151.960.48545.060.49439.0845.530.50951.04
Nanonets-OCR23B64.830.25444.9874.940.37761.230.30745.4068.970.40849.030.43535.5055.090.46858.36
DeepSeek-OCR-23B55.530.35446.0056.010.46649.410.41240.7848.670.49343.600.48637.3042.060.53349.51
OCRFlux-3B3B47.140.45438.3548.460.42441.820.48631.9042.170.43737.210.55932.6534.870.49142.06
DeepSeek-OCR3B53.500.41945.3957.060.51446.950.47839.9948.640.54840.480.53734.0441.120.57546.98
dots.ocr2.9B72.010.24861.3779.510.37965.950.30756.6771.860.41755.680.40347.7059.630.46764.55
FireRed-OCR2B70.810.28763.8677.230.39668.490.31962.6474.770.42257.420.41551.6062.160.47465.57
HunyuanOCR1B65.610.26955.7468.020.38261.490.30851.6263.680.40054.580.42148.3057.540.45960.56
UniRec-0.1B0.1B58.910.42251.3167.600.52652.420.50148.3759.040.57834.440.65830.9738.160.68548.59
OpenDoc-0.1B0.1B60.280.41153.0968.860.51952.460.50148.4159.040.57744.270.54738.4649.060.60352.00
General VLMs: Qwen3.5
Qwen3.5-397B-A17B397B/17B69.120.23365.2665.400.36668.340.24463.9165.530.37662.700.28760.7056.120.39966.72
Qwen3.5-122B-A10B122B/10B76.140.22667.9683.030.37576.340.22067.8283.210.36669.850.28162.1975.440.40174.11
Qwen3.5-35B-A3B35B/3B68.400.23264.9463.450.37468.040.24564.7863.860.37960.590.31059.6853.070.41965.68
Qwen3.5-27B27B72.070.22766.3672.510.36270.730.23664.6171.170.36765.920.28361.2364.820.39069.57
Qwen3.5-9B9B73.870.25467.6079.390.38873.340.26067.0079.010.39665.450.33260.9168.590.43770.89
Qwen3.5-4B4B73.450.27669.9678.020.41072.530.28168.8876.780.41263.470.38061.2767.170.47769.82
Qwen3.5-2B2B66.240.34862.8470.700.47365.220.35058.3072.360.47755.920.44050.9960.790.52162.46
Qwen3.5-0.8B0.8B60.770.37654.3965.540.50059.280.38654.2262.220.51047.930.49844.6048.980.55755.99
General VLMs: Qwen3-VL
Qwen3-VL-8B8B72.440.26165.1078.350.41172.030.26664.8877.820.40962.730.34255.5566.810.44869.07
Qwen3-VL-4B4B72.040.26265.1077.170.41870.840.27263.5476.130.42559.610.37855.1561.470.48067.50
Qwen3-VL-2B2B66.370.30059.0470.030.43965.810.31460.2568.520.44854.090.42851.0553.990.51162.09
General VLMs: Other
Kimi K2.61T/32B72.320.30366.9380.300.46669.950.32264.6977.310.47568.020.33562.4475.140.48170.10
Step3-VL10B53.650.49653.4157.160.50952.740.51653.6256.150.52945.060.57945.4247.660.57350.48
MiniCPM-V-4.58B51.810.43945.9753.360.48149.380.46142.7951.500.48937.590.58332.0139.060.55246.26
Gemini-3.1-Pro---70.040.30665.6375.080.40969.280.32265.8174.240.41771.980.30068.6277.260.38670.43

* OvisOCR2 scores are author-reported post-publication results from its official model card. Only per-track Overall and Avg3 were published; unreported component metrics are shown as —.

NaviDC-OCR scores are author-reported post-publication results from its official repository. The source reports TextEdit, FormulaCDM, and TableTEDS for each track, but not ROEdit; unreported ROEdit metrics are shown as —.

Bold marks the best reported score in each column; underlined marks the runner-up. GitHub README tables cannot run sorting scripts, so use the interactive leaderboard to sort any metric in either direction.

Diagnostics

The diagnostic panel shows where current systems still have headroom. Formula recognition is the largest single bottleneck, and real degradation changes rankings more sharply than digital degradation.

Diagnostic panels

Case Studies

The four case studies below are all taken from the paper. They show failures that aggregate scores can hide: notation loss, reading-order mistakes, annotation contamination, table-structure errors, character-level corruption, and missing visual authentication cues.

Case 1: Academic

Case study 1: academic structured lab report

Case 2: Business

Case study 2: business product specification table

Case 3: Finance

Case study 3: finance actuarial valuation report

Case 4: Certificate

Case study 4: Chinese product quality certificate

Appendix Highlights

The appendix documents the degradation design, per-category behavior, and source-validity checks used to make the benchmark reproducible.

Degradation operations

Degradation scenarios

Per-category overview

Source-validity dashboard

Download

The full image/GT/HTML release is hosted on Hugging Face:

# After downloading all files from Hugging Face:
shasum -a 256 -c SHA256SUMS.txt
cat pdb_full.tar.part-* | tar -xf -

Verify the split archive and reconstructed release:

python scripts/verify_split_archive.py /path/to/downloaded/files

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv

GT Review

Stable public identifiers: puredocbench-gt-latest for the GT archive and puredocbench-review-v1 for this review UI. Exact update timestamps remain in the Hugging Face metadata and changelog rather than public URLs or directory names. Use the review app to inspect annotations and export correction patches.

Local launch:

mkdir -p review/assets
ln -s /path/to/puredocbench/images/clean review/assets/images
python3 -m http.server 8767 --directory review

Open:

http://127.0.0.1:8767/index.html

Static app URL:

https://zhihengli-casia.github.io/PureDocBench/review/

The GitHub repository does not include the full image release. For visual review on GitHub Pages, click Load Images and select the downloaded images/clean folder. Local launch can also use the symlink above.

GT Coordinates

If you need spatial labels, regenerate clean-render coordinates from the HTML/CSS sources:

python scripts/add_gt_coordinates.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --in-place \
  --include-bbox \
  --include-coordinate-system \
  --report coordinate_report.json

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --require-coordinates \
  --require-bbox

The script follows the OmniDocBench GT convention and adds a rectangular poly field to each layout_dets item. poly is a flat list of clean-image pixel coordinates in top-left, top-right, bottom-right, bottom-left order: [x1, y1, x2, y1, x2, y2, x1, y2]. A derived bbox: [x1, y1, x2, y2] can also be written with --include-bbox, but poly is the primary coordinate field. Run playwright install chromium first if the Playwright browser is not installed, or pass --browser-channel chrome to use a local Chrome installation.

Inference And Scoring

PureDocBench includes a public CLI for model-agnostic inference, lightweight scoring, and OmniDocBench export:

pip install -e .

puredocbench infer \
  --images /path/to/puredocbench/images/clean \
  --output-dir predictions/my_model_clean \
  --command-template 'python my_model_infer.py --image {image} --out {output}'

puredocbench score \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --pred-dir predictions/my_model_clean \
  --track clean \
  --out-dir scores/my_model_clean

See docs/INFERENCE_SCORING.md for the full interface and OmniDocBench export path.

Repository Contents

manifests/                         Release and sample manifests
metadata/                          Dataset card and Croissant metadata
scripts/                           Rendering, degradation, validation, leaderboard tools
puredocbench/                      Public inference, scoring, and OmniDocBench export CLI
model_inference/                   Sanitized model inference configs and runners
supplemental_inference_scoring/    API/local inference and scoring utilities
assets/figures/                    Figures from the paper
paper/                             Paper PDF

Quick Start

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium

Render one HTML page:

python scripts/render_single_image.py \
  --html /path/to/page.html \
  --out /path/to/page.png \
  --dpi 300

Apply a deterministic degradation profile:

python scripts/apply_degradation_ablation.py \
  --input /path/to/clean_images \
  --output /path/to/degraded_images \
  --profile full_medium

License

  • Dataset assets are released under CC BY 4.0; see LICENSE_DATA.
  • Code in this repository is released under the license in LICENSE.
  • Model weights are not redistributed.

Citation

@misc{puredocbench,
  title        = {How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings},
  author       = {Li, Zhiheng and collaborators},
  year         = {2026},
  howpublished = {\url{https://github.com/zhihengli-casia/puredocbench}},
  note         = {Dataset and benchmark release}
}

Contributors

zhihengli-casia

27 commits

Languages

Python

81.5%

Shell

7.8%

JavaScript

4.8%

HTML

4.3%

CSS

1.6%