zhihengli-casia/puredocbench

Dataset

Main Leaderboard

3

17 commits

1 linked in READMEs

updated Sep 17, 2026

See the code
benchmark
document
document-parsing
image
multimodal

README

Main Leaderboard

42 models · 3 tracks · 🏆 Search & sort the leaderboard →

RankModelModel typeParamsAvg₃ ↑Clean ↑Digital ↑Real ↑
1TeleOCRPipeline / Multi-stage1.2B78.4186.9077.4770.85
2OvisOCR2End-to-End0.8B75.0681.5577.0966.56
3Qwen3.5-122B-A10BGeneral VLM122B/10B74.1176.1476.3469.85
4FD-RLEnd-to-End4B73.9278.3876.3367.04
5Logics-Parsing-v2End-to-End4B72.6176.3573.8567.64
6Qwen3.5-9BGeneral VLM9B70.8973.8773.3465.45
7Gemini-3.1-ProGeneral VLM70.4370.0469.2871.98
8DotsMOCRPipeline / Multi-stage3B70.3976.2773.1661.73
9Kimi K2.6General VLM1T/32B70.1072.3269.9568.02
10MinerU2.5-ProPipeline / Multi-stage1.2B70.0775.8771.7762.56

Top 10 by Avg₃, highest first. Bold = best; underlined = runner-up. ‡ = author-reported result.

CSV · JSON · Evaluation protocol · Submit results

Show all 42 models
RankModelModel typeParamsAvg₃ ↑Clean ↑Digital ↑Real ↑
1TeleOCRPipeline / Multi-stage1.2B78.4186.9077.4770.85
2OvisOCR2End-to-End0.8B75.0681.5577.0966.56
3Qwen3.5-122B-A10BGeneral VLM122B/10B74.1176.1476.3469.85
4FD-RLEnd-to-End4B73.9278.3876.3367.04
5Logics-Parsing-v2End-to-End4B72.6176.3573.8567.64
6Qwen3.5-9BGeneral VLM9B70.8973.8773.3465.45
7Gemini-3.1-ProGeneral VLM70.4370.0469.2871.98
8DotsMOCRPipeline / Multi-stage3B70.3976.2773.1661.73
9Kimi K2.6General VLM1T/32B70.1072.3269.9568.02
10MinerU2.5-ProPipeline / Multi-stage1.2B70.0775.8771.7762.56
11Qwen3.5-4BGeneral VLM4B69.8273.4572.5363.47
12Qwen3.5-27BGeneral VLM27B69.5772.0770.7365.92
13OCRVerseEnd-to-End4B69.4073.1871.3663.66
14Qwen3-VL-8BGeneral VLM8B69.0772.4472.0362.73
15YouTu-ParsingPipeline / Multi-stage2B68.3275.0269.6660.29
16MinerU2.5Pipeline / Multi-stage1.2B67.6674.9068.9259.15
17Qwen3-VL-4BGeneral VLM4B67.5072.0470.8459.61
18PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B66.7573.0166.7360.50
19Qwen3.5-397B-A17BGeneral VLM397B/17B66.7269.1268.3462.70
20Qwen3.5-35B-A3BGeneral VLM35B/3B65.6868.4068.0460.59
21FireRed-OCREnd-to-End2B65.5770.8168.4957.42
22dots.ocrEnd-to-End2.9B64.5572.0165.9555.68
23olmOCR-2-7BEnd-to-End7B63.7869.3665.8756.10
24GLM-OCRPipeline / Multi-stage0.9B63.3468.6563.0658.31
25Qwen3.5-2BGeneral VLM2B62.4666.2465.2255.92
26Qwen3-VL-2BGeneral VLM2B62.0966.3765.8154.09
27HunyuanOCREnd-to-End1B60.5665.6161.4954.58
28Nanonets-OCR2End-to-End3B58.3664.8361.2349.03
29Dolphin-v2Pipeline / Multi-stage3B57.0265.9060.2444.92
30Qwen3.5-0.8BGeneral VLM0.8B55.9960.7759.2847.93
31olmOCR-7BEnd-to-End7B55.9062.5657.8447.30
32MonkeyOCR-pro-3BPipeline / Multi-stage3B55.3762.2357.4046.49
33MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B53.5461.0955.7243.82
34OpenDoc-0.1B †End-to-End0.1B52.0060.2852.4644.27
35Qianfan-OCREnd-to-End4B51.0457.2250.8545.06
36Step3-VLGeneral VLM10B50.4853.6552.7445.06
37DeepSeek-OCR-2End-to-End3B49.5155.5349.4143.60
38UniRec-0.1BEnd-to-End0.1B48.5958.9152.4234.44
39DeepSeek-OCREnd-to-End3B46.9853.5046.9540.48
40MiniCPM-V-4.5General VLM8B46.2651.8149.3837.59
41OCRFlux-3BEnd-to-End3B42.0647.1441.8237.21
42OpenOCRPipeline / Multi-stage0.1B29.4932.7030.0325.73
Clean: full metrics for all 42 models
RankModelModel typeParamsOverall ↑TextEdit ↓FormulaCDM ↑TableTEDS ↑ROEdit ↓
1TeleOCRPipeline / Multi-stage1.2B86.900.11181.0191.09
2OvisOCR2End-to-End0.8B81.55
3FD-RLEnd-to-End4B78.380.19368.2186.220.334
4Logics-Parsing-v2End-to-End4B76.350.21367.6782.670.342
5DotsMOCRPipeline / Multi-stage3B76.270.15166.2377.650.273
6Qwen3.5-122B-A10BGeneral VLM122B/10B76.140.22667.9683.030.375
7MinerU2.5-ProPipeline / Multi-stage1.2B75.870.22265.1484.680.346
8YouTu-ParsingPipeline / Multi-stage2B75.020.23067.3480.740.358
9MinerU2.5Pipeline / Multi-stage1.2B74.900.18462.0881.040.327
10Qwen3.5-9BGeneral VLM9B73.870.25467.6079.390.388
11Qwen3.5-4BGeneral VLM4B73.450.27669.9678.020.410
12OCRVerseEnd-to-End4B73.180.27363.7883.090.393
13PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B73.010.26663.5382.120.428
14Qwen3-VL-8BGeneral VLM8B72.440.26165.1078.350.411
15Kimi K2.6General VLM1T/32B72.320.30366.9380.300.466
16Qwen3.5-27BGeneral VLM27B72.070.22766.3672.510.362
17Qwen3-VL-4BGeneral VLM4B72.040.26265.1077.170.418
18dots.ocrEnd-to-End2.9B72.010.24861.3779.510.379
19FireRed-OCREnd-to-End2B70.810.28763.8677.230.396
20Gemini-3.1-ProGeneral VLM70.040.30665.6375.080.409
21olmOCR-2-7BEnd-to-End7B69.360.28456.8979.590.358
22Qwen3.5-397B-A17BGeneral VLM397B/17B69.120.23365.2665.400.366
23GLM-OCRPipeline / Multi-stage0.9B68.650.31457.8979.440.470
24Qwen3.5-35B-A3BGeneral VLM35B/3B68.400.23264.9463.450.374
25Qwen3-VL-2BGeneral VLM2B66.370.30059.0470.030.439
26Qwen3.5-2BGeneral VLM2B66.240.34862.8470.700.473
27Dolphin-v2Pipeline / Multi-stage3B65.900.34259.8072.120.429
28HunyuanOCREnd-to-End1B65.610.26955.7468.020.382
29Nanonets-OCR2End-to-End3B64.830.25444.9874.940.377
30olmOCR-7BEnd-to-End7B62.560.38858.6967.770.466
31MonkeyOCR-pro-3BPipeline / Multi-stage3B62.230.34648.4672.830.492
32MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B61.090.35847.4371.600.498
33Qwen3.5-0.8BGeneral VLM0.8B60.770.37654.3965.540.500
34OpenDoc-0.1B †End-to-End0.1B60.280.41153.0968.860.519
35UniRec-0.1BEnd-to-End0.1B58.910.42251.3167.600.526
36Qianfan-OCREnd-to-End4B57.220.37049.7958.830.443
37DeepSeek-OCR-2End-to-End3B55.530.35446.0056.010.466
38Step3-VLGeneral VLM10B53.650.49653.4157.160.509
39DeepSeek-OCREnd-to-End3B53.500.41945.3957.060.514
40MiniCPM-V-4.5General VLM8B51.810.43945.9753.360.481
41OCRFlux-3BEnd-to-End3B47.140.45438.3548.460.424
42OpenOCRPipeline / Multi-stage0.1B32.700.35433.500.000.507
Digital-degraded: full metrics for all 42 models
RankModelModel typeParamsOverall ↑TextEdit ↓FormulaCDM ↑TableTEDS ↑ROEdit ↓
1TeleOCRPipeline / Multi-stage1.2B77.470.20672.5980.45
2OvisOCR2End-to-End0.8B77.09
3Qwen3.5-122B-A10BGeneral VLM122B/10B76.340.22067.8283.210.366
4FD-RLEnd-to-End4B76.330.21467.1683.220.350
5Logics-Parsing-v2End-to-End4B73.850.24867.3379.020.375
6Qwen3.5-9BGeneral VLM9B73.340.26067.0079.010.396
7DotsMOCRPipeline / Multi-stage3B73.160.19864.3274.950.309
8Qwen3.5-4BGeneral VLM4B72.530.28168.8876.780.412
9Qwen3-VL-8BGeneral VLM8B72.030.26664.8877.820.409
10MinerU2.5-ProPipeline / Multi-stage1.2B71.770.27261.7980.730.378
11OCRVerseEnd-to-End4B71.360.30263.9580.360.415
12Qwen3-VL-4BGeneral VLM4B70.840.27263.5476.130.425
13Qwen3.5-27BGeneral VLM27B70.730.23664.6171.170.367
14Kimi K2.6General VLM1T/32B69.950.32264.6977.310.475
15YouTu-ParsingPipeline / Multi-stage2B69.660.27061.4474.490.388
16Gemini-3.1-ProGeneral VLM69.280.32265.8174.240.417
17MinerU2.5Pipeline / Multi-stage1.2B68.920.24556.9974.240.374
18FireRed-OCREnd-to-End2B68.490.31962.6474.770.422
19Qwen3.5-397B-A17BGeneral VLM397B/17B68.340.24463.9165.530.376
20Qwen3.5-35B-A3BGeneral VLM35B/3B68.040.24564.7863.860.379
21PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B66.730.33958.0376.070.478
22dots.ocrEnd-to-End2.9B65.950.30756.6771.860.417
23olmOCR-2-7BEnd-to-End7B65.870.31854.5774.810.378
24Qwen3-VL-2BGeneral VLM2B65.810.31460.2568.520.448
25Qwen3.5-2BGeneral VLM2B65.220.35058.3072.360.477
26GLM-OCRPipeline / Multi-stage0.9B63.060.38353.2374.210.520
27HunyuanOCREnd-to-End1B61.490.30851.6263.680.400
28Nanonets-OCR2End-to-End3B61.230.30745.4068.970.408
29Dolphin-v2Pipeline / Multi-stage3B60.240.39352.2067.860.461
30Qwen3.5-0.8BGeneral VLM0.8B59.280.38654.2262.220.510
31olmOCR-7BEnd-to-End7B57.840.43655.4461.660.499
32MonkeyOCR-pro-3BPipeline / Multi-stage3B57.400.39745.5766.320.526
33MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B55.720.41643.9164.830.529
34Step3-VLGeneral VLM10B52.740.51653.6256.150.529
35OpenDoc-0.1B †End-to-End0.1B52.460.50148.4159.040.577
36UniRec-0.1BEnd-to-End0.1B52.420.50148.3759.040.578
37Qianfan-OCREnd-to-End4B50.850.43844.4151.960.485
38DeepSeek-OCR-2End-to-End3B49.410.41240.7848.670.493
39MiniCPM-V-4.5General VLM8B49.380.46142.7951.500.489
40DeepSeek-OCREnd-to-End3B46.950.47839.9948.640.548
41OCRFlux-3BEnd-to-End3B41.820.48631.9042.170.437
42OpenOCRPipeline / Multi-stage0.1B30.030.41031.090.000.541
Real-degraded: full metrics for all 42 models
RankModelModel typeParamsOverall ↑TextEdit ↓FormulaCDM ↑TableTEDS ↑ROEdit ↓
1Gemini-3.1-ProGeneral VLM71.980.30068.6277.260.386
2TeleOCRPipeline / Multi-stage1.2B70.850.30265.1177.66
3Qwen3.5-122B-A10BGeneral VLM122B/10B69.850.28162.1975.440.401
4Kimi K2.6General VLM1T/32B68.020.33562.4475.140.481
5Logics-Parsing-v2End-to-End4B67.640.30461.6571.640.416
6FD-RLEnd-to-End4B67.040.29858.8272.080.391
7OvisOCR2End-to-End0.8B66.56
8Qwen3.5-27BGeneral VLM27B65.920.28361.2364.820.390
9Qwen3.5-9BGeneral VLM9B65.450.33260.9168.590.437
10OCRVerseEnd-to-End4B63.660.36357.0370.300.452
11Qwen3.5-4BGeneral VLM4B63.470.38061.2767.170.477
12Qwen3-VL-8BGeneral VLM8B62.730.34255.5566.810.448
13Qwen3.5-397B-A17BGeneral VLM397B/17B62.700.28760.7056.120.399
14MinerU2.5-ProPipeline / Multi-stage1.2B62.560.37552.7072.470.446
15DotsMOCRPipeline / Multi-stage3B61.730.31254.3961.970.393
16Qwen3.5-35B-A3BGeneral VLM35B/3B60.590.31059.6853.070.419
17PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B60.500.39854.0067.330.510
18YouTu-ParsingPipeline / Multi-stage2B60.290.36052.2064.690.430
19Qwen3-VL-4BGeneral VLM4B59.610.37855.1561.470.480
20MinerU2.5Pipeline / Multi-stage1.2B59.150.37049.0165.410.446
21GLM-OCRPipeline / Multi-stage0.9B58.310.43350.3467.830.543
22FireRed-OCREnd-to-End2B57.420.41551.6062.160.474
23olmOCR-2-7BEnd-to-End7B56.100.41748.7961.250.439
24Qwen3.5-2BGeneral VLM2B55.920.44050.9960.790.521
25dots.ocrEnd-to-End2.9B55.680.40347.7059.630.467
26HunyuanOCREnd-to-End1B54.580.42148.3057.540.459
27Qwen3-VL-2BGeneral VLM2B54.090.42851.0553.990.511
28Nanonets-OCR2End-to-End3B49.030.43535.5055.090.468
29Qwen3.5-0.8BGeneral VLM0.8B47.930.49844.6048.980.557
30olmOCR-7BEnd-to-End7B47.300.54246.2649.800.568
31MonkeyOCR-pro-3BPipeline / Multi-stage3B46.490.51138.1852.430.600
32Qianfan-OCREnd-to-End4B45.060.49439.0845.530.509
33Step3-VLGeneral VLM10B45.060.57945.4247.660.573
34Dolphin-v2Pipeline / Multi-stage3B44.920.55339.9850.040.558
35OpenDoc-0.1B †End-to-End0.1B44.270.54738.4649.060.603
36MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B43.820.55636.9450.070.609
37DeepSeek-OCR-2End-to-End3B43.600.48637.3042.060.533
38DeepSeek-OCREnd-to-End3B40.480.53734.0441.120.575
39MiniCPM-V-4.5General VLM8B37.590.58332.0139.060.552
40OCRFlux-3BEnd-to-End3B37.210.55932.6534.870.491
41UniRec-0.1BEnd-to-End0.1B34.440.65830.9738.160.685
42OpenOCRPipeline / Multi-stage0.1B25.730.48625.810.000.591

‡ Author-reported results added after the paper: TeleOCR, OvisOCR2. TeleOCR was previously named NaviDC-OCR. TeleOCR's Avg₃ (78.41) is the mean of its reported track scores. Unreported component metrics are shown as .

Source and evaluation notes document the added results, rounding, and evaluation disclosures.

† OpenDoc-0.1B retains the published Avg3 of 52.00. The mean of its displayed track scores is 52.34; this transcription preserves the paper value pending an erratum.

Original paper Table 2 (40 baselines)

Original paper Table 2

Evaluation protocol

  • Overall ↑ = (100 × (1 − TextEdit) + FormulaCDM + TableTEDS) / 3.
  • Avg₃ ↑ is the arithmetic mean of the three track Overall scores, with the published values preserved above.
  • TextEdit ↓ and ROEdit ↓ use the 0–1 scale; FormulaCDM ↑, TableTEDS ↑, Overall ↑, and Avg₃ ↑ use the 0–100 scale. Reading order is reported separately from Overall.
  • The benchmark has 1,475 pages per track. The 40 paper baselines retain their original scores and protocol. TeleOCR and OvisOCR2 are transcribed from the linked author reports; their exact GT revisions are not specified in those public score tables. The current GT download has continued to receive annotation updates. New submissions should identify their GT revision and page manifest.
  • For paper-compatible scoring, follow the OmniDocBench export and evaluation instructions. The lightweight public scorer is for development checks and uses a different scoring implementation.

Source: paper Table 2 and versioned published table.

Submit results

Open a discussion or a dataset pull request with the following information:

  1. Model name, exact checkpoint or API version, parameter count, and model/code link.
  2. GT revision, page manifest, and results for Clean, Digital-degraded, and Real-degraded, including all five metrics per track.
  3. Inference settings, evaluator version/configuration, and prediction or evaluation-log links where available.
  4. A clear distinction between author-reported results and results reproduced by the benchmark maintainers.

Maintainers review submissions before adding them to the published leaderboard. Evaluation revisions and result sources are recorded explicitly.

PureDocBench

How far is document parsing from solved?
A source-traceable benchmark for OCR and document parsing across clean, digitally degraded, and real-degraded document settings.

Hugging Face Dataset Data License Code License Paper

中文说明 | Leaderboard | Dataset | Paper | GT Review & Corrections

Benchmark Overview

PureDocBench uses HTML/CSS document sources as hidden anchors: each page is rendered into images and annotated from the same structured source. This gives a benchmark where text, tables, formulas, captions, and reading order can be scored with less post-hoc annotation noise.

PureDocBench 是一个源可追踪的 OCR / 文档解析 benchmark。数据由 HTML/CSS 源文件渲染而来,GT 标注从同源结构中抽取,覆盖 clean、digital-degraded、real-degraded 三条图像轨道。

Updates

  • 2026-09-16: Published the Hugging Face leaderboard with 40 paper baselines and author-reported TeleOCR and OvisOCR2 results, plus CSV/JSON downloads.

  • Current GT: The stable alias is puredocbench-gt-latest; download gt/puredocbench_gt_latest.tar.gz and check gt/latest.json for the exact revision and timestamp.

  • 2026-06-14: Updated GT annotations and opened the GT Review app for community corrections.

  • 2026-05-08: Initial public release of PureDocBench, including the paper PDF and full dataset on Hugging Face.

GT Annotation Examples

The examples below show colored coordinate boxes over clean rendered pages from an academic paper, a patent form, and a tuition invoice.

PureDocBench GT coordinate annotation examples

PureDocBench overview

At A Glance

ItemCount
Official pages1,475
Official images4,425
Top-level domains10
Fine-grained subcategories66
Image tracksclean, digital-degraded, real-degraded
Scored structurestext, formulas, tables, reading order

Diagnostics

The diagnostic panel shows where current systems still have headroom. Formula recognition is the largest single bottleneck, and real degradation changes rankings more sharply than digital degradation.

Diagnostic panels

Case Studies

The four case studies below are all taken from the paper. They show failures that aggregate scores can hide: notation loss, reading-order mistakes, annotation contamination, table-structure errors, character-level corruption, and missing visual authentication cues.

Case 1: Academic

Case study 1: academic structured lab report

Case 2: Business

Case study 2: business product specification table

Case 3: Finance

Case study 3: finance actuarial valuation report

Case 4: Certificate

Case study 4: Chinese product quality certificate

Appendix Highlights

The appendix documents the degradation design, per-category behavior, and source-validity checks used to make the benchmark reproducible.

Degradation operations

Degradation scenarios

Per-category overview

Source-validity dashboard

Download

The full image/GT/HTML release is hosted on Hugging Face:

# After downloading all files from Hugging Face:
shasum -a 256 -c SHA256SUMS.txt
cat pdb_full.tar.part-* | tar -xf -

Latest GT-only annotations are available separately:

  • gt/puredocbench_gt_latest.tar.gz
  • gt/puredocbench_gt_latest.sha256
  • gt/latest.json
  • gt/changed_cases.json

If you already downloaded an older GT archive, replace it with gt/puredocbench_gt_latest.tar.gz.

Verify the split archive and reconstructed release:

python scripts/verify_split_archive.py /path/to/downloaded/files

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv

GT Review

Stable GT alias: puredocbench-gt-latest (gt/puredocbench_gt_latest.tar.gz). The exact revision is recorded in gt/latest.json; dates are kept as provenance metadata rather than embedded in public URLs or directory names. Use the review app to inspect annotations and export correction patches.

Local launch:

mkdir -p review/assets
ln -s /path/to/puredocbench/images/clean review/assets/images
python3 -m http.server 8767 --directory review

Open:

http://127.0.0.1:8767/index.html

Static app URL:

https://zhihengli-casia.github.io/PureDocBench/review/

The GitHub repository does not include the full image release. For visual review on GitHub Pages, click Load Images and select the downloaded images/clean folder. Local launch can also use the symlink above.

GT Coordinates

If you need spatial labels, regenerate clean-render coordinates from the HTML/CSS sources:

python scripts/add_gt_coordinates.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --in-place \
  --include-bbox \
  --include-coordinate-system \
  --report coordinate_report.json

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --require-coordinates \
  --require-bbox

The script follows the OmniDocBench GT convention and adds a rectangular poly field to each layout_dets item. poly is a flat list of clean-image pixel coordinates in top-left, top-right, bottom-right, bottom-left order: [x1, y1, x2, y1, x2, y2, x1, y2]. A derived bbox: [x1, y1, x2, y2] can also be written with --include-bbox, but poly is the primary coordinate field. Run playwright install chromium first if the Playwright browser is not installed, or pass --browser-channel chrome to use a local Chrome installation.

Inference And Scoring

PureDocBench includes a public CLI for model-agnostic inference, lightweight scoring, and OmniDocBench export:

pip install -e .

puredocbench infer \
  --images /path/to/puredocbench/images/clean \
  --output-dir predictions/my_model_clean \
  --command-template 'python my_model_infer.py --image {image} --out {output}'

puredocbench score \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --pred-dir predictions/my_model_clean \
  --track clean \
  --out-dir scores/my_model_clean

See docs/INFERENCE_SCORING.md for the full interface and OmniDocBench export path.

Repository Contents

manifests/                         Release and sample manifests
metadata/                          Dataset card and Croissant metadata
scripts/                           Rendering, degradation, validation, leaderboard tools
puredocbench/                      Public inference, scoring, and OmniDocBench export CLI
model_inference/                   Sanitized model inference configs and runners
supplemental_inference_scoring/    API/local inference and scoring utilities
assets/figures/                    Figures from the paper
paper/                             Paper PDF

Quick Start

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium

Render one HTML page:

python scripts/render_single_image.py \
  --html /path/to/page.html \
  --out /path/to/page.png \
  --dpi 300

Apply a deterministic degradation profile:

python scripts/apply_degradation_ablation.py \
  --input /path/to/clean_images \
  --output /path/to/degraded_images \
  --profile full_medium

License

  • Dataset assets are released under CC BY 4.0; see LICENSE_DATA.
  • Code in this repository is released under the license in LICENSE.
  • Model weights are not redistributed.

Citation

@misc{puredocbench,
  title        = {How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings},
  author       = {Li, Zhiheng and collaborators},
  year         = {2026},
  howpublished = {\url{https://github.com/zhihengli-casia/puredocbench}},
  note         = {Dataset and benchmark release}
}

Contributors

zhihengli-casia

13 commits

RO
rosetulip

2 commits

ZL
Zhiheng Li

1 commits

zhihengli-casia/puredocbench

Dataset

Main Leaderboard

3

17 commits

1 linked in READMEs

updated Sep 17, 2026

See the code
benchmark
document
document-parsing
image
multimodal

README

Main Leaderboard

42 models · 3 tracks · 🏆 Search & sort the leaderboard →

RankModelModel typeParamsAvg₃ ↑Clean ↑Digital ↑Real ↑
1TeleOCRPipeline / Multi-stage1.2B78.4186.9077.4770.85
2OvisOCR2End-to-End0.8B75.0681.5577.0966.56
3Qwen3.5-122B-A10BGeneral VLM122B/10B74.1176.1476.3469.85
4FD-RLEnd-to-End4B73.9278.3876.3367.04
5Logics-Parsing-v2End-to-End4B72.6176.3573.8567.64
6Qwen3.5-9BGeneral VLM9B70.8973.8773.3465.45
7Gemini-3.1-ProGeneral VLM70.4370.0469.2871.98
8DotsMOCRPipeline / Multi-stage3B70.3976.2773.1661.73
9Kimi K2.6General VLM1T/32B70.1072.3269.9568.02
10MinerU2.5-ProPipeline / Multi-stage1.2B70.0775.8771.7762.56

Top 10 by Avg₃, highest first. Bold = best; underlined = runner-up. ‡ = author-reported result.

CSV · JSON · Evaluation protocol · Submit results

Show all 42 models
RankModelModel typeParamsAvg₃ ↑Clean ↑Digital ↑Real ↑
1TeleOCRPipeline / Multi-stage1.2B78.4186.9077.4770.85
2OvisOCR2End-to-End0.8B75.0681.5577.0966.56
3Qwen3.5-122B-A10BGeneral VLM122B/10B74.1176.1476.3469.85
4FD-RLEnd-to-End4B73.9278.3876.3367.04
5Logics-Parsing-v2End-to-End4B72.6176.3573.8567.64
6Qwen3.5-9BGeneral VLM9B70.8973.8773.3465.45
7Gemini-3.1-ProGeneral VLM70.4370.0469.2871.98
8DotsMOCRPipeline / Multi-stage3B70.3976.2773.1661.73
9Kimi K2.6General VLM1T/32B70.1072.3269.9568.02
10MinerU2.5-ProPipeline / Multi-stage1.2B70.0775.8771.7762.56
11Qwen3.5-4BGeneral VLM4B69.8273.4572.5363.47
12Qwen3.5-27BGeneral VLM27B69.5772.0770.7365.92
13OCRVerseEnd-to-End4B69.4073.1871.3663.66
14Qwen3-VL-8BGeneral VLM8B69.0772.4472.0362.73
15YouTu-ParsingPipeline / Multi-stage2B68.3275.0269.6660.29
16MinerU2.5Pipeline / Multi-stage1.2B67.6674.9068.9259.15
17Qwen3-VL-4BGeneral VLM4B67.5072.0470.8459.61
18PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B66.7573.0166.7360.50
19Qwen3.5-397B-A17BGeneral VLM397B/17B66.7269.1268.3462.70
20Qwen3.5-35B-A3BGeneral VLM35B/3B65.6868.4068.0460.59
21FireRed-OCREnd-to-End2B65.5770.8168.4957.42
22dots.ocrEnd-to-End2.9B64.5572.0165.9555.68
23olmOCR-2-7BEnd-to-End7B63.7869.3665.8756.10
24GLM-OCRPipeline / Multi-stage0.9B63.3468.6563.0658.31
25Qwen3.5-2BGeneral VLM2B62.4666.2465.2255.92
26Qwen3-VL-2BGeneral VLM2B62.0966.3765.8154.09
27HunyuanOCREnd-to-End1B60.5665.6161.4954.58
28Nanonets-OCR2End-to-End3B58.3664.8361.2349.03
29Dolphin-v2Pipeline / Multi-stage3B57.0265.9060.2444.92
30Qwen3.5-0.8BGeneral VLM0.8B55.9960.7759.2847.93
31olmOCR-7BEnd-to-End7B55.9062.5657.8447.30
32MonkeyOCR-pro-3BPipeline / Multi-stage3B55.3762.2357.4046.49
33MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B53.5461.0955.7243.82
34OpenDoc-0.1B †End-to-End0.1B52.0060.2852.4644.27
35Qianfan-OCREnd-to-End4B51.0457.2250.8545.06
36Step3-VLGeneral VLM10B50.4853.6552.7445.06
37DeepSeek-OCR-2End-to-End3B49.5155.5349.4143.60
38UniRec-0.1BEnd-to-End0.1B48.5958.9152.4234.44
39DeepSeek-OCREnd-to-End3B46.9853.5046.9540.48
40MiniCPM-V-4.5General VLM8B46.2651.8149.3837.59
41OCRFlux-3BEnd-to-End3B42.0647.1441.8237.21
42OpenOCRPipeline / Multi-stage0.1B29.4932.7030.0325.73
Clean: full metrics for all 42 models
RankModelModel typeParamsOverall ↑TextEdit ↓FormulaCDM ↑TableTEDS ↑ROEdit ↓
1TeleOCRPipeline / Multi-stage1.2B86.900.11181.0191.09
2OvisOCR2End-to-End0.8B81.55
3FD-RLEnd-to-End4B78.380.19368.2186.220.334
4Logics-Parsing-v2End-to-End4B76.350.21367.6782.670.342
5DotsMOCRPipeline / Multi-stage3B76.270.15166.2377.650.273
6Qwen3.5-122B-A10BGeneral VLM122B/10B76.140.22667.9683.030.375
7MinerU2.5-ProPipeline / Multi-stage1.2B75.870.22265.1484.680.346
8YouTu-ParsingPipeline / Multi-stage2B75.020.23067.3480.740.358
9MinerU2.5Pipeline / Multi-stage1.2B74.900.18462.0881.040.327
10Qwen3.5-9BGeneral VLM9B73.870.25467.6079.390.388
11Qwen3.5-4BGeneral VLM4B73.450.27669.9678.020.410
12OCRVerseEnd-to-End4B73.180.27363.7883.090.393
13PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B73.010.26663.5382.120.428
14Qwen3-VL-8BGeneral VLM8B72.440.26165.1078.350.411
15Kimi K2.6General VLM1T/32B72.320.30366.9380.300.466
16Qwen3.5-27BGeneral VLM27B72.070.22766.3672.510.362
17Qwen3-VL-4BGeneral VLM4B72.040.26265.1077.170.418
18dots.ocrEnd-to-End2.9B72.010.24861.3779.510.379
19FireRed-OCREnd-to-End2B70.810.28763.8677.230.396
20Gemini-3.1-ProGeneral VLM70.040.30665.6375.080.409
21olmOCR-2-7BEnd-to-End7B69.360.28456.8979.590.358
22Qwen3.5-397B-A17BGeneral VLM397B/17B69.120.23365.2665.400.366
23GLM-OCRPipeline / Multi-stage0.9B68.650.31457.8979.440.470
24Qwen3.5-35B-A3BGeneral VLM35B/3B68.400.23264.9463.450.374
25Qwen3-VL-2BGeneral VLM2B66.370.30059.0470.030.439
26Qwen3.5-2BGeneral VLM2B66.240.34862.8470.700.473
27Dolphin-v2Pipeline / Multi-stage3B65.900.34259.8072.120.429
28HunyuanOCREnd-to-End1B65.610.26955.7468.020.382
29Nanonets-OCR2End-to-End3B64.830.25444.9874.940.377
30olmOCR-7BEnd-to-End7B62.560.38858.6967.770.466
31MonkeyOCR-pro-3BPipeline / Multi-stage3B62.230.34648.4672.830.492
32MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B61.090.35847.4371.600.498
33Qwen3.5-0.8BGeneral VLM0.8B60.770.37654.3965.540.500
34OpenDoc-0.1B †End-to-End0.1B60.280.41153.0968.860.519
35UniRec-0.1BEnd-to-End0.1B58.910.42251.3167.600.526
36Qianfan-OCREnd-to-End4B57.220.37049.7958.830.443
37DeepSeek-OCR-2End-to-End3B55.530.35446.0056.010.466
38Step3-VLGeneral VLM10B53.650.49653.4157.160.509
39DeepSeek-OCREnd-to-End3B53.500.41945.3957.060.514
40MiniCPM-V-4.5General VLM8B51.810.43945.9753.360.481
41OCRFlux-3BEnd-to-End3B47.140.45438.3548.460.424
42OpenOCRPipeline / Multi-stage0.1B32.700.35433.500.000.507
Digital-degraded: full metrics for all 42 models
RankModelModel typeParamsOverall ↑TextEdit ↓FormulaCDM ↑TableTEDS ↑ROEdit ↓
1TeleOCRPipeline / Multi-stage1.2B77.470.20672.5980.45
2OvisOCR2End-to-End0.8B77.09
3Qwen3.5-122B-A10BGeneral VLM122B/10B76.340.22067.8283.210.366
4FD-RLEnd-to-End4B76.330.21467.1683.220.350
5Logics-Parsing-v2End-to-End4B73.850.24867.3379.020.375
6Qwen3.5-9BGeneral VLM9B73.340.26067.0079.010.396
7DotsMOCRPipeline / Multi-stage3B73.160.19864.3274.950.309
8Qwen3.5-4BGeneral VLM4B72.530.28168.8876.780.412
9Qwen3-VL-8BGeneral VLM8B72.030.26664.8877.820.409
10MinerU2.5-ProPipeline / Multi-stage1.2B71.770.27261.7980.730.378
11OCRVerseEnd-to-End4B71.360.30263.9580.360.415
12Qwen3-VL-4BGeneral VLM4B70.840.27263.5476.130.425
13Qwen3.5-27BGeneral VLM27B70.730.23664.6171.170.367
14Kimi K2.6General VLM1T/32B69.950.32264.6977.310.475
15YouTu-ParsingPipeline / Multi-stage2B69.660.27061.4474.490.388
16Gemini-3.1-ProGeneral VLM69.280.32265.8174.240.417
17MinerU2.5Pipeline / Multi-stage1.2B68.920.24556.9974.240.374
18FireRed-OCREnd-to-End2B68.490.31962.6474.770.422
19Qwen3.5-397B-A17BGeneral VLM397B/17B68.340.24463.9165.530.376
20Qwen3.5-35B-A3BGeneral VLM35B/3B68.040.24564.7863.860.379
21PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B66.730.33958.0376.070.478
22dots.ocrEnd-to-End2.9B65.950.30756.6771.860.417
23olmOCR-2-7BEnd-to-End7B65.870.31854.5774.810.378
24Qwen3-VL-2BGeneral VLM2B65.810.31460.2568.520.448
25Qwen3.5-2BGeneral VLM2B65.220.35058.3072.360.477
26GLM-OCRPipeline / Multi-stage0.9B63.060.38353.2374.210.520
27HunyuanOCREnd-to-End1B61.490.30851.6263.680.400
28Nanonets-OCR2End-to-End3B61.230.30745.4068.970.408
29Dolphin-v2Pipeline / Multi-stage3B60.240.39352.2067.860.461
30Qwen3.5-0.8BGeneral VLM0.8B59.280.38654.2262.220.510
31olmOCR-7BEnd-to-End7B57.840.43655.4461.660.499
32MonkeyOCR-pro-3BPipeline / Multi-stage3B57.400.39745.5766.320.526
33MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B55.720.41643.9164.830.529
34Step3-VLGeneral VLM10B52.740.51653.6256.150.529
35OpenDoc-0.1B †End-to-End0.1B52.460.50148.4159.040.577
36UniRec-0.1BEnd-to-End0.1B52.420.50148.3759.040.578
37Qianfan-OCREnd-to-End4B50.850.43844.4151.960.485
38DeepSeek-OCR-2End-to-End3B49.410.41240.7848.670.493
39MiniCPM-V-4.5General VLM8B49.380.46142.7951.500.489
40DeepSeek-OCREnd-to-End3B46.950.47839.9948.640.548
41OCRFlux-3BEnd-to-End3B41.820.48631.9042.170.437
42OpenOCRPipeline / Multi-stage0.1B30.030.41031.090.000.541
Real-degraded: full metrics for all 42 models
RankModelModel typeParamsOverall ↑TextEdit ↓FormulaCDM ↑TableTEDS ↑ROEdit ↓
1Gemini-3.1-ProGeneral VLM71.980.30068.6277.260.386
2TeleOCRPipeline / Multi-stage1.2B70.850.30265.1177.66
3Qwen3.5-122B-A10BGeneral VLM122B/10B69.850.28162.1975.440.401
4Kimi K2.6General VLM1T/32B68.020.33562.4475.140.481
5Logics-Parsing-v2End-to-End4B67.640.30461.6571.640.416
6FD-RLEnd-to-End4B67.040.29858.8272.080.391
7OvisOCR2End-to-End0.8B66.56
8Qwen3.5-27BGeneral VLM27B65.920.28361.2364.820.390
9Qwen3.5-9BGeneral VLM9B65.450.33260.9168.590.437
10OCRVerseEnd-to-End4B63.660.36357.0370.300.452
11Qwen3.5-4BGeneral VLM4B63.470.38061.2767.170.477
12Qwen3-VL-8BGeneral VLM8B62.730.34255.5566.810.448
13Qwen3.5-397B-A17BGeneral VLM397B/17B62.700.28760.7056.120.399
14MinerU2.5-ProPipeline / Multi-stage1.2B62.560.37552.7072.470.446
15DotsMOCRPipeline / Multi-stage3B61.730.31254.3961.970.393
16Qwen3.5-35B-A3BGeneral VLM35B/3B60.590.31059.6853.070.419
17PaddleOCR-VL-1.5Pipeline / Multi-stage0.9B60.500.39854.0067.330.510
18YouTu-ParsingPipeline / Multi-stage2B60.290.36052.2064.690.430
19Qwen3-VL-4BGeneral VLM4B59.610.37855.1561.470.480
20MinerU2.5Pipeline / Multi-stage1.2B59.150.37049.0165.410.446
21GLM-OCRPipeline / Multi-stage0.9B58.310.43350.3467.830.543
22FireRed-OCREnd-to-End2B57.420.41551.6062.160.474
23olmOCR-2-7BEnd-to-End7B56.100.41748.7961.250.439
24Qwen3.5-2BGeneral VLM2B55.920.44050.9960.790.521
25dots.ocrEnd-to-End2.9B55.680.40347.7059.630.467
26HunyuanOCREnd-to-End1B54.580.42148.3057.540.459
27Qwen3-VL-2BGeneral VLM2B54.090.42851.0553.990.511
28Nanonets-OCR2End-to-End3B49.030.43535.5055.090.468
29Qwen3.5-0.8BGeneral VLM0.8B47.930.49844.6048.980.557
30olmOCR-7BEnd-to-End7B47.300.54246.2649.800.568
31MonkeyOCR-pro-3BPipeline / Multi-stage3B46.490.51138.1852.430.600
32Qianfan-OCREnd-to-End4B45.060.49439.0845.530.509
33Step3-VLGeneral VLM10B45.060.57945.4247.660.573
34Dolphin-v2Pipeline / Multi-stage3B44.920.55339.9850.040.558
35OpenDoc-0.1B †End-to-End0.1B44.270.54738.4649.060.603
36MonkeyOCR-pro-1.2BPipeline / Multi-stage1.2B43.820.55636.9450.070.609
37DeepSeek-OCR-2End-to-End3B43.600.48637.3042.060.533
38DeepSeek-OCREnd-to-End3B40.480.53734.0441.120.575
39MiniCPM-V-4.5General VLM8B37.590.58332.0139.060.552
40OCRFlux-3BEnd-to-End3B37.210.55932.6534.870.491
41UniRec-0.1BEnd-to-End0.1B34.440.65830.9738.160.685
42OpenOCRPipeline / Multi-stage0.1B25.730.48625.810.000.591

‡ Author-reported results added after the paper: TeleOCR, OvisOCR2. TeleOCR was previously named NaviDC-OCR. TeleOCR's Avg₃ (78.41) is the mean of its reported track scores. Unreported component metrics are shown as .

Source and evaluation notes document the added results, rounding, and evaluation disclosures.

† OpenDoc-0.1B retains the published Avg3 of 52.00. The mean of its displayed track scores is 52.34; this transcription preserves the paper value pending an erratum.

Original paper Table 2 (40 baselines)

Original paper Table 2

Evaluation protocol

  • Overall ↑ = (100 × (1 − TextEdit) + FormulaCDM + TableTEDS) / 3.
  • Avg₃ ↑ is the arithmetic mean of the three track Overall scores, with the published values preserved above.
  • TextEdit ↓ and ROEdit ↓ use the 0–1 scale; FormulaCDM ↑, TableTEDS ↑, Overall ↑, and Avg₃ ↑ use the 0–100 scale. Reading order is reported separately from Overall.
  • The benchmark has 1,475 pages per track. The 40 paper baselines retain their original scores and protocol. TeleOCR and OvisOCR2 are transcribed from the linked author reports; their exact GT revisions are not specified in those public score tables. The current GT download has continued to receive annotation updates. New submissions should identify their GT revision and page manifest.
  • For paper-compatible scoring, follow the OmniDocBench export and evaluation instructions. The lightweight public scorer is for development checks and uses a different scoring implementation.

Source: paper Table 2 and versioned published table.

Submit results

Open a discussion or a dataset pull request with the following information:

  1. Model name, exact checkpoint or API version, parameter count, and model/code link.
  2. GT revision, page manifest, and results for Clean, Digital-degraded, and Real-degraded, including all five metrics per track.
  3. Inference settings, evaluator version/configuration, and prediction or evaluation-log links where available.
  4. A clear distinction between author-reported results and results reproduced by the benchmark maintainers.

Maintainers review submissions before adding them to the published leaderboard. Evaluation revisions and result sources are recorded explicitly.

PureDocBench

How far is document parsing from solved?
A source-traceable benchmark for OCR and document parsing across clean, digitally degraded, and real-degraded document settings.

Hugging Face Dataset Data License Code License Paper

中文说明 | Leaderboard | Dataset | Paper | GT Review & Corrections

Benchmark Overview

PureDocBench uses HTML/CSS document sources as hidden anchors: each page is rendered into images and annotated from the same structured source. This gives a benchmark where text, tables, formulas, captions, and reading order can be scored with less post-hoc annotation noise.

PureDocBench 是一个源可追踪的 OCR / 文档解析 benchmark。数据由 HTML/CSS 源文件渲染而来,GT 标注从同源结构中抽取,覆盖 clean、digital-degraded、real-degraded 三条图像轨道。

Updates

  • 2026-09-16: Published the Hugging Face leaderboard with 40 paper baselines and author-reported TeleOCR and OvisOCR2 results, plus CSV/JSON downloads.

  • Current GT: The stable alias is puredocbench-gt-latest; download gt/puredocbench_gt_latest.tar.gz and check gt/latest.json for the exact revision and timestamp.

  • 2026-06-14: Updated GT annotations and opened the GT Review app for community corrections.

  • 2026-05-08: Initial public release of PureDocBench, including the paper PDF and full dataset on Hugging Face.

GT Annotation Examples

The examples below show colored coordinate boxes over clean rendered pages from an academic paper, a patent form, and a tuition invoice.

PureDocBench GT coordinate annotation examples

PureDocBench overview

At A Glance

ItemCount
Official pages1,475
Official images4,425
Top-level domains10
Fine-grained subcategories66
Image tracksclean, digital-degraded, real-degraded
Scored structurestext, formulas, tables, reading order

Diagnostics

The diagnostic panel shows where current systems still have headroom. Formula recognition is the largest single bottleneck, and real degradation changes rankings more sharply than digital degradation.

Diagnostic panels

Case Studies

The four case studies below are all taken from the paper. They show failures that aggregate scores can hide: notation loss, reading-order mistakes, annotation contamination, table-structure errors, character-level corruption, and missing visual authentication cues.

Case 1: Academic

Case study 1: academic structured lab report

Case 2: Business

Case study 2: business product specification table

Case 3: Finance

Case study 3: finance actuarial valuation report

Case 4: Certificate

Case study 4: Chinese product quality certificate

Appendix Highlights

The appendix documents the degradation design, per-category behavior, and source-validity checks used to make the benchmark reproducible.

Degradation operations

Degradation scenarios

Per-category overview

Source-validity dashboard

Download

The full image/GT/HTML release is hosted on Hugging Face:

# After downloading all files from Hugging Face:
shasum -a 256 -c SHA256SUMS.txt
cat pdb_full.tar.part-* | tar -xf -

Latest GT-only annotations are available separately:

  • gt/puredocbench_gt_latest.tar.gz
  • gt/puredocbench_gt_latest.sha256
  • gt/latest.json
  • gt/changed_cases.json

If you already downloaded an older GT archive, replace it with gt/puredocbench_gt_latest.tar.gz.

Verify the split archive and reconstructed release:

python scripts/verify_split_archive.py /path/to/downloaded/files

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv

GT Review

Stable GT alias: puredocbench-gt-latest (gt/puredocbench_gt_latest.tar.gz). The exact revision is recorded in gt/latest.json; dates are kept as provenance metadata rather than embedded in public URLs or directory names. Use the review app to inspect annotations and export correction patches.

Local launch:

mkdir -p review/assets
ln -s /path/to/puredocbench/images/clean review/assets/images
python3 -m http.server 8767 --directory review

Open:

http://127.0.0.1:8767/index.html

Static app URL:

https://zhihengli-casia.github.io/PureDocBench/review/

The GitHub repository does not include the full image release. For visual review on GitHub Pages, click Load Images and select the downloaded images/clean folder. Local launch can also use the symlink above.

GT Coordinates

If you need spatial labels, regenerate clean-render coordinates from the HTML/CSS sources:

python scripts/add_gt_coordinates.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --in-place \
  --include-bbox \
  --include-coordinate-system \
  --report coordinate_report.json

python scripts/validate_release_manifest.py \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --require-coordinates \
  --require-bbox

The script follows the OmniDocBench GT convention and adds a rectangular poly field to each layout_dets item. poly is a flat list of clean-image pixel coordinates in top-left, top-right, bottom-right, bottom-left order: [x1, y1, x2, y1, x2, y2, x1, y2]. A derived bbox: [x1, y1, x2, y2] can also be written with --include-bbox, but poly is the primary coordinate field. Run playwright install chromium first if the Playwright browser is not installed, or pass --browser-channel chrome to use a local Chrome installation.

Inference And Scoring

PureDocBench includes a public CLI for model-agnostic inference, lightweight scoring, and OmniDocBench export:

pip install -e .

puredocbench infer \
  --images /path/to/puredocbench/images/clean \
  --output-dir predictions/my_model_clean \
  --command-template 'python my_model_infer.py --image {image} --out {output}'

puredocbench score \
  --release-root /path/to/puredocbench \
  --manifest manifests/release_manifest_candidate_1475.csv \
  --pred-dir predictions/my_model_clean \
  --track clean \
  --out-dir scores/my_model_clean

See docs/INFERENCE_SCORING.md for the full interface and OmniDocBench export path.

Repository Contents

manifests/                         Release and sample manifests
metadata/                          Dataset card and Croissant metadata
scripts/                           Rendering, degradation, validation, leaderboard tools
puredocbench/                      Public inference, scoring, and OmniDocBench export CLI
model_inference/                   Sanitized model inference configs and runners
supplemental_inference_scoring/    API/local inference and scoring utilities
assets/figures/                    Figures from the paper
paper/                             Paper PDF

Quick Start

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium

Render one HTML page:

python scripts/render_single_image.py \
  --html /path/to/page.html \
  --out /path/to/page.png \
  --dpi 300

Apply a deterministic degradation profile:

python scripts/apply_degradation_ablation.py \
  --input /path/to/clean_images \
  --output /path/to/degraded_images \
  --profile full_medium

License

  • Dataset assets are released under CC BY 4.0; see LICENSE_DATA.
  • Code in this repository is released under the license in LICENSE.
  • Model weights are not redistributed.

Citation

@misc{puredocbench,
  title        = {How Far Is Document Parsing from Solved? PureDocBench: A Source-Traceable Benchmark across Clean, Degraded, and Real-World Settings},
  author       = {Li, Zhiheng and collaborators},
  year         = {2026},
  howpublished = {\url{https://github.com/zhihengli-casia/puredocbench}},
  note         = {Dataset and benchmark release}
}

Contributors

zhihengli-casia

13 commits

RO
rosetulip

2 commits

ZL
Zhiheng Li

1 commits