PaddlePaddle/Real5-OmniDocBench

Dataset

37

stars

31

commits

2

linked in READMEs

Aug 27, 2026

updated

benchmark
document
document-parsing
image
multimodal
ocr

README

Real5-OmniDocBench

A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

ECCV 2026 arXiv Dataset Base Benchmark License

Leaderboard | Overview | Dataset | Evaluation | Submit Results | Citation

Real5-OmniDocBench measures the robustness of document parsing systems under five physical acquisition conditions: Scanning, Warping, Screen-Photography, Illumination, and Skew. It reconstructs the same 1,355 pages from OmniDocBench v1.5 in every condition, producing 6,775 images in total. The one-to-one page correspondence and shared evaluation protocol isolate the effect of acquisition conditions from changes in document content.

Original pages and corresponding Real5-OmniDocBench reconstructions

Original pages and their corresponding reconstructions under five physical acquisition conditions.

News

  • 2026-08-27: Added evaluation results for NaviDC-OCR.
  • 2026-08-08: Added evaluation results for Kimi-K2.5、Kimi-K2.6、Doubao-Seed-2.1-Pro、MonkeyOCRv2-S-Parsing、MonkeyOCRv2-B-Parsing and OvisOCR2.
  • 2026-06-18: Real5-OmniDocBench was accepted to ECCV 2026. 🎉
Earlier updates
  • 2026-05-28: Added evaluation results for PaddleOCR-VL-1.6 and MinerU2.5-Pro.
  • 2026-03-05: Released the paper and added results for DeepSeek-OCR 2 and GLM-OCR.
  • 2026-01-28: Released the dataset and benchmark.

Leaderboard

The leaderboard reports performance over all five acquisition conditions. All metrics follow OmniDocBench: Overall↑, TextEdit↓, FormulaCDM↑, TableTEDS↑, and Reading OrderEdit↓.

Higher is better for ↑ metrics and lower is better for ↓ metrics. Best results in each column are shown in bold, and second-best results are underlined. Tables are sorted by Overall score in descending order.

1. Overall

MethodsModel TypeParametersOverall↑Scanning↑Warping↑Screen-Photography↑Illumination↑Skew↑
PaddleOCR-VL-1.6Specialized VLMs0.9B93.1994.7492.4892.7893.2892.66
OvisOCR2Specialized VLMs0.9B92.2993.7791.4093.0992.8890.33
PaddleOCR-VL-1.5Specialized VLMs0.9B92.0593.4391.2591.7692.1691.66
NaviDC-OCRSpecialized VLMs1.2B90.7292.3390.6389.0990.9090.63
GLM-OCRSpecialized VLMs0.9B90.3292.6790.6891.7591.1285.39
Kimi-K2.6General VLMs1.1T89.7690.0889.6289.5889.9189.61
Gemini-3 ProGeneral VLMs-89.2489.4788.9088.8689.5389.45
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.2289.4989.7088.4088.5489.97
Kimi-K2.5General VLMs1.1T89.0989.6788.8688.3989.6688.86
Doubao-Seed-2.1-ProGeneral VLMs-89.0288.8589.3688.9989.1388.79
MinerU2.5-proSpecialized VLMs1.2B88.9492.1188.7291.2991.3181.26
Qwen3-VL-235BGeneral VLMs235B88.9089.4389.9989.2789.2786.56
Gemini-2.5 ProGeneral VLMs-88.2189.2587.6387.1187.9789.07
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B87.9088.8788.1787.6486.7588.09
Qwen2.5-VL-72BGeneral VLMs72B86.9286.1987.7786.4887.2586.90
dots.ocrSpecialized VLMs3B86.3886.8786.0187.1887.5784.27
MinerU2.5Specialized VLMs1.2B85.6190.0683.7689.4189.5775.24
PaddleOCR-VLSpecialized VLMs0.9B85.5492.1185.9782.5489.6177.47
Nanonets-OCR-sSpecialized VLMs3B84.1985.5283.5684.8685.0181.98
MonkeyOCR-pro-3BSpecialized VLMs3.7B79.4986.9478.9082.4484.7164.47
GPT-5.2General VLMs-78.6684.4376.2676.7580.8875.00
MonkeyOCR-3BSpecialized VLMs3.7B78.2984.6577.2780.7183.1665.67
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B77.1584.6476.5980.2482.1162.18
MinerU2-VLMSpecialized VLMs0.9B76.9583.6073.7378.7780.5168.16
Deepseek-OCRSpecialized VLMs3B73.9986.1767.2075.3178.1063.01
Deepseek-OCR 2Specialized VLMs3B73.0189.5966.5371.6576.0261.28
PP-StructureV3Pipeline Tools-64.4584.6859.3466.8973.3837.98
DolphinSpecialized VLMs322M61.7872.1660.3564.2967.2944.83
Dolphin-1.5Specialized VLMs0.3B61.4883.3950.5069.7675.6128.16
Marker-1.8.2Pipeline Tools-60.1070.2758.9863.6566.3141.27

Overall is the mean score across all five scenarios.

View per-scenario leaderboards

2. Scanning

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B94.740.03593.6594.120.042
OvisOCR2Specialized VLMs0.9B93.770.04192.1393.300.039
PaddleOCR-VL-1.5Specialized VLMs0.9B93.430.03793.0490.970.045
GLM-OCRSpecialized VLMs0.9B92.670.05491.1092.280.061
NaviDC-OCRSpecialized VLMs1.2B92.330.04588.2693.260.039
PaddleOCR-VLSpecialized VLMs0.9B92.110.03990.3589.900.048
MinerU2.5-proSpecialized VLMs1.2B92.110.04089.7790.570.043
Kimi-K2.6General VLMs1.1T90.080.06289.0285.770.072
MinerU2.5Specialized VLMs1.2B90.060.05288.2287.160.050
Kimi-K2.5General VLMs1.1T89.670.06088.6786.310.079
Deepseek-OCR 2Specialized VLMs3B89.590.05588.5585.720.056
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.490.05386.2987.530.051
Gemini-3 ProGeneral VLMs-89.470.07188.1687.370.078
Qwen3-VL-235BGeneral VLMs235B89.430.05989.0185.190.066
Gemini-2.5 ProGeneral VLMs-89.250.07387.4487.620.098
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B88.870.05886.3286.090.053
Doubao-Seed-2.1-ProGeneral VLMs-88.850.08486.5688.330.093
MonkeyOCR-pro-3BSpecialized VLMs3.7B86.940.10386.2984.860.141
dots.ocrSpecialized VLMs3B86.870.08383.2785.680.081
Qwen2.5-VL-72BGeneral VLMs72B86.190.11086.1483.410.114
Deepseek-OCRSpecialized VLMs3B86.170.07883.5982.690.085
Nanonets-OCR-sSpecialized VLMs3B85.520.10688.0979.110.106
PP-StructureV3Pipeline Tools-84.680.09484.3479.060.092
MonkeyOCR-3BSpecialized VLMs3.7B84.650.10084.1679.810.143
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B84.640.12384.1782.130.145
GPT-5.2General VLMs-84.430.14285.6881.780.109
MinerU2-VLMSpecialized VLMs0.9B83.600.09479.7680.440.091
Dolphin-1.5Specialized VLMs0.3B83.390.09776.2583.650.090
DolphinSpecialized VLMs322M72.160.15464.5867.270.130
Marker-1.8.2Pipeline Tools-70.270.22377.0356.050.238

3. Warping

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B92.480.04891.6390.660.061
OvisOCR2Specialized VLMs0.9B91.400.05890.9489.080.058
PaddleOCR-VL-1.5Specialized VLMs0.9B91.250.05390.9488.100.063
GLM-OCRSpecialized VLMs0.9B90.680.07190.3088.780.100
NaviDC-OCRSpecialized VLMs1.2B90.630.06289.1488.930.051
Qwen3-VL-235BGeneral VLMs235B89.990.05189.0685.950.064
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.700.05887.5087.390.052
Kimi-K2.6General VLMs1.1T89.620.07391.1885.020.079
Doubao-Seed-2.1-ProGeneral VLMs-89.360.08787.7489.030.095
Gemini-3 ProGeneral VLMs-88.900.08688.1087.200.087
Kimi-K2.5General VLMs1.1T88.860.06989.7783.710.084
MinerU2.5-proSpecialized VLMs1.2B88.720.10087.8188.360.076
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B88.170.06986.2085.170.056
Qwen2.5-VL-72BGeneral VLMs72B87.770.08688.8583.060.102
Gemini-2.5 ProGeneral VLMs-87.630.09286.5085.590.109
dots.ocrSpecialized VLMs3B86.010.08785.0381.740.093
PaddleOCR-VLSpecialized VLMs0.9B85.970.09385.4581.770.092
MinerU2.5Specialized VLMs1.2B83.760.15485.9280.710.104
Nanonets-OCR-sSpecialized VLMs3B83.560.12186.2476.570.124
MonkeyOCR-pro-3BSpecialized VLMs3.7B78.900.16879.5573.940.212
MonkeyOCR-3BSpecialized VLMs3.7B77.270.16479.0869.180.211
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B76.590.19678.8570.520.221
GPT-5.2General VLMs-76.260.23980.9071.800.165
MinerU2-VLMSpecialized VLMs0.9B73.730.20277.7263.650.173
Deepseek-OCRSpecialized VLMs3B67.200.32873.5960.800.226
Deepseek-OCR 2Specialized VLMs3B66.530.29370.4258.440.209
DolphinSpecialized VLMs322M60.350.31661.0651.580.247
PP-StructureV3Pipeline Tools-59.340.37668.2247.400.261
Marker-1.8.2Pipeline Tools-58.980.34972.7139.080.390
Dolphin-1.5Specialized VLMs0.3B50.500.38347.2442.520.309

4. Screen-Photography

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
OvisOCR2Specialized VLMs0.9B93.090.05491.8092.830.046
PaddleOCR-VL-1.6Specialized VLMs0.9B92.780.04590.6492.190.054
PaddleOCR-VL-1.5Specialized VLMs0.9B91.760.05090.8889.380.059
GLM-OCRSpecialized VLMs0.9B91.750.06389.8391.660.070
MinerU2.5-proSpecialized VLMs1.2B91.290.05087.4191.440.044
Kimi-K2.6General VLMs1.1T89.580.07390.1085.490.077
MinerU2.5Specialized VLMs1.2B89.410.06287.5586.830.053
Qwen3-VL-235BGeneral VLMs235B89.270.06888.7285.850.071
NaviDC-OCRSpecialized VLMs1.2B89.090.07885.8089.300.057
Doubao-Seed-2.1-ProGeneral VLMs-88.990.09887.4789.310.102
Gemini-3 ProGeneral VLMs-88.860.08487.3387.650.087
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B88.400.07087.0785.120.050
Kimi-K2.5General VLMs1.1T88.390.07087.4184.770.078
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B87.640.06884.9684.720.049
dots.ocrSpecialized VLMs3B87.180.08185.3484.260.079
Gemini-2.5 ProGeneral VLMs-87.110.10385.3086.310.117
Qwen2.5-VL-72BGeneral VLMs72B86.480.10087.4682.000.102
Nanonets-OCR-sSpecialized VLMs3B84.860.11286.6579.090.117
PaddleOCR-VLSpecialized VLMs0.9B82.540.10383.5874.360.107
MonkeyOCR-pro-3BSpecialized VLMs3.7B82.440.12481.5578.130.177
MonkeyOCR-3BSpecialized VLMs3.7B80.710.12281.3373.040.177
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B80.240.14880.7874.740.179
MinerU2-VLMSpecialized VLMs0.9B78.770.13979.0271.170.123
GPT-5.2General VLMs-76.750.20879.2771.730.148
Deepseek-OCRSpecialized VLMs3B75.310.22077.6870.260.169
Deepseek-OCR 2Specialized VLMs3B71.650.20173.4961.540.157
Dolphin-1.5Specialized VLMs0.3B69.760.20561.8068.000.177
PP-StructureV3Pipeline Tools-66.890.20473.2647.820.165
DolphinSpecialized VLMs322M64.290.23258.6657.380.195
Marker-1.8.2Pipeline Tools-63.650.29072.7347.210.325

5. Illumination

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B93.280.04292.6191.480.051
OvisOCR2Specialized VLMs0.9B92.880.04391.1891.720.040
PaddleOCR-VL-1.5Specialized VLMs0.9B92.160.04691.8089.330.051
MinerU2.5-proSpecialized VLMs1.2B91.310.05088.1590.760.052
GLM-OCRSpecialized VLMs0.9B91.120.05991.0288.200.071
NaviDC-OCRSpecialized VLMs1.2B90.900.05587.2790.910.045
Kimi-K2.6General VLMs1.1T89.910.06289.7686.200.072
Kimi-K2.5General VLMs1.1T89.660.06489.8385.530.077
PaddleOCR-VLSpecialized VLMs0.9B89.610.04986.6687.020.055
MinerU2.5Specialized VLMs1.2B89.570.06588.3686.870.062
Gemini-3 ProGeneral VLMs-89.530.07387.7888.140.080
Qwen3-VL-235BGeneral VLMs235B89.270.06087.8186.050.070
Doubao-Seed-2.1-ProGeneral VLMs-89.130.08587.2988.590.091
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B88.540.06686.7285.530.059
Gemini-2.5 ProGeneral VLMs-87.970.08386.1386.110.103
dots.ocrSpecialized VLMs3B87.570.06885.0784.440.076
Qwen2.5-VL-72BGeneral VLMs72B87.250.08786.4484.030.097
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B86.750.07484.1683.490.057
Nanonets-OCR-sSpecialized VLMs3B85.010.09987.9476.960.112
MonkeyOCR-pro-3BSpecialized VLMs3.7B84.710.12084.1382.020.171
MonkeyOCR-3BSpecialized VLMs3.7B83.160.11883.6377.620.168
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B82.110.14482.0778.670.172
GPT-5.2General VLMs-80.880.19184.4177.370.134
MinerU2-VLMSpecialized VLMs0.9B80.510.13580.7274.290.123
Deepseek-OCRSpecialized VLMs3B78.100.19281.7171.810.156
Deepseek-OCR 2Specialized VLMs3B76.020.16877.8367.010.122
Dolphin-1.5Specialized VLMs0.3B75.610.15970.0472.690.133
PP-StructureV3Pipeline Tools-73.380.15877.7558.190.126
DolphinSpecialized VLMs322M67.290.19761.4260.100.173
Marker-1.8.2Pipeline Tools-66.310.25974.8050.030.337

6. Skew

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B92.660.04591.4491.040.058
PaddleOCR-VL-1.5Specialized VLMs0.9B91.660.04791.0088.690.061
NaviDC-OCRSpecialized VLMs1.2B90.630.05487.1590.100.042
OvisOCR2Specialized VLMs0.9B90.330.04890.5885.230.048
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.970.05188.1086.880.051
Kimi-K2.6General VLMs1.1T89.610.06090.2584.270.078
Gemini-3 ProGeneral VLMs-89.450.08088.3388.060.092
Gemini-2.5 ProGeneral VLMs-89.070.07787.8986.990.104
Kimi-K2.5General VLMs1.1T88.860.06089.6283.000.080
Doubao-Seed-2.1-ProGeneral VLMs-88.790.08088.2386.110.095
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B88.090.06384.5486.070.055
Qwen2.5-VL-72BGeneral VLMs72B86.900.07787.2681.140.091
Qwen3-VL-235BGeneral VLMs235B86.560.07783.9683.410.091
GLM-OCRSpecialized VLMs0.9B85.390.09985.7880.280.156
dots.ocrSpecialized VLMs3B84.270.08785.7375.740.094
Nanonets-OCR-sSpecialized VLMs3B81.980.12185.7872.220.133
MinerU2.5-proSpecialized VLMs1.2B81.260.20283.9281.070.107
PaddleOCR-VLSpecialized VLMs0.9B77.470.19278.8172.830.193
MinerU2.5Specialized VLMs1.2B75.240.30581.7874.390.151
GPT-5.2General VLMs-75.000.25780.2770.470.167
MinerU2-VLMSpecialized VLMs0.9B68.160.23074.4553.070.191
MonkeyOCR-3BSpecialized VLMs3.7B65.670.24869.2352.590.300
MonkeyOCR-pro-3BSpecialized VLMs3.7B64.470.25169.0649.420.301
Deepseek-OCRSpecialized VLMs3B63.010.32773.2748.480.231
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B62.180.29266.2549.460.317
Deepseek-OCR 2Specialized VLMs3B61.280.29566.1647.180.221
DolphinSpecialized VLMs322M44.830.50051.3433.220.321
Marker-1.8.2Pipeline Tools-41.270.53660.1617.230.543
PP-StructureV3Pipeline Tools-37.980.55744.3725.270.417
Dolphin-1.5Specialized VLMs0.3B28.160.55325.6014.180.419

Benchmark Overview

Real5-OmniDocBench evaluates the same document content under controlled changes to the physical acquisition process. Except for the scanning subset, images were manually captured with handheld mobile devices.

ScenarioAcquisition conditionRepresentative artifacts
ScanningDocuments captured with scanning devicesScanner characteristics and clean planar capture
WarpingCurved or non-planar pages photographed by handPage curvature, folding, and local deformation
Screen-PhotographyScreens displaying documents photographed by handMoiré patterns, reflections, and display artifacts
IlluminationDocuments photographed under varied lightingShadows, glare, and uneven exposure
SkewDocuments photographed from oblique viewpointsPerspective distortion and geometric skew

This design provides:

  • Controlled comparison: every scenario contains the same 1,355 source pages.
  • Physical realism: acquisition artifacts are produced by real devices and environments rather than synthetic transformations.
  • Protocol compatibility: page identities, annotations, prediction format, and metrics follow OmniDocBench v1.5.

Dataset

The five scenario directories each contain 1,355 images. The complete download is approximately 16 GB.

Real5-OmniDocBench/
├── Real5-OmniDocBench-Scanning/
├── Real5-OmniDocBench-Warping/
├── Real5-OmniDocBench-Screen-Photography/
├── Real5-OmniDocBench-Illumination/
└── Real5-OmniDocBench-Skew/

Install the current Hugging Face Hub CLI and download the complete dataset:

pip install -U huggingface_hub
hf download PaddlePaddle/Real5-OmniDocBench \
  --repo-type dataset \
  --local-dir ./Real5-OmniDocBench

To download a single condition, use a file pattern:

hf download PaddlePaddle/Real5-OmniDocBench \
  --repo-type dataset \
  --include "Real5-OmniDocBench-Warping/*" \
  --local-dir ./Real5-OmniDocBench

Download the matching OmniDocBench v1.5 ground-truth annotations separately:

hf download opendatalab/OmniDocBench OmniDocBench.json \
  --repo-type dataset \
  --revision v1_5 \
  --local-dir ./OmniDocBench-v1.5

Evaluation

Real5-OmniDocBench does not introduce a new prediction schema. It reuses the OmniDocBench v1.5 annotation format and evaluation pipeline so that performance can be compared across acquisition conditions.

  1. Run the model independently on all 1,355 images in each scenario.
  2. Export predictions in the OmniDocBench end-to-end parsing format.
  3. Match each image to its corresponding OmniDocBench v1.5 ground-truth annotation.
  4. Apply the same preprocessing, evaluator, and metric settings to every scenario.

Refer to the version-pinned OmniDocBench v1.5 evaluation guide for environment setup, prediction formats, and evaluation commands.

Metrics

MetricDirectionDefinition
Overall((1 - TextEdit) * 100 + TableTEDS + FormulaCDM) / 3
TextEditNormalized edit distance for plain-text content
FormulaCDMCharacter Detection Matching score for formulas
TableTEDSTree-Edit-Distance-based Similarity for table structure
Reading OrderEditNormalized edit distance for the reading-order sequence

Submit Results

Model results can appear in the Hugging Face Hub leaderboard through the Hub evaluation-results workflow. Add an evaluation result file under .eval_results/ in the model repository, set evaluation_framework: real5-omnidocbench, and follow the Hugging Face evaluation results documentation for the supported schema and submission process.

Citation

If you use Real5-OmniDocBench in your research, please cite the following paper. Please also cite OmniDocBench when using its annotations or evaluation pipeline.

@misc{zhou2026real5omnidocbench,
  title         = {Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild},
  author        = {Changda Zhou and Ziyue Gao and Xueqing Wang and Tingquan Gao and Cheng Cui and Jing Tang and Yi Liu},
  year          = {2026},
  eprint        = {2603.04205},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2603.04205},
  url           = {https://arxiv.org/abs/2603.04205}
}

Acknowledgements

Real5-OmniDocBench is built on OmniDocBench v1.5 and adopts its annotations and evaluation protocol. We thank the OmniDocBench authors for making their benchmark and evaluation tools available to the community.

License

The Real5-OmniDocBench dataset and repository materials are released under the Apache License 2.0. This benchmark inherits annotations and source-page correspondence from OmniDocBench v1.5; use of those materials remains subject to the applicable OmniDocBench and source-document terms.

Contributors

ZH
zhouchangda

28 commits

ChengCui

3 commits

PaddlePaddle/Real5-OmniDocBench

Dataset

37

stars

31

commits

2

linked in READMEs

Aug 27, 2026

updated

benchmark
document
document-parsing
image
multimodal
ocr

README

Real5-OmniDocBench

A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

ECCV 2026 arXiv Dataset Base Benchmark License

Leaderboard | Overview | Dataset | Evaluation | Submit Results | Citation

Real5-OmniDocBench measures the robustness of document parsing systems under five physical acquisition conditions: Scanning, Warping, Screen-Photography, Illumination, and Skew. It reconstructs the same 1,355 pages from OmniDocBench v1.5 in every condition, producing 6,775 images in total. The one-to-one page correspondence and shared evaluation protocol isolate the effect of acquisition conditions from changes in document content.

Original pages and corresponding Real5-OmniDocBench reconstructions

Original pages and their corresponding reconstructions under five physical acquisition conditions.

News

  • 2026-08-27: Added evaluation results for NaviDC-OCR.
  • 2026-08-08: Added evaluation results for Kimi-K2.5、Kimi-K2.6、Doubao-Seed-2.1-Pro、MonkeyOCRv2-S-Parsing、MonkeyOCRv2-B-Parsing and OvisOCR2.
  • 2026-06-18: Real5-OmniDocBench was accepted to ECCV 2026. 🎉
Earlier updates
  • 2026-05-28: Added evaluation results for PaddleOCR-VL-1.6 and MinerU2.5-Pro.
  • 2026-03-05: Released the paper and added results for DeepSeek-OCR 2 and GLM-OCR.
  • 2026-01-28: Released the dataset and benchmark.

Leaderboard

The leaderboard reports performance over all five acquisition conditions. All metrics follow OmniDocBench: Overall↑, TextEdit↓, FormulaCDM↑, TableTEDS↑, and Reading OrderEdit↓.

Higher is better for ↑ metrics and lower is better for ↓ metrics. Best results in each column are shown in bold, and second-best results are underlined. Tables are sorted by Overall score in descending order.

1. Overall

MethodsModel TypeParametersOverall↑Scanning↑Warping↑Screen-Photography↑Illumination↑Skew↑
PaddleOCR-VL-1.6Specialized VLMs0.9B93.1994.7492.4892.7893.2892.66
OvisOCR2Specialized VLMs0.9B92.2993.7791.4093.0992.8890.33
PaddleOCR-VL-1.5Specialized VLMs0.9B92.0593.4391.2591.7692.1691.66
NaviDC-OCRSpecialized VLMs1.2B90.7292.3390.6389.0990.9090.63
GLM-OCRSpecialized VLMs0.9B90.3292.6790.6891.7591.1285.39
Kimi-K2.6General VLMs1.1T89.7690.0889.6289.5889.9189.61
Gemini-3 ProGeneral VLMs-89.2489.4788.9088.8689.5389.45
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.2289.4989.7088.4088.5489.97
Kimi-K2.5General VLMs1.1T89.0989.6788.8688.3989.6688.86
Doubao-Seed-2.1-ProGeneral VLMs-89.0288.8589.3688.9989.1388.79
MinerU2.5-proSpecialized VLMs1.2B88.9492.1188.7291.2991.3181.26
Qwen3-VL-235BGeneral VLMs235B88.9089.4389.9989.2789.2786.56
Gemini-2.5 ProGeneral VLMs-88.2189.2587.6387.1187.9789.07
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B87.9088.8788.1787.6486.7588.09
Qwen2.5-VL-72BGeneral VLMs72B86.9286.1987.7786.4887.2586.90
dots.ocrSpecialized VLMs3B86.3886.8786.0187.1887.5784.27
MinerU2.5Specialized VLMs1.2B85.6190.0683.7689.4189.5775.24
PaddleOCR-VLSpecialized VLMs0.9B85.5492.1185.9782.5489.6177.47
Nanonets-OCR-sSpecialized VLMs3B84.1985.5283.5684.8685.0181.98
MonkeyOCR-pro-3BSpecialized VLMs3.7B79.4986.9478.9082.4484.7164.47
GPT-5.2General VLMs-78.6684.4376.2676.7580.8875.00
MonkeyOCR-3BSpecialized VLMs3.7B78.2984.6577.2780.7183.1665.67
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B77.1584.6476.5980.2482.1162.18
MinerU2-VLMSpecialized VLMs0.9B76.9583.6073.7378.7780.5168.16
Deepseek-OCRSpecialized VLMs3B73.9986.1767.2075.3178.1063.01
Deepseek-OCR 2Specialized VLMs3B73.0189.5966.5371.6576.0261.28
PP-StructureV3Pipeline Tools-64.4584.6859.3466.8973.3837.98
DolphinSpecialized VLMs322M61.7872.1660.3564.2967.2944.83
Dolphin-1.5Specialized VLMs0.3B61.4883.3950.5069.7675.6128.16
Marker-1.8.2Pipeline Tools-60.1070.2758.9863.6566.3141.27

Overall is the mean score across all five scenarios.

View per-scenario leaderboards

2. Scanning

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B94.740.03593.6594.120.042
OvisOCR2Specialized VLMs0.9B93.770.04192.1393.300.039
PaddleOCR-VL-1.5Specialized VLMs0.9B93.430.03793.0490.970.045
GLM-OCRSpecialized VLMs0.9B92.670.05491.1092.280.061
NaviDC-OCRSpecialized VLMs1.2B92.330.04588.2693.260.039
PaddleOCR-VLSpecialized VLMs0.9B92.110.03990.3589.900.048
MinerU2.5-proSpecialized VLMs1.2B92.110.04089.7790.570.043
Kimi-K2.6General VLMs1.1T90.080.06289.0285.770.072
MinerU2.5Specialized VLMs1.2B90.060.05288.2287.160.050
Kimi-K2.5General VLMs1.1T89.670.06088.6786.310.079
Deepseek-OCR 2Specialized VLMs3B89.590.05588.5585.720.056
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.490.05386.2987.530.051
Gemini-3 ProGeneral VLMs-89.470.07188.1687.370.078
Qwen3-VL-235BGeneral VLMs235B89.430.05989.0185.190.066
Gemini-2.5 ProGeneral VLMs-89.250.07387.4487.620.098
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B88.870.05886.3286.090.053
Doubao-Seed-2.1-ProGeneral VLMs-88.850.08486.5688.330.093
MonkeyOCR-pro-3BSpecialized VLMs3.7B86.940.10386.2984.860.141
dots.ocrSpecialized VLMs3B86.870.08383.2785.680.081
Qwen2.5-VL-72BGeneral VLMs72B86.190.11086.1483.410.114
Deepseek-OCRSpecialized VLMs3B86.170.07883.5982.690.085
Nanonets-OCR-sSpecialized VLMs3B85.520.10688.0979.110.106
PP-StructureV3Pipeline Tools-84.680.09484.3479.060.092
MonkeyOCR-3BSpecialized VLMs3.7B84.650.10084.1679.810.143
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B84.640.12384.1782.130.145
GPT-5.2General VLMs-84.430.14285.6881.780.109
MinerU2-VLMSpecialized VLMs0.9B83.600.09479.7680.440.091
Dolphin-1.5Specialized VLMs0.3B83.390.09776.2583.650.090
DolphinSpecialized VLMs322M72.160.15464.5867.270.130
Marker-1.8.2Pipeline Tools-70.270.22377.0356.050.238

3. Warping

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B92.480.04891.6390.660.061
OvisOCR2Specialized VLMs0.9B91.400.05890.9489.080.058
PaddleOCR-VL-1.5Specialized VLMs0.9B91.250.05390.9488.100.063
GLM-OCRSpecialized VLMs0.9B90.680.07190.3088.780.100
NaviDC-OCRSpecialized VLMs1.2B90.630.06289.1488.930.051
Qwen3-VL-235BGeneral VLMs235B89.990.05189.0685.950.064
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.700.05887.5087.390.052
Kimi-K2.6General VLMs1.1T89.620.07391.1885.020.079
Doubao-Seed-2.1-ProGeneral VLMs-89.360.08787.7489.030.095
Gemini-3 ProGeneral VLMs-88.900.08688.1087.200.087
Kimi-K2.5General VLMs1.1T88.860.06989.7783.710.084
MinerU2.5-proSpecialized VLMs1.2B88.720.10087.8188.360.076
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B88.170.06986.2085.170.056
Qwen2.5-VL-72BGeneral VLMs72B87.770.08688.8583.060.102
Gemini-2.5 ProGeneral VLMs-87.630.09286.5085.590.109
dots.ocrSpecialized VLMs3B86.010.08785.0381.740.093
PaddleOCR-VLSpecialized VLMs0.9B85.970.09385.4581.770.092
MinerU2.5Specialized VLMs1.2B83.760.15485.9280.710.104
Nanonets-OCR-sSpecialized VLMs3B83.560.12186.2476.570.124
MonkeyOCR-pro-3BSpecialized VLMs3.7B78.900.16879.5573.940.212
MonkeyOCR-3BSpecialized VLMs3.7B77.270.16479.0869.180.211
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B76.590.19678.8570.520.221
GPT-5.2General VLMs-76.260.23980.9071.800.165
MinerU2-VLMSpecialized VLMs0.9B73.730.20277.7263.650.173
Deepseek-OCRSpecialized VLMs3B67.200.32873.5960.800.226
Deepseek-OCR 2Specialized VLMs3B66.530.29370.4258.440.209
DolphinSpecialized VLMs322M60.350.31661.0651.580.247
PP-StructureV3Pipeline Tools-59.340.37668.2247.400.261
Marker-1.8.2Pipeline Tools-58.980.34972.7139.080.390
Dolphin-1.5Specialized VLMs0.3B50.500.38347.2442.520.309

4. Screen-Photography

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
OvisOCR2Specialized VLMs0.9B93.090.05491.8092.830.046
PaddleOCR-VL-1.6Specialized VLMs0.9B92.780.04590.6492.190.054
PaddleOCR-VL-1.5Specialized VLMs0.9B91.760.05090.8889.380.059
GLM-OCRSpecialized VLMs0.9B91.750.06389.8391.660.070
MinerU2.5-proSpecialized VLMs1.2B91.290.05087.4191.440.044
Kimi-K2.6General VLMs1.1T89.580.07390.1085.490.077
MinerU2.5Specialized VLMs1.2B89.410.06287.5586.830.053
Qwen3-VL-235BGeneral VLMs235B89.270.06888.7285.850.071
NaviDC-OCRSpecialized VLMs1.2B89.090.07885.8089.300.057
Doubao-Seed-2.1-ProGeneral VLMs-88.990.09887.4789.310.102
Gemini-3 ProGeneral VLMs-88.860.08487.3387.650.087
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B88.400.07087.0785.120.050
Kimi-K2.5General VLMs1.1T88.390.07087.4184.770.078
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B87.640.06884.9684.720.049
dots.ocrSpecialized VLMs3B87.180.08185.3484.260.079
Gemini-2.5 ProGeneral VLMs-87.110.10385.3086.310.117
Qwen2.5-VL-72BGeneral VLMs72B86.480.10087.4682.000.102
Nanonets-OCR-sSpecialized VLMs3B84.860.11286.6579.090.117
PaddleOCR-VLSpecialized VLMs0.9B82.540.10383.5874.360.107
MonkeyOCR-pro-3BSpecialized VLMs3.7B82.440.12481.5578.130.177
MonkeyOCR-3BSpecialized VLMs3.7B80.710.12281.3373.040.177
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B80.240.14880.7874.740.179
MinerU2-VLMSpecialized VLMs0.9B78.770.13979.0271.170.123
GPT-5.2General VLMs-76.750.20879.2771.730.148
Deepseek-OCRSpecialized VLMs3B75.310.22077.6870.260.169
Deepseek-OCR 2Specialized VLMs3B71.650.20173.4961.540.157
Dolphin-1.5Specialized VLMs0.3B69.760.20561.8068.000.177
PP-StructureV3Pipeline Tools-66.890.20473.2647.820.165
DolphinSpecialized VLMs322M64.290.23258.6657.380.195
Marker-1.8.2Pipeline Tools-63.650.29072.7347.210.325

5. Illumination

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B93.280.04292.6191.480.051
OvisOCR2Specialized VLMs0.9B92.880.04391.1891.720.040
PaddleOCR-VL-1.5Specialized VLMs0.9B92.160.04691.8089.330.051
MinerU2.5-proSpecialized VLMs1.2B91.310.05088.1590.760.052
GLM-OCRSpecialized VLMs0.9B91.120.05991.0288.200.071
NaviDC-OCRSpecialized VLMs1.2B90.900.05587.2790.910.045
Kimi-K2.6General VLMs1.1T89.910.06289.7686.200.072
Kimi-K2.5General VLMs1.1T89.660.06489.8385.530.077
PaddleOCR-VLSpecialized VLMs0.9B89.610.04986.6687.020.055
MinerU2.5Specialized VLMs1.2B89.570.06588.3686.870.062
Gemini-3 ProGeneral VLMs-89.530.07387.7888.140.080
Qwen3-VL-235BGeneral VLMs235B89.270.06087.8186.050.070
Doubao-Seed-2.1-ProGeneral VLMs-89.130.08587.2988.590.091
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B88.540.06686.7285.530.059
Gemini-2.5 ProGeneral VLMs-87.970.08386.1386.110.103
dots.ocrSpecialized VLMs3B87.570.06885.0784.440.076
Qwen2.5-VL-72BGeneral VLMs72B87.250.08786.4484.030.097
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B86.750.07484.1683.490.057
Nanonets-OCR-sSpecialized VLMs3B85.010.09987.9476.960.112
MonkeyOCR-pro-3BSpecialized VLMs3.7B84.710.12084.1382.020.171
MonkeyOCR-3BSpecialized VLMs3.7B83.160.11883.6377.620.168
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B82.110.14482.0778.670.172
GPT-5.2General VLMs-80.880.19184.4177.370.134
MinerU2-VLMSpecialized VLMs0.9B80.510.13580.7274.290.123
Deepseek-OCRSpecialized VLMs3B78.100.19281.7171.810.156
Deepseek-OCR 2Specialized VLMs3B76.020.16877.8367.010.122
Dolphin-1.5Specialized VLMs0.3B75.610.15970.0472.690.133
PP-StructureV3Pipeline Tools-73.380.15877.7558.190.126
DolphinSpecialized VLMs322M67.290.19761.4260.100.173
Marker-1.8.2Pipeline Tools-66.310.25974.8050.030.337

6. Skew

MethodsModel TypeParametersOverall↑TextEditFormulaCDMTableTEDSReading OrderEdit
PaddleOCR-VL-1.6Specialized VLMs0.9B92.660.04591.4491.040.058
PaddleOCR-VL-1.5Specialized VLMs0.9B91.660.04791.0088.690.061
NaviDC-OCRSpecialized VLMs1.2B90.630.05487.1590.100.042
OvisOCR2Specialized VLMs0.9B90.330.04890.5885.230.048
MonkeyOCRv2-B-ParsingSpecialized VLMs0.7B89.970.05188.1086.880.051
Kimi-K2.6General VLMs1.1T89.610.06090.2584.270.078
Gemini-3 ProGeneral VLMs-89.450.08088.3388.060.092
Gemini-2.5 ProGeneral VLMs-89.070.07787.8986.990.104
Kimi-K2.5General VLMs1.1T88.860.06089.6283.000.080
Doubao-Seed-2.1-ProGeneral VLMs-88.790.08088.2386.110.095
MonkeyOCRv2-S-ParsingSpecialized VLMs0.6B88.090.06384.5486.070.055
Qwen2.5-VL-72BGeneral VLMs72B86.900.07787.2681.140.091
Qwen3-VL-235BGeneral VLMs235B86.560.07783.9683.410.091
GLM-OCRSpecialized VLMs0.9B85.390.09985.7880.280.156
dots.ocrSpecialized VLMs3B84.270.08785.7375.740.094
Nanonets-OCR-sSpecialized VLMs3B81.980.12185.7872.220.133
MinerU2.5-proSpecialized VLMs1.2B81.260.20283.9281.070.107
PaddleOCR-VLSpecialized VLMs0.9B77.470.19278.8172.830.193
MinerU2.5Specialized VLMs1.2B75.240.30581.7874.390.151
GPT-5.2General VLMs-75.000.25780.2770.470.167
MinerU2-VLMSpecialized VLMs0.9B68.160.23074.4553.070.191
MonkeyOCR-3BSpecialized VLMs3.7B65.670.24869.2352.590.300
MonkeyOCR-pro-3BSpecialized VLMs3.7B64.470.25169.0649.420.301
Deepseek-OCRSpecialized VLMs3B63.010.32773.2748.480.231
MonkeyOCR-pro-1.2BSpecialized VLMs1.9B62.180.29266.2549.460.317
Deepseek-OCR 2Specialized VLMs3B61.280.29566.1647.180.221
DolphinSpecialized VLMs322M44.830.50051.3433.220.321
Marker-1.8.2Pipeline Tools-41.270.53660.1617.230.543
PP-StructureV3Pipeline Tools-37.980.55744.3725.270.417
Dolphin-1.5Specialized VLMs0.3B28.160.55325.6014.180.419

Benchmark Overview

Real5-OmniDocBench evaluates the same document content under controlled changes to the physical acquisition process. Except for the scanning subset, images were manually captured with handheld mobile devices.

ScenarioAcquisition conditionRepresentative artifacts
ScanningDocuments captured with scanning devicesScanner characteristics and clean planar capture
WarpingCurved or non-planar pages photographed by handPage curvature, folding, and local deformation
Screen-PhotographyScreens displaying documents photographed by handMoiré patterns, reflections, and display artifacts
IlluminationDocuments photographed under varied lightingShadows, glare, and uneven exposure
SkewDocuments photographed from oblique viewpointsPerspective distortion and geometric skew

This design provides:

  • Controlled comparison: every scenario contains the same 1,355 source pages.
  • Physical realism: acquisition artifacts are produced by real devices and environments rather than synthetic transformations.
  • Protocol compatibility: page identities, annotations, prediction format, and metrics follow OmniDocBench v1.5.

Dataset

The five scenario directories each contain 1,355 images. The complete download is approximately 16 GB.

Real5-OmniDocBench/
├── Real5-OmniDocBench-Scanning/
├── Real5-OmniDocBench-Warping/
├── Real5-OmniDocBench-Screen-Photography/
├── Real5-OmniDocBench-Illumination/
└── Real5-OmniDocBench-Skew/

Install the current Hugging Face Hub CLI and download the complete dataset:

pip install -U huggingface_hub
hf download PaddlePaddle/Real5-OmniDocBench \
  --repo-type dataset \
  --local-dir ./Real5-OmniDocBench

To download a single condition, use a file pattern:

hf download PaddlePaddle/Real5-OmniDocBench \
  --repo-type dataset \
  --include "Real5-OmniDocBench-Warping/*" \
  --local-dir ./Real5-OmniDocBench

Download the matching OmniDocBench v1.5 ground-truth annotations separately:

hf download opendatalab/OmniDocBench OmniDocBench.json \
  --repo-type dataset \
  --revision v1_5 \
  --local-dir ./OmniDocBench-v1.5

Evaluation

Real5-OmniDocBench does not introduce a new prediction schema. It reuses the OmniDocBench v1.5 annotation format and evaluation pipeline so that performance can be compared across acquisition conditions.

  1. Run the model independently on all 1,355 images in each scenario.
  2. Export predictions in the OmniDocBench end-to-end parsing format.
  3. Match each image to its corresponding OmniDocBench v1.5 ground-truth annotation.
  4. Apply the same preprocessing, evaluator, and metric settings to every scenario.

Refer to the version-pinned OmniDocBench v1.5 evaluation guide for environment setup, prediction formats, and evaluation commands.

Metrics

MetricDirectionDefinition
Overall((1 - TextEdit) * 100 + TableTEDS + FormulaCDM) / 3
TextEditNormalized edit distance for plain-text content
FormulaCDMCharacter Detection Matching score for formulas
TableTEDSTree-Edit-Distance-based Similarity for table structure
Reading OrderEditNormalized edit distance for the reading-order sequence

Submit Results

Model results can appear in the Hugging Face Hub leaderboard through the Hub evaluation-results workflow. Add an evaluation result file under .eval_results/ in the model repository, set evaluation_framework: real5-omnidocbench, and follow the Hugging Face evaluation results documentation for the supported schema and submission process.

Citation

If you use Real5-OmniDocBench in your research, please cite the following paper. Please also cite OmniDocBench when using its annotations or evaluation pipeline.

@misc{zhou2026real5omnidocbench,
  title         = {Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild},
  author        = {Changda Zhou and Ziyue Gao and Xueqing Wang and Tingquan Gao and Cheng Cui and Jing Tang and Yi Liu},
  year          = {2026},
  eprint        = {2603.04205},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2603.04205},
  url           = {https://arxiv.org/abs/2603.04205}
}

Acknowledgements

Real5-OmniDocBench is built on OmniDocBench v1.5 and adopts its annotations and evaluation protocol. We thank the OmniDocBench authors for making their benchmark and evaluation tools available to the community.

License

The Real5-OmniDocBench dataset and repository materials are released under the Apache License 2.0. This benchmark inherits annotations and source-page correspondence from OmniDocBench v1.5; use of those materials remains subject to the applicable OmniDocBench and source-document terms.

Contributors

ZH
zhouchangda

28 commits

ChengCui

3 commits