VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
| Rank | Model | ELO | 95% CI | Wins | Losses | Ties | Win% |
|---|---|---|---|---|---|---|---|
| 1 | lightonai/LightOnOCR-2-1B | 1559 | 1497–1630 | 39 | 25 | 0 | 61% |
| 2 | zai-org/GLM-OCR | 1535 | 1471–1591 | 48 | 35 | 1 | 57% |
| 3 | rednote-hilab/dots.ocr | 1453 | 1385–1515 | 26 | 37 | 0 | 41% |
| 4 | deepseek-ai/DeepSeek-OCR | 1452 | 1388–1514 | 33 | 49 | 1 | 40% |
davanstrien/bpl-ocr-benchload_dataset("davanstrien/bpl-ocr-bench-results") — leaderboard tableload_dataset("davanstrien/bpl-ocr-bench-results", name="comparisons") — full pairwise comparison logload_dataset("davanstrien/bpl-ocr-bench-results", name="metadata") — evaluation run historyGenerated by ocr-bench
21 commits
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
| Rank | Model | ELO | 95% CI | Wins | Losses | Ties | Win% |
|---|---|---|---|---|---|---|---|
| 1 | lightonai/LightOnOCR-2-1B | 1559 | 1497–1630 | 39 | 25 | 0 | 61% |
| 2 | zai-org/GLM-OCR | 1535 | 1471–1591 | 48 | 35 | 1 | 57% |
| 3 | rednote-hilab/dots.ocr | 1453 | 1385–1515 | 26 | 37 | 0 | 41% |
| 4 | deepseek-ai/DeepSeek-OCR | 1452 | 1388–1514 | 33 | 49 | 1 | 40% |
davanstrien/bpl-ocr-benchload_dataset("davanstrien/bpl-ocr-bench-results") — leaderboard tableload_dataset("davanstrien/bpl-ocr-bench-results", name="comparisons") — full pairwise comparison logload_dataset("davanstrien/bpl-ocr-bench-results", name="metadata") — evaluation run historyGenerated by ocr-bench
21 commits