mx262/OCR-Reasoning

Dataset

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

7

10 commits

2 linked in READMEs

updated Jun 15, 2026

See the code

README

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Project Page | Paper | GitHub

OCR-Reasoning is a comprehensive benchmark designed to systematically assess Multimodal Large Language Models (MLLMs) on text-rich image reasoning tasks. The benchmark comprises 1,069 human-annotated examples spanning 6 core reasoning abilities: spatial reasoning, numerical analysis, mathematical reasoning, enumerative reasoning, logical reasoning, and multidisciplinary knowledge.

Unlike existing text-rich image understanding benchmarks that only provide final answers, OCR-Reasoning also provides detailed step-by-step reasoning processes, enabling a holistic evaluation of a model's problem-solving logic.

Evaluation

The evaluation of OCR-Reasoning is supported in VLMEvalKit. To evaluate a model (e.g., Qwen2.5-VL), you can use the following commands from the official repository:

git clone https://github.com/SCUT-DLVCLab/OCR-Reasoning
cd OCR_Reasoning
python run.py --data OCR_Reasoning --model Qwen2.5-VL-7B-Instruct --verbose

Citation

If you find OCR-Reasoning helpful in your research, please cite the following paper:

@article{huang2025ocreasoning,
  title={OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning},
  author={Huang, Mingxin and Shi, Yongxin and Peng, Dezhi and Lai, Songxuan and Xie, Zecheng and Jin, Lianwen},
  journal={arXiv preprint arXiv:2505.17163},
  year={2025}
}
Multimodal Large Language Models
Reasoning

mx262/OCR-Reasoning

Dataset

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

7

10 commits

2 linked in READMEs

updated Jun 15, 2026

See the code

README

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Project Page | Paper | GitHub

OCR-Reasoning is a comprehensive benchmark designed to systematically assess Multimodal Large Language Models (MLLMs) on text-rich image reasoning tasks. The benchmark comprises 1,069 human-annotated examples spanning 6 core reasoning abilities: spatial reasoning, numerical analysis, mathematical reasoning, enumerative reasoning, logical reasoning, and multidisciplinary knowledge.

Unlike existing text-rich image understanding benchmarks that only provide final answers, OCR-Reasoning also provides detailed step-by-step reasoning processes, enabling a holistic evaluation of a model's problem-solving logic.

Evaluation

The evaluation of OCR-Reasoning is supported in VLMEvalKit. To evaluate a model (e.g., Qwen2.5-VL), you can use the following commands from the official repository:

git clone https://github.com/SCUT-DLVCLab/OCR-Reasoning
cd OCR_Reasoning
python run.py --data OCR_Reasoning --model Qwen2.5-VL-7B-Instruct --verbose

Citation

If you find OCR-Reasoning helpful in your research, please cite the following paper:

@article{huang2025ocreasoning,
  title={OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning},
  author={Huang, Mingxin and Shi, Yongxin and Peng, Dezhi and Lai, Songxuan and Xie, Zecheng and Jin, Lianwen},
  journal={arXiv preprint arXiv:2505.17163},
  year={2025}
}
Multimodal Large Language Models
Reasoning