Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks mainly focus on the final diagnostic answer, providing limited insight into whether models engage in clinically meaningful reasoning.
To address this, we present CheXStruct and CXReasonBench:
CheXStruct: A fully automated pipeline extracting structured clinical information directly from chest X-rays. It performs anatomical segmentation, derives anatomical landmarks and diagnostic measurements, computes diagnostic indices, and applies clinical thresholds based on expert guidelines.
CXReasonBench: A multi-path, multi-stage evaluation framework that assesses a model’s ability to perform structured diagnostic reasoning. The benchmark includes 18,988 QA pairs across 12 diagnostic tasks and 1,200 cases, each with up to 4 visual inputs, enabling detailed evaluation of reasoning steps including visual grounding and diagnostic measurements.
Even the strongest LVLMs evaluated struggle with structured reasoning and generalization, showing the importance of this benchmark.
If you use this code or dataset in your research, please cite our work:
@inproceedings{leecxreasonbench,
title={CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays},
author={Lee, Hyungyung and Choi, Geon and Lee, Jung-Oh and Yoon, Hangyul and Hong, Hyuk Gi and Choi, Edward},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track}
}
If you have any questions, feedback, or issues regarding this project, please reach out to us via email: ttumyche@kaist.ac.kr
3 commits
Python
100.0%
Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks mainly focus on the final diagnostic answer, providing limited insight into whether models engage in clinically meaningful reasoning.
To address this, we present CheXStruct and CXReasonBench:
CheXStruct: A fully automated pipeline extracting structured clinical information directly from chest X-rays. It performs anatomical segmentation, derives anatomical landmarks and diagnostic measurements, computes diagnostic indices, and applies clinical thresholds based on expert guidelines.
CXReasonBench: A multi-path, multi-stage evaluation framework that assesses a model’s ability to perform structured diagnostic reasoning. The benchmark includes 18,988 QA pairs across 12 diagnostic tasks and 1,200 cases, each with up to 4 visual inputs, enabling detailed evaluation of reasoning steps including visual grounding and diagnostic measurements.
Even the strongest LVLMs evaluated struggle with structured reasoning and generalization, showing the importance of this benchmark.
If you use this code or dataset in your research, please cite our work:
@inproceedings{leecxreasonbench,
title={CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays},
author={Lee, Hyungyung and Choi, Geon and Lee, Jung-Oh and Yoon, Hangyul and Hong, Hyuk Gi and Choi, Edward},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track}
}
If you have any questions, feedback, or issues regarding this project, please reach out to us via email: ttumyche@kaist.ac.kr
3 commits
Python
100.0%