ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
2
4 commits
1 linked in READMEs
updated Jun 16, 2025

ClinBench-HPB is a comprehensive benchmark dataset specifically designed to evaluate the performance of Large Language Models (LLMs) in the field of hepatobiliary and pancreatic diseases.
This dataset consists of five carefully constructed subsets covering different data sources and question types:
Comprehensiveness: Covers multiple subspecialties of HPB diseases
Diversity: Includes both multiple-choice questions and real clinical cases
High Quality: All data reviewed by medical experts
Multilingual: Supports both Chinese and English evaluation
This dataset is suitable for:
Evaluating LLMs' medical knowledge in HPB diseases
Comparing model performance on medical QA tasks
Developing medical education applications
Evaluating clinical decision support systems
This benchmark is for research purposes only and does not constitute medical advice. Always consult healthcare professionals for clinical decisions. This dataset is licensed for academic research only. Commercial use is strictly prohibited. No modification, distribution, or use in commercial products is permitted without explicit authorization. If you believe this dataset infringes your rights, please contact us via: Email: yuchong.li@connect.polyu.hk
If you use ClinBench-HPB in your research, please cite our work:
@misc{li2025clinbenchhpbclinicalbenchmarkevaluating,
title={ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases},
author={Yuchong Li and Xiaojun Zeng and Chihua Fang and Jian Yang and Fucang Jia and Lei Zhang},
year={2025},
eprint={2506.00095},
archivePrefix={arXiv},
primaryClass={cs.CY},
url={https://arxiv.org/abs/2506.00095},
}
4 commits
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
2
4 commits
1 linked in READMEs
updated Jun 16, 2025

ClinBench-HPB is a comprehensive benchmark dataset specifically designed to evaluate the performance of Large Language Models (LLMs) in the field of hepatobiliary and pancreatic diseases.
This dataset consists of five carefully constructed subsets covering different data sources and question types:
Comprehensiveness: Covers multiple subspecialties of HPB diseases
Diversity: Includes both multiple-choice questions and real clinical cases
High Quality: All data reviewed by medical experts
Multilingual: Supports both Chinese and English evaluation
This dataset is suitable for:
Evaluating LLMs' medical knowledge in HPB diseases
Comparing model performance on medical QA tasks
Developing medical education applications
Evaluating clinical decision support systems
This benchmark is for research purposes only and does not constitute medical advice. Always consult healthcare professionals for clinical decisions. This dataset is licensed for academic research only. Commercial use is strictly prohibited. No modification, distribution, or use in commercial products is permitted without explicit authorization. If you believe this dataset infringes your rights, please contact us via: Email: yuchong.li@connect.polyu.hk
If you use ClinBench-HPB in your research, please cite our work:
@misc{li2025clinbenchhpbclinicalbenchmarkevaluating,
title={ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases},
author={Yuchong Li and Xiaojun Zeng and Chihua Fang and Jian Yang and Fucang Jia and Lei Zhang},
year={2025},
eprint={2506.00095},
archivePrefix={arXiv},
primaryClass={cs.CY},
url={https://arxiv.org/abs/2506.00095},
}
4 commits