yezanting/Med-VLM-Bench-Summary

A Curated Benchmark Repository for Medical Vision-Language Models

200

53 commits

updated Jan 21, 2026

See the code

README

🧠 Med-VLM-Bench: A Curated Benchmark Repository for Medical Vision-Language Models

📚 A comprehensive summary of recent benchmarks for evaluating and training Medical Vision-Language Models (Med-VLMs)


👨‍💻 Contributors

  • 🧑‍🔬 Zanting Ye
    Southern Medical University
    📧 yzt2861252880@gmail.com

  • 🧑‍🔬 Xu Han
    Shanghai Jiao Tong University
    📧 hanxv8826@gmail.com

  • 🧑‍🔬 Xiaolong Niu
    Southern Medical University

  • 🧑‍🔬 Zian Wang
    Shanghai Jiao Tong University

  • 🧑‍🔬 Shengyuan Liu
    The Chinese University of Hong Kong
    📧 liushengyuan@link.cuhk.edu.hk

  • 🧑‍🔬 Xin Liu
    Southern Medical University
    📧 lx10230114@gmail.com

  • 👨‍🏫 Lijun Lu
    Southern Medical University


📊 GitHub Stats

Stars Forks License Last Commit


🔍 Project Overview

With the continuous advancement of research on Medical Vision-Language Models (Med-VLMs) and their reasoning capabilities, a number of high-quality, publicly available datasets focusing on medical reasoning have been released between March and May 2025. These datasets provide a solid foundation for the development of multimodal medical AI systems.

Med-VLM-Bench is a curated, continuously updated repository of the latest and most important datasets for training and evaluating medical LLMs and VLMs. This project focuses on:

  • ✅ Reasoning-centric multimodal benchmarks
  • 📅 Latest datasets published in Mar 2025–2026
  • 🧠 Foundational datasets from 2023–2024
  • 🔗 Direct access to dataset links or HuggingFace/GitHub repositories

💡 Our knowledge is limited to public sources. We welcome community contributions — feel free to open an issue to share new datasets, and we will update promptly.

📌Note: The annotation time of the dataset is based on the publication time of the corresponding article.


📢 News

🌟 Latest Updates

  • 2026-01-21: 🎉 Added some recent datasets and benchmarks!Check it out for detailed information and download links!
  • 2025-06-29: 🎉 Added new datasets/benchmarks AbdomenAtlas 3.0 (ICCV2025), Derm1M(ICCV2025), MedTVT-R1, , GEMeX(ICCV2025) and HIE-Reasoning(ICML2025). Check it out for detailed information and download links!
  • 2025-06-18: 🎉 Added new datasets/benchmarks Lingshu, ReasonMed. Check it out for detailed information and download links!
  • 2025-06-11: 🎉 Added some recent datasets and benchmarks!
  • 2025-06-11: 🎉 Create our github project!

📊 Dataset Summary Table

Dataset NamePaper TitleYear / VenueData ModalityTask TypeSizeDownload Link
Multi-RADSMulti-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models2026.1.6Text (Synthetic Reports)RADS Classification1,600 synthetic reports covering 17 distinct imaging findings.Github
Bones and Joints (B&J) BenchmarkThe Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision Language Models for Clinical Competency2025.12.25Text & Image (X-ray, CT, MRI)VQA & Treatment Planning1,245 question-answer pairs spanning 7 clinical competency tasks. Hugging Face
MediEvalMediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs2025.12.23Text (EHR & Notes)Natural Language Inference37,144 medical statements derived from 2,015 hospital admissions.GitHub
TCM-BEST4SDTA benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models2025.12.2Text (TCM Case Reports)Syndrome Differentiation600 total questions including 300 clinical syndrome differentiation cases.GitHub
SurgMLLMBenchSurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding2025.11.26Video & TextSurgical Scene Understanding10,652 frames with annotations integrated from 5 surgical datasets.project page
MedVisionMedVision: Dataset and Benchmark for Quantitative Medical Image Analysis2025.11.24Image (CT, MRI, X-ray, PET)Detection & Measurement30.8 million image-annotation pairs across 22 public datasets.project page
EHRStructEHRStruct: A Comprehensive Benchmark Framework for Evaluating Large Language Models on Structured Electronic Health Record Tasks2025.12.1Structured EHR (Tables)Relational Data Reasoning2,200 task-specific samples across 11 data and knowledge tasks.GitHub
TCM-EvalTCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine2025.12.26Text (TCM Knowledge)Professional MCQ6,099 questions from expert-level Chinese medical examinations.dataset
RxSafeBenchRxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated ConsultationBIBM2025Text (Dialogue & MCQ)Medication Safety QA2,443 consultation scenarios including 1,063 contraindication cases.GitHub
SemBenchSemBench: A Benchmark for Semantic Query Processing Engines2025.11.3Text (Knowledge Graph)Semantic Query Evaluation1,400+ SPARQL templates for evaluating medical query engines.GitHub
XBenchXBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography2025.10.22Image (X-ray) & TextGrounding & Explanation12,601 chest X-ray cases with localization and textual explanations.GitHub
IMBIMB: An Italian Medical Benchmark for Question AnsweringCLIC-it 2025Text (Italian)Medical QA & MCQA808,506 items featuring 782,644 clinical Italian conversations.GitHub
ViPET-ReportGenToward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report GenerationNeurIPS 2025Image (3D PET/CT) & TextReport Generation & VQA1.5 million slices paired with 2,757 Vietnamese clinical reports.GitHub
Neural-MedBenchBeyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks2025.12.13Text & Image (MRI/CT)Differential Diagnosis120 expert cases resulting in 200 depth-of-reasoning tasks.Project page
MedQARoMedQARo: A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian2025.12.31Text (Romanian)Multilingual QA102,646 Romanian QA pairs covering 1,011 clinical patients.GitHub
AnesSuiteAnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs2025.12.25Text (Anesthesiology)Specialized Knowledge QA4,427 anesthesiology MCQ items focused on complex decision-making.GitHub
TracSumTracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domainthe 2025 Conference on Empirical Methods in Natural Language ProcessingText (Abstracts/Notes)Aspect-Based Summarization500 abstracts resulting in 3,500 summary-citation traceable pairs.GitHub
MedAgentBoardMedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks2025.10.30Text, Image & EHRMulti-Agent Collaboration8 benchmark categories designed for multi-agent reasoning tasks.Project page
HEAL-MedVQALocalizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs2025IJCAIImage & TextGrounded Medical VQA11,000+ samples requiring localization prior to medical answering.Project page
BRIDGEBRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text2025.10.28Text (EHR & Notes)Multitask Evaluation1.4 million samples covering 87 tasks in 9 different languages.Project page
LLMEval-MedLLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation2025.8.31TextClinical QA Validation~1,000 real-world cases validated via physician-in-the-loop audits.GitHub
CSEDBA Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains2025.8.13Text (Clinical Context)Safety-Effectiveness Eval30 criteria across 26 specialties based on expert physician consensus.GitHub
SSG-VQAChallenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study2025.7.8Image & TextSurgical VQA1,300+ scene graph samples focused on instrument-tissue interaction.GitHub
M$^3$-MedM$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding2025.7.6Video & TextMulti-hop Reasoning3,748 instructional videos with 12,747 reasoning-intensive QA pairs.Project page
PET2RepPET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography2025.08.06Text + Image (PET/CT)Report Generation565 whole-body paired pet/ct data combinations with detailed radiology reportGithub
MedTVT-QAMedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis2025.06.23Text + Time Series (ECG) + Image (CXR) + Tabular (Lab Test)Multimodal Medical Reasoning, Multi-disease Diagnosis, Report Generation8,706 multimodal data combinations used to generate QA pairsGithub
HIE-ReasoningVisual and Domain Knowledge for Professional-level Graph-of-Thought Medical Reasoning2025.06.18(ICML2025)Text + Image (MRI) + Clinical DataProfessional-level Medical Reasoning, Neurocognitive Outcome Prediction, Lesion Analysis133 unique MRIs, 749 professional QA pairs, 133 interpretation summariesGithub
ReasonMedReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning2025.06.11Text (Multi-agent CoT, Summary, QA)Medical Reasoning, QA, CoT Fine-tuning370K high-quality samples distilled from 1.75M CoT paths, based on 195K questions from 4 benchmarksHF
Lingshu (Train)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08Text + Image (Multimodal Instruction, VQA, Report)Multimodal Medical QA, Reasoning, Consultation, Report Generation~9.3M training samples from 60+ datasetsProject Page
MedEvalKit (Linshu Test)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08Text + Image (Multimodal Benchmarks)Benchmarking: VQA, Report Generation, Medical Text QA152,066 evaluation samples from 16 benchmarksGithub
MIRIADMIRIAD: Augmenting LLMs with millions of medical query-response pairs2025.06.09Text (Instruction-Response)Medical QA, Retrieval-Augmented Generation (RAG), Hallucination Detection5.8M / 4.4M QA pairsHF
ClinBench-HPBClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases2025.06.04Text (Multiple-choice Questions, Clinical Cases)Medical Question Answering, Hepato-Pancreato-Biliary Clinical Case Diagnosis3,535 MCQs & 337 clinical cases, covering 465+ Hepato-Pancreato-Biliary diseasesProject Page, HF
SurgVLM-DBSurgVLM-DB: A Large-scale Multimodal Surgical Database Comprising Over 1.81 Million Frames with 7.79 Million Conversations2025.06Video + TextMultimodal QA1.81M frames, 7.79M QAsGitHub
EndoBenchEndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis2025.05.29Image + Text (Visual QA, Multimodal Tasks, Multi-level Visual Prompts)Endoscopy Analysis, Medical Imaging, Multimodal Model EvaluationCovers 4 endoscopy scenarios (Gastroscopy, Colonoscopy, Capsule Endoscopy, Surgical Endoscopy); includes 12 clinical tasks and 12 subtasks; 5 levels of visual prompt granularity; 6832 clinically validated VQA samplesHF
MedXpertQAMedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingICML2025Text + Image (Multimodal MCQs)Expert-level Medical QA, Clinical Reasoning, Multimodal Understanding4,460 questions (2,455 text / 2,005 image)HF
MedCaseReasoningMedCaseReasoning: Evaluating and Learning Diagnostic Reasoning from Clinical Case Reports2025.05.20TextDiagnostic Reasoning14,489 QA casesGitHub
vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-2025.06MRI Image + BBox + Multilingual QA (Including 7 languages: vi, en, fr, de, zh, ko, ja)VQA, Lesion Detection, Clinical Reasoning (Tree-of-Thought)12.3k samplesHF
DrVD-BenchDrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?2025.05.30Medical Images + TextVQA, Reasoning, Report Gen.7,789 image–QA pairsGitHub, HF
MedS-InsTowards Evaluating and Building Versatile LLMs for Medicine2025.05TextInstruction Tuning5M instances, 19K instructionsHF
MedS-Bench-2025.05TextClinical Task Benchmark11 task typesHF
MM-SkinMM-Skin:Enhancing Dermatology VLM with an Image-Text Dataset Derived from Textbooks2025.05.09Image + TextOpen-ended VQA (no reasoning)-GitHub
AlphaMed19KBeyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL2025.05.23TextQA Reasoning19K QAsHF
Derm1MDerm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology2025.04.13(ICCV2025)Text + Image (Dermatology)Skin Disease Classification, Concept Identification, Cross-modal Retrieval1,029,761 image-text pairsGithub
GEMeXGEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis2025.03.23 (ICCV2025)Text + Image (Chest X-ray)Medical Visual Question Answering (VQA) for Chest X-ray Diagnosis151,025 images and 1,605,575 QA pairsHF, Github
Surg-396KEndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery2025.03.15Image + Text (Multimodal Instruction, VQA, Grounding, Description)Endoscopic Surgery, Surgical Scene Understanding, Visual QA, Grounded Dialogue396K instruction-image pairs from 41.4K images across 3 datasets (EndoVis, CoPESD, Cholec80) with 5 conversation types and 7 scene understanding tasksGitHub, Data Link
AbdomenAtlas 3.0RadGPT: Constructing 3D Image-Text Tumor Datasets2025.01.08(ICCV2025)Text + 3D Image (Abdominal CT)3D Abdominal CT Report Generation, Tumor Segmentation, Staging, and Analysis9,262 3D CT scans with paired reports, detailing 8,562 tumor instancesGithub, HF
HuatuoGPT-o1 DatasetHuatuoGPT-o1,Towards Medical Complex Reasoning with LLMs2024.12.25Text (Complex CoT, Medical Verifiable Problems, Multi-Step Reasoning)Medical Complex Reasoning, CoT Fine-tuning, Reinforcement LearningContains 40K high-quality medical complex reasoning problems filtered by a medical verifier, based on MedQA-USMLE and MedMCQA medical exam training setsGitHub, HF
PubMedVisionHuatuoGPT-Visionn, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale2024.09.30Image + Text (Multimodal)Medical VQA (Alignment VQA, Instruction-Tuning VQA), Captioning, Summarization1.3M VQA samples from 914,960 filtered PubMed medical images & text (647K + 647K)Hugging Face
PMC-VQAPMC-VQA: Visual Instruction Tuning for Medical VQA2024.08.08Image + TextVQA226,946 QA pairsHF
VQARad-2023.08.07RadiographyVQA315 images, 3,515 QAsOSF
AsclepiusAsclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsACL 2025Image + TextVQA3232 VQA pairs encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasksGitHub
MedTrinity-25MMedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for MedicineICLR 2025Image + TextVQA25M VQA pairs, covering over 25 million images across 10 modalitiesHF
MediConfusionMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsICLR 2025Image + TextVQA176 confusing pairs, a set of two images that share the same question and corresponding answer options, but the correct answer is different for the images.HF
GMAI-MMBenchGMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AINeurIPS 2024Multi-modal (38 types)VQA26K QA pairsHF
PathMMUPathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology2024.03.20Pathology Image + TextMulti-choice, Reasoning33,428 QAs, 24,067 imagesHF
OmniMedVQAOmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMCVPR 2024Multi-modal (12 types)VQA118,010 images, 127,995 QAOpenXLab
CARESA Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsNeurIPS 2024Medical Images + QAOpen/Closed QA41K QA pairsGitHub
MultiMedEvalMultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models2024.02.16Image + TextMulti-task Evaluation6 tasks, 23 datasetsGitHub
medical-o1-reasoning-SFTHuatuoGPT-o1, Towards Medical Complex Reasoning with LLMsACL 2025Medical VQA、ReasoningMedical VQA19.7k QA pairsHF

🇨🇳 中文版本

🔍 项目简介

随着医学视觉语言模型(Med-VLM)及其推理能力研究的持续推进,尤其在 2025 年 3 月至 5 月期间,陆续发布了众多高质量、聚焦于医学推理能力的新型公开数据集,为多模态医疗人工智能的发展提供了坚实的数据基础。为此,我们希望尽可能汇总这一些数据,期待能为该社区提供更便捷的数据访问方式。我们发布了Med-VLM-Bench:

Med-VLM-Bench 致力于汇总并整理这些模型训练与评估的关键资源:

  • ✅ 聚焦 2025 年 3月–2026年发布的新数据集
  • 🧠 重点强调推理能力、多模态理解和问答能力的数据集
  • 🧪 同时覆盖 2023–2024 年的经典Med LLM/VLM benchmark datasets
  • 🔗 提供直接可用的下载链接和开源地址

💡 我们的知识来源有限,欢迎大家通过 Issue 或 PR 推荐更多数据集,我们会第一时间更新!

📌Note: 此外我们的数据集标注时间以相应文章发表时间为准


📊 数据集汇总表

数据集名称论文标题年份 / 会议数据模态任务类型数据规模下载链接
Multi-RADSMulti-RADS Synthetic Radiology Report Dataset...2026.1.6文本 (合成报告)RADS 分类1,600 份合成报告,涵盖 17 种影像发现。Github
Bones and Joints (B&J) BenchmarkThe Illusion of Clinical Reasoning...2025.12.25文本与图像 (X光, CT, MRI)视觉问答 (VQA) 与治疗计划1,245 个问答对,涵盖 7 项临床能力任务。 Hugging Face
MediEvalMediEval: A Unified Medical Benchmark...2025.12.23文本 (电子健康记录与笔记)自然语言推理 (NLI)37,144 条医疗陈述,源自 2,015 次入院记录。GitHub
TCM-BEST4SDTA benchmark dataset for evaluating Syndrome...2025.12.2文本 (中医病历报告)辨证论治总计 600 道题目,包含 300 例临床辨证案例。GitHub
SurgMLLMBenchSurgMLLMBench: A Multimodal Large Language Model...2025.11.26视频与文本手术场景理解10,652 帧标注图像,整合自 5 个手术数据集。project page
MedVisionMedVision: Dataset and Benchmark for Quantitative...2025.11.24图像 (CT, MRI, X光, PET)检测与测量3,080 万个图像-标注对,涵盖 22 个公共数据集。project page
EHRStructEHRStruct: A Comprehensive Benchmark Framework...2025.12.1结构化电子健康记录 (表格)关系数据推理2,200 个针对 11 项任务的特定评估样本。GitHub
TCM-EvalTCM-Eval: An Expert-Level Dynamic and Extensible...2025.12.26文本 (中医知识)专业选择题 (MCQ)6,099 道来自专家级中医考试的题目。dataset
RxSafeBenchRxSafeBench: Identifying Medication Safety Issues...BIBM2025文本 (对话与选择题)用药安全问答2,443 个咨询场景,包含 1,063 例禁忌症案例。GitHub
SemBenchSemBench: A Benchmark for Semantic Query Processing...2025.11.3文本 (知识图谱)语义查询评估1,400 多个用于评估医疗查询引擎的模板。GitHub
XBenchXBench: A Comprehensive Benchmark for Visual...2025.10.22图像 (X光) 与文本定位与解释12,601 例附带定位与文本解释的胸片案例。GitHub
IMBIMB: An Italian Medical Benchmark for Question AnsweringCLIC-it 2025文本 (意大利语)医疗问答与选择题808,506 条数据,包含 78 万条临床对话。GitHub
ViPET-ReportGenToward a Vision-Language Foundation Model for Medical...NeurIPS 2025图像 (3D PET/CT) 与文本报告生成与视觉问答150 万张切片,配对 2,757 份越南语临床报告。GitHub
Neural-MedBenchBeyond Classification Accuracy: Neural-MedBench...2025.12.13文本与图像 (MRI/CT)鉴别诊断120 个专家病例,衍生出 200 个深度推理任务。Project page
AnesSuiteAnesSuite: A Comprehensive Benchmark and Dataset...2025.12.25文本 (麻醉学)专业知识问答4,427 个专注于复杂决策的麻醉学选择题。GitHub
MedQARoMedQARo: A Large-Scale Benchmark for Evaluating...2025.12.31文本 (罗马尼亚语)多语言问答102,646 个罗马尼亚语问答对,涵盖 1,011 名患者。GitHub
TracSumTracSum: A New Benchmark for Aspect-Based Summarization...EMNLP 2025文本 (摘要/笔记)基于维度的医疗摘要500 篇摘要,生成 3,500 个可追溯的摘要-引用对。GitHub
MedAgentBoardMedAgentBoard: Benchmarking Multi-Agent Collaboration...2025.10.30文本, 图像与记录多智能体协作8 个基准类别,专为多智能体推理任务设计。Project page
HEAL-MedVQALocalizing Before Answering: A Hallucination Evaluation...IJCAI 2025图像与文本基于定位的医疗 VQA11,000 多个要求在回答前先定位病理区域的样本。Project page
BRIDGEBRIDGE: Benchmarking Large Language Models for...2025.10.28文本 (电子健康记录与笔记)多任务评估140 万个样本,涵盖 9 种语言的 87 项任务。Project page
LLMEval-MedLLMEval-Med: A Real-world Clinical Benchmark for Medical...2025.8.31文本临床问答验证约 1,000 例通过医师参与审计验证的真实案例。GitHub
CSEDBA Novel Evaluation Benchmark for Medical LLMs...2025.8.13文本 (临床语境)安全性-有效性评估基于专家共识的 30 项标准,涵盖 26 个专科。GitHub
SSG-VQAChallenging Vision-Language Models with Surgical Data...2025.7.8图像与文本手术视觉问答 (VQA)1,300 多个专注于器械-组织交互的场景图样本。GitHub
M$^3$-MedM$^3$-Med: A Benchmark for Multi-lingual, Multi-modal...2025.7.6视频与文本多跳推理3,748 段教学视频,包含 12,747 个重推理问答对。Project page
PET2RepPET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography2025.08.06Text + Image (PET/CT)报告生成565全身配对PET/CT数据组合与详细的放射学报告Github
MedTVT-QAMedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis2025.06.23文本 + 时间序列 (心电图) + 图像 (胸部X光) + 表格 (血液检测)多模态医疗推理、多病种诊断、报告生成用于生成QA对的8,706组多模态数据组合Github
HIE-ReasoningVisual and Domain Knowledge for Professional-level Graph-of-Thought Medical Reasoning2025.06.18(ICML2025)文本 + 图像 (MRI) + 临床数据专业级医疗推理、神经认知结局预测、病灶分析133个独立MRI、749个专业问答对、133份解读摘要Github
ReasonMedReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning2025.06.11文本(多代理推理、多步总结、医学问答)医学推理、问答、链式思维微调从 175 万条 CoT 路径中精炼出的 37 万高质量样本,覆盖来自 4 个基准的 19.5 万问题HF
Lingshu(Train)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08文本 + 图像(多模态指令、VQA、报告)多模态医学问答、推理、问诊、报告生成约 930 万训练样本,来自 60+ 数据集Project Page
MedEvalKit(Linshu test)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08文本 + 图像(多模态评测基准)基准评测:VQA、报告生成、医学文本问答共 152,066 个评估样本,来自 16 个基准数据集Github
MIRIADMIRIAD: Augmenting LLMs with millions of medical query-response pairs2025.06.09文本(指令-回答对)医学问答、RAG 检索增强、幻觉检测582 万 / 448 万HF
ClinBench-HPBClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases2025.06.04文本 (选择题, 临床病例)医学问答, 临床病例诊断3,535道选择题和337个临床病例, 覆盖465+种肝胆胰疾病Project Page, HF
SurgVLM-DBSurgVLM-DB: A Large-Scale Multimodal Surgical Database2025.06视频 + 文本多模态问答1.81M帧, 7.79M对话GitHub
EndoBenchEndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis2025.05.29图像+文本(视觉问答、多模态任务、多层次视觉提示)内镜分析、医学影像、多模态模型评估覆盖胃镜、结肠镜、胶囊内镜和手术内镜 4 大场景;包含 12 个临床任务及 12 个次任务;5 种视觉提示粒度;6832 个经过临床验证的 VQA 样本HF
MedXpertQAMedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingICML2025文本 + 图像(多模态选择题)专家级医学问答、临床推理、多模态理解共 4,460 题(文本 2,455 / 图像 2,005)HF
MedCaseReasoningMedCaseReasoning: Evaluating and Learning Diagnostic Reasoning from Clinical Case Reports2025.05.20文本诊断推理14,489问答对GitHub
vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-2025.06MRI 图像 + BBox + 多语种问答(vi, en, fr, de, zh, ko, ja)医学VQA、病灶检测、临床推理(Tree-of-Thought)12,325 条样本HF
DrVD-BenchDrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?2025.05.30医学图像 + 文本医学VQA、推理、报告生成7,789图文QA对GitHub, HF
MedS-InsTowards Evaluating and Building Versatile LLMs for Medicine2025.05文本指令微调5M样本, 19K指令HF
MedS-Bench-2025.05文本临床任务评估11大类任务HF
MM-SkinMM-Skin:Enhancing Dermatology VLM with an Image-Text Dataset Derived from Textbooks2025.05.09图像 + 文本开放式VQA(无推理)-GitHub
AlphaMed19KBeyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL2025.05.23文本推理问答19K问答对HF
Derm1MDerm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology2025.04.13(ICCV2025)文本 + 图像(皮肤病学)皮肤病分类、概念识别、跨模态检索1,029,761个图文对Github
GEMeXGEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis2025.03.23 (ICCV2025)文本 + 图像(胸部X光片)用于胸部X光诊断的医疗视觉问答(VQA)151,025张图片和1,605,575个问答HF, Github
Surg-396KEndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery2025.03.15图像 + 文本(多模态指令、问答、目标定位、详细描述)内窥镜外科、手术场景理解、视觉问答、定位对话来自 EndoVis、CoPESD 和 Cholec80 的 41,400 张图像,生成 396,000 图文对,覆盖 5 种对话类型与 7 类手术理解任务GitHub Data link
AbdomenAtlas 3.0RadGPT: Constructing 3D Image-Text Tumor Datasets2025.01.08 (ICCV2025)文本 + 3D图像 (腹部CT)3D腹部CT报告生成、肿瘤分割、分期与分析9,262组3D CT扫描及配对报告,包含8,562个肿瘤实例Github, HF
HuatuoGPT-o1 DatasetHuatuoGPT-o1,Towards Medical Complex Reasoning with LLMs2024.12.25文本(复杂链式思维、医学验证题、多步推理)医学复杂推理、链式思维微调、强化学习包含 40K 经医学验证器筛选的高质量医学复杂推理问题,基于 MedQA-USMLE 和 MedMCQA 医学考试训练集GitHub, HF
PubMedVisionHuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale2024.09.30图像 + 文本(多模态)医学视觉问答(VQA)、图文对齐、指令微调、描述生成等130 万 VQA 样本,来自 PubMed 中筛选的 91.5 万医学图像与上下文(647K + 647K)HF
PMC-VQAPMC-VQA: Visual Instruction Tuning for Medical VQA2024.08.08图像 + 文本医学VQA226,946问答对HF
VQARad-2023.08.07放射图像VQA315图像, 3515问答OSF
AsclepiusAsclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsACL 2025图像 + 文本医学VQA3232条问答对,涵盖 15 个医学专业,分为 3 个主要类别和 8 个子类别的临床任务GitHub
MedTrinity-25MMedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for MedicineICLR 2025图像 + 文本医学VQA大规模医学多模态数据集,涵盖 10 种模态的2500万张图像,为65种疾病提供多粒度注释HF
MediConfusionMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsICLR 2025图像 + 文本医学VQA由 176 个令人困惑的对组成。混淆对是一组两张图像,它们共享相同的问题和相应的答案选项,但图像的正确答案不同。HF
GMAI-MMBenchGMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AINeurIPS 2024多模态(38种)医学VQA26K问答对HF
PathMMUPathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Patholog2024.03.20病理图像 + 文本选择题+推理33,428问答, 24,067图像HF
OmniMedVQAOmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMCVPR 2024多模态(12种)医学VQA118,010图像, 127,995问答OpenXLab
CARESA Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsNeurIPS 2024医学图像+问答开放与封闭问答41K问答对GitHub
MultiMedEvalMultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models2024.02.16图文多模态多任务评估6任务, 23数据集GitHub
medical-o1-reasoning-SFTHuatuoGPT-o1, Towards Medical Complex Reasoning with LLMsACL 2025医学VQA、推理医学VQA19.7k问答对HF

📬 联系我们 / Issues & Contact

如有任何问题欢迎提交 Issue 或通过邮件联系:


👥 合作者(Contributors)


⭐ Star 趋势图 (Star History)

Star History Chart



⭐ Star 本项目支持我们持续更新,欢迎 PR 和建议交流!

Contributors

yezanting

46 commits

Saint-lsy

3 commits

lailainan

1 commits

WangRongsheng

1 commits

yezanting/Med-VLM-Bench-Summary

A Curated Benchmark Repository for Medical Vision-Language Models

200

53 commits

updated Jan 21, 2026

See the code

README

🧠 Med-VLM-Bench: A Curated Benchmark Repository for Medical Vision-Language Models

📚 A comprehensive summary of recent benchmarks for evaluating and training Medical Vision-Language Models (Med-VLMs)


👨‍💻 Contributors

  • 🧑‍🔬 Zanting Ye
    Southern Medical University
    📧 yzt2861252880@gmail.com

  • 🧑‍🔬 Xu Han
    Shanghai Jiao Tong University
    📧 hanxv8826@gmail.com

  • 🧑‍🔬 Xiaolong Niu
    Southern Medical University

  • 🧑‍🔬 Zian Wang
    Shanghai Jiao Tong University

  • 🧑‍🔬 Shengyuan Liu
    The Chinese University of Hong Kong
    📧 liushengyuan@link.cuhk.edu.hk

  • 🧑‍🔬 Xin Liu
    Southern Medical University
    📧 lx10230114@gmail.com

  • 👨‍🏫 Lijun Lu
    Southern Medical University


📊 GitHub Stats

Stars Forks License Last Commit


🔍 Project Overview

With the continuous advancement of research on Medical Vision-Language Models (Med-VLMs) and their reasoning capabilities, a number of high-quality, publicly available datasets focusing on medical reasoning have been released between March and May 2025. These datasets provide a solid foundation for the development of multimodal medical AI systems.

Med-VLM-Bench is a curated, continuously updated repository of the latest and most important datasets for training and evaluating medical LLMs and VLMs. This project focuses on:

  • ✅ Reasoning-centric multimodal benchmarks
  • 📅 Latest datasets published in Mar 2025–2026
  • 🧠 Foundational datasets from 2023–2024
  • 🔗 Direct access to dataset links or HuggingFace/GitHub repositories

💡 Our knowledge is limited to public sources. We welcome community contributions — feel free to open an issue to share new datasets, and we will update promptly.

📌Note: The annotation time of the dataset is based on the publication time of the corresponding article.


📢 News

🌟 Latest Updates

  • 2026-01-21: 🎉 Added some recent datasets and benchmarks!Check it out for detailed information and download links!
  • 2025-06-29: 🎉 Added new datasets/benchmarks AbdomenAtlas 3.0 (ICCV2025), Derm1M(ICCV2025), MedTVT-R1, , GEMeX(ICCV2025) and HIE-Reasoning(ICML2025). Check it out for detailed information and download links!
  • 2025-06-18: 🎉 Added new datasets/benchmarks Lingshu, ReasonMed. Check it out for detailed information and download links!
  • 2025-06-11: 🎉 Added some recent datasets and benchmarks!
  • 2025-06-11: 🎉 Create our github project!

📊 Dataset Summary Table

Dataset NamePaper TitleYear / VenueData ModalityTask TypeSizeDownload Link
Multi-RADSMulti-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models2026.1.6Text (Synthetic Reports)RADS Classification1,600 synthetic reports covering 17 distinct imaging findings.Github
Bones and Joints (B&J) BenchmarkThe Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision Language Models for Clinical Competency2025.12.25Text & Image (X-ray, CT, MRI)VQA & Treatment Planning1,245 question-answer pairs spanning 7 clinical competency tasks. Hugging Face
MediEvalMediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs2025.12.23Text (EHR & Notes)Natural Language Inference37,144 medical statements derived from 2,015 hospital admissions.GitHub
TCM-BEST4SDTA benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models2025.12.2Text (TCM Case Reports)Syndrome Differentiation600 total questions including 300 clinical syndrome differentiation cases.GitHub
SurgMLLMBenchSurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding2025.11.26Video & TextSurgical Scene Understanding10,652 frames with annotations integrated from 5 surgical datasets.project page
MedVisionMedVision: Dataset and Benchmark for Quantitative Medical Image Analysis2025.11.24Image (CT, MRI, X-ray, PET)Detection & Measurement30.8 million image-annotation pairs across 22 public datasets.project page
EHRStructEHRStruct: A Comprehensive Benchmark Framework for Evaluating Large Language Models on Structured Electronic Health Record Tasks2025.12.1Structured EHR (Tables)Relational Data Reasoning2,200 task-specific samples across 11 data and knowledge tasks.GitHub
TCM-EvalTCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine2025.12.26Text (TCM Knowledge)Professional MCQ6,099 questions from expert-level Chinese medical examinations.dataset
RxSafeBenchRxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated ConsultationBIBM2025Text (Dialogue & MCQ)Medication Safety QA2,443 consultation scenarios including 1,063 contraindication cases.GitHub
SemBenchSemBench: A Benchmark for Semantic Query Processing Engines2025.11.3Text (Knowledge Graph)Semantic Query Evaluation1,400+ SPARQL templates for evaluating medical query engines.GitHub
XBenchXBench: A Comprehensive Benchmark for Visual-Language Explanations in Chest Radiography2025.10.22Image (X-ray) & TextGrounding & Explanation12,601 chest X-ray cases with localization and textual explanations.GitHub
IMBIMB: An Italian Medical Benchmark for Question AnsweringCLIC-it 2025Text (Italian)Medical QA & MCQA808,506 items featuring 782,644 clinical Italian conversations.GitHub
ViPET-ReportGenToward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report GenerationNeurIPS 2025Image (3D PET/CT) & TextReport Generation & VQA1.5 million slices paired with 2,757 Vietnamese clinical reports.GitHub
Neural-MedBenchBeyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks2025.12.13Text & Image (MRI/CT)Differential Diagnosis120 expert cases resulting in 200 depth-of-reasoning tasks.Project page
MedQARoMedQARo: A Large-Scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian2025.12.31Text (Romanian)Multilingual QA102,646 Romanian QA pairs covering 1,011 clinical patients.GitHub
AnesSuiteAnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs2025.12.25Text (Anesthesiology)Specialized Knowledge QA4,427 anesthesiology MCQ items focused on complex decision-making.GitHub
TracSumTracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domainthe 2025 Conference on Empirical Methods in Natural Language ProcessingText (Abstracts/Notes)Aspect-Based Summarization500 abstracts resulting in 3,500 summary-citation traceable pairs.GitHub
MedAgentBoardMedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks2025.10.30Text, Image & EHRMulti-Agent Collaboration8 benchmark categories designed for multi-agent reasoning tasks.Project page
HEAL-MedVQALocalizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs2025IJCAIImage & TextGrounded Medical VQA11,000+ samples requiring localization prior to medical answering.Project page
BRIDGEBRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text2025.10.28Text (EHR & Notes)Multitask Evaluation1.4 million samples covering 87 tasks in 9 different languages.Project page
LLMEval-MedLLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation2025.8.31TextClinical QA Validation~1,000 real-world cases validated via physician-in-the-loop audits.GitHub
CSEDBA Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains2025.8.13Text (Clinical Context)Safety-Effectiveness Eval30 criteria across 26 specialties based on expert physician consensus.GitHub
SSG-VQAChallenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study2025.7.8Image & TextSurgical VQA1,300+ scene graph samples focused on instrument-tissue interaction.GitHub
M$^3$-MedM$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding2025.7.6Video & TextMulti-hop Reasoning3,748 instructional videos with 12,747 reasoning-intensive QA pairs.Project page
PET2RepPET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography2025.08.06Text + Image (PET/CT)Report Generation565 whole-body paired pet/ct data combinations with detailed radiology reportGithub
MedTVT-QAMedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis2025.06.23Text + Time Series (ECG) + Image (CXR) + Tabular (Lab Test)Multimodal Medical Reasoning, Multi-disease Diagnosis, Report Generation8,706 multimodal data combinations used to generate QA pairsGithub
HIE-ReasoningVisual and Domain Knowledge for Professional-level Graph-of-Thought Medical Reasoning2025.06.18(ICML2025)Text + Image (MRI) + Clinical DataProfessional-level Medical Reasoning, Neurocognitive Outcome Prediction, Lesion Analysis133 unique MRIs, 749 professional QA pairs, 133 interpretation summariesGithub
ReasonMedReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning2025.06.11Text (Multi-agent CoT, Summary, QA)Medical Reasoning, QA, CoT Fine-tuning370K high-quality samples distilled from 1.75M CoT paths, based on 195K questions from 4 benchmarksHF
Lingshu (Train)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08Text + Image (Multimodal Instruction, VQA, Report)Multimodal Medical QA, Reasoning, Consultation, Report Generation~9.3M training samples from 60+ datasetsProject Page
MedEvalKit (Linshu Test)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08Text + Image (Multimodal Benchmarks)Benchmarking: VQA, Report Generation, Medical Text QA152,066 evaluation samples from 16 benchmarksGithub
MIRIADMIRIAD: Augmenting LLMs with millions of medical query-response pairs2025.06.09Text (Instruction-Response)Medical QA, Retrieval-Augmented Generation (RAG), Hallucination Detection5.8M / 4.4M QA pairsHF
ClinBench-HPBClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases2025.06.04Text (Multiple-choice Questions, Clinical Cases)Medical Question Answering, Hepato-Pancreato-Biliary Clinical Case Diagnosis3,535 MCQs & 337 clinical cases, covering 465+ Hepato-Pancreato-Biliary diseasesProject Page, HF
SurgVLM-DBSurgVLM-DB: A Large-scale Multimodal Surgical Database Comprising Over 1.81 Million Frames with 7.79 Million Conversations2025.06Video + TextMultimodal QA1.81M frames, 7.79M QAsGitHub
EndoBenchEndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis2025.05.29Image + Text (Visual QA, Multimodal Tasks, Multi-level Visual Prompts)Endoscopy Analysis, Medical Imaging, Multimodal Model EvaluationCovers 4 endoscopy scenarios (Gastroscopy, Colonoscopy, Capsule Endoscopy, Surgical Endoscopy); includes 12 clinical tasks and 12 subtasks; 5 levels of visual prompt granularity; 6832 clinically validated VQA samplesHF
MedXpertQAMedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingICML2025Text + Image (Multimodal MCQs)Expert-level Medical QA, Clinical Reasoning, Multimodal Understanding4,460 questions (2,455 text / 2,005 image)HF
MedCaseReasoningMedCaseReasoning: Evaluating and Learning Diagnostic Reasoning from Clinical Case Reports2025.05.20TextDiagnostic Reasoning14,489 QA casesGitHub
vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-2025.06MRI Image + BBox + Multilingual QA (Including 7 languages: vi, en, fr, de, zh, ko, ja)VQA, Lesion Detection, Clinical Reasoning (Tree-of-Thought)12.3k samplesHF
DrVD-BenchDrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?2025.05.30Medical Images + TextVQA, Reasoning, Report Gen.7,789 image–QA pairsGitHub, HF
MedS-InsTowards Evaluating and Building Versatile LLMs for Medicine2025.05TextInstruction Tuning5M instances, 19K instructionsHF
MedS-Bench-2025.05TextClinical Task Benchmark11 task typesHF
MM-SkinMM-Skin:Enhancing Dermatology VLM with an Image-Text Dataset Derived from Textbooks2025.05.09Image + TextOpen-ended VQA (no reasoning)-GitHub
AlphaMed19KBeyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL2025.05.23TextQA Reasoning19K QAsHF
Derm1MDerm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology2025.04.13(ICCV2025)Text + Image (Dermatology)Skin Disease Classification, Concept Identification, Cross-modal Retrieval1,029,761 image-text pairsGithub
GEMeXGEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis2025.03.23 (ICCV2025)Text + Image (Chest X-ray)Medical Visual Question Answering (VQA) for Chest X-ray Diagnosis151,025 images and 1,605,575 QA pairsHF, Github
Surg-396KEndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery2025.03.15Image + Text (Multimodal Instruction, VQA, Grounding, Description)Endoscopic Surgery, Surgical Scene Understanding, Visual QA, Grounded Dialogue396K instruction-image pairs from 41.4K images across 3 datasets (EndoVis, CoPESD, Cholec80) with 5 conversation types and 7 scene understanding tasksGitHub, Data Link
AbdomenAtlas 3.0RadGPT: Constructing 3D Image-Text Tumor Datasets2025.01.08(ICCV2025)Text + 3D Image (Abdominal CT)3D Abdominal CT Report Generation, Tumor Segmentation, Staging, and Analysis9,262 3D CT scans with paired reports, detailing 8,562 tumor instancesGithub, HF
HuatuoGPT-o1 DatasetHuatuoGPT-o1,Towards Medical Complex Reasoning with LLMs2024.12.25Text (Complex CoT, Medical Verifiable Problems, Multi-Step Reasoning)Medical Complex Reasoning, CoT Fine-tuning, Reinforcement LearningContains 40K high-quality medical complex reasoning problems filtered by a medical verifier, based on MedQA-USMLE and MedMCQA medical exam training setsGitHub, HF
PubMedVisionHuatuoGPT-Visionn, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale2024.09.30Image + Text (Multimodal)Medical VQA (Alignment VQA, Instruction-Tuning VQA), Captioning, Summarization1.3M VQA samples from 914,960 filtered PubMed medical images & text (647K + 647K)Hugging Face
PMC-VQAPMC-VQA: Visual Instruction Tuning for Medical VQA2024.08.08Image + TextVQA226,946 QA pairsHF
VQARad-2023.08.07RadiographyVQA315 images, 3,515 QAsOSF
AsclepiusAsclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsACL 2025Image + TextVQA3232 VQA pairs encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasksGitHub
MedTrinity-25MMedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for MedicineICLR 2025Image + TextVQA25M VQA pairs, covering over 25 million images across 10 modalitiesHF
MediConfusionMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsICLR 2025Image + TextVQA176 confusing pairs, a set of two images that share the same question and corresponding answer options, but the correct answer is different for the images.HF
GMAI-MMBenchGMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AINeurIPS 2024Multi-modal (38 types)VQA26K QA pairsHF
PathMMUPathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology2024.03.20Pathology Image + TextMulti-choice, Reasoning33,428 QAs, 24,067 imagesHF
OmniMedVQAOmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMCVPR 2024Multi-modal (12 types)VQA118,010 images, 127,995 QAOpenXLab
CARESA Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsNeurIPS 2024Medical Images + QAOpen/Closed QA41K QA pairsGitHub
MultiMedEvalMultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models2024.02.16Image + TextMulti-task Evaluation6 tasks, 23 datasetsGitHub
medical-o1-reasoning-SFTHuatuoGPT-o1, Towards Medical Complex Reasoning with LLMsACL 2025Medical VQA、ReasoningMedical VQA19.7k QA pairsHF

🇨🇳 中文版本

🔍 项目简介

随着医学视觉语言模型(Med-VLM)及其推理能力研究的持续推进,尤其在 2025 年 3 月至 5 月期间,陆续发布了众多高质量、聚焦于医学推理能力的新型公开数据集,为多模态医疗人工智能的发展提供了坚实的数据基础。为此,我们希望尽可能汇总这一些数据,期待能为该社区提供更便捷的数据访问方式。我们发布了Med-VLM-Bench:

Med-VLM-Bench 致力于汇总并整理这些模型训练与评估的关键资源:

  • ✅ 聚焦 2025 年 3月–2026年发布的新数据集
  • 🧠 重点强调推理能力、多模态理解和问答能力的数据集
  • 🧪 同时覆盖 2023–2024 年的经典Med LLM/VLM benchmark datasets
  • 🔗 提供直接可用的下载链接和开源地址

💡 我们的知识来源有限,欢迎大家通过 Issue 或 PR 推荐更多数据集,我们会第一时间更新!

📌Note: 此外我们的数据集标注时间以相应文章发表时间为准


📊 数据集汇总表

数据集名称论文标题年份 / 会议数据模态任务类型数据规模下载链接
Multi-RADSMulti-RADS Synthetic Radiology Report Dataset...2026.1.6文本 (合成报告)RADS 分类1,600 份合成报告,涵盖 17 种影像发现。Github
Bones and Joints (B&J) BenchmarkThe Illusion of Clinical Reasoning...2025.12.25文本与图像 (X光, CT, MRI)视觉问答 (VQA) 与治疗计划1,245 个问答对,涵盖 7 项临床能力任务。 Hugging Face
MediEvalMediEval: A Unified Medical Benchmark...2025.12.23文本 (电子健康记录与笔记)自然语言推理 (NLI)37,144 条医疗陈述,源自 2,015 次入院记录。GitHub
TCM-BEST4SDTA benchmark dataset for evaluating Syndrome...2025.12.2文本 (中医病历报告)辨证论治总计 600 道题目,包含 300 例临床辨证案例。GitHub
SurgMLLMBenchSurgMLLMBench: A Multimodal Large Language Model...2025.11.26视频与文本手术场景理解10,652 帧标注图像,整合自 5 个手术数据集。project page
MedVisionMedVision: Dataset and Benchmark for Quantitative...2025.11.24图像 (CT, MRI, X光, PET)检测与测量3,080 万个图像-标注对,涵盖 22 个公共数据集。project page
EHRStructEHRStruct: A Comprehensive Benchmark Framework...2025.12.1结构化电子健康记录 (表格)关系数据推理2,200 个针对 11 项任务的特定评估样本。GitHub
TCM-EvalTCM-Eval: An Expert-Level Dynamic and Extensible...2025.12.26文本 (中医知识)专业选择题 (MCQ)6,099 道来自专家级中医考试的题目。dataset
RxSafeBenchRxSafeBench: Identifying Medication Safety Issues...BIBM2025文本 (对话与选择题)用药安全问答2,443 个咨询场景,包含 1,063 例禁忌症案例。GitHub
SemBenchSemBench: A Benchmark for Semantic Query Processing...2025.11.3文本 (知识图谱)语义查询评估1,400 多个用于评估医疗查询引擎的模板。GitHub
XBenchXBench: A Comprehensive Benchmark for Visual...2025.10.22图像 (X光) 与文本定位与解释12,601 例附带定位与文本解释的胸片案例。GitHub
IMBIMB: An Italian Medical Benchmark for Question AnsweringCLIC-it 2025文本 (意大利语)医疗问答与选择题808,506 条数据,包含 78 万条临床对话。GitHub
ViPET-ReportGenToward a Vision-Language Foundation Model for Medical...NeurIPS 2025图像 (3D PET/CT) 与文本报告生成与视觉问答150 万张切片,配对 2,757 份越南语临床报告。GitHub
Neural-MedBenchBeyond Classification Accuracy: Neural-MedBench...2025.12.13文本与图像 (MRI/CT)鉴别诊断120 个专家病例,衍生出 200 个深度推理任务。Project page
AnesSuiteAnesSuite: A Comprehensive Benchmark and Dataset...2025.12.25文本 (麻醉学)专业知识问答4,427 个专注于复杂决策的麻醉学选择题。GitHub
MedQARoMedQARo: A Large-Scale Benchmark for Evaluating...2025.12.31文本 (罗马尼亚语)多语言问答102,646 个罗马尼亚语问答对,涵盖 1,011 名患者。GitHub
TracSumTracSum: A New Benchmark for Aspect-Based Summarization...EMNLP 2025文本 (摘要/笔记)基于维度的医疗摘要500 篇摘要,生成 3,500 个可追溯的摘要-引用对。GitHub
MedAgentBoardMedAgentBoard: Benchmarking Multi-Agent Collaboration...2025.10.30文本, 图像与记录多智能体协作8 个基准类别,专为多智能体推理任务设计。Project page
HEAL-MedVQALocalizing Before Answering: A Hallucination Evaluation...IJCAI 2025图像与文本基于定位的医疗 VQA11,000 多个要求在回答前先定位病理区域的样本。Project page
BRIDGEBRIDGE: Benchmarking Large Language Models for...2025.10.28文本 (电子健康记录与笔记)多任务评估140 万个样本,涵盖 9 种语言的 87 项任务。Project page
LLMEval-MedLLMEval-Med: A Real-world Clinical Benchmark for Medical...2025.8.31文本临床问答验证约 1,000 例通过医师参与审计验证的真实案例。GitHub
CSEDBA Novel Evaluation Benchmark for Medical LLMs...2025.8.13文本 (临床语境)安全性-有效性评估基于专家共识的 30 项标准,涵盖 26 个专科。GitHub
SSG-VQAChallenging Vision-Language Models with Surgical Data...2025.7.8图像与文本手术视觉问答 (VQA)1,300 多个专注于器械-组织交互的场景图样本。GitHub
M$^3$-MedM$^3$-Med: A Benchmark for Multi-lingual, Multi-modal...2025.7.6视频与文本多跳推理3,748 段教学视频,包含 12,747 个重推理问答对。Project page
PET2RepPET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography2025.08.06Text + Image (PET/CT)报告生成565全身配对PET/CT数据组合与详细的放射学报告Github
MedTVT-QAMedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis2025.06.23文本 + 时间序列 (心电图) + 图像 (胸部X光) + 表格 (血液检测)多模态医疗推理、多病种诊断、报告生成用于生成QA对的8,706组多模态数据组合Github
HIE-ReasoningVisual and Domain Knowledge for Professional-level Graph-of-Thought Medical Reasoning2025.06.18(ICML2025)文本 + 图像 (MRI) + 临床数据专业级医疗推理、神经认知结局预测、病灶分析133个独立MRI、749个专业问答对、133份解读摘要Github
ReasonMedReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning2025.06.11文本(多代理推理、多步总结、医学问答)医学推理、问答、链式思维微调从 175 万条 CoT 路径中精炼出的 37 万高质量样本,覆盖来自 4 个基准的 19.5 万问题HF
Lingshu(Train)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08文本 + 图像(多模态指令、VQA、报告)多模态医学问答、推理、问诊、报告生成约 930 万训练样本,来自 60+ 数据集Project Page
MedEvalKit(Linshu test)Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning2025.06.08文本 + 图像(多模态评测基准)基准评测:VQA、报告生成、医学文本问答共 152,066 个评估样本,来自 16 个基准数据集Github
MIRIADMIRIAD: Augmenting LLMs with millions of medical query-response pairs2025.06.09文本(指令-回答对)医学问答、RAG 检索增强、幻觉检测582 万 / 448 万HF
ClinBench-HPBClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases2025.06.04文本 (选择题, 临床病例)医学问答, 临床病例诊断3,535道选择题和337个临床病例, 覆盖465+种肝胆胰疾病Project Page, HF
SurgVLM-DBSurgVLM-DB: A Large-Scale Multimodal Surgical Database2025.06视频 + 文本多模态问答1.81M帧, 7.79M对话GitHub
EndoBenchEndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis2025.05.29图像+文本(视觉问答、多模态任务、多层次视觉提示)内镜分析、医学影像、多模态模型评估覆盖胃镜、结肠镜、胶囊内镜和手术内镜 4 大场景;包含 12 个临床任务及 12 个次任务;5 种视觉提示粒度;6832 个经过临床验证的 VQA 样本HF
MedXpertQAMedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingICML2025文本 + 图像(多模态选择题)专家级医学问答、临床推理、多模态理解共 4,460 题(文本 2,455 / 图像 2,005)HF
MedCaseReasoningMedCaseReasoning: Evaluating and Learning Diagnostic Reasoning from Clinical Case Reports2025.05.20文本诊断推理14,489问答对GitHub
vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-2025.06MRI 图像 + BBox + 多语种问答(vi, en, fr, de, zh, ko, ja)医学VQA、病灶检测、临床推理(Tree-of-Thought)12,325 条样本HF
DrVD-BenchDrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?2025.05.30医学图像 + 文本医学VQA、推理、报告生成7,789图文QA对GitHub, HF
MedS-InsTowards Evaluating and Building Versatile LLMs for Medicine2025.05文本指令微调5M样本, 19K指令HF
MedS-Bench-2025.05文本临床任务评估11大类任务HF
MM-SkinMM-Skin:Enhancing Dermatology VLM with an Image-Text Dataset Derived from Textbooks2025.05.09图像 + 文本开放式VQA(无推理)-GitHub
AlphaMed19KBeyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL2025.05.23文本推理问答19K问答对HF
Derm1MDerm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology2025.04.13(ICCV2025)文本 + 图像(皮肤病学)皮肤病分类、概念识别、跨模态检索1,029,761个图文对Github
GEMeXGEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis2025.03.23 (ICCV2025)文本 + 图像(胸部X光片)用于胸部X光诊断的医疗视觉问答(VQA)151,025张图片和1,605,575个问答HF, Github
Surg-396KEndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery2025.03.15图像 + 文本(多模态指令、问答、目标定位、详细描述)内窥镜外科、手术场景理解、视觉问答、定位对话来自 EndoVis、CoPESD 和 Cholec80 的 41,400 张图像,生成 396,000 图文对,覆盖 5 种对话类型与 7 类手术理解任务GitHub Data link
AbdomenAtlas 3.0RadGPT: Constructing 3D Image-Text Tumor Datasets2025.01.08 (ICCV2025)文本 + 3D图像 (腹部CT)3D腹部CT报告生成、肿瘤分割、分期与分析9,262组3D CT扫描及配对报告,包含8,562个肿瘤实例Github, HF
HuatuoGPT-o1 DatasetHuatuoGPT-o1,Towards Medical Complex Reasoning with LLMs2024.12.25文本(复杂链式思维、医学验证题、多步推理)医学复杂推理、链式思维微调、强化学习包含 40K 经医学验证器筛选的高质量医学复杂推理问题,基于 MedQA-USMLE 和 MedMCQA 医学考试训练集GitHub, HF
PubMedVisionHuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale2024.09.30图像 + 文本(多模态)医学视觉问答(VQA)、图文对齐、指令微调、描述生成等130 万 VQA 样本,来自 PubMed 中筛选的 91.5 万医学图像与上下文(647K + 647K)HF
PMC-VQAPMC-VQA: Visual Instruction Tuning for Medical VQA2024.08.08图像 + 文本医学VQA226,946问答对HF
VQARad-2023.08.07放射图像VQA315图像, 3515问答OSF
AsclepiusAsclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsACL 2025图像 + 文本医学VQA3232条问答对,涵盖 15 个医学专业,分为 3 个主要类别和 8 个子类别的临床任务GitHub
MedTrinity-25MMedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for MedicineICLR 2025图像 + 文本医学VQA大规模医学多模态数据集,涵盖 10 种模态的2500万张图像,为65种疾病提供多粒度注释HF
MediConfusionMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation modelsICLR 2025图像 + 文本医学VQA由 176 个令人困惑的对组成。混淆对是一组两张图像,它们共享相同的问题和相应的答案选项,但图像的正确答案不同。HF
GMAI-MMBenchGMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AINeurIPS 2024多模态(38种)医学VQA26K问答对HF
PathMMUPathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Patholog2024.03.20病理图像 + 文本选择题+推理33,428问答, 24,067图像HF
OmniMedVQAOmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMCVPR 2024多模态(12种)医学VQA118,010图像, 127,995问答OpenXLab
CARESA Comprehensive Benchmark of Trustworthiness in Medical Vision Language ModelsNeurIPS 2024医学图像+问答开放与封闭问答41K问答对GitHub
MultiMedEvalMultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models2024.02.16图文多模态多任务评估6任务, 23数据集GitHub
medical-o1-reasoning-SFTHuatuoGPT-o1, Towards Medical Complex Reasoning with LLMsACL 2025医学VQA、推理医学VQA19.7k问答对HF

📬 联系我们 / Issues & Contact

如有任何问题欢迎提交 Issue 或通过邮件联系:


👥 合作者(Contributors)


⭐ Star 趋势图 (Star History)

Star History Chart



⭐ Star 本项目支持我们持续更新,欢迎 PR 和建议交流!

Contributors

yezanting

46 commits

Saint-lsy

3 commits

lailainan

1 commits

WangRongsheng

1 commits