PRAISELab-PicusLab/ExtendedMMMED

🩺 Extended MMMED is a benchmark dataset for evaluating Vision-Language Models (VLMs) on medical multiple-choice question answering (MCQA) tasks. 🏥It features 955 real-world medical questions from Spanish MIR and Italian SSM exams, available in 🇪🇸 Spanish, 🇬🇧 English, and 🇮🇹 Italian.

0

stars

4

commits

Jupyter Notebook

primary language

Apr 15, 2026

updated

huggingface.co/datasets/praiselab-picuslab/MMMED/viewer/extended/
benchmarking
dataset
multimodal
vision-language-model
visual
visual-question-answering

README

🏥 Extended Multilingual Multimodal Medical Exam Dataset for Visual Question Answering in Healthcare

GitHub Repository MMMED License

The Extended Multilingual Multimodal Medical Exam Dataset (Extended MMMED) is the new, larger release of MMMED for evaluating Vision-Language Models (VLMs) on medical multiple-choice question answering (MCQA) tasks.

Compared to the original benchmark, this extension substantially increases the number of questions and updates the benchmark with 28 tested VLMs (general-purpose, medical-specialized, and closed-source) across Spanish, English, and Italian.

The dataset includes challenging, real-world medical content from Médico Interno Residente (MIR - Spain) and Scuole di Specializzazione in Medicina (SSM - Italy) exam settings, with heterogeneous diagnostic images and clinically grounded questions.

🔓 How to Access the Dataset

You can access the dataset via Hugging Face. Follow these steps to download it:

⚠️ Disclaimer: This dataset contains medical images that may be sensitive for some users. Viewer discretion is advised, especially if the content may evoke strong emotional reactions or be distressing.

Login using e.g. huggingface-cli login to access this dataset

from datasets import load_dataset

# Login using e.g. `huggingface-cli login` to access this dataset
ds = load_dataset("praiselab-picuslab/MMMED", split="extended")

🌟 Key Features:

  • Languages: 🇪🇸 Spanish, 🇬🇧 English, 🇮🇹 Italian
  • Scale: 955 questions per language (2,865 total samples)
  • Medical Content: Questions based on Spanish residency exam material
  • Image Types: Diagnostic medical images (e.g., CT scans, X-rays)
  • Categories: 26 medical categories per language
  • Multimodal: Each question comes with a medical image 📸
  • Benchmarking: 28 VLMs evaluated in multilingual settings

🔄 Dataset Workflow

Here is the general workflow for building the MMMED dataset for Vision-Language Model (VLM) evaluation:

image

📊 Dataset Overview

The Extended MMMED benchmark contains 955 questions for each language and is organized into 26 medical categories per language. The table below reports updated corpus statistics used in the new study.

Statistic🇪🇸 Spanish🇬🇧 English🇮🇹 Italian
# Questions955955955
# Categories262626
Last Update202620262026
Avg. Option Length4.203.874.03
Max. Option Length737674
Total Question Tokens*42,26241,32739,716
Avg. Question Length43.4540.8140.78
Max. Question Length264258254

* Token counts are computed with the preprocessing pipeline used in this repository (SpaCy-based analysis notebooks).

image

🖼️ Image Types

Categorization of Image Types in the Extended MMMED Dataset. This figure presents the four main categories of images included in the dataset and their respective distributions.

image

Example MMCQA

Each multimodal multiple-choice question-answer (MMCQA) pair integrates three essential components with the following structure:

  • Category: $C$
  • Question: $Q$
  • Image URL: $I$
  • Answer Options: $O$
  • Correct Answer: 💡

Here’s an illustrative example of multimodal QA in three languages:

image

🔍 VLMs Evaluated in the Extended Benchmark (28 Models)

The following table reports architecture details for all tested models.

ModelTypeParam (B)Language ModelVision Model
medvlm-r1Medical2Qwen2-2BQwenViT
maira-2Medical7Vicuna-7B-v1.5RAD-DINO-MAIRA-2
medgemma-4b-itMedical4Gemma-3-4BMedSigLIP-448
llava-med-v1.5-7bMedical7Mistral-7BCLIP ViT-L/14
chexagent-8bMedical8Phi-2-2BSigLIP-Large
medgemma-27b-itMedical27Gemma-3-27BMedSigLIP-448
minicpm-v-2.6General2.6Qwen2-7BSigLip-400M
paligemma-3b-mix-448General3Gemma-2BSigLIP-So400m/14
paligemma2-3b-mix-448General3Gemma-2-2BSigLIP-So400m/14
deepseek-vl2-tinyGeneral3DeepSeekMoE-3BSigLIP-400M
qwen2.5-vl-3bGeneral3Qwen2.5-3BQwenViT
phi-3.5-visionGeneral4Phi-3.5CLIP ViT-L/14
gemma-3-4b-itGeneral4Gemma-3-4BSigLIP
llava-v1.5-7bGeneral7Vicuna-7B-v1.5CLIP ViT-L/14
deepseek-vl-7bGeneral7DeepSeek-LLM-7BSigLIP + SAM
qwen2.5-vl-7bGeneral7Qwen2.5-7BQwenViT
qwen2-vl-7bGeneral8Qwen2-7BQwenViT
qwen3-vl-8bGeneral8Qwen3-8BQwenViT
internvl2.5-8bGeneral8InternLM2.5-7BInternViT-300M
paligemma2-10b-mix-448General10Gemma-2-9BSigLIP-So400m/14
pixtral-12bGeneral12Mistral-Nemo-12BPixtral ViT
gemma-3-27b-itGeneral27Gemma-3-27BSigLIP
qwen3-vl-30bGeneral30Qwen3-30BQwenViT
qwen2.5-vl-32bGeneral32Qwen2.5-32BQwenViT
qwen2.5-vl-72bGeneral72Qwen2.5-72BQwenViT
claude-4-sonnetClosedUnknownClosed-SourceClosed-Source
gpt-5-miniClosedUnknownClosed-SourceClosed-Source
gemini-2.5-flashClosedUnknownClosed-SourceClosed-Source

📈 VLM Performance on MMMED

The following figure presents the overall multilingual performance trend.

image

For complete analysis outputs (tables and publication-quality figures), see:

  • Analysis/analysis_output/tables/accuracy_table.csv
  • Analysis/analysis_output/tables/summary_table.csv
  • Analysis/analysis_output/figures/

🖋️ Original MMMED Citation

Please cite also the original work as follows:

@inproceedings{riccio2025multilingual,
  title={A Multilingual Multimodal Medical Examination Dataset for Visual Question Answering in Healthcare},
  author={Riccio, Giuseppe and Romano, Antonio and Barone, Mariano and Orlando, Gian Marco and Russo, Diego and Postiglione, Marco and La Gatta, Valerio and Moscato, Vincenzo},
  booktitle={2025 IEEE 38th International Symposium on Computer-Based Medical Systems (CBMS)},
  pages={435--440},
  year={2025},
  organization={IEEE Computer Society}
}

🌐 Notes

Dataset Usage: The dataset is intended for academic and research purposes only. It is not recommended for clinical decision-making or commercial use.

👨‍💻 This project was developed by Mariano Barone, Francesco Di Serio, Giuseppe Riccio, Antonio Romano, Vincenzo Moscato, and Marco Postiglione University of Naples, Federico II

📜 License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

CC BY-NC 4.0

Contributors

LaErre9

2 commits

PRAISELab-PicusLab/ExtendedMMMED

🩺 Extended MMMED is a benchmark dataset for evaluating Vision-Language Models (VLMs) on medical multiple-choice question answering (MCQA) tasks. 🏥It features 955 real-world medical questions from Spanish MIR and Italian SSM exams, available in 🇪🇸 Spanish, 🇬🇧 English, and 🇮🇹 Italian.

0

stars

4

commits

Jupyter Notebook

primary language

Apr 15, 2026

updated

huggingface.co/datasets/praiselab-picuslab/MMMED/viewer/extended/
benchmarking
dataset
multimodal
vision-language-model
visual
visual-question-answering

README

🏥 Extended Multilingual Multimodal Medical Exam Dataset for Visual Question Answering in Healthcare

GitHub Repository MMMED License

The Extended Multilingual Multimodal Medical Exam Dataset (Extended MMMED) is the new, larger release of MMMED for evaluating Vision-Language Models (VLMs) on medical multiple-choice question answering (MCQA) tasks.

Compared to the original benchmark, this extension substantially increases the number of questions and updates the benchmark with 28 tested VLMs (general-purpose, medical-specialized, and closed-source) across Spanish, English, and Italian.

The dataset includes challenging, real-world medical content from Médico Interno Residente (MIR - Spain) and Scuole di Specializzazione in Medicina (SSM - Italy) exam settings, with heterogeneous diagnostic images and clinically grounded questions.

🔓 How to Access the Dataset

You can access the dataset via Hugging Face. Follow these steps to download it:

⚠️ Disclaimer: This dataset contains medical images that may be sensitive for some users. Viewer discretion is advised, especially if the content may evoke strong emotional reactions or be distressing.

Login using e.g. huggingface-cli login to access this dataset

from datasets import load_dataset

# Login using e.g. `huggingface-cli login` to access this dataset
ds = load_dataset("praiselab-picuslab/MMMED", split="extended")

🌟 Key Features:

  • Languages: 🇪🇸 Spanish, 🇬🇧 English, 🇮🇹 Italian
  • Scale: 955 questions per language (2,865 total samples)
  • Medical Content: Questions based on Spanish residency exam material
  • Image Types: Diagnostic medical images (e.g., CT scans, X-rays)
  • Categories: 26 medical categories per language
  • Multimodal: Each question comes with a medical image 📸
  • Benchmarking: 28 VLMs evaluated in multilingual settings

🔄 Dataset Workflow

Here is the general workflow for building the MMMED dataset for Vision-Language Model (VLM) evaluation:

image

📊 Dataset Overview

The Extended MMMED benchmark contains 955 questions for each language and is organized into 26 medical categories per language. The table below reports updated corpus statistics used in the new study.

Statistic🇪🇸 Spanish🇬🇧 English🇮🇹 Italian
# Questions955955955
# Categories262626
Last Update202620262026
Avg. Option Length4.203.874.03
Max. Option Length737674
Total Question Tokens*42,26241,32739,716
Avg. Question Length43.4540.8140.78
Max. Question Length264258254

* Token counts are computed with the preprocessing pipeline used in this repository (SpaCy-based analysis notebooks).

image

🖼️ Image Types

Categorization of Image Types in the Extended MMMED Dataset. This figure presents the four main categories of images included in the dataset and their respective distributions.

image

Example MMCQA

Each multimodal multiple-choice question-answer (MMCQA) pair integrates three essential components with the following structure:

  • Category: $C$
  • Question: $Q$
  • Image URL: $I$
  • Answer Options: $O$
  • Correct Answer: 💡

Here’s an illustrative example of multimodal QA in three languages:

image

🔍 VLMs Evaluated in the Extended Benchmark (28 Models)

The following table reports architecture details for all tested models.

ModelTypeParam (B)Language ModelVision Model
medvlm-r1Medical2Qwen2-2BQwenViT
maira-2Medical7Vicuna-7B-v1.5RAD-DINO-MAIRA-2
medgemma-4b-itMedical4Gemma-3-4BMedSigLIP-448
llava-med-v1.5-7bMedical7Mistral-7BCLIP ViT-L/14
chexagent-8bMedical8Phi-2-2BSigLIP-Large
medgemma-27b-itMedical27Gemma-3-27BMedSigLIP-448
minicpm-v-2.6General2.6Qwen2-7BSigLip-400M
paligemma-3b-mix-448General3Gemma-2BSigLIP-So400m/14
paligemma2-3b-mix-448General3Gemma-2-2BSigLIP-So400m/14
deepseek-vl2-tinyGeneral3DeepSeekMoE-3BSigLIP-400M
qwen2.5-vl-3bGeneral3Qwen2.5-3BQwenViT
phi-3.5-visionGeneral4Phi-3.5CLIP ViT-L/14
gemma-3-4b-itGeneral4Gemma-3-4BSigLIP
llava-v1.5-7bGeneral7Vicuna-7B-v1.5CLIP ViT-L/14
deepseek-vl-7bGeneral7DeepSeek-LLM-7BSigLIP + SAM
qwen2.5-vl-7bGeneral7Qwen2.5-7BQwenViT
qwen2-vl-7bGeneral8Qwen2-7BQwenViT
qwen3-vl-8bGeneral8Qwen3-8BQwenViT
internvl2.5-8bGeneral8InternLM2.5-7BInternViT-300M
paligemma2-10b-mix-448General10Gemma-2-9BSigLIP-So400m/14
pixtral-12bGeneral12Mistral-Nemo-12BPixtral ViT
gemma-3-27b-itGeneral27Gemma-3-27BSigLIP
qwen3-vl-30bGeneral30Qwen3-30BQwenViT
qwen2.5-vl-32bGeneral32Qwen2.5-32BQwenViT
qwen2.5-vl-72bGeneral72Qwen2.5-72BQwenViT
claude-4-sonnetClosedUnknownClosed-SourceClosed-Source
gpt-5-miniClosedUnknownClosed-SourceClosed-Source
gemini-2.5-flashClosedUnknownClosed-SourceClosed-Source

📈 VLM Performance on MMMED

The following figure presents the overall multilingual performance trend.

image

For complete analysis outputs (tables and publication-quality figures), see:

  • Analysis/analysis_output/tables/accuracy_table.csv
  • Analysis/analysis_output/tables/summary_table.csv
  • Analysis/analysis_output/figures/

🖋️ Original MMMED Citation

Please cite also the original work as follows:

@inproceedings{riccio2025multilingual,
  title={A Multilingual Multimodal Medical Examination Dataset for Visual Question Answering in Healthcare},
  author={Riccio, Giuseppe and Romano, Antonio and Barone, Mariano and Orlando, Gian Marco and Russo, Diego and Postiglione, Marco and La Gatta, Valerio and Moscato, Vincenzo},
  booktitle={2025 IEEE 38th International Symposium on Computer-Based Medical Systems (CBMS)},
  pages={435--440},
  year={2025},
  organization={IEEE Computer Society}
}

🌐 Notes

Dataset Usage: The dataset is intended for academic and research purposes only. It is not recommended for clinical decision-making or commercial use.

👨‍💻 This project was developed by Mariano Barone, Francesco Di Serio, Giuseppe Riccio, Antonio Romano, Vincenzo Moscato, and Marco Postiglione University of Naples, Federico II

📜 License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

CC BY-NC 4.0

Contributors

LaErre9

2 commits

Languages

Jupyter Notebook

74.6%

Python

25.4%