shiym2000/MedM-VL-2D-3B-en

Model

MedM-VL-2D-3B-en

0

20 commits

1 linked in READMEs

updated Oct 11, 2025

See the code

README

MedM-VL-2D-3B-en

Introduction

A 2D medical LVLM trained on 2D medical images and English medical texts, enabling tasks such as report generation, VQA, referring expression comprehension (REC), referring expression generation (REG) and image classification.

Config
Image encodergoogle/siglip-base-patch16-256-multilingual
ConnectorMLP (2-layer)
LLMQwen/Qwen2.5-3B-Instruct
Image resolution256*256
Sequence length2048

Evaluation

BenchmarkMed-FlamingoLLaVA-MedRadFMMedM-VL-2D-3B-en
MedMNISTderma0.0120.2580.0510.810
MedMNISTorgan0.0890.6680.1890.791
MedPix0.0810.151-0.087
MIMIC-CXR0.2330.2040.0680.222
PathVQA0.3340.3780.2480.634
SAMedidentify-0.458-0.637
SAMedrefer-0.086-0.225
SLAKEidentify-0.272-0.349
SLAKErefer-0.041-0.261
SLAKEvqa0.2150.3370.8170.812

Quickstart

Please refer to MedM-VL.

Citation

@inproceedings{shi2025medm,
  title={Medm-vl: What makes a good medical lvlm?},
  author={Shi, Yiming and Yang, Shaoshuai and Zhu, Xun and Wang, Haoyu and Fu, Xiangling and Li, Miao and Wu, Ji},
  booktitle={International Workshop on Agentic AI for Medicine},
  pages={290--299},
  year={2025},
  organization={Springer}
}
2D_Medical_LVLMs
conversational
image-text-to-text
lvlm
safetensors

Contributors

shiym2000

20 commits

shiym2000/MedM-VL-2D-3B-en

Model

MedM-VL-2D-3B-en

0

20 commits

1 linked in READMEs

updated Oct 11, 2025

See the code

README

MedM-VL-2D-3B-en

Introduction

A 2D medical LVLM trained on 2D medical images and English medical texts, enabling tasks such as report generation, VQA, referring expression comprehension (REC), referring expression generation (REG) and image classification.

Config
Image encodergoogle/siglip-base-patch16-256-multilingual
ConnectorMLP (2-layer)
LLMQwen/Qwen2.5-3B-Instruct
Image resolution256*256
Sequence length2048

Evaluation

BenchmarkMed-FlamingoLLaVA-MedRadFMMedM-VL-2D-3B-en
MedMNISTderma0.0120.2580.0510.810
MedMNISTorgan0.0890.6680.1890.791
MedPix0.0810.151-0.087
MIMIC-CXR0.2330.2040.0680.222
PathVQA0.3340.3780.2480.634
SAMedidentify-0.458-0.637
SAMedrefer-0.086-0.225
SLAKEidentify-0.272-0.349
SLAKErefer-0.041-0.261
SLAKEvqa0.2150.3370.8170.812

Quickstart

Please refer to MedM-VL.

Citation

@inproceedings{shi2025medm,
  title={Medm-vl: What makes a good medical lvlm?},
  author={Shi, Yiming and Yang, Shaoshuai and Zhu, Xun and Wang, Haoyu and Fu, Xiangling and Li, Miao and Wu, Ji},
  booktitle={International Workshop on Agentic AI for Medicine},
  pages={290--299},
  year={2025},
  organization={Springer}
}
2D_Medical_LVLMs
conversational
image-text-to-text
lvlm
safetensors

Contributors

shiym2000

20 commits