shiym2000/MedM-VL-CT-Chest-3B-en

Model

MedM-VL-CT-Chest-3B-en

0

20 commits

1 linked in READMEs

updated Oct 12, 2025

See the code

README

MedM-VL-CT-Chest-3B-en

Introduction

A 3D medical LVLM trained on 3D chest CT volumes and English medical texts (CT-RATE), enabling tasks such as report generation and medical VQA.

Config
Image encodergoogle/siglip-base-patch16-256-multilingual
ConnectorCross-Attention + MLP (2-layer)
LLMQwen/Qwen2.5-3B-Instruct
Image resolution32*256*256
Sequence length2048

Evaluation

TaskCT-CHATMedM-VL-CT-Chest (3D)MedM-VL-CT-Chest (2D+Avg)MedM-VL-CT-Chest (2D+Attn)
Long answer0.4820.6190.6220.623
Short answer0.2740.6580.6640.667
Multiple choice0.8380.9240.9200.925
Report generation0.3950.4190.4410.439

Quickstart

Please refer to MedM-VL.

Citation

@inproceedings{shi2025medm,
  title={Medm-vl: What makes a good medical lvlm?},
  author={Shi, Yiming and Yang, Shaoshuai and Zhu, Xun and Wang, Haoyu and Fu, Xiangling and Li, Miao and Wu, Ji},
  booktitle={International Workshop on Agentic AI for Medicine},
  pages={290--299},
  year={2025},
  organization={Springer}
}
3D_Medical_LVLMs
conversational
image-text-to-text
lvlm
safetensors

Contributors

shiym2000

20 commits

shiym2000/MedM-VL-CT-Chest-3B-en

Model

MedM-VL-CT-Chest-3B-en

0

20 commits

1 linked in READMEs

updated Oct 12, 2025

See the code

README

MedM-VL-CT-Chest-3B-en

Introduction

A 3D medical LVLM trained on 3D chest CT volumes and English medical texts (CT-RATE), enabling tasks such as report generation and medical VQA.

Config
Image encodergoogle/siglip-base-patch16-256-multilingual
ConnectorCross-Attention + MLP (2-layer)
LLMQwen/Qwen2.5-3B-Instruct
Image resolution32*256*256
Sequence length2048

Evaluation

TaskCT-CHATMedM-VL-CT-Chest (3D)MedM-VL-CT-Chest (2D+Avg)MedM-VL-CT-Chest (2D+Attn)
Long answer0.4820.6190.6220.623
Short answer0.2740.6580.6640.667
Multiple choice0.8380.9240.9200.925
Report generation0.3950.4190.4410.439

Quickstart

Please refer to MedM-VL.

Citation

@inproceedings{shi2025medm,
  title={Medm-vl: What makes a good medical lvlm?},
  author={Shi, Yiming and Yang, Shaoshuai and Zhu, Xun and Wang, Haoyu and Fu, Xiangling and Li, Miao and Wu, Ji},
  booktitle={International Workshop on Agentic AI for Medicine},
  pages={290--299},
  year={2025},
  organization={Springer}
}
3D_Medical_LVLMs
conversational
image-text-to-text
lvlm
safetensors

Contributors

shiym2000

20 commits