Welcome to the official code repository for GMAI-VL & GMAI-VL-5.5M (AAAI 2026). This repository provides the training code needed for reproducing the results and furthering research in medical AI through vision-language models.
GMAI-VL-5.5M is systematically constructed from 219 professional medical data sources using an Annotation-Guided Data Generation pipeline β all text is grounded in expert annotations, not hallucinated.
| Subset | Size | Type | Description |
|---|---|---|---|
| GMAI-MM-Caption | 1.7M | Multimodal | High-quality medical image captions |
| GMAI-MM-Percept | 1.3M | Multimodal | Medical image classification & segmentation labels |
| GMAI-MM-Instrunct | 0.9M | Multimodal | Medical image analysis instruction QA |
| GMAI-Text-Single | 1.0M | Text-only | Single-turn medical text QA |
| GMAI-Text-Multi | 0.7M | Text-only | Multi-turn medical text QA |
| Dataset | Scale | Modalities | Languages | Traceable | Source |
|---|---|---|---|---|---|
| PathVQA | 32.7K | Pathology | EN | β | Textbooks |
| MIMIC-CXR | 227K | X-Ray | EN | β | Hospital |
| PMC-OA | 1.65M | Multi | EN | β | PubMed |
| PubMedVision | 1.29M | Multi | EN&CN | β | PubMed |
| GMAI-VL-5.5M | 5.5M | Multi (13) | EN&CN | β | 219 medical datasets |
π Download: π€ Hugging Face
GMAI-VL is built on the LLaVA architecture with InternLM2.5-7B (LLM) + CLIP Vision Encoder + MLP Projector, trained using a three-stage progressive strategy:
| Stage | Strategy | Trainable Components | Learning Rate |
|---|---|---|---|
| Stage I | Shallow Alignment | Projector only | 1e-3 |
| Stage II | Deep Alignment | Projector + Vision Encoder | 1e-4 |
| Stage III | Instruction Tuning | Full model | 1e-5 |
| Model | Params | OmniMedVQA | GMAI-MMBench | MMMU H&M | VQA-RAD |
|---|---|---|---|---|---|
| InternVL2 | 40B | 78.70 | β | β | β |
| HuatuoGPT-Vision | 34B | 73.23 | β | 50.3 | β |
| medgemma | 4B | 81.92 | β | 43.3 | β |
| GMAI-VL | 7B | 88.48 | 62.43 | 51.3 | 66.3 |
With only 7B parameters, GMAI-VL outperforms models with 34B+ parameters on multiple benchmarks, demonstrating the value of high-quality data + progressive training.
A Gradio-based visualization demo is available in demo/gmai_vl_demo. It supports medical image upload, preset multimodal examples, task/body-system controls, and screenshot-friendly model responses.
cd demo/gmai_vl_demo
pip install -r requirements-gmai-vl-demo.txt
export GMAI_VL_MODEL_PATH=./model_weight
PORT=10083 bash run_demo.sh
To set up the environment for training, please follow the instructions in the xtuner repository. Ensure that all dependencies are correctly installed.
git clone https://github.com/uni-medical/GMAI-VL
cd GMAI-VL
pip install -e .
Prepare your training data in the format shown in examples/example_data.json. To support multiple JSON datasets with different sampling ratios, use the format defined in examples/example_list.json.
Example structure for example_list.json:
{
"FILE1": {
"image_dir": "",
"annotations": "examples/example_data.json",
"sample_ratio": 1.0,
"length": 38276,
"data_augment": true
},
...
}
Here is a reference script to start training:
SRUN_ARGS=${SRUN_ARGS:-""}
export PYTHONPATH="$(pwd):$(pwd)/../"
export MASTER_PORT=34220
export NPROC_PER_NODE=${KUBERNETES_CONTAINER_RESOURCE_GPU}
export PORT=${MASTER_PORT}
export NNODES=${WORLD_SIZE}
export NODE_RANK=${RANK}
export ADDR=${MASTER_ADDR}
export HF_HOME=~/.cache
export USE_TRITON_KERNEL=1
torchrun --nnodes=${WORLD_SIZE} --node_rank=${RANK} --nproc_per_node=${NPROC_PER_NODE} --master_addr=${MASTER_ADDR} --master_port=${MASTER_PORT} \
tools/llava/llava_sft.py \
--freeze-llm \
--freeze-vit \
--llava $base_model_path/ \
--chat-template internlm2 \
--datasets examples/examples_list.json \
--num-workers 6 \
--mirco-batch-size $MIRCO_BATCH_SIZE \
--global-batch-size $GLOBAL_BATCH_SIZE \
--lr 1e-5 \
--wd 0 \
--dset-pack-level soft \
--shard-strategy 'zero2' \
--group-by-length \
--resume \
--max-length 2048 \
--checkpoint-interval 500 \
--log-interval 10 \
--work-dir $work_dir/ \
--dset-cache-dir $work_dir/cache/ \
--dset-from-cache
π‘ Note: Our training follows a multi-stage strategy. At each stage, different components (e.g., LLM or ViT) may be frozen or fine-tuned. Please adjust flags such as --freeze-llm, --freeze-vit, and learning rates accordingly, as described in the paper.
For evaluation, we use VLMEvalKit. GMAI-VL has been evaluated on 7 mainstream medical multimodal benchmarks: OmniMedVQA, GMAI-MMBench, MMMU (Health & Medicine), VQA-RAD, SLAKE, PMC-VQA, and PathVQA. See the paper for full results.
| Resource | Link |
|---|---|
| Paper (AAAI 2026) | Proceedings |
| Paper (arXiv) | arXiv:2411.14522 |
| GMAI-VL-5.5M Dataset | π€ Hugging Face |
| Training Code | This repository |
| Training Framework | XTuner |
| Evaluation Toolkit | VLMEvalKit |
For inquiries, collaboration opportunities, or access requests, feel free to reach out via email or open a GitHub issue.
Thank you for your interest and support!
We would like to express our sincere gratitude to the open-source community. Our work is built upon the excellent contributions of the following toolkits:
We deeply appreciate the efforts of the developers and contributors behind these projects.
If you find our work helpful in your research, please consider citing us:
@inproceedings{li2026gmai,
title={Gmai-vl \& gmai-vl-5.5 m: A large vision-language model and a comprehensive multimodal dataset towards general medical ai},
author={Li, Tianbin and Su, Yanzhou and Li, Wei and Fu, Bin and Chen, Zhe and Huang, Ziyan and Wang, Guoan and Ma, Chenglong and Chen, Ying and Hu, Ming and others},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={40},
number={28},
pages={23177--23185},
year={2026}
}
Python
100.0%
Welcome to the official code repository for GMAI-VL & GMAI-VL-5.5M (AAAI 2026). This repository provides the training code needed for reproducing the results and furthering research in medical AI through vision-language models.
GMAI-VL-5.5M is systematically constructed from 219 professional medical data sources using an Annotation-Guided Data Generation pipeline β all text is grounded in expert annotations, not hallucinated.
| Subset | Size | Type | Description |
|---|---|---|---|
| GMAI-MM-Caption | 1.7M | Multimodal | High-quality medical image captions |
| GMAI-MM-Percept | 1.3M | Multimodal | Medical image classification & segmentation labels |
| GMAI-MM-Instrunct | 0.9M | Multimodal | Medical image analysis instruction QA |
| GMAI-Text-Single | 1.0M | Text-only | Single-turn medical text QA |
| GMAI-Text-Multi | 0.7M | Text-only | Multi-turn medical text QA |
| Dataset | Scale | Modalities | Languages | Traceable | Source |
|---|---|---|---|---|---|
| PathVQA | 32.7K | Pathology | EN | β | Textbooks |
| MIMIC-CXR | 227K | X-Ray | EN | β | Hospital |
| PMC-OA | 1.65M | Multi | EN | β | PubMed |
| PubMedVision | 1.29M | Multi | EN&CN | β | PubMed |
| GMAI-VL-5.5M | 5.5M | Multi (13) | EN&CN | β | 219 medical datasets |
π Download: π€ Hugging Face
GMAI-VL is built on the LLaVA architecture with InternLM2.5-7B (LLM) + CLIP Vision Encoder + MLP Projector, trained using a three-stage progressive strategy:
| Stage | Strategy | Trainable Components | Learning Rate |
|---|---|---|---|
| Stage I | Shallow Alignment | Projector only | 1e-3 |
| Stage II | Deep Alignment | Projector + Vision Encoder | 1e-4 |
| Stage III | Instruction Tuning | Full model | 1e-5 |
| Model | Params | OmniMedVQA | GMAI-MMBench | MMMU H&M | VQA-RAD |
|---|---|---|---|---|---|
| InternVL2 | 40B | 78.70 | β | β | β |
| HuatuoGPT-Vision | 34B | 73.23 | β | 50.3 | β |
| medgemma | 4B | 81.92 | β | 43.3 | β |
| GMAI-VL | 7B | 88.48 | 62.43 | 51.3 | 66.3 |
With only 7B parameters, GMAI-VL outperforms models with 34B+ parameters on multiple benchmarks, demonstrating the value of high-quality data + progressive training.
A Gradio-based visualization demo is available in demo/gmai_vl_demo. It supports medical image upload, preset multimodal examples, task/body-system controls, and screenshot-friendly model responses.
cd demo/gmai_vl_demo
pip install -r requirements-gmai-vl-demo.txt
export GMAI_VL_MODEL_PATH=./model_weight
PORT=10083 bash run_demo.sh
To set up the environment for training, please follow the instructions in the xtuner repository. Ensure that all dependencies are correctly installed.
git clone https://github.com/uni-medical/GMAI-VL
cd GMAI-VL
pip install -e .
Prepare your training data in the format shown in examples/example_data.json. To support multiple JSON datasets with different sampling ratios, use the format defined in examples/example_list.json.
Example structure for example_list.json:
{
"FILE1": {
"image_dir": "",
"annotations": "examples/example_data.json",
"sample_ratio": 1.0,
"length": 38276,
"data_augment": true
},
...
}
Here is a reference script to start training:
SRUN_ARGS=${SRUN_ARGS:-""}
export PYTHONPATH="$(pwd):$(pwd)/../"
export MASTER_PORT=34220
export NPROC_PER_NODE=${KUBERNETES_CONTAINER_RESOURCE_GPU}
export PORT=${MASTER_PORT}
export NNODES=${WORLD_SIZE}
export NODE_RANK=${RANK}
export ADDR=${MASTER_ADDR}
export HF_HOME=~/.cache
export USE_TRITON_KERNEL=1
torchrun --nnodes=${WORLD_SIZE} --node_rank=${RANK} --nproc_per_node=${NPROC_PER_NODE} --master_addr=${MASTER_ADDR} --master_port=${MASTER_PORT} \
tools/llava/llava_sft.py \
--freeze-llm \
--freeze-vit \
--llava $base_model_path/ \
--chat-template internlm2 \
--datasets examples/examples_list.json \
--num-workers 6 \
--mirco-batch-size $MIRCO_BATCH_SIZE \
--global-batch-size $GLOBAL_BATCH_SIZE \
--lr 1e-5 \
--wd 0 \
--dset-pack-level soft \
--shard-strategy 'zero2' \
--group-by-length \
--resume \
--max-length 2048 \
--checkpoint-interval 500 \
--log-interval 10 \
--work-dir $work_dir/ \
--dset-cache-dir $work_dir/cache/ \
--dset-from-cache
π‘ Note: Our training follows a multi-stage strategy. At each stage, different components (e.g., LLM or ViT) may be frozen or fine-tuned. Please adjust flags such as --freeze-llm, --freeze-vit, and learning rates accordingly, as described in the paper.
For evaluation, we use VLMEvalKit. GMAI-VL has been evaluated on 7 mainstream medical multimodal benchmarks: OmniMedVQA, GMAI-MMBench, MMMU (Health & Medicine), VQA-RAD, SLAKE, PMC-VQA, and PathVQA. See the paper for full results.
| Resource | Link |
|---|---|
| Paper (AAAI 2026) | Proceedings |
| Paper (arXiv) | arXiv:2411.14522 |
| GMAI-VL-5.5M Dataset | π€ Hugging Face |
| Training Code | This repository |
| Training Framework | XTuner |
| Evaluation Toolkit | VLMEvalKit |
For inquiries, collaboration opportunities, or access requests, feel free to reach out via email or open a GitHub issue.
Thank you for your interest and support!
We would like to express our sincere gratitude to the open-source community. Our work is built upon the excellent contributions of the following toolkits:
We deeply appreciate the efforts of the developers and contributors behind these projects.
If you find our work helpful in your research, please consider citing us:
@inproceedings{li2026gmai,
title={Gmai-vl \& gmai-vl-5.5 m: A large vision-language model and a comprehensive multimodal dataset towards general medical ai},
author={Li, Tianbin and Su, Yanzhou and Li, Wei and Fu, Bin and Chen, Zhe and Huang, Ziyan and Wang, Guoan and Ma, Chenglong and Chen, Ying and Hu, Ming and others},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={40},
number={28},
pages={23177--23185},
year={2026}
}
Python
100.0%