A Vision-Centric Multimodal LLM for Holistic Medical Understanding (8B Parameter Version)
ClinFusion-8B is the highly efficient, lightweight version of the ClinFusion multimodal large language model series. Built upon the Qwen3-VL-8B-Instruct base, ClinFusion-8B integrates a composition and cascaded vision encoder framework (featuring DINOv2 and ConvNeXt) to deliver state-of-the-art medical and clinical understanding.
Despite its compact size, ClinFusion-8B strikes an optimal balance between exceptional clinical reasoning capabilities and manageable resource consumption, making it ideal for edge deployment, real-time clinical assistants, and institutions with limited computational budgets.
.nii.gz CT/MRI files).To set up the environment, run inference, or evaluate ClinFusion-8B on your own data, please refer directly to our official GitHub Repository.
The GitHub repository provides comprehensive, step-by-step instructions for:
uv.On our newly introduced MedIF-Bench (instruction-following evaluation) and clinical report generation benchmarks, ClinFusion-8B outperforms prominent open-source counterparts (e.g., Hulu-Med, Lingshu-8B) and exhibits competitive performance compared to proprietary medical systems.
Refer to the main ClinFusion Paper for complete evaluation metrics.
If you find ClinFusion useful in your research, please consider citing:
@article{yuan2026clinfusion,
title={ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding},
author={Yuan, Hangjie and Qian, Yichen and Tang, Zhiwei and Xu, Xianzhe and Wu, Lirong and Yang, Sicheng and Wang, Jinwang and Wang, Pengju and Zeng, Zhitao and Han, Yizeng and Xing, Yan and Luo, Shengxuan and Feng, Tao and Xie, Qing and Yao, Weigen and Yang, Yi and Liu, Zuozhu and Tang, Jiasheng and Wang, Shaocheng and Wang, Jitao and Dong, Jiahong and Chen, Weihua and Xu, Feng and Wang, Fan},
journal={arXiv preprint arXiv:2607.24743},
year={2026}
}
Built with β€οΈ by Alibaba DAMO Academy. We acknowledge the authors of Qwen3-VL, DINOv2, and OpenCLIP for their excellent foundation models.
25 commits
A Vision-Centric Multimodal LLM for Holistic Medical Understanding (8B Parameter Version)
ClinFusion-8B is the highly efficient, lightweight version of the ClinFusion multimodal large language model series. Built upon the Qwen3-VL-8B-Instruct base, ClinFusion-8B integrates a composition and cascaded vision encoder framework (featuring DINOv2 and ConvNeXt) to deliver state-of-the-art medical and clinical understanding.
Despite its compact size, ClinFusion-8B strikes an optimal balance between exceptional clinical reasoning capabilities and manageable resource consumption, making it ideal for edge deployment, real-time clinical assistants, and institutions with limited computational budgets.
.nii.gz CT/MRI files).To set up the environment, run inference, or evaluate ClinFusion-8B on your own data, please refer directly to our official GitHub Repository.
The GitHub repository provides comprehensive, step-by-step instructions for:
uv.On our newly introduced MedIF-Bench (instruction-following evaluation) and clinical report generation benchmarks, ClinFusion-8B outperforms prominent open-source counterparts (e.g., Hulu-Med, Lingshu-8B) and exhibits competitive performance compared to proprietary medical systems.
Refer to the main ClinFusion Paper for complete evaluation metrics.
If you find ClinFusion useful in your research, please consider citing:
@article{yuan2026clinfusion,
title={ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding},
author={Yuan, Hangjie and Qian, Yichen and Tang, Zhiwei and Xu, Xianzhe and Wu, Lirong and Yang, Sicheng and Wang, Jinwang and Wang, Pengju and Zeng, Zhitao and Han, Yizeng and Xing, Yan and Luo, Shengxuan and Feng, Tao and Xie, Qing and Yao, Weigen and Yang, Yi and Liu, Zuozhu and Tang, Jiasheng and Wang, Shaocheng and Wang, Jitao and Dong, Jiahong and Chen, Weihua and Xu, Feng and Wang, Fan},
journal={arXiv preprint arXiv:2607.24743},
year={2026}
}
Built with β€οΈ by Alibaba DAMO Academy. We acknowledge the authors of Qwen3-VL, DINOv2, and OpenCLIP for their excellent foundation models.
25 commits