A Vision-Centric Multimodal LLM for Holistic Medical Understanding (32B Parameter Flagship)
ClinFusion-32B is the flagship, high-performance model of the ClinFusion multimodal large language model series. Built on Qwen3-VL-32B-Instruct, ClinFusion-32B is engineered for complex clinical reasoning, high-fidelity medical report generation, and fine-grained visual-textual instruction-following.
By fusing DINOv2 and ConvNeXt vision encoders with a 32-billion parameter language backbone through our custom Cascade Spatial-Aware Locality Fusion operator, ClinFusion-32B delivers unmatched comprehension of medical cases, approaching and sometimes exceeding proprietary models like GPT-5.2 and Gemini-3-Flash.
.nii.gz CT/MRI files).To set up the environment, run inference, or evaluate ClinFusion-32B on your own data, please refer directly to our official GitHub Repository.
The GitHub repository provides comprehensive, step-by-step instructions for:
uv.ClinFusion-32B sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks.
If you find ClinFusion useful in your research, please consider citing:
@article{yuan2026ClinFusion,
title={ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding},
author={Yuan, Hangjie and Qian, Yichen and Tang, Zhiwei and Xu, Xianzhe and Wu, Lirong and Yang, Sicheng and Wang, Jinwang and Wang, Pengju and Zeng, Zhitao and Han, Yizeng and Xing, Yan and Luo, Shengxuan and Feng, Tao and Xie, Qing and Yao, Weigen and Yang, Yi and Liu, Zuozhu and Tang, Jiasheng and Wang, Shaocheng and Wang, Jitao and Dong, Jiahong and Chen, Weihua and Xu, Feng and Wang, Fan},
journal={arXiv preprint arXiv:2607.24743},
year={2026}
}
Built with ❤️ by Alibaba DAMO Academy. Special thanks to the open-source community behind Qwen3-VL, DINOv2, and OpenCLIP.
34 commits
A Vision-Centric Multimodal LLM for Holistic Medical Understanding (32B Parameter Flagship)
ClinFusion-32B is the flagship, high-performance model of the ClinFusion multimodal large language model series. Built on Qwen3-VL-32B-Instruct, ClinFusion-32B is engineered for complex clinical reasoning, high-fidelity medical report generation, and fine-grained visual-textual instruction-following.
By fusing DINOv2 and ConvNeXt vision encoders with a 32-billion parameter language backbone through our custom Cascade Spatial-Aware Locality Fusion operator, ClinFusion-32B delivers unmatched comprehension of medical cases, approaching and sometimes exceeding proprietary models like GPT-5.2 and Gemini-3-Flash.
.nii.gz CT/MRI files).To set up the environment, run inference, or evaluate ClinFusion-32B on your own data, please refer directly to our official GitHub Repository.
The GitHub repository provides comprehensive, step-by-step instructions for:
uv.ClinFusion-32B sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks.
If you find ClinFusion useful in your research, please consider citing:
@article{yuan2026ClinFusion,
title={ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding},
author={Yuan, Hangjie and Qian, Yichen and Tang, Zhiwei and Xu, Xianzhe and Wu, Lirong and Yang, Sicheng and Wang, Jinwang and Wang, Pengju and Zeng, Zhitao and Han, Yizeng and Xing, Yan and Luo, Shengxuan and Feng, Tao and Xie, Qing and Yao, Weigen and Yang, Yi and Liu, Zuozhu and Tang, Jiasheng and Wang, Shaocheng and Wang, Jitao and Dong, Jiahong and Chen, Weihua and Xu, Feng and Wang, Fan},
journal={arXiv preprint arXiv:2607.24743},
year={2026}
}
Built with ❤️ by Alibaba DAMO Academy. Special thanks to the open-source community behind Qwen3-VL, DINOv2, and OpenCLIP.
34 commits