Medical Vision-Language Models

32 repos

Multimodal AI systems that combine vision and language processing for medical imaging and clinical applications, particularly in radiology. These repositories feature models like MedVLThinker at various scales (3B, 7B, 32B parameters) trained with reinforcement learning and supervised fine-tuning on medical imaging datasets, designed to perform visual reasoning and question-answering over medical images. The cluster represents work at the intersection of computer vision, natural language processing, and medical AI—enabling models to understand and reason about diagnostic imagery alongside clinical text.

medical ·561
multimodal ·558
radiology ·464
vision-language ·401
chest-ct ·301
diagnostic-imaging ·301
healthcare ·301
foundation-model ·301
huggingscience ·301
3d-medical-imaging ·301