32 repos
Multimodal AI systems that combine vision and language processing for medical imaging and clinical applications, particularly in radiology. These repositories feature models like MedVLThinker at various scales (3B, 7B, 32B parameters) trained with reinforcement learning and supervised fine-tuning on medical imaging datasets, designed to perform visual reasoning and question-answering over medical images. The cluster represents work at the intersection of computer vision, natural language processing, and medical AI—enabling models to understand and reason about diagnostic imagery alongside clinical text.