20 repos
Fine-tuned and optimized implementations of medical vision-language models built on Qwen2.5-VL, a multimodal foundation model. These repositories focus on adapting large vision-language models for medical imaging and clinical applications through reinforcement learning and other optimization techniques. Researchers exploring this cluster will find model variants at different scales (3B and 7B parameters), training datasets, and inference implementations designed to improve medical image understanding and report generation.