10 repos
Large language models and multimodal systems adapted for image segmentation tasks, particularly in medical imaging. The cluster combines vision-language model integration (DINO, SAM2, Qwen) with specialized segmentation frameworks like LISA, which leverages LLMs to understand segmentation queries and improve annotation efficiency. Projects here explore how language understanding can enhance traditional computer vision pipelines.