19 repos
Vision-language models and systems for understanding images and videos through combined visual and textual representations. This cluster focuses on VLM architectures, multimodal AI systems, and their applications in image and video understanding tasks. The central repositories are medical-domain specialized models (Hulu-Med series), representing domain adaptation of vision-language techniques to healthcare contexts.