16 repos
Large-scale multimodal models that combine vision and language capabilities, enabling AI systems to understand and reason about both images/videos and text together. This cluster covers instruction-tuning approaches for vision-language models, video understanding frameworks, and techniques for integrating visual information with language model reasoning. Repositories here span foundational model architectures, training methodologies, and practical applications of multimodal AI across image and video domains.