16 repos
Multimodal AI models that combine visual understanding with language processing, enabling systems to reason about images and videos using natural language. The cluster centers on video-language pretraining architectures and cross-modal fusion techniques, with VideoLLaMA representing a key family of models that extend large language models to process and understand video content alongside text. These repositories contain model implementations, training frameworks, and utilities for building systems that ground language understanding in visual and temporal information.