Vision-Language and Video-Language Models

16 repos

Multimodal AI models that combine visual understanding with language processing, enabling systems to reason about images and videos using natural language. The cluster centers on video-language pretraining architectures and cross-modal fusion techniques, with VideoLLaMA representing a key family of models that extend large language models to process and understand video content alongside text. These repositories contain model implementations, training frameworks, and utilities for building systems that ground language understanding in visual and temporal information.

Python · 3
video-language-pretraining ·3,141
llama ·3,141
multi-modal-chatgpt ·3,141
minigpt4 ·3,141
blip2 ·3,141
vision-language-pretraining ·3,141
large-language-models ·3,141
cross-modal-pretraining ·3,141
text-generation ·121
transformers ·121