Vision-Language Video Understanding

14 repos

Multimodal AI models that combine vision and language capabilities for understanding video content, including temporal reasoning, action recognition, and video question-answering tasks. The cluster centers on specialized vision-language models like the Cambrian-S and OpenSpatial series that are optimized for spatial and temporal reasoning across video frames, enabling systems to track multi-object interactions and reason causally about sequences of events.

Python · 1
vision-language ·232
video-understanding ·204
videoqa ·191
video-question-answering ·191
multi-object-interaction ·191
causal-temporal-action-reasoning ·191
multimodal ·41
spatial-reasoning ·39
remyx ·28
quantitative-spatial-reasoning ·28