Multimodal Video Understanding & Segmentation

13 repos

Video object segmentation and understanding systems that combine visual and language modalities to enable referring expression segmentation and video-language reasoning tasks. The cluster centers on methods for localizing and segmenting objects in video based on textual descriptions, with datasets and model implementations spanning Python and Jupyter notebook-based research. Core projects like GLUS and variants explore grounding and segmentation approaches for connecting language queries to dynamic video content.

Python · 7
Jupyter Notebook · 3
referring-video-object-segmentation ·951
referring-expression-segmentation ·595
video-understanding ·595
multimodal-learning ·525
referring-expression-comprehension ·525
mevis-dataset ·525
mose-dataset ·525
video-language ·356
multi-modality ·70
video-segmentation ·70