13 repos
Video object segmentation and understanding systems that combine visual and language modalities to enable referring expression segmentation and video-language reasoning tasks. The cluster centers on methods for localizing and segmenting objects in video based on textual descriptions, with datasets and model implementations spanning Python and Jupyter notebook-based research. Core projects like GLUS and variants explore grounding and segmentation approaches for connecting language queries to dynamic video content.