0
stars
23
commits
Python
primary language
Jul 28, 2026
updated
Lin, Lang, et al. "Glus: Global-local reasoning unified into a single large language model for video segmentation." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\ ↩
Parikh, Chirag, et al. "Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from social video narratives." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\ ↩
Yuan, Haobo, et al. "Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos." arXiv preprint arXiv:2501.04001 (2025).\ ↩
Ravi, Nikhila, et al. "Sam 2: Segment anything in images and videos." International Conference on Learning Representations. Vol. 2025. 2025.\ ↩
Zhang, Boqiang, et al. "Videollama 3: Frontier multimodal foundation models for image and video understanding." arXiv preprint arXiv:2501.13106 (2025). ↩
23 commits
Python
100.0%
0
stars
23
commits
Python
primary language
Jul 28, 2026
updated
Lin, Lang, et al. "Glus: Global-local reasoning unified into a single large language model for video segmentation." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\ ↩
Parikh, Chirag, et al. "Roadsocial: A diverse videoqa dataset and benchmark for road event understanding from social video narratives." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\ ↩
Yuan, Haobo, et al. "Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos." arXiv preprint arXiv:2501.04001 (2025).\ ↩
Ravi, Nikhila, et al. "Sam 2: Segment anything in images and videos." International Conference on Learning Representations. Vol. 2025. 2025.\ ↩
Zhang, Boqiang, et al. "Videollama 3: Frontier multimodal foundation models for image and video understanding." arXiv preprint arXiv:2501.13106 (2025). ↩
23 commits
Python
100.0%