Dense-World/Sa2VA-Training

Dataset

8

stars

16

commits

2

linked in READMEs

Jul 28, 2026

updated

README

This repository contains the code and data for the paper "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos".

🏠 Project Page
📜 arXiv 🧑‍💻 GitHub

Sa2VA is the first unified model for the dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation, with minimal one-shot instruction tuning. Sa2VA combines SAM-2, a foundation video segmentation model, with LLaVA, an advanced vision-language model, and unifies text, image, and video into a shared LLM token space.

Contributors

HarborYuan

15 commits

nielsr

1 commits

Dense-World/Sa2VA-Training

Dataset

8

stars

16

commits

2

linked in READMEs

Jul 28, 2026

updated

README

This repository contains the code and data for the paper "Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos".

🏠 Project Page
📜 arXiv 🧑‍💻 GitHub

Sa2VA is the first unified model for the dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation, with minimal one-shot instruction tuning. Sa2VA combines SAM-2, a foundation video segmentation model, with LLaVA, an advanced vision-language model, and unifies text, image, and video into a shared LLM token space.

Contributors

HarborYuan

15 commits

nielsr

1 commits