A ComfyUI node implementation for ByteDance's Sa2VA
Python
95
9 commits
updated Dec 22, 2025
A ComfyUI node implementation for ByteDance's Sa2VA (Segment Anything 2 Video Assistant) models, enabling advanced multimodal image and video understanding with precise segmentation capabilities. This repo only implements the image portion of the model.
Sa2VA is a state-of-the-art multimodal large language model (MLLM) that combines SAM2 (Segment Anything Model 2) with VLLMs for grounded understanding of images and videos. It achieves comparable performance to SOTA MLLMs like Qwen2.5-VL and InternVL3 on question-answering benchmarks while adding advanced visual prompt understanding and dense object segmentation capabilities.
This Sa2VA node can be thought of as a more advanced version of neverbiasu's ComfyUI-SAM2 node that allows for segmentation of objects in an image using natural langauge. Unlike that node which is based on Grounded SAM/Grounding DINO, Sa2VA uses a full VLLM trained to output SAM2 segmentation masks, which means it can handle significantly longer and more descriptive text. This allows Sa2VA to be better for uses cases where simple phrases like "woman on right" isn't sufficient to completely disambiguate between objects.
It outperforms Grounding DINO on short prompts:

And can follow longer instructions quite well, such as describing a character in general, rather than their position or traits in the image itself. This lends itself well to auto-generated or agentic segmentation prompts:

It can also segment more than one mask at a time, but the prompt needs to be precise:

cd ComfyUI/custom_nodes
git clone https://github.com/adambarbato/ComfyUI-Sa2VA.git
cd ComfyUI-Sa2VA
pip install -r requirements.txt
Important: Sa2VA models require:
Older transformers versions will fail with "No module named 'transformers.models.qwen3_vl'" error.
ByteDance/Sa2VA-Qwen3-VL-4B (recommended - 4B parameters)ByteDance/Sa2VA-Qwen2_5-VL-7B (7B parameters)ByteDance/Sa2VA-InternVL3-8B (8B parameters)ByteDance/Sa2VA-InternVL3-14B (14B parameters)ByteDance/Sa2VA-Qwen2_5-VL-3B (3B parameters)ByteDance/Sa2VA-InternVL3-2B (2B parameters)A single, comprehensive node that provides:
model_name and mask_threshold as neededsegmentation_prompt: "Please describe the image in detail."model_name and mask_threshold as neededsegmentation_prompt: "Please provide segmentation masks for all objects."masks output to mask-compatible nodes or mask_images to Preview ImageSa2VA models use bfloat16 precision by default with the option to quantize to 8 bits using bits-and-bytes.
"No module named 'transformers.models.qwen3_vl'"
pip install transformers>=4.57.0 --upgrade
This is the most common issue - your transformers version is too old.
"No module named 'qwen_vl_utils'"
pip install qwen_vl_utils
This dependency is required for Sa2VA model utilities.
"'NoneType' object is not subscriptable"
CUDA Out of Memory
Model Loading Errors
torch.cuda.empty_cache() to clear VRAMPoor Segmentation Quality
For detailed troubleshooting, see TROUBLESHOOTING.md.
# Object-specific segmentation
"Please segment the person in the image"
"Identify and segment all vehicles"
# Multi-object segmentation
"Create separate masks for all distinct objects"
"Segment foreground and background separately"
Contributions welcome! Areas for improvement:
MIT
Python
100.0%
A ComfyUI node implementation for ByteDance's Sa2VA
Python
95
9 commits
updated Dec 22, 2025
A ComfyUI node implementation for ByteDance's Sa2VA (Segment Anything 2 Video Assistant) models, enabling advanced multimodal image and video understanding with precise segmentation capabilities. This repo only implements the image portion of the model.
Sa2VA is a state-of-the-art multimodal large language model (MLLM) that combines SAM2 (Segment Anything Model 2) with VLLMs for grounded understanding of images and videos. It achieves comparable performance to SOTA MLLMs like Qwen2.5-VL and InternVL3 on question-answering benchmarks while adding advanced visual prompt understanding and dense object segmentation capabilities.
This Sa2VA node can be thought of as a more advanced version of neverbiasu's ComfyUI-SAM2 node that allows for segmentation of objects in an image using natural langauge. Unlike that node which is based on Grounded SAM/Grounding DINO, Sa2VA uses a full VLLM trained to output SAM2 segmentation masks, which means it can handle significantly longer and more descriptive text. This allows Sa2VA to be better for uses cases where simple phrases like "woman on right" isn't sufficient to completely disambiguate between objects.
It outperforms Grounding DINO on short prompts:

And can follow longer instructions quite well, such as describing a character in general, rather than their position or traits in the image itself. This lends itself well to auto-generated or agentic segmentation prompts:

It can also segment more than one mask at a time, but the prompt needs to be precise:

cd ComfyUI/custom_nodes
git clone https://github.com/adambarbato/ComfyUI-Sa2VA.git
cd ComfyUI-Sa2VA
pip install -r requirements.txt
Important: Sa2VA models require:
Older transformers versions will fail with "No module named 'transformers.models.qwen3_vl'" error.
ByteDance/Sa2VA-Qwen3-VL-4B (recommended - 4B parameters)ByteDance/Sa2VA-Qwen2_5-VL-7B (7B parameters)ByteDance/Sa2VA-InternVL3-8B (8B parameters)ByteDance/Sa2VA-InternVL3-14B (14B parameters)ByteDance/Sa2VA-Qwen2_5-VL-3B (3B parameters)ByteDance/Sa2VA-InternVL3-2B (2B parameters)A single, comprehensive node that provides:
model_name and mask_threshold as neededsegmentation_prompt: "Please describe the image in detail."model_name and mask_threshold as neededsegmentation_prompt: "Please provide segmentation masks for all objects."masks output to mask-compatible nodes or mask_images to Preview ImageSa2VA models use bfloat16 precision by default with the option to quantize to 8 bits using bits-and-bytes.
"No module named 'transformers.models.qwen3_vl'"
pip install transformers>=4.57.0 --upgrade
This is the most common issue - your transformers version is too old.
"No module named 'qwen_vl_utils'"
pip install qwen_vl_utils
This dependency is required for Sa2VA model utilities.
"'NoneType' object is not subscriptable"
CUDA Out of Memory
Model Loading Errors
torch.cuda.empty_cache() to clear VRAMPoor Segmentation Quality
For detailed troubleshooting, see TROUBLESHOOTING.md.
# Object-specific segmentation
"Please segment the person in the image"
"Identify and segment all vehicles"
# Multi-object segmentation
"Create separate masks for all distinct objects"
"Segment foreground and background separately"
Contributions welcome! Areas for improvement:
MIT
Python
100.0%