Awesome-MLLM-Segmentation
If you find this project helpful, please consider giving it a star ⭐.
Last Updated: 2026-09-26
New updates are added directly to their corresponding sections.
Contents
Image Segmentation
Referring Expression Segmentation
- [LISA] | CVPR'24 | LISA: Reasoning Segmentation via Large Language Model |
[pdf] | [code]
- [GLaMM] | CVPR'24 | GLaMM: Pixel Grounding Large Multimodal Model |
[pdf] | [code]
- [PixelLM] | CVPR'24 | PixelLM: Pixel Reasoning with Large Multimodal Model |
[pdf] | [code]
- [GSVA] | CVPR'24 | GSVA: Generalized Segmentation via Multimodal Large Language Models |
[pdf] | [code]
- [AnyRef] | CVPR'24 | Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception |
[pdf] | [code]
- [GROUNDHOG] | CVPR'24 | GROUNDHOG: Grounding Large Language Models to Holistic Segmentation |
[pdf] | [code]
- [NExT-Chat] | ICML'24 | NExT-Chat: An LMM for Chat, Detection and Segmentation |
[pdf] | [code]
- [PSALM] | ECCV'24 | PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model |
[pdf] | [code]
- [SAM4MLLM] | ECCV'24 | SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation |
[pdf] | [code]
- [CoReS] | ECCV'24 | CoReS: Orchestrating the Dance of Reasoning and Segmentation |
[pdf] | [code]
- [VISA] | ECCV'24 | VISA: Reasoning Video Object Segmentation via Large Language Models |
[pdf] | [code]
- [OMG-LLaVA] | NeurIPS'24 | OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding |
[pdf] | [code]
- [VITRON ] | NeurIPS'24 | Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing |
[pdf] | [code]
- [VisionLLM v2] | NeurIPS'24 | VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks |
[pdf] | [code]
- [VLTP] | WACV'25 | VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation |
[pdf] | [code]
- [LaVASeg] | ArXiv'2403 | Empowering Segmentation Ability to Multi-modal Large Language Models |
[pdf]
- [LaSagnA] | ArXiv'2404 | LaSagnA: Language-based Segmentation Assistant for Complex Queries |
[pdf] | [code]
- [F-LMM] | CVPR'25 | F-LMM: Grounding Frozen Large Multimodal Models |
[pdf] | [code]
- [SETOKIM] | ICLR'25 | Towards Semantic Equivalence of Tokenization in Multimodal LLM |
[pdf] | [code]
- [u-LLaVA] | ArXiv'2408 | u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model |
[pdf] | [code]
- [UnifiedMLLM] | ArXiv'2408 | UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model |
[pdf] | [code]
- [DIFFLMM] | ArXiv'2410 | Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision |
[pdf] | [code]
- [SegLLM] | ICLR'25 | SegLLM: Multi-round Reasoning Segmentation |
[pdf] | [code]
- [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
[pdf] | [code]
- [InstructSeg] | ArXiv'2412 | InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models |
[pdf] | [code]
- [PRIMA] | ArXiv'2412 | PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation |
[pdf] | [code]
- [GeoGround] | ArXiv'2411 | GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding |
[pdf] | [code]
- [RSUniVLM] | ArXiv'2412 | RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts |
[pdf] | [code]
- [Text4Seg] | ICLR'25 | Text4Seg: Reimagining Image Segmentation as Text Generation |
[pdf] | [code]
- [Sa2VA] | ArXiv'2501 | Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos |
[pdf] | [code]
- [MIRAS] | ArXiv'2502 | Pixel-Level Reasoning Segmentation via Multi-turn Conversations |
[pdf] | [code]
- [GeoPix] | ArXiv'2501 | GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing |
[pdf]
- [UFO] | NeurIPS'25 | UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface |
[pdf] | [code]
- [GroundingSuite] | ArXiv'2503 | GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding |
[pdf] | [code]
- [HiMTok] | ICCV'25 | HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model |
[pdf] | [code]
- [Seg-Zero] | ArXiv'2503 | Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement |
[pdf] | [code]
- [POPEN] | CVPR'25 | POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation |
[pdf] | [code]
- [MMSA] | ICLR'25 | MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation |
[pdf] | [code]
- [SegEarth-R1] | ArXiv'2504 | SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model |
[pdf] | [code]
- [Pixel-SAIL] | ArXiv'2504 | Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding |
[pdf] | [code]
- [Ground-V] | CVPR'25 | Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels |
[pdf]
- [VistaLLM] | CVPR'24 | Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model |
[pdf]
- [SegAgent] | CVPR'25 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories |
[pdf] | [code]
- [ALTo] | ArXiv'2505 | ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation |
[pdf] | [code]
- [SAM-R1] | NeurIPS'25 | SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning |
[pdf]
- [RSVP] | ACL'25 | RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought |
[pdf] | [code]
- [VRS-HQ] | CVPR'25 | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation |
[pdf] | [code]
- [VisionReasoner] | ArXiv'2505 | VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning |
[pdf] | [code]
- [PixelThink] | ArXiv'2505 | PixelThink: Towards Efficient Chain-of-Pixel Reasoning |
[pdf] | [code]
- [Seg-R1] | ArXiv'2506 | Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning |
[pdf] | [code]
- [HRSeg] | ArXiv'2507 | HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation |
[pdf] | [code]
- [OmniAVS] | ICCV'25 | Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation |
[pdf] | [code]
- [LENS] | ArXiv'2508 | LENS: Learning to Segment Anything with Unified Reinforced Reasoning |
[pdf] | [code]
- [X-SAM] | ArXiv'2508 | X-SAM: From Segment Anything to Any Segmentation |
[pdf] | [code]
- [Text4Seg++] | ArXiv'2509 | Text4Seg++: Advancing Image Segmentation via Generative Language Modeling |
[pdf] | [code]
- [SVP] | ArXiv'2509 | Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation |
[pdf] |
- [CoPRS] | ArXiv'2510 | CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation |
[pdf] | [code]
- [PaDT] | ArXiv'2510 | Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs |
[pdf] | [code]
- [LENS] | ArXiv'2510 | Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs |
[pdf]
- [ARGenSeg] | NeurIPS'25 | ARGenSeg: Image Segmentation with Autoregressive Image Generation Model |
[pdf]
- [UniPixel] | NeurIPS'25 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning |
[pdf] | [code]
- [READ] | CVPR'25 | Reasoning to Attend: Try to Understand How Token Works |
[pdf] | [code]
- [UniGeoSeg] | ArXiv'2511 | UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes |
[pdf] | [code]
- [STAMP] | ArXiv'2511 | Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction |
[pdf] | [code]
- [GETok] | ArXiv'2512 | Grounding Everything in Tokens for Multimodal Large Language Models |
[pdf] | [code]
- [Tarot-SAM3] | ArXiv'2604 | Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation |
[pdf]
- [WISE] | ArXiv'2604 | Efficient Reasoning via Thought Compression for Language Segmentation |
[pdf] | [code]
- [DPAD] | CVPR'26 | Discriminative Perception via Anchored Description for Reasoning Segmentation |
[pdf] | [code]
- [WeatherReasonSeg] | ArXiv'2603 | WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models |
[pdf]
- [GroundedSurg] | ArXiv'2603 | GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation |
[pdf] | [code]
- [PRS-Med] | ArXiv'2505 | PRS-Med: Position Reasoning Segmentation for Medical Images with Vision-Language Model |
[pdf] | [code]
- [LlamaSeg] | ArXiv'2505 | LlamaSeg: Image Segmentation via Autoregressive Mask Generation |
[pdf]
- [MedSeg-R] | ArXiv'2506 | MedSeg-R: Medical Image Segmentation with Domain Alignment and RL for Reasoning |
[pdf]
- [SegEarth-R2] | ArXiv'2512 | SegEarth-R2: Geospatial Multimodal Large Language Model for Reasoning Segmentation |
[pdf] | [code]
- [DR^2Seg] | ArXiv'2601 | DR^2Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models |
[pdf]
- [PixDLM] | CVPR'26 | PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation |
[pdf]
- [UGround] | ICML'26 | UGround: Towards Unified Visual Grounding with Unrolled Transformers |
[pdf] | [code]
- [Dr. Seg] | CVPR'26 | Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design |
[pdf] | [code]
- [MTRS] | ArXiv'2606 | An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation |
[pdf]
- [AnchorSeg] | ACL'26 | AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation |
[pdf] | [code]
- [SetCon] | ArXiv'2605 | SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction |
[pdf]
- [IC-Seg] | ArXiv'2605 | Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification |
[pdf] | [code]
- [SegCompass] | CVPR'26 | SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation |
[pdf] | [code]
- [FlowSeg] | ICML'26 | FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation |
[pdf] | [code]
- [MedVol-R1] | ArXiv'2605 | MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation |
[pdf]
- [CR-Seg] | ArXiv'2606 | CR-Seg: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation |
[pdf]
- [StAR] | ArXiv'2603 | StAR: Segment Anything Reasoner |
[pdf] | [code]
- [InstructSAM] | ArXiv'2605 | InstructSAM: Segment Any Instance with Any Instructions |
[pdf] | [code]
- [B-GRTO] | ArXiv'2605 | B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation |
[pdf]
- [VASA] | ArXiv'2605 | Vision Harnessing Agent for Open Ad-hoc Segmentation |
[pdf]
- [Qwen3-VL-Seg] | ArXiv'2605 | Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding |
[pdf]
- [Rea2Seg] | ArXiv'2606 | Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning |
[pdf] | [code]
- [CCRC] | ArXiv'2606 | CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation |
[pdf]
- [GEAR-Seg] | ArXiv'2607 | GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine |
[pdf]
- [PanoSeeker] | ECCV'26 | Seek to Segment: Active Perception for Panoramic Referring Segmentation |
[pdf] | [code]
- [DGSeg] | ECCV'26 | DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation |
[pdf] | [code]
- [SegAnswer] | ArXiv'2607 | Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning |
[pdf]
- [CycleGRPO] | ECCV'26 | Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO |
[pdf] | [code]
- [OP-HRG] | ECCV'26 | Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning |
[pdf]
- [VespaSeg] | ArXiv'2608 | VespaSeg: A Resource-Aware Ground-then-Segment Pipeline for Referring Expression Segmentation |
[pdf] | [code]
- [PixVL] | ArXiv'2608 | PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask-Text Consistency Cycle |
[pdf]
- [STAMPlus] | ArXiv'2608 | Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation |
[pdf]
- [EgoAfford] | ArXiv'2608 | EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation |
[pdf] | [code]
- [MedPixel] | ArXiv'2608 | MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation |
[pdf] | [code]
- [MedUP] | ArXiv'2608 | MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models |
[pdf]
- [FIRM] | ArXiv'2608 | FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation |
[pdf]
- [FloodReasonBench] | ArXiv'2608 | FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge |
[pdf]
- [DRAgent] | ArXiv'2608 | DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation |
[pdf]
- [PAYN] | ICML'26 | Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation |
[pdf] | [code]
- [MedREAL] | ECCV'26 | From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation |
[pdf]
- [SeGDeP] | ArXiv'2609 | SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation |
[pdf]
- [AgriScope] | ArXiv'2609 | AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images |
[pdf] | [project]
- [MLLM Point Prompts] | ArXiv'2609 | Adapting Open-Weight MLLMs to Generate Point Prompts for Electron Microscopy Segmentation |
[pdf]
- [PSMA PET/CT VLM] | ArXiv'2609 | A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation |
[pdf]
Open-Vocabulary Semantic Segmentation
- [PSALM] | ECCV'24 | PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model |
[pdf] | [code]
- [LLMFormer] | IJCV'24 | LLMFormer: Large Language Model for Open-Vocabulary Semantic Segmentation |
[pdf]
- [LaSagnA] | ArXiv'2404 | LaSagnA: Language-based Segmentation Assistant for Complex Queries |
[pdf] | [code]
- [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
[pdf] | [code]
- [Text4Seg] | ICLR'25 | Text4Seg: Reimagining Image Segmentation as Text Generation |
[pdf] | [code]
- [HiMTok] | ICCV'25 | HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model |
[pdf] | [code]
- [ALTo] | ArXiv'2505 | ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation |
[pdf] | [code]
- [X-SAM] | ArXiv'2508 | X-SAM: From Segment Anything to Any Segmentation |
[pdf] | [code]
- [OV-Stitcher] | ArXiv'2604 | OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation |
[pdf]
- [SPAR] | CVPR'26 | SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation |
[pdf] | [code]
- [MM-OVSeg] | CVPR'26 | MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing |
[pdf] | [code]
- [DSLO] | CVPR'26 | Direct Segmentation without Logits Optimization for Open-Vocabulary Semantic Segmentation |
[pdf]
- [ProC-SAM3] | GRSL'26 | Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation |
[pdf] | [code]
- [STAMPlus] | ArXiv'2608 | Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation |
[pdf]
- [DAF] | ArXiv'2608 | Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation |
[pdf] | [code]
- [SatOV] | ArXiv'2609 | SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery |
[pdf]
- [HyperCLIP++] | ArXiv'2609 | HyperCLIP++: Fine-tuning CLIP for Open-Vocabulary Semantic Segmentation in Hyperbolic Space |
[pdf]
Video Segmentation
- [VISA] | ECCV'24 | VISA: Reasoning Video Object Segmentation via Large Language Models |
[pdf] | [code]
- [VideoLISA] | NeurIPS'24 | One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos |
[pdf] | [code]
- [VITRON ] | NeurIPS'24 | Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing |
[pdf] | [code]
- [ViLLa] | NeurIPS'25 | ViLLa: Video Reasoning Segmentation with Large Language Model |
[pdf] | [code]
- [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model |
[pdf] | [code]
- [InstructSeg] | ArXiv'2412 | InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models |
[pdf] | [code]
- [Sa2VA] | ArXiv'2501 | Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos |
[pdf] | [code]
- [VRS-HQ] | CVPR'25 | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation |
[pdf] | [code]
- [GLUS] | CVPR'25 | GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation |
[pdf] | [code]
- [DeSa2VA] | ArXiv'2506 | Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder |
[pdf] | [code]
- [OmniAVS] | ICCV'25 | Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation |
[pdf] | [code]
- [Veason-R1] | CVPR'26 | Reinforcing Video Reasoning Segmentation to Think Before It Segments |
[pdf] | [code]
- [VoCap] | ArXiv'2508 | VoCap: Video Object Captioning and Segmentation from Any Prompt |
[pdf] | [code]
- [PixFoundation 2.0] | ArXiv'2508 | PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? |
[pdf] | [code]
- [DecAF] | ICLR'26 | DecAF: Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation |
[pdf] | [code]
- [UniPixel] | NeurIPS'25 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning |
[pdf] | [code]
- [TrajSeg] | ArXiv'2603 | Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation |
[pdf] | [code]
- [VIRST] | CVPR'26 | VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation |
[pdf] | [code]
- [AgentRVOS] | ArXiv'2603 | AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation |
[pdf] | [code]
- [SDAM] | AAAI'26 | Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory |
[pdf]
- [SPARROW] | CVPR'26 | SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs |
[pdf] | [code]
- [MAR3] | ArXiv'2603 | MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation |
[pdf]
- [SteerSeg] | ArXiv'2605 | SteerSeg: Attention Steering for Reasoning Video Segmentation |
[pdf] | [code]
- [Seg-ReSearch] | ArXiv'2602 | Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search |
[pdf] | [code]
- [RCoT-Seg] | ArXiv'2605 | RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation |
[pdf] | [code]
- [FeVOS] | ECCV'26 | FeVOS: Foresight Expression Video Object Segmentation |
[pdf] | [code]
- [STAC] | ArXiv'2607 | STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation |
[pdf] | [code]
- [ReflexTrack] | ArXiv'2607 | ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation |
[pdf]
- [PhysMLLMs] | ArXiv'2608 | PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos |
[pdf] | [code]
- [MLLM-Assisted Audio VOS] | ArXiv'2608 | MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge |
[pdf]
- [MoVISA] | ArXiv'2609 | MoVISA: Multi-Token Reasoning for Video Object Segmentation |
[pdf]
Feedback
If you have any suggestions or find missing papers, please feel free to reach out at lanm0002@e.ntu.edu.sg.