mc-lan/Awesome-MLLM-Segmentation

A curated list of publications on image and video segmentation leveraging Multimodal Large Language Models (MLLMs), highlighting state-of-the-art methods, innovative applications, and key advancements in the field.

236

65 commits

updated Sep 26, 2026

See the code

README

Awesome-MLLM-Segmentation

If you find this project helpful, please consider giving it a star ⭐.

Last Updated: 2026-09-26

New updates are added directly to their corresponding sections.

Contents

Image Segmentation

Referring Expression Segmentation

  1. [LISA] | CVPR'24 | LISA: Reasoning Segmentation via Large Language Model | [pdf] | [code]
  2. [GLaMM] | CVPR'24 | GLaMM: Pixel Grounding Large Multimodal Model | [pdf] | [code]
  3. [PixelLM] | CVPR'24 | PixelLM: Pixel Reasoning with Large Multimodal Model | [pdf] | [code]
  4. [GSVA] | CVPR'24 | GSVA: Generalized Segmentation via Multimodal Large Language Models | [pdf] | [code]
  5. [AnyRef] | CVPR'24 | Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception | [pdf] | [code]
  6. [GROUNDHOG] | CVPR'24 | GROUNDHOG: Grounding Large Language Models to Holistic Segmentation | [pdf] | [code]
  7. [NExT-Chat] | ICML'24 | NExT-Chat: An LMM for Chat, Detection and Segmentation | [pdf] | [code]
  8. [PSALM] | ECCV'24 | PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model | [pdf] | [code]
  9. [SAM4MLLM] | ECCV'24 | SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation | [pdf] | [code]
  10. [CoReS] | ECCV'24 | CoReS: Orchestrating the Dance of Reasoning and Segmentation | [pdf] | [code]
  11. [VISA] | ECCV'24 | VISA: Reasoning Video Object Segmentation via Large Language Models | [pdf] | [code]
  12. [OMG-LLaVA] | NeurIPS'24 | OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding | [pdf] | [code]
  13. [VITRON ] | NeurIPS'24 | Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing | [pdf] | [code]
  14. [VisionLLM v2] | NeurIPS'24 | VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks | [pdf] | [code]
  15. [VLTP] | WACV'25 | VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation | [pdf] | [code]
  16. [LaVASeg] | ArXiv'2403 | Empowering Segmentation Ability to Multi-modal Large Language Models | [pdf]
  17. [LaSagnA] | ArXiv'2404 | LaSagnA: Language-based Segmentation Assistant for Complex Queries | [pdf] | [code]
  18. [F-LMM] | CVPR'25 | F-LMM: Grounding Frozen Large Multimodal Models | [pdf] | [code]
  19. [SETOKIM] | ICLR'25 | Towards Semantic Equivalence of Tokenization in Multimodal LLM | [pdf] | [code]
  20. [u-LLaVA] | ArXiv'2408 | u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model | [pdf] | [code]
  21. [UnifiedMLLM] | ArXiv'2408 | UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model | [pdf] | [code]
  22. [DIFFLMM] | ArXiv'2410 | Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision | [pdf] | [code]
  23. [SegLLM] | ICLR'25 | SegLLM: Multi-round Reasoning Segmentation | [pdf] | [code]
  24. [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model | [pdf] | [code]
  25. [InstructSeg] | ArXiv'2412 | InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models | [pdf] | [code]
  26. [PRIMA] | ArXiv'2412 | PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation | [pdf] | [code]
  27. [GeoGround] | ArXiv'2411 | GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding | [pdf] | [code]
  28. [RSUniVLM] | ArXiv'2412 | RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts | [pdf] | [code]
  29. [Text4Seg] | ICLR'25 | Text4Seg: Reimagining Image Segmentation as Text Generation | [pdf] | [code]
  30. [Sa2VA] | ArXiv'2501 | Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos | [pdf] | [code]
  31. [MIRAS] | ArXiv'2502 | Pixel-Level Reasoning Segmentation via Multi-turn Conversations | [pdf] | [code]
  32. [GeoPix] | ArXiv'2501 | GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing | [pdf]
  33. [UFO] | NeurIPS'25 | UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface | [pdf] | [code]
  34. [GroundingSuite] | ArXiv'2503 | GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding | [pdf] | [code]
  35. [HiMTok] | ICCV'25 | HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model | [pdf] | [code]
  36. [Seg-Zero] | ArXiv'2503 | Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement | [pdf] | [code]
  37. [POPEN] | CVPR'25 | POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation | [pdf] | [code]
  38. [MMSA] | ICLR'25 | MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation | [pdf] | [code]
  39. [SegEarth-R1] | ArXiv'2504 | SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model | [pdf] | [code]
  40. [Pixel-SAIL] | ArXiv'2504 | Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding | [pdf] | [code]
  41. [Ground-V] | CVPR'25 | Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels | [pdf]
  42. [VistaLLM] | CVPR'24 | Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model | [pdf]
  43. [SegAgent] | CVPR'25 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories | [pdf] | [code]
  44. [ALTo] | ArXiv'2505 | ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation | [pdf] | [code]
  45. [SAM-R1] | NeurIPS'25 | SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning | [pdf]
  46. [RSVP] | ACL'25 | RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought | [pdf] | [code]
  47. [VRS-HQ] | CVPR'25 | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | [pdf] | [code]
  48. [VisionReasoner] | ArXiv'2505 | VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning | [pdf] | [code]
  49. [PixelThink] | ArXiv'2505 | PixelThink: Towards Efficient Chain-of-Pixel Reasoning | [pdf] | [code]
  50. [Seg-R1] | ArXiv'2506 | Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning | [pdf] | [code]
  51. [HRSeg] | ArXiv'2507 | HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation | [pdf] | [code]
  52. [OmniAVS] | ICCV'25 | Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation | [pdf] | [code]
  53. [LENS] | ArXiv'2508 | LENS: Learning to Segment Anything with Unified Reinforced Reasoning | [pdf] | [code]
  54. [X-SAM] | ArXiv'2508 | X-SAM: From Segment Anything to Any Segmentation | [pdf] | [code]
  55. [Text4Seg++] | ArXiv'2509 | Text4Seg++: Advancing Image Segmentation via Generative Language Modeling | [pdf] | [code]
  56. [SVP] | ArXiv'2509 | Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation | [pdf] |
  57. [CoPRS] | ArXiv'2510 | CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation | [pdf] | [code]
  58. [PaDT] | ArXiv'2510 | Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs | [pdf] | [code]
  59. [LENS] | ArXiv'2510 | Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs | [pdf]
  60. [ARGenSeg] | NeurIPS'25 | ARGenSeg: Image Segmentation with Autoregressive Image Generation Model | [pdf]
  61. [UniPixel] | NeurIPS'25 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | [pdf] | [code]
  62. [READ] | CVPR'25 | Reasoning to Attend: Try to Understand How Token Works | [pdf] | [code]
  63. [UniGeoSeg] | ArXiv'2511 | UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes | [pdf] | [code]
  64. [STAMP] | ArXiv'2511 | Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction | [pdf] | [code]
  65. [GETok] | ArXiv'2512 | Grounding Everything in Tokens for Multimodal Large Language Models | [pdf] | [code]
  66. [Tarot-SAM3] | ArXiv'2604 | Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation | [pdf]
  67. [WISE] | ArXiv'2604 | Efficient Reasoning via Thought Compression for Language Segmentation | [pdf] | [code]
  68. [DPAD] | CVPR'26 | Discriminative Perception via Anchored Description for Reasoning Segmentation | [pdf] | [code]
  69. [WeatherReasonSeg] | ArXiv'2603 | WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models | [pdf]
  70. [GroundedSurg] | ArXiv'2603 | GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation | [pdf] | [code]
  71. [PRS-Med] | ArXiv'2505 | PRS-Med: Position Reasoning Segmentation for Medical Images with Vision-Language Model | [pdf] | [code]
  72. [LlamaSeg] | ArXiv'2505 | LlamaSeg: Image Segmentation via Autoregressive Mask Generation | [pdf]
  73. [MedSeg-R] | ArXiv'2506 | MedSeg-R: Medical Image Segmentation with Domain Alignment and RL for Reasoning | [pdf]
  74. [SegEarth-R2] | ArXiv'2512 | SegEarth-R2: Geospatial Multimodal Large Language Model for Reasoning Segmentation | [pdf] | [code]
  75. [DR^2Seg] | ArXiv'2601 | DR^2Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models | [pdf]
  76. [PixDLM] | CVPR'26 | PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation | [pdf]
  77. [UGround] | ICML'26 | UGround: Towards Unified Visual Grounding with Unrolled Transformers | [pdf] | [code]
  78. [Dr. Seg] | CVPR'26 | Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design | [pdf] | [code]
  79. [MTRS] | ArXiv'2606 | An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation | [pdf]
  80. [AnchorSeg] | ACL'26 | AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation | [pdf] | [code]
  81. [SetCon] | ArXiv'2605 | SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction | [pdf]
  82. [IC-Seg] | ArXiv'2605 | Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification | [pdf] | [code]
  83. [SegCompass] | CVPR'26 | SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation | [pdf] | [code]
  84. [FlowSeg] | ICML'26 | FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation | [pdf] | [code]
  85. [MedVol-R1] | ArXiv'2605 | MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation | [pdf]
  86. [CR-Seg] | ArXiv'2606 | CR-Seg: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation | [pdf]
  87. [StAR] | ArXiv'2603 | StAR: Segment Anything Reasoner | [pdf] | [code]
  88. [InstructSAM] | ArXiv'2605 | InstructSAM: Segment Any Instance with Any Instructions | [pdf] | [code]
  89. [B-GRTO] | ArXiv'2605 | B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation | [pdf]
  90. [VASA] | ArXiv'2605 | Vision Harnessing Agent for Open Ad-hoc Segmentation | [pdf]
  91. [Qwen3-VL-Seg] | ArXiv'2605 | Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding | [pdf]
  92. [Rea2Seg] | ArXiv'2606 | Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning | [pdf] | [code]
  93. [CCRC] | ArXiv'2606 | CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation | [pdf]
  94. [GEAR-Seg] | ArXiv'2607 | GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine | [pdf]
  95. [PanoSeeker] | ECCV'26 | Seek to Segment: Active Perception for Panoramic Referring Segmentation | [pdf] | [code]
  96. [DGSeg] | ECCV'26 | DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation | [pdf] | [code]
  97. [SegAnswer] | ArXiv'2607 | Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning | [pdf]
  98. [CycleGRPO] | ECCV'26 | Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO | [pdf] | [code]
  99. [OP-HRG] | ECCV'26 | Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning | [pdf]
  100. [VespaSeg] | ArXiv'2608 | VespaSeg: A Resource-Aware Ground-then-Segment Pipeline for Referring Expression Segmentation | [pdf] | [code]
  101. [PixVL] | ArXiv'2608 | PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask-Text Consistency Cycle | [pdf]
  102. [STAMPlus] | ArXiv'2608 | Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation | [pdf]
  103. [EgoAfford] | ArXiv'2608 | EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation | [pdf] | [code]
  104. [MedPixel] | ArXiv'2608 | MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation | [pdf] | [code]
  105. [MedUP] | ArXiv'2608 | MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models | [pdf]
  106. [FIRM] | ArXiv'2608 | FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation | [pdf]
  107. [FloodReasonBench] | ArXiv'2608 | FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge | [pdf]
  108. [DRAgent] | ArXiv'2608 | DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation | [pdf]
  109. [PAYN] | ICML'26 | Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation | [pdf] | [code]
  110. [MedREAL] | ECCV'26 | From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation | [pdf]
  111. [SeGDeP] | ArXiv'2609 | SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation | [pdf]
  112. [AgriScope] | ArXiv'2609 | AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images | [pdf] | [project]
  113. [MLLM Point Prompts] | ArXiv'2609 | Adapting Open-Weight MLLMs to Generate Point Prompts for Electron Microscopy Segmentation | [pdf]
  114. [PSMA PET/CT VLM] | ArXiv'2609 | A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation | [pdf]

Open-Vocabulary Semantic Segmentation

  1. [PSALM] | ECCV'24 | PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model | [pdf] | [code]
  2. [LLMFormer] | IJCV'24 | LLMFormer: Large Language Model for Open-Vocabulary Semantic Segmentation | [pdf]
  3. [LaSagnA] | ArXiv'2404 | LaSagnA: Language-based Segmentation Assistant for Complex Queries | [pdf] | [code]
  4. [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model | [pdf] | [code]
  5. [Text4Seg] | ICLR'25 | Text4Seg: Reimagining Image Segmentation as Text Generation | [pdf] | [code]
  6. [HiMTok] | ICCV'25 | HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model | [pdf] | [code]
  7. [ALTo] | ArXiv'2505 | ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation | [pdf] | [code]
  8. [X-SAM] | ArXiv'2508 | X-SAM: From Segment Anything to Any Segmentation | [pdf] | [code]
  9. [OV-Stitcher] | ArXiv'2604 | OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation | [pdf]
  10. [SPAR] | CVPR'26 | SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation | [pdf] | [code]
  11. [MM-OVSeg] | CVPR'26 | MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing | [pdf] | [code]
  12. [DSLO] | CVPR'26 | Direct Segmentation without Logits Optimization for Open-Vocabulary Semantic Segmentation | [pdf]
  13. [ProC-SAM3] | GRSL'26 | Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation | [pdf] | [code]
  14. [STAMPlus] | ArXiv'2608 | Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation | [pdf]
  15. [DAF] | ArXiv'2608 | Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation | [pdf] | [code]
  16. [SatOV] | ArXiv'2609 | SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery | [pdf]
  17. [HyperCLIP++] | ArXiv'2609 | HyperCLIP++: Fine-tuning CLIP for Open-Vocabulary Semantic Segmentation in Hyperbolic Space | [pdf]

Video Segmentation

  1. [VISA] | ECCV'24 | VISA: Reasoning Video Object Segmentation via Large Language Models | [pdf] | [code]
  2. [VideoLISA] | NeurIPS'24 | One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos | [pdf] | [code]
  3. [VITRON ] | NeurIPS'24 | Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing | [pdf] | [code]
  4. [ViLLa] | NeurIPS'25 | ViLLa: Video Reasoning Segmentation with Large Language Model | [pdf] | [code]
  5. [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model | [pdf] | [code]
  6. [InstructSeg] | ArXiv'2412 | InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models | [pdf] | [code]
  7. [Sa2VA] | ArXiv'2501 | Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos | [pdf] | [code]
  8. [VRS-HQ] | CVPR'25 | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | [pdf] | [code]
  9. [GLUS] | CVPR'25 | GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation | [pdf] | [code]
  10. [DeSa2VA] | ArXiv'2506 | Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder | [pdf] | [code]
  11. [OmniAVS] | ICCV'25 | Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation | [pdf] | [code]
  12. [Veason-R1] | CVPR'26 | Reinforcing Video Reasoning Segmentation to Think Before It Segments | [pdf] | [code]
  13. [VoCap] | ArXiv'2508 | VoCap: Video Object Captioning and Segmentation from Any Prompt | [pdf] | [code]
  14. [PixFoundation 2.0] | ArXiv'2508 | PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? | [pdf] | [code]
  15. [DecAF] | ICLR'26 | DecAF: Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation | [pdf] | [code]
  16. [UniPixel] | NeurIPS'25 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | [pdf] | [code]
  17. [TrajSeg] | ArXiv'2603 | Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation | [pdf] | [code]
  18. [VIRST] | CVPR'26 | VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation | [pdf] | [code]
  19. [AgentRVOS] | ArXiv'2603 | AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation | [pdf] | [code]
  20. [SDAM] | AAAI'26 | Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory | [pdf]
  21. [SPARROW] | CVPR'26 | SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs | [pdf] | [code]
  22. [MAR3] | ArXiv'2603 | MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation | [pdf]
  23. [SteerSeg] | ArXiv'2605 | SteerSeg: Attention Steering for Reasoning Video Segmentation | [pdf] | [code]
  24. [Seg-ReSearch] | ArXiv'2602 | Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search | [pdf] | [code]
  25. [RCoT-Seg] | ArXiv'2605 | RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation | [pdf] | [code]
  26. [FeVOS] | ECCV'26 | FeVOS: Foresight Expression Video Object Segmentation | [pdf] | [code]
  27. [STAC] | ArXiv'2607 | STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation | [pdf] | [code]
  28. [ReflexTrack] | ArXiv'2607 | ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation | [pdf]
  29. [PhysMLLMs] | ArXiv'2608 | PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos | [pdf] | [code]
  30. [MLLM-Assisted Audio VOS] | ArXiv'2608 | MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge | [pdf]
  31. [MoVISA] | ArXiv'2609 | MoVISA: Multi-Token Reasoning for Video Object Segmentation | [pdf]

Feedback

If you have any suggestions or find missing papers, please feel free to reach out at lanm0002@e.ntu.edu.sg.

Contributors

mc-lan

63 commits

ildar-idrisov

2 commits

mc-lan/Awesome-MLLM-Segmentation

A curated list of publications on image and video segmentation leveraging Multimodal Large Language Models (MLLMs), highlighting state-of-the-art methods, innovative applications, and key advancements in the field.

236

65 commits

updated Sep 26, 2026

See the code

README

Awesome-MLLM-Segmentation

If you find this project helpful, please consider giving it a star ⭐.

Last Updated: 2026-09-26

New updates are added directly to their corresponding sections.

Contents

Image Segmentation

Referring Expression Segmentation

  1. [LISA] | CVPR'24 | LISA: Reasoning Segmentation via Large Language Model | [pdf] | [code]
  2. [GLaMM] | CVPR'24 | GLaMM: Pixel Grounding Large Multimodal Model | [pdf] | [code]
  3. [PixelLM] | CVPR'24 | PixelLM: Pixel Reasoning with Large Multimodal Model | [pdf] | [code]
  4. [GSVA] | CVPR'24 | GSVA: Generalized Segmentation via Multimodal Large Language Models | [pdf] | [code]
  5. [AnyRef] | CVPR'24 | Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception | [pdf] | [code]
  6. [GROUNDHOG] | CVPR'24 | GROUNDHOG: Grounding Large Language Models to Holistic Segmentation | [pdf] | [code]
  7. [NExT-Chat] | ICML'24 | NExT-Chat: An LMM for Chat, Detection and Segmentation | [pdf] | [code]
  8. [PSALM] | ECCV'24 | PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model | [pdf] | [code]
  9. [SAM4MLLM] | ECCV'24 | SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation | [pdf] | [code]
  10. [CoReS] | ECCV'24 | CoReS: Orchestrating the Dance of Reasoning and Segmentation | [pdf] | [code]
  11. [VISA] | ECCV'24 | VISA: Reasoning Video Object Segmentation via Large Language Models | [pdf] | [code]
  12. [OMG-LLaVA] | NeurIPS'24 | OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding | [pdf] | [code]
  13. [VITRON ] | NeurIPS'24 | Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing | [pdf] | [code]
  14. [VisionLLM v2] | NeurIPS'24 | VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks | [pdf] | [code]
  15. [VLTP] | WACV'25 | VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation | [pdf] | [code]
  16. [LaVASeg] | ArXiv'2403 | Empowering Segmentation Ability to Multi-modal Large Language Models | [pdf]
  17. [LaSagnA] | ArXiv'2404 | LaSagnA: Language-based Segmentation Assistant for Complex Queries | [pdf] | [code]
  18. [F-LMM] | CVPR'25 | F-LMM: Grounding Frozen Large Multimodal Models | [pdf] | [code]
  19. [SETOKIM] | ICLR'25 | Towards Semantic Equivalence of Tokenization in Multimodal LLM | [pdf] | [code]
  20. [u-LLaVA] | ArXiv'2408 | u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model | [pdf] | [code]
  21. [UnifiedMLLM] | ArXiv'2408 | UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model | [pdf] | [code]
  22. [DIFFLMM] | ArXiv'2410 | Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision | [pdf] | [code]
  23. [SegLLM] | ICLR'25 | SegLLM: Multi-round Reasoning Segmentation | [pdf] | [code]
  24. [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model | [pdf] | [code]
  25. [InstructSeg] | ArXiv'2412 | InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models | [pdf] | [code]
  26. [PRIMA] | ArXiv'2412 | PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation | [pdf] | [code]
  27. [GeoGround] | ArXiv'2411 | GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding | [pdf] | [code]
  28. [RSUniVLM] | ArXiv'2412 | RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts | [pdf] | [code]
  29. [Text4Seg] | ICLR'25 | Text4Seg: Reimagining Image Segmentation as Text Generation | [pdf] | [code]
  30. [Sa2VA] | ArXiv'2501 | Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos | [pdf] | [code]
  31. [MIRAS] | ArXiv'2502 | Pixel-Level Reasoning Segmentation via Multi-turn Conversations | [pdf] | [code]
  32. [GeoPix] | ArXiv'2501 | GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing | [pdf]
  33. [UFO] | NeurIPS'25 | UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface | [pdf] | [code]
  34. [GroundingSuite] | ArXiv'2503 | GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding | [pdf] | [code]
  35. [HiMTok] | ICCV'25 | HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model | [pdf] | [code]
  36. [Seg-Zero] | ArXiv'2503 | Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement | [pdf] | [code]
  37. [POPEN] | CVPR'25 | POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation | [pdf] | [code]
  38. [MMSA] | ICLR'25 | MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation | [pdf] | [code]
  39. [SegEarth-R1] | ArXiv'2504 | SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model | [pdf] | [code]
  40. [Pixel-SAIL] | ArXiv'2504 | Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding | [pdf] | [code]
  41. [Ground-V] | CVPR'25 | Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels | [pdf]
  42. [VistaLLM] | CVPR'24 | Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model | [pdf]
  43. [SegAgent] | CVPR'25 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories | [pdf] | [code]
  44. [ALTo] | ArXiv'2505 | ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation | [pdf] | [code]
  45. [SAM-R1] | NeurIPS'25 | SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning | [pdf]
  46. [RSVP] | ACL'25 | RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought | [pdf] | [code]
  47. [VRS-HQ] | CVPR'25 | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | [pdf] | [code]
  48. [VisionReasoner] | ArXiv'2505 | VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning | [pdf] | [code]
  49. [PixelThink] | ArXiv'2505 | PixelThink: Towards Efficient Chain-of-Pixel Reasoning | [pdf] | [code]
  50. [Seg-R1] | ArXiv'2506 | Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning | [pdf] | [code]
  51. [HRSeg] | ArXiv'2507 | HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation | [pdf] | [code]
  52. [OmniAVS] | ICCV'25 | Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation | [pdf] | [code]
  53. [LENS] | ArXiv'2508 | LENS: Learning to Segment Anything with Unified Reinforced Reasoning | [pdf] | [code]
  54. [X-SAM] | ArXiv'2508 | X-SAM: From Segment Anything to Any Segmentation | [pdf] | [code]
  55. [Text4Seg++] | ArXiv'2509 | Text4Seg++: Advancing Image Segmentation via Generative Language Modeling | [pdf] | [code]
  56. [SVP] | ArXiv'2509 | Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation | [pdf] |
  57. [CoPRS] | ArXiv'2510 | CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation | [pdf] | [code]
  58. [PaDT] | ArXiv'2510 | Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs | [pdf] | [code]
  59. [LENS] | ArXiv'2510 | Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs | [pdf]
  60. [ARGenSeg] | NeurIPS'25 | ARGenSeg: Image Segmentation with Autoregressive Image Generation Model | [pdf]
  61. [UniPixel] | NeurIPS'25 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | [pdf] | [code]
  62. [READ] | CVPR'25 | Reasoning to Attend: Try to Understand How Token Works | [pdf] | [code]
  63. [UniGeoSeg] | ArXiv'2511 | UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes | [pdf] | [code]
  64. [STAMP] | ArXiv'2511 | Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction | [pdf] | [code]
  65. [GETok] | ArXiv'2512 | Grounding Everything in Tokens for Multimodal Large Language Models | [pdf] | [code]
  66. [Tarot-SAM3] | ArXiv'2604 | Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation | [pdf]
  67. [WISE] | ArXiv'2604 | Efficient Reasoning via Thought Compression for Language Segmentation | [pdf] | [code]
  68. [DPAD] | CVPR'26 | Discriminative Perception via Anchored Description for Reasoning Segmentation | [pdf] | [code]
  69. [WeatherReasonSeg] | ArXiv'2603 | WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models | [pdf]
  70. [GroundedSurg] | ArXiv'2603 | GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation | [pdf] | [code]
  71. [PRS-Med] | ArXiv'2505 | PRS-Med: Position Reasoning Segmentation for Medical Images with Vision-Language Model | [pdf] | [code]
  72. [LlamaSeg] | ArXiv'2505 | LlamaSeg: Image Segmentation via Autoregressive Mask Generation | [pdf]
  73. [MedSeg-R] | ArXiv'2506 | MedSeg-R: Medical Image Segmentation with Domain Alignment and RL for Reasoning | [pdf]
  74. [SegEarth-R2] | ArXiv'2512 | SegEarth-R2: Geospatial Multimodal Large Language Model for Reasoning Segmentation | [pdf] | [code]
  75. [DR^2Seg] | ArXiv'2601 | DR^2Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models | [pdf]
  76. [PixDLM] | CVPR'26 | PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation | [pdf]
  77. [UGround] | ICML'26 | UGround: Towards Unified Visual Grounding with Unrolled Transformers | [pdf] | [code]
  78. [Dr. Seg] | CVPR'26 | Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design | [pdf] | [code]
  79. [MTRS] | ArXiv'2606 | An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation | [pdf]
  80. [AnchorSeg] | ACL'26 | AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation | [pdf] | [code]
  81. [SetCon] | ArXiv'2605 | SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction | [pdf]
  82. [IC-Seg] | ArXiv'2605 | Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification | [pdf] | [code]
  83. [SegCompass] | CVPR'26 | SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation | [pdf] | [code]
  84. [FlowSeg] | ICML'26 | FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation | [pdf] | [code]
  85. [MedVol-R1] | ArXiv'2605 | MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation | [pdf]
  86. [CR-Seg] | ArXiv'2606 | CR-Seg: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation | [pdf]
  87. [StAR] | ArXiv'2603 | StAR: Segment Anything Reasoner | [pdf] | [code]
  88. [InstructSAM] | ArXiv'2605 | InstructSAM: Segment Any Instance with Any Instructions | [pdf] | [code]
  89. [B-GRTO] | ArXiv'2605 | B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation | [pdf]
  90. [VASA] | ArXiv'2605 | Vision Harnessing Agent for Open Ad-hoc Segmentation | [pdf]
  91. [Qwen3-VL-Seg] | ArXiv'2605 | Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding | [pdf]
  92. [Rea2Seg] | ArXiv'2606 | Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning | [pdf] | [code]
  93. [CCRC] | ArXiv'2606 | CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation | [pdf]
  94. [GEAR-Seg] | ArXiv'2607 | GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine | [pdf]
  95. [PanoSeeker] | ECCV'26 | Seek to Segment: Active Perception for Panoramic Referring Segmentation | [pdf] | [code]
  96. [DGSeg] | ECCV'26 | DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation | [pdf] | [code]
  97. [SegAnswer] | ArXiv'2607 | Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning | [pdf]
  98. [CycleGRPO] | ECCV'26 | Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO | [pdf] | [code]
  99. [OP-HRG] | ECCV'26 | Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning | [pdf]
  100. [VespaSeg] | ArXiv'2608 | VespaSeg: A Resource-Aware Ground-then-Segment Pipeline for Referring Expression Segmentation | [pdf] | [code]
  101. [PixVL] | ArXiv'2608 | PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask-Text Consistency Cycle | [pdf]
  102. [STAMPlus] | ArXiv'2608 | Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation | [pdf]
  103. [EgoAfford] | ArXiv'2608 | EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation | [pdf] | [code]
  104. [MedPixel] | ArXiv'2608 | MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation | [pdf] | [code]
  105. [MedUP] | ArXiv'2608 | MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models | [pdf]
  106. [FIRM] | ArXiv'2608 | FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation | [pdf]
  107. [FloodReasonBench] | ArXiv'2608 | FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge | [pdf]
  108. [DRAgent] | ArXiv'2608 | DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation | [pdf]
  109. [PAYN] | ICML'26 | Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation | [pdf] | [code]
  110. [MedREAL] | ECCV'26 | From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation | [pdf]
  111. [SeGDeP] | ArXiv'2609 | SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation | [pdf]
  112. [AgriScope] | ArXiv'2609 | AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images | [pdf] | [project]
  113. [MLLM Point Prompts] | ArXiv'2609 | Adapting Open-Weight MLLMs to Generate Point Prompts for Electron Microscopy Segmentation | [pdf]
  114. [PSMA PET/CT VLM] | ArXiv'2609 | A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation | [pdf]

Open-Vocabulary Semantic Segmentation

  1. [PSALM] | ECCV'24 | PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model | [pdf] | [code]
  2. [LLMFormer] | IJCV'24 | LLMFormer: Large Language Model for Open-Vocabulary Semantic Segmentation | [pdf]
  3. [LaSagnA] | ArXiv'2404 | LaSagnA: Language-based Segmentation Assistant for Complex Queries | [pdf] | [code]
  4. [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model | [pdf] | [code]
  5. [Text4Seg] | ICLR'25 | Text4Seg: Reimagining Image Segmentation as Text Generation | [pdf] | [code]
  6. [HiMTok] | ICCV'25 | HiMTok: Learning Hierarchical Mask Tokens for Image Segmentation with Large Multimodal Model | [pdf] | [code]
  7. [ALTo] | ArXiv'2505 | ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation | [pdf] | [code]
  8. [X-SAM] | ArXiv'2508 | X-SAM: From Segment Anything to Any Segmentation | [pdf] | [code]
  9. [OV-Stitcher] | ArXiv'2604 | OV-Stitcher: A Global Context-Aware Framework for Training-Free Open-Vocabulary Semantic Segmentation | [pdf]
  10. [SPAR] | CVPR'26 | SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation | [pdf] | [code]
  11. [MM-OVSeg] | CVPR'26 | MM-OVSeg: Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing | [pdf] | [code]
  12. [DSLO] | CVPR'26 | Direct Segmentation without Logits Optimization for Open-Vocabulary Semantic Segmentation | [pdf]
  13. [ProC-SAM3] | GRSL'26 | Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation | [pdf] | [code]
  14. [STAMPlus] | ArXiv'2608 | Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation | [pdf]
  15. [DAF] | ArXiv'2608 | Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation | [pdf] | [code]
  16. [SatOV] | ArXiv'2609 | SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery | [pdf]
  17. [HyperCLIP++] | ArXiv'2609 | HyperCLIP++: Fine-tuning CLIP for Open-Vocabulary Semantic Segmentation in Hyperbolic Space | [pdf]

Video Segmentation

  1. [VISA] | ECCV'24 | VISA: Reasoning Video Object Segmentation via Large Language Models | [pdf] | [code]
  2. [VideoLISA] | NeurIPS'24 | One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos | [pdf] | [code]
  3. [VITRON ] | NeurIPS'24 | Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing | [pdf] | [code]
  4. [ViLLa] | NeurIPS'25 | ViLLa: Video Reasoning Segmentation with Large Language Model | [pdf] | [code]
  5. [HyperSeg] | ArXiv'2411 | HyperSeg: Towards Universal Visual Segmentation with Large Language Model | [pdf] | [code]
  6. [InstructSeg] | ArXiv'2412 | InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models | [pdf] | [code]
  7. [Sa2VA] | ArXiv'2501 | Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos | [pdf] | [code]
  8. [VRS-HQ] | CVPR'25 | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | [pdf] | [code]
  9. [GLUS] | CVPR'25 | GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation | [pdf] | [code]
  10. [DeSa2VA] | ArXiv'2506 | Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder | [pdf] | [code]
  11. [OmniAVS] | ICCV'25 | Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation | [pdf] | [code]
  12. [Veason-R1] | CVPR'26 | Reinforcing Video Reasoning Segmentation to Think Before It Segments | [pdf] | [code]
  13. [VoCap] | ArXiv'2508 | VoCap: Video Object Captioning and Segmentation from Any Prompt | [pdf] | [code]
  14. [PixFoundation 2.0] | ArXiv'2508 | PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? | [pdf] | [code]
  15. [DecAF] | ICLR'26 | DecAF: Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation | [pdf] | [code]
  16. [UniPixel] | NeurIPS'25 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | [pdf] | [code]
  17. [TrajSeg] | ArXiv'2603 | Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation | [pdf] | [code]
  18. [VIRST] | CVPR'26 | VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation | [pdf] | [code]
  19. [AgentRVOS] | ArXiv'2603 | AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation | [pdf] | [code]
  20. [SDAM] | AAAI'26 | Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory | [pdf]
  21. [SPARROW] | CVPR'26 | SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs | [pdf] | [code]
  22. [MAR3] | ArXiv'2603 | MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation | [pdf]
  23. [SteerSeg] | ArXiv'2605 | SteerSeg: Attention Steering for Reasoning Video Segmentation | [pdf] | [code]
  24. [Seg-ReSearch] | ArXiv'2602 | Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search | [pdf] | [code]
  25. [RCoT-Seg] | ArXiv'2605 | RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation | [pdf] | [code]
  26. [FeVOS] | ECCV'26 | FeVOS: Foresight Expression Video Object Segmentation | [pdf] | [code]
  27. [STAC] | ArXiv'2607 | STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation | [pdf] | [code]
  28. [ReflexTrack] | ArXiv'2607 | ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation | [pdf]
  29. [PhysMLLMs] | ArXiv'2608 | PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos | [pdf] | [code]
  30. [MLLM-Assisted Audio VOS] | ArXiv'2608 | MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge | [pdf]
  31. [MoVISA] | ArXiv'2609 | MoVISA: Multi-Token Reasoning for Video Object Segmentation | [pdf]

Feedback

If you have any suggestions or find missing papers, please feel free to reach out at lanm0002@e.ntu.edu.sg.

Contributors

mc-lan

63 commits

ildar-idrisov

2 commits