This repository is for the first comprehensive survey on Meta AI's Segment Anything Model (SAM).
1,222
1,367 commits
updated Sep 23, 2026
The First Comprehensive SAM Survey: A Comprehensive Survey on Segment Anything Model for Vision and Beyond. Chunhui Zhang, Li Liu, Yawen Cui, Guanjie Huang, Weilin Lin, Yiqian Yang, Yuehong Hu. [paper] [homepage][中文解读]
Abstract: Artificial intelligence (AI) is evolving towards artificial general intelligence, which refers to the ability of an AI system to perform a wide range of tasks and exhibit a level of intelligence similar to that of a human being. This is in contrast to narrow or specialized AI, which is designed to perform specific tasks with a high degree of efficiency. Therefore, it is urgent to design a general class of models, which we term foundation models, trained on broad data that can be adapted to various downstream tasks. The recently proposed segment anything model (SAM) has made significant progress in breaking the boundaries of segmentation, greatly promoting the development of foundation models for computer vision. To fully comprehend SAM, we conduct a survey study. As the first to comprehensively review the progress of segmenting anything task for vision and beyond based on the foundation model of SAM, this work focuses on its applications to various tasks and data types by discussing its historical development, recent progress, and profound impact on broad applications. We first introduce the background and terminology for foundation models including SAM, as well as state-of-the-art methods contemporaneous with SAM that are significant for segmenting anything task. Then, we analyze and summarize the advantages and limitations of SAM across various image processing applications, including software scenes, real-world scenes, and complex scenes. Importantly, many insights are drawn to guide future research to develop more versatile foundation models and improve the architecture of SAM. We also summarize massive other amazing applications of SAM in vision and beyond. Finally, we maintain a continuously updated paper list and an open-source project summary for foundation model SAM at here.
Awesome Segment Anything Models: A curated list of awesome segment anything models in computer vision and beyond. This repository supplements our survey paper. We intend to continuously update it.
:boom:SAM 3.1: ''SAM 3.1 Object Multiplex'' was released.
:boom:SAM Audio: ''SAM Audio: Segment Anything in Audio'' was released.
:boom:SAM 3D: ''SAM 3D: 3Dfy Anything in Images'' was released.
:boom:SAM 3: ''SAM 3: Segment Anything with Concepts'' was released.
:boom:SAM 2: ''Segment Anything in Images and Videos'' was released.
:boom:SAM: ''Segment Anything'' was released.
:boom:SAM & SAM2 for videos: The first survey on Segment Anything for Videos: A Systematic Survey was online.
- 2026.06.05: SAM 3D won the CVPR 2026 Best Paper Honorable Mention.
- 2026.03.27: SAM 3.1 Object Multiplex was released.
- 2025.12.15: SAM Audio was released.
- 2025.11.19: SAM 3 and SAM 3D were released.
- 2025.10.11: SAM 3 arrives! Officially announced and set to launch.
- 2025.04.22: SAM 2 won the ICLR 2025 Best Paper Honorable Mention.
- 2024.07.31: The first survey on SAM & SAM2 for Videos was online.
- 2024.07.29: The SAM 2 was released.
- 2023.07.14: "Segment Anything" was accepted by ICCV 2023 (Best Paper Honorable Mention).
- 2023.05.16: An initial version of this Awesome-Segment-Anything project.
- 2023.05.14: The first comprehensive SAM survey was online.
- 2023.04.05: The paper of "Segment Anything" was online.
If you find our work useful in your research, please consider citing:
@article{zhang2023comprehensive,
title={A Comprehensive Survey on Segment Anything Model for Vision and Beyond},
author={Zhang, Chunhui and Liu, Li and Cui, Yawen and Huang, Guanjie and Lin, Weilin and Yang, Yiqian and Hu, Yuehong},
journal={arXiv preprint arXiv:2305.08196},
year={2023}
}
@article{zhang2024segment,
title={Segment Anything for Videos: A Systematic Survey},
author={Zhang, Chunhui and Cui, Yawen and Lin, Weilin and Huang, Guanjie and Rong, Yan and Liu, Li and Shan, Shiguang},
journal={arXiv preprint arXiv:2408.08315},
year={2024}
}
The First Comprehensive SAM Survey: Chunhui Zhang, Li Liu, Yawen Cui, Guanjie Huang, Weilin Lin, Yiqian Yang, Yuehong Hu.
"A Comprehensive Survey on Segment Anything Model for Vision and Beyond." ArXiv (2024).
[paper]
[homepage]
[中文解读]
[2023.05]
The First Survey on SAM & SAM2 for Videos: Chunhui Zhang, Yawen Cui, Weilin Lin, Guanjie Huang, Yan Rong, Li Liu, Shiguang Shan.
"Segment Anything for Videos: A Systematic Survey." ArXiv (2024).
[ArXiv]
[ChinaXiv]
[ResearchGate]
[Project]
[中文解读]
[2024.07]
SAM4MIS: Yichi Zhang, Rushi Jiao.
"Towards Segment Anything Model (SAM) for Medical Image Segmentation: A Survey." CBM (2024).
[paper]
[project]
[2023.05]
Yichi Zhang, Zhenrong Shen.
"Unleashing the Potential of SAM2 for Biomedical Images and Videos: A Survey." ArXiv (2024).
[paper]
[code]
[2024.08]
Tianfei Zhou, Fei Zhang, Boyu Chang, Wenguan Wang, Ye Yuan, Ender Konukoglu, Daniel Cremers.
"Image Segmentation in Foundation Model Era: A Survey." ArXiv (2024).
[paper]
[2024.08]
Chaoning Zhang, Fachrina Dewi Puspitasari, Sheng Zheng, Chenghao Li, Yu Qiao, Taegoo Kang, Xinru Shan, Chenshuang Zhang, Caiyan Qin, Francois Rameau, Lik-Hang Lee, Sung-Ho Bae, Choong Seon Hong.
"A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering." ArXiv (2024).
[paper]
[2023.05]
Xiaorui Sun, Jun Liu, Heng Tao Shen, Xiaofeng Zhu, Ping Hu.
"On Efficient Variants of Segment Anything Model: A Survey." IJCV (2025).
[paper]
[2024.10]
Mudassar Ali and Tong Wu and Haoji Hu and Qiong Luo and Dong Xu and Weizeng Zheng and Neng Jin and Chen Yang and Jincao Yao.
"A review of the Segment Anything Model (SAM) for medical image analysis: Accomplishments and perspectives." Computerized Medical Imaging and Graphics (2024).
[paper]
[2024.12]
Zhang Jiaxing, Tang Hao.
"SAM2 for Image and Video Segmentation: A Comprehensive Survey." ArXiv (2025).
[paper]
[2025.03]
Kang Wang.
"A survey on SAM-based methods for medical image segmentation." IS-AII (2025).
[paper]
[2025.07]
Guoping Xu, Jayaram K. Udupa, Yajun Yu, Hua-Chieh Shao, Songlin Zhao, Wei Liu, You Zhang.
"Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future." ArXiv (2025).
[paper]
[2025.07]
WanSAM4RS-Tracker: Zhipeng Wan and Sheng Wang and Wei Han and Yuewei Wang and Xiaohui Huang and Xiaohan Zhang and Xiaodao Chen and Yunliang Chen.
"A systematic survey and meta-analysis of the segment anything model in remote sensing image processing: Challenges, advances, applications, and opportunities." ISPRS Journal of Photogrammetry and Remote Sensing (2025).
[paper]
[project]
[2025.09]
Yang, Yizai and Cheng, Lechao and Wang, Yaxiong and Hui, Tianrui and Li, Wenjing and Zhong, Zhun.
"A Survey for Point Prompt of Segment Anything Model." MMAsia Workshops (2025).
[paper]
[2025.12]
SAM: Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, Ross Girshick.
"Segment Anything." ICCV (2023) Best Paper Honorable Mention.
[paper]
[homepage]
[code]
[Zhihu]
[Reddit]
[2023.04]
SAM 2: Nikhila Ravi∗,†, Valentin Gabeur∗, Yuan-Ting Hu∗, Ronghang Hu∗, Chaitanya Ryali∗, Tengyu Ma∗, Haitham Khedr∗, Roman Rädle∗ Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár†, Christoph Feichtenhofer∗,†.
"SAM 2: Segment Anything in Images and Videos." ICLR (2025) Best Paper Honorable Mention.
[paper]
[demo]]
[code]
[project]]
[dataset]
[blog]
[2024.07]
SAM 3: Nicolas Carion*, Laura Gustafson*, Yuan-Ting Hu*, Shoubhik Debnath*, Ronghang Hu*, Didac Suris*, Chaitanya Ryali*, Kalyan Vasudev Alwala*, Haitham Khedr*, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu°, Tsung-Han Wu°, Yu Zhou°, Liliane Momeni°, Rishi Hazra°, Shuangrui Ding°, Sagar Vaze°, Francois Porcher°, Feng Li°, Siyuan Li°, Aishwarya Kamath°, Ho Kei Cheng°, Piotr Dollar†, Nikhila Ravi†, Kate Saenko†, Pengchuan Zhang†, Christoph Feichtenhofer†.
"SAM 3: Segment Anything with Concepts." ICLR (2026).
[paper]
[arXiv]
[code]
[homepage]
[中文解读]
[2025.10]
SAM 3D: SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Dollár, Georgia Gkioxari, MattFeiszli, Jitendra Malik.
"SAM 3D: 3Dfy Anything in Images." CVPR (2026). CVPR (2026) Best Paper Honorable Mention.
[paper]
[code]
[project]
[demo]
[blog]
[中文解读]
[2025.11]
SAM 3D Body: Xitong Yang⋆, Devansh Kukreja⋆, Don Pinkus⋆, Anushka Sagar, Taosha Fan, Jinhyung Park◦, Soyong Shin◦, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli†, Jitendra Malik†, Piotr Dollar†, Kris Kitani†.
"SAM 3D Body: Robust Full-Body Human Mesh Recovery." ArXiv (2025).
[paper]
[code]
[project]
[2025.11]
SAM Audio: Bowen Shi∗, Andros Tjandra∗, John Hoffman∗, Helin Wang∗, Yi-Chiao Wu∗, Luya Gao∗, Julius Richter†,Matt Le†, Apoorv Vyas†, Sanyuan Chen†, Christoph Feichtenhofer‡, Piotr Dollár‡, Wei-Ning Hsu‡, Ann Lee‡.
"SAM Audio: Segment Anything in Audio." ArXiv (2025).
[paper]
[code]
[project]
[demo]
[2025.12]
GPT-4V: OpenAI.
"GPT-4V(ision) System Card." ArXiv (2023).
[paper]
[homepage]
[2023.09]
Gemini: Gemini Team, Google.
"Gemini: A Family of Highly Capable Multimodal Models." ArXiv (2023).
[paper]
[homepage]
[blog]
[2023.12]
SEEM: Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, Yong Jae Lee.
"Segment Everything Everywhere All at Once." NeurIPS (2023).
[paper]
[code]
[2023.04]
SegGPT: Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, Tiejun Huang.
"SegGPT: Segmenting Everything In Context." ICCV (2023).
[paper]
[code]
[2023.04]
Grounding DINO: Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang.
"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection." ArXiv (2023).
[paper]
[code]
[2023.04]
ImageBind: Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra.
"ImageBind: One Embedding Space To Bind Them All." CVPR (2023).
[paper]
[homepage]
[code]
[2023.05]
LanguageBind: Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Wancai Zhang, Zhifeng Li, Wei Liu, Li Yuan.
"LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment." ArXiv (2023).
[paper]
[code]
Meta-Transformer: Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, Xiangyu Yue.
"Meta-Transformer: A Unified Framework for Multimodal Learning." ArXiv (2023).
[paper]
[homepage]
[code]
[中文解读]
[2023.07]
OpenSeeD: Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianfeng Gao, Jianwei Yang, Lei Zhang.
"A Simple Framework for Open-Vocabulary Segmentation and Detection." ICCV (2023).
[paper]
[code]
[2023.03]
RAM: Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, Lei Zhang.
"Recognize Anything: A Strong Image Tagging Model." ArXiv (2023).
[paper]
[homepage]
[code]
[2023.06]
PACGen: Yuheng Li, Haotian Liu, Yangming Wen, Yong Jae Lee.
"Generate Anything Anywhere in Any Scene." ArXiv (2023).
[paper]
[homepage]
[code]
[2023.06]
ASM: Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen, Hao Li, Xizhou Zhu, Zhiguo Cao, Yushi Chen, Tong Lu, Jifeng Dai, Yu Qiao.
"The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World." ArXiv (2023).
[paper]
[homepage]
[demo]
[2023.08]
OneFormer: Jitesh Jain, Jiachen Li, MangTik Chiu, Ali Hassani, Nikita Orlov, Humphrey Shi.
"OneFormer: One Transformer to Rule Universal Image Segmentation." CVPR (2023).
[paper]
[homepage]
[code]
[2022.11]
OVSeg: Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, Diana Marculescu.
"Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP." CVPR (2023).
[paper]
[homepage]
[code]
[2022.10]
WAM: Tom Sander, Pierre Fernandez, Alain Durmus, Teddy Furon, Matthijs Douze.
"Watermark Anything with Localized Messages." ArXiv (2024).
[paper]
[code]
[2024.11]
Sa2VA: Haobo Yuan, Xiangtai Li, Tao Zhang, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang.
"Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos." ArXiv (2025).
[paper]
[code]
[project]
[hugging face]
[2025.01]
SAMTok: Yikang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu, Ye Tian, Haochen Wang, Haobo Yuan, Jiacong Wang, Lu Qi, Hao Fei, Anran Wang, Zhuochen Wang, Yujing Wang, Cheng Chen, Shunping Ji, Xiangtai Li.
"SAMTok: Representing Any Mask with Two Words." ArXiv (2026).
[paper]
[code]
[project]
[hugging face]
[demo]
[2026.01]
DAM: Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui.
"Describe Anything: Detailed Localized Image and Video Captioning." ArXiv (2025).
[paper]
[code]
[project]
[huggingface]
[2025.04]
DINOv2: Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski.
"DINOv2: Learning Robust Visual Features without Supervision." TMLR (2024).
[paper]
[code]
[project]
[2023.04]
DINOv3: Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Mairal, Hervé Jégou, Patrick Labatut, Piotr Bojanowski.
"DINOv3." ArXiv (2025).
[paper]
[code]
[2025.08]
Rex-Omni: Qing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong, Zhaoyang Zeng, Yihao Chen, Tianhe Ren, Junzhi Yu, Lei Zhang.
"Detect Anything via Next Point Prediction." ArXiv (2025).
[paper]
[project]
[code]
[2025.10]
Mamba-3: Anonymous authors.
"Mamba-3: Improved Sequence Modeling using State Space Principles." ICLR (2026).
[paper]
[2025.11]
Depth Anything 3: Haotong Lin, Sili Chen, Junhao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, Bingyi Kang.
"Depth Anything 3: Recovering the Visual Space from Any Views." ICLR (2026).
[paper]
[code]
[2025.11]
Vision Banana: Valentin Gabeur, Shangbang Long, Songyou Peng, Paul Voigtlaender, Shuyang Sun, Yanan Bao, Karen Truong, Zhicheng Wang, Wenlei Zhou, Jonathan T. Barron, Kyle Genova, Nithish Kannen, Sherry Ben, Yandong Li, Mandy Guo, Suhas Yogin, Yiming Gu, Huizhong Chen, Oliver Wang, Saining Xie, Howard Zhou, Kaiming He, Thomas Funkhouser, Jean-Baptiste Alayrac, Radu Soricut.
"Image Generators are Generalist Vision Learners." ArXiv (2026).
[paper]
[code]
[2026.04]
RelateAnything: Maëlic Neau.
"RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs." ArXiv (2026).
[paper]
[code]
[models]
[dataset]
[2026.09]
:boom:AgentDSM: Wentao Sun, Zhengsen Xu, Yiping Chen, John S. Zelek, Jonathan Li.
"Agentic Building-Aware Satellite Gaussian Splatting for Auditable Urban DSM Reconstruction." ArXiv (2026).
[paper]
[code]
[2026.09]
:boom:SAM-V: Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem.
"SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
:boom:Sepideh Gohari, Goodarz Mehr, Azim Eskandarian.
"Real-World Perception for Autonomous Driving in Adverse Weather: Enhancing Standard Detectors via Foundation-Guided Auto-Annotation." TITS (2026).
[paper]
[2026.09]
:boom:SAFe: Anja Delić, Jurica Runtas, Marin Oršić, Ivan Marković, Ivan Petrović.
"SAFe: Segment-guided Aggregation of Feature Densities for Anomaly-aware Segmentation." ArXiv (2026).
[paper]
[2026.09]
:boom:PhysReflect: Shuheng Ge, Hongwei Ren, Li Zhang, Xiangqian Wu.
"PhysReflect: Geometry and Perception Guided Diffusion for Physically-Plausible Mirror Reflections." ArXiv (2026).
[paper]
[2026.09]
:boom:Dhruv Gamdha, James Afful, Shambhavi Joshi, Ulrike Passe, Adarsh Krishnamurthy, Baskar Ganapathysubramanian.
"Semi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms." ArXiv (2026).
[paper]
[2026.09]
:boom:SRPR-Net: Lufei Liu, Guojie Li, Suncheng Xiang, Fan Zhang.
"SRPR-Net: Semantic and Relational Prompt Refinement for Automated SAM-based Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
:boom:RoboDistill: Ziying Song, Lin Liu, Hongyu Pan, Shaoqing Xu, Lei Yang, Mingzhe Guo, Caiyan Jia.
"Towards robust multimodal 3D object detection via visual foundation models." ArXiv (2026).
[paper]
[2026.09]
:boom:CODY-SAM3: Laura Cif, Zohra Souei, Diane Demailly, Mayte Castro-Jimenez, Juan Dario Ortigoza-Escobar, Muhammad Mushhood Ur Rehman, Morgan Dornadic, Sophie Huby, Gun-Marie Hariz, Cecile Hubsch, Nathalie Dorison, Eduardo M. Moraud, Jocelyne Bloch, Gabriella Horvath, Olivier Oullier, Xavier Vasques.
"Foundation-model-based multi-label phenotyping of combined hyperkinetic movement disorders." ArXiv (2026).
[paper]
[code]
[data]
[2026.09]
:boom:FOM-SAM3: Haolong Meng, Fangbo Qin, Mengchen Bai, Houwu Wang, Cirong Liu, Shan Yu.
"Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning." ICRA (2026).
[paper]
[code]
[2026.09]
:boom:AgenticSwarm: Muhammad Ahsan Mustafa, Yasheerah Yaqoot, Faryal Batool, Roohan Ahmed Khan, Valerii Serpiva, Dzmitry Tsetserukou.
"AgenticSwarm: Semantic Perception and Adaptive Task Allocation for Heterogeneous Multi-UAV Missions." ArXiv (2026).
[paper]
[2026.09]
:boom:P3-SAM: Qian Xu, Hang Xiong, Anpeng Wang, Sam Kwong, Cong Zhang, Runmin Cong.
"P3-SAM: SAM with Perceptual Parallel Prompt for Few-Shot Strip Steel Surface Defect Segmentation." ICME (2026).
[paper]
[2026.09]
:boom:FedSAM-3D: Xinran Wu, Rencheng Zheng, Yuxiang Dai, Hui Zhang, Xueqin Xia, Yu Cheng, Chengyan Wang, He Wang.
"FedSAM-3D: Adapter-Constrained Federated Adaptation for Transferable Medical Segmentation Foundation Models." TBME (2026).
[paper]
[code]
[2026.09]
:boom:Shujun Lv, Bo Fang, Yongfei Wu, Kun Li, Qiankun Li, Junxin Chen.
"From SAM 1 to SAM 3: Benchmarking Zero-Shot Cross-Domain Medical Image Segmentation." Expert Systems (2026).
[paper]
[2026.09]
:boom:PVPSAM: Han, Fangzhou and Li, Xiaoci and Gu, Li and Li, Li and Mi, Ke and Gu, Shenming and Zhang, Hailong.
"PVPSAM: Method and Benchmark for Weakly Supervised Object-level Photovoltaic Panels Extraction in Remote Sensing Imagery." JSTARS(2026).
[paper]
[2026.09]
:boom:PBP-SAM: Liangchao Chen, Guanying Huo, Weifeng Kong, Ziheng Cao, and Jiaying Chen.
"PBP-SAM: polarization-driven boundary prompt SAM for camouflaged object detection." ArXiv (2026).
[paper]
[2026.09]
:boom:GeoFuse-SAM: Pengtao Ren, et al.
"GeoFuse-SAM: A Multimodal Data Fusion Frameworkfor Boundary-Aware Foundation Model Adaptation inMedical Image Segmentation." ArXiv (2026).
[paper]
[2026.09]
:boom:DPSAM2: Wenbo Lei, Long Yu & Shengwei Tian.
"DPSAM2: memory-guided dual-path adaptation of SAM2 for boundary-aware low-contrast segmentation." The Visual Computer (2026).
[paper]
[code]
[2026.09]
:boom:Marcel Hudcovič, et al.
"From Prompt to Plot: Proof-of-Concept forSegmentation of Agricultural Landscapes in AerialImagery Using SAM 3 Agent." IGARSS (2026).
[paper]
[2026.09]
:boom:Sampaio, Filipe A. and Astudillo, Carlos A. and Souza, Alan and Miranda, Daniel and Borin, Edson.
"Improving SAM-Based Seismic Facies Segmentation With Logits Feedback." LGRS (2026).
[paper]
[2026.09]
:boom:PE-MedSAM2: Yuan, Xuejia and Yang, Zongjian and Guo, Yu and Kong, Fanhui and Ma, Jiquan.
"PE-MedSAM2: Parameter-Efficient Adaptation of MedSAM2 for 2D Medical Image Segmentation." TBME (2026).
[paper]
[code]
[2026.09]
:boom:ReliefSAM: Yihang Chen, Xiang Lyu, Rui Xu, Jiao Pan, Fadjar Ibnu Thufail, Brahmantara, Jiaqing Liu, Satoshi Tanaka & Liang Li.
"ReliefSAM: A Geometry-Augmented Multi-prior Adapter for Bas-Relief Segmentation." ECCV (2026).
[paper]
[code]
[2026.09]
:boom:ODG-SAM2-Morph: Yang, Dongxu, Xirui Xu, Shengmao Zhang, Zuli Wu, Tianfei Cheng, Jianglong Que, Siyao Wu, and Fei Wang.
"Morphometric Information for Yangtze Finless Porpoises Using Detection-Guided SAM2 Segmentation with UAV Imagery." Fishes (2026).
[paper]
[2026.09]
:boom:SnakeSAM: Jingwen Li, et al.
"SnakeSAM: A topology-preserving foundation model for medical curvilinear segmentation." Array(2026).
[paper]
[2026.09]
:boom:Chen, Xuan, and Shaolong Chen.
"Dynamic Consistency-Aware Multi-View Learning with SAM3 for 3D Medical Image Segmentation." Sensors (2026).
[paper]
[2026.09]
:boom:ASAM2-UNet: Xie, Caiyun, Linfeng Zhang, Zhaokun Chen, and Junyun Wu.
"ASAM2-UNet: An Attention-Enhanced SAM2 U-Net for Polyp Segmentation." Electronics (2026).
[paper]
[2026.09]
:boom:Shaghayegh Chavoshian, Ali Barzegar Khanghah & Atena Roshan Fekr.
"Transfer Learning on Segment Anything Model for Footwear Outsole Segmentation to Predict Footwear Slip Resistance." Annals of Biomedical Engineering (2026).
[paper]
[2026.09]
:boom:FST-SAM3: Guanhao Wu, Guilian Chen, Huisi Wu, and Jin Qin.
"FST-SAM3: Taming SAM 3 with Frequency-Spatio-Temporal Refinement for Video Polyp Segmentation." ECCV (2026).
[paper]
[code]
[2026.09]
:boom:FWSAM-Net: Shuchi Chen, Shengbing Chen, Qian Chen.
"FWSAM-Net: Wavelet-enhanced SAM2-based framework with frequency-aware adapter for Infrared Small Target Detection." Infrared Physics & Technology (2026).
[paper]
[2026.09]
:boom:SAM3_Remote_Sensing_LoRA: Nermeen Abou Baker.
"Parameter-Efficient Adaptation of SAM3 for Remote Sensing Segmentation Beyond Single-Domain Prompting." ICANN (2026).
[paper]
[code]
[2026.09]
:boom:MTGF-SAM: Zhang, Liangdong and Liu, Xiaohui and Zhang, Junxiao and Shao, Qinglong and Xing, Huaqiao and Zhu, Qing.
"MTGF-SAM: Multi-Level Terrain-Gated Fusion of Segment Anything Model for Landslide Detection in Remote Sensing Imagery." JSTARS (2026).
[paper]
[2026.09]
:boom:YLSAM2: Jinghui Yang and Liang Wang and Shuyin Hu and Bohao Zhang and Huiyuan Pang and Longqin Xu and Meng Cui and Shuangyin Liu.
"YLSAM2: Attention-guided LoRA enhanced underwater multi-scene fish segmentation and counting based on YOLO11 prompting SAM2." Aquacultural Engineering (2026).
[paper]
[2026.09]
:boom:Busra Aslan.
"YOLO–SAM-Guided ROI-Based Deep Learning for Non-Invasive Neonatal Jaundice Detection." BALKAN JOURNAL OF ELECTRICAL & COMPUTER ENGINEERING(2026).
[paper]
[2026.09]
:boom:ES-SAM: Xudong Yang, Xinnan Fan, Peiyu Zhao, Qi Sun, Pengfei Shi.
"ES-SAM: An Enhanced Semantic-SAM for semantic segmentation." PR (2026).
[paper]
[2026.09]
:boom:WOFT-SAM: Jonáš Šerých ⋅ Jiri Matas.
"Segmentation-Guided Homography Estimation for Long-Term Planar Tracking." ECCV (2026).
[paper]
[code]
[2026.09]
Ömer Faruk Deniz, Mustafa Taha Koçyiğit.
"Open-vocabulary 3D object detection with promptable segmentation." ArXiv (2026).
[paper]
[2026.09]
Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan.
"Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings." ArXiv (2026).
[paper]
[2026.09]
HYDRA: Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Damith Ranasinghe.
"Queries Knew More Than We Thought: Uncovering Latent Knowledge in Segmentation Models." ArXiv (2026).
[paper]
[2026.09]
Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins.
"Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions." ArXiv (2026).
[paper]
[code]
[2026.09]
SetPlanner: Dawei Yan, Yuezhe Yang, Menglan Ruan, Chunfeng Yang, Yudong Zhang.
"SetPlanner: A Lightweight Plug-in Point-Set Planner for Frozen SAM." ArXiv (2026).
[paper]
[code]
[2026.09]
PSMP-CLIP: Xuezhi Xiang, Guanghao Wu, Heqi Xiang, Jiayao Liu, Xiaoheng Li, Yiming Chen, Shanjun Zhang.
"PSMP-CLIP: Patch-Prompt SAM and Multi-Semantic Prompting for CLIP-Based Zero-Shot Anomaly Detection." ArXiv (2026).
[paper]
[2026.09]
VPRef: Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang.
"VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
ReliSAM: Shen Jiang and Xiaoyan Kui and Zhipeng Hu and Yangyang Shi and Ziwei Zou and Zexin Ji and Zeru Hai and Qinsong Li and Zuheng Ming and Haodong Xu and Beiji Zou.
"Reliability-guided dual-view learning with SAM distillation for semi-supervised medical image segmentation." Expert Systems with Applications (2026).
[paper]
[2026.09]
ViCo-SAM3: Qiangqiang Zhou, Wenjun Tang, Yong Chen, Dandan Zhu, Jiawei Xu.
"ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation." ArXiv (2026).
[paper]
[2026.09]
PEFT-SAM-Liver-CT: Ramtin Mojtahedi, Mohammad Hamghalam, Jacob J. Peoples, Richard K. G. Do, Amber L. Simpson.
"Parameter-Efficient Fine-Tuning of Foundation Models for Liver Tumor Segmentation in CT." SPIE Medical Imaging(2026).
[paper]
[code]
[2026.09]
PRR: Weihong Qi, Chen Ling.
"Perceive, Refine, Reason: A Calibrated Pipeline for Measuring Indicators in Strategic Visual Communication on Social Media." ArXiv (2026).
[paper]
[2026.09]
WorldMem: Aditi Tiwari, Akshit Bhalla, Darshan Prasad, Heng Ji.
"Does Video Memory Use What It Retrieves? A Causal Audit of Memory Specificity." ArXiv (2026).
[paper]
[2026.09]
ResoSeg: Chunkai Li, Junhao Yin, Ke Li, Jingde Chen.
"ResoSeg: Resonance Tagger using Transformer and Segment Model." ArXiv (2026).
[paper]
[code]
[2026.09]
MIE-SAM: Ze Li, Ying Ying Zhang, Shuai Zhang, Zhi Peng Wang.
"Multi-modal interaction enhanced segment anything model (MIE-SAM) for RGB-T salient object detection." Neural Networks (2026).
[paper]
[code]
[2026.09]
Silas Kwabla Gah, Ebenezer Owusu.
"Beyond Argmax: A Mechanistic Study of Semantic Retention in Frozen Foundation-Model Composition for Generalized Few-Shot 3D Segmentation." ArXiv (2026).
[paper]
[2026.09]
SAMV-DUSt3R: Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao.
"SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views." ArXiv (2026).
[paper]
[2026.09]
DINO-Med: Boya Wang, Ruizhe Li, Chao Chen, Xin Chen.
"DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging." ArXiv (2026).
[paper]
[2026.09]
BruNet: Qiming Wang, Richard J. Motley, Ebube E. Obi, Xianfang Sun, Paul L. Rosin.
"BruNet: A Cross-Domain Transfer Framework for Bruise Segmentation." ArXiv (2026).
[paper]
[2026.09]
DiSECT: Ramtin Mojtahedi, Mohammad Hamghalam, Jacob J. Peoples, Natalie Gangai, Mithat Gonen, Yun Shin Chun, HyunSeon Christine Kang, Richard K. G. Do, Amber L. Simpson.
"Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed Tomography." ArXiv (2026).
[paper]
[2026.09]
LeCor: Yi Luo, Yike Guo, Wenxuan Li, Zongwei Zhou, Rui Zhang, Kai Ding.
"LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour Segmentation." ArXiv (2026).
[paper]
[2026.09]
Guided SAM 3D: Jerred Chen, Simon Weber, Ronald Clark.
"Guiding Image-to-3D Generation with Test-Time Partial Observations." ArXiv (2026).
[paper]
[2026.09]
PBP-SAM: Liangchao Chen, Guanying Huo, Weifeng Kong, Ziheng Cao, and Jiaying Chen.
"PBP-SAM: polarization-driven boundary prompt SAM for camouflaged object detection." Applied Optics (2026).
[paper]
[2026.09]
LSVOS: Chang Liu, Henghui Ding, Lingyi Hong, Ning Xu, Linjie Yang, Yuchen Fan, Canyang Wu, Jinrong Zhang, Xusheng He, Ce Bian, Xianjing Han, Jianlong Wu, Mingqi Gao, Sijie Li, Jungong Han, JeongRae Kim, Chaehyun Kim, Changwon Lim, Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim, Pranjal Aggarwal, Sean Welleck, Yiwen Ren, Jianing Liu, Yingxin Wang, Kexin Zhang, Licheng Jiao, Lingling Li, Xu Liu, Jinxing Zhou, Suiyi Zhao, Yanghao Zhou, Ruohao Guo, Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu.
"Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation." ECCVW (2026).
[paper]
[code]
[2026.09]
MorphoOrgaAgent: Hanyi Zhang, Maximilian Hoermann, Lion J. Gleiter, Yiling Xu, Bettina Katalin Budai, Hans-Ulrich Kauczor, Carsten Marr, Tingying Peng.
"MorphoOrgaAgent: A Foundation-Model-Based Multi-Agent System for Autonomous Organoid Analysis." MICCAI workshop (2026).
[paper]
[code]
[2026.09]
SAM3-O2D2: Lucas Görnhardt, Timo Bartels, Tim Fingscheidt.
"SAM3-O2D2: Zero-Shot Object Out-of-Distribution Detection by Object Class Prompting of the SAM3-Image Model." ArXiv (2026).
[paper]
[2026.09]
MR-RS-SDFR: Quanxin Zheng, Shuai Zhao.
"MRI-Guided Reslice-Refined Cross-Slice SDF Reconstruction of the Left Ventricle from Cardiac MRI with Sparse Axial Supervision." ArXiv (2026).
[paper]
[2026.09]
Andreas Gilson, Laura Hennig, Peter Pietrzyk.
"Zero-Shot 3D Plant Organ Segmentation with SAM3 and Semantic NeRFs." ECCVW (2026).
[paper]
[2026.09]
CoRe-SAM3: Shipeng Liu, Liang Zhao, Dengfeng Chen.
"CoRe-SAM3: Conditional Semantic--Visual Reconciliation for SAM3 Crack Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
DriveZero: Hao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang, Kuangye Chen, Xiongjie Zhang, Wentao Dai, Hengchen Dai, Siyuan Liu, Zehao Huang, Naiyan Wang.
"DriveZero: End-to-End Driving Beyond Human Demonstrations." ArXiv (2026).
[paper]
[code]
[2026.09]
CrACK: Feifei Liu, Jintao Cheng, Chi Man Vong, Xiaoyu Tang.
"CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models." ArXiv (2026).
[paper]
[2026.09]
SAM-Radar: Jue Wang, Xuan Wang, Hao Zhou, Ruixiang Zhou, Yixuan Zhou, Tianshuo Yuan, Jieming Ma, Jie Zhang, Fei Luo.
"Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking." ArXiv (2026).
[paper]
[2026.09]
SeGDeP: Linnan Zhao, Xu Liu, Lingling Li, Licheng Jiao, Fang Liu, Wenping Ma.
"SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation." ArXiv (2026).
[paper]
[2026.09]
Diffuse2Seg: Christoph Hümmer, Joachim Sicking, Fabian Hüger, Hanno Gottschalk.
"Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision." ArXiv (2026).
[paper]
[2026.09]
Hesham, S.A.S., Liu, Y., Sun, G. et al.
"Evaluating SAM2 for Video Semantic Segmentation." Mach. Intell. Res.(2026).
[paper]
[2026.09]
Carballo Pérez, Áurea Genoveva, Pablo Magariños-Docampo, Pedro Orgeira-Crespo, and Fernando Aguado-Agelet.
"Weakly Supervised Segmentation of Macroalgae Through Gradient Analysis in Convolutional Neural Networks and Segment Anything Model." Applied Sciences(2026).
[paper]
[2026.09]
STCM: Dong, Wuzhou, Yulin Chen, Zhipan Wang, and Qingling Zhang.
"Semantic–Texture Complementation and Prediction-Guided SAM Fusion for Plastic Mulch Segmentation in GF-7 Imagery." Remote Sensing (2026).
[paper]
[2026.09]
SarSAM: Fatih Fehmi Ş İMŞEK, Melih ALTAY, Saygin ABDIKAN.
"SarSAM: An Integrated Framework for Agricultural Field Boundary Delineation and Mapping Using High-Resolution SAR (PAZ) Imagery and the Segment Anything Model." Advances in Space Research (2026).
[paper]
[2026.09]
SMART: Gao, Fang and Shi, Lei and Jin, Yan and Zheng, Hanbo and Huang, Qingbao and Yu, Jun.
"Say, Move, and Remember: Enhancing SAM2 for Referring Video Segmentation via Motion Modeling and Key Action-Aware Memory." TMM (2026).
[paper]
[code]
[2026.09]
TeF-SAM: Jinxin Liang, Xiaoming Liu, Zhiyuan Zang, Xin Wang, Yicheng Qi, Xiang Li.
"TeF-SAM: Prototype memory for medical lesion segmentation with text-free inference." Computerized Medical Imaging and Graphics (2026).
[paper]
[code]
[2026.09]
FastSAM-GAL: Zhu, Zhifu and Yuan, Xiping and Gan, Shu and Luo, Weidong and Chen, Cheng and Li, Xuan.
"FastSAM Guided Adversarial Learning for Unsupervised Multimodal Remote Sensing Change Detection." TGRS (2026).
[paper]
[2026.09]
A2SAM: Yong Chen and Renyi Chen and Qingbo Kang and Rui Wang and He Lyu and Zekun Jiang and Hongqiu Wang and Huan Song and Kang Li.
"Informativeness-driven active adaptation of SAM: Structural prompts and contrastive parameter selection for medical tubular segmentation." MIA (2026).
[paper]
[code]
[2026.09]
Yuxiang Huang, et al.
"Automated Zero-Shot Video Segmentation of Indoor Walls via SAM2 with Pretrained Semantic Prompts." Construction Research Congress(2026).
[paper]
[2026.09]
SAM3TGNet: Zhang, Jiayin, Nan Mo, Gege Ma, and Bangyan Tang.
"SAM3TGNet: A SAM3 Feature Encoding and Global Context Spatiotemporal Attention-Enhanced Change Detection Method for Optical Remote Sensing Images." Sensors (2026).
[paper]
[2026.09]
SAM-AUT: Amir-M. Naddaf-Sh, Vinay S Baburao, Hassan Zargarzadeh.
"SAM for Weld Defect Detection in Ultrasonic B-Scans." ArXiv (2026).
[paper]
[code]
[2026.09]
SWIFT: Lucie Bracq, Romain Guiet, Sandra Offner, Béatrice Kunz, Gisou van der Goot & Nathalie Brandenberg.
"A Single-organoid Workflow for quantitative Imaging classiFication and Tracking (SWIFT)." Communications Biology (2026).
[paper]
[2026.09]
SarSAM: Fatih Fehmi ŞİMŞEK and Melih ALTAY and Saygin ABDIKAN.
"SarSAM: An Integrated Framework for Agricultural Field Boundary Delineation and Mapping Using High-Resolution SAR (PAZ) Imagery and the Segment Anything Model." Advances in Space Research (2026).
[paper]
[2026.09]
ReliSAM: Shen Jiang and Xiaoyan Kui and Zhipeng Hu and Yangyang Shi and Ziwei Zou and Zexin Ji and Zeru Hai and Qinsong Li and Zuheng Ming and Haodong Xu and Beiji Zou.
"Reliability-guided dual-view learning with SAM distillation for semi-supervised medical image segmentation." Expert Systems with Applications (2026).
[paper]
[2026.09]
SAMLoRA: Wang, Xuewu, Wenlu Zhao, Cai Wang, Xu Chen, Yan Xu, Zuoman Zhang, Xirui Qiao, Bing Cao, Huifang Wang, and Hao Liu.
"High-Resolution Mapping and Spatial Pattern Analysis of Areca Palm Plantations in Sanya, China, Using SAMLoRA." Forests (2026).
[paper]
[2026.09]
Kai Zhao and Chenchen Kang and Suzy Rogiers and Oula Ghannoum and Yi Guo.
"Adapting SAM3 for 3D fruit counting with cross-view contrastive learning and Hough voting." Computers and Electronics in Agriculture (2026).
[paper]
[2026.09]
SAM-FSYOLO: Zhongyi Wang, Luohua Zhang, Changning Wei, Richu Jin, Dongjun Zhang, Tijun Bie, and Yonghui Yang.
"SAM-FSYOLO: An Integrated Framework for Miniature Covert Imaging Device Detection in Hotel Environments." ArXiv (2026).
[paper]
[2026.09]
CLON: Seojin Ji, Yoojin Kwon, Hyung-Sin Kim.
"CLON: Cue-Calibrated Linguistic Object Onboarding for Zero-Shot 6D Pose Front-Ends." ArXiv (2026).
[paper]
[2026.09]
WireSeg-32K: Zilin Dai, Lehong Wang, Yi Yang, Xiang Fei.
"WireSeg-32K: A Physics-Grounded Synthetic Dataset for Wire Instance Segmentation." CVPRW (2026).
[paper]
[code]
[2026.09]
Hailong Ning, Hao Wang, Yimeng Wang, Tao Lei, Renwei Dian, Asoke K. Nandi.
"Progressive Pseudo-Label Optimization for Point-Supervised Change Detection." ArXiv (2026).
[paper]
[2026.09]
FreNet: Yinan Liu, Jiankang Hong, Zhen Gao, Ye Lu.
"Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation." ArXiv (2026).
[paper]
[2026.09]
ENEAS: Javier del Pino, Salvador Rodríguez, Alejandro Garabito, Javier Álvarez, Chema Garabito.
"ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
SAM3-LoRA: P. Malaisree, S. Youwai, S. Janrungautai, D. Amorndechaphon, P. Rojanavasu, W. Songkitti.
"SAM3-LoRA: Parameter-Efficient Adaptation of a Concept-Promptable Foundation Model for Multi-Class Structural Defect Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl.
"Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting." ECML-PKDD(2026).
[paper]
[2026.09]
NeuSOGA: Qingde Li, Qingqi Hong, Jie Tian.
"Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations." ArXiv (2026).
[paper]
[code]
[2026.09]
AcrossVAM1.0: Yafei Zhang, Nan Wu.
"AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction." ArXiv (2026).
[paper]
[2026.08]
OPUS: Xiaoyan Wei, Zhimin Yao, Ruilin Yang, Wei Zhang, Yong Dai, Yi Zhang, Wei Ge.
"OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection." ArXiv (2026).
[paper]
[2026.08]
Marin Maletic, Goran Vasiljevic.
"Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring." ICASSP (2026).
[paper]
[2026.08]
MemMTL: Yangyang Xu, Haobo Yuan, Yuzhu Wang, Duo Su, Xi Ye, Yibo Yang, Jun Zhu.
"Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction." ArXiv (2026).
[paper]
[2026.08]
MariSat: Amir Abbes, Ines Harrabi, Lucas Justin Yirepoa Kinda, Rim Trabelsi, Adnane Cabani, Fatma Abdelkefi.
"MariSat: A Maritime Dataset for Instance Segmentation of Objects in Satellite and Aerial Images." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM-3D-Body: R. James Cotton, J. D. Peiffer, Lucinda Williamson, John Leske, Georgios Pavlakos.
"Biomechanical 3D Body: Self-Supervised Distillation of Biomechanical Pose from a 3D Body Foundation Model." ECCVW (2026).
[paper]
[2026.08]
FoundYou: Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano, Gabriele Berton, Luis Payá, Carlo Masone.
"FoundYou: A Unified Model for Personalized Segmentation and Retrieval." ECCV (2026).
[paper]
[code]
[2026.08]
SAM-STIR: Hu, Xiao and Zhou, Yun and Lv, Jian and Lai, Wenjie and Jiang, Yadong.
"SAM-STIR: Detection-Guided Memory Update for SAM 3-Based Infrared Small and Tiny Object Tracking." TGRS (2026).
[paper]
[2026.08]
RASP-SAM: Peng Zhang and Chen Liu and Zeyu Liu and Guanglei Zhang and Hongming Shan and Wenjian Wang.
"Adapting foundation models to weakly-supervised few-shot medical image segmentation via retrieval-augmented semantic prompting." Pattern Recognition (2026).
[paper]
[2026.08]
AffectOmni: Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua.
"AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes." ArXiv (2026).
[paper]
[code]
[2026.08]
SD-SAM 3: Ali Lesani, Chul Min Yeum, Su-Min Kang.
"Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery." ArXiv (2026).
[paper]
[2026.08]
T2S: Kumju Jo, Heesun Jung, Sungyong Baik.
"Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator." KBS (2026).
[paper]
[2026.08]
FAN-LoRA: Ziquan Liu, Zhewei Zhu, Xuyang Shi.
"FAN-LoRA: A Fourier-Adaptive Nonlinear Low-Rank Adaptor for Medical Foundation Model Domain Adaptation." ArXiv (2026).
[paper]
[2026.08]
LiDAR-SAM2: Jihun Kim, Hyun-Kurl Jang, Hyemin Yang, Jinnyeong Yang, Hyeokjun Kweon, Kuk-Jin Yoon.
"Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models." ECCV Workshop(2026).
[paper]
[2026.08]
HPMA: Xinning Yao, Jingjing Wang, Jinghua Yue, Xiaoyan Luo, Fugen Zhou, Bo Liu.
"Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation." ArXiv (2026).
[paper]
[2026.08]
ReGround-Surg: Jiaxin Wen, Ming Yin, Lu Liu, Zeyu Fu.
"ReGround-Surg: Reliability-Guided Anchor Grounding for Referring Surgical Video Segmentation." PRCV (2026).
[paper]
[code]
[2026.08]
OptiSight: Alperen Avan, Jordi Sanchez-Riera.
"OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation." ArXiv (2026).
[paper]
[code]
[2026.08]
Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu.
"MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
SAM3Dual: JeongRae Kim, Chaehyun Kim, Changwon Lim.
"SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
Mingqi Gao, Sijie Li, Jungong Han.
"Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
RAVP: Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto, Daniela Giordano, Concetto Spampinato, Federica Proietto Salanitri.
"Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging." ArXiv (2026).
[paper]
[2026.08]
SPARK-SAM: Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu.
"SPARK-SAM: Self-Prompt Adaptation with Response Knowledge for SAM in Infrared Small Target Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
GAP-SAM: Haozhen Yan, Siyuan Shan, Zijian Yu, Youqi Wang, Yan Hong, Jun Lan, Jianfu Zhang.
"GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization." ArXiv (2026).
[paper]
[2026.08]
Manani, Hiren.
"From Detection to Segmentation: A Foundation Model Approach to Organoid Brightfield Image Analysis Using SAM 3." ArXiv (2026).
[paper]
[2026.08]
DPD-SAM: Ma, Yize and Liang, Zhenyang and Yin, Feiyu and Song, Pengfei and Song, Junyi and Yang, Shengxiang and Wu, Guoqing and Ren, Yan and Yu, Jinhua.
"Latent Domain-Specific Prompt-Driven SAM via Conditional Diffusion Refinement for Postoperative Glioma Segmentation." TII (2026).
[paper]
[2026.08]
Hilla Fred and Mogens Agerbo Krogh and Britt {Bang Jensen} and Laura Ruotsalainen and Jouni Vielma and Matti Pastell.
"Automatic visual detection of fish in Recirculated Aquaculture Systems using the Segment Anything Model." Aquaculture (2026).
[paper]
[2026.08]
SNFusion: Yang Fang and Jingjing Chen and Dahang Wan and Xianli Lang and Shuangbao Shu and Rongsheng Lu and Qianqian Wu.
"SNFusion: A SAMv2-guided boundary-aware fusion framework for visible-infrared maritime vessel detection." Ocean Engineering (2026).
[paper]
[code]
[2026.08]
ACRIS-SAM2: Jiaxiang Luo, Weiwen Chen.
"ACRIS-SAM2: Attribute-driven cross-modal interaction and dual semantic prompting for few-shot segmentation." Neurocomputing (2026).
[paper]
[2026.08]
SAM2 R-CNN: Mehdi Gharbage, Céline Teulière, Pierre Bouges & Thierry Chateau .
"SAM2 R-CNN: Transferring SAM 2 Knowledge for Data Efficient Instance Segmentation." ICPR (2026).
[paper]
[code]
[2026.08]
Yang, Chenzheng, Shenhua Yang, Pu Wang, Weijun Wang, and Zeyang Huang.
"Constrained Boundary Enhancement for SAM 2-Based Ship Segmentation in UAV Berthing and Unberthing Videos." Applied Sciences (2026).
[paper]
[2026.08]
Swapnil Biswas.
"Enhancing Skin Lesion Classification in Teledermatology via SAM-Based Segmentation and Wavelet-Based Convolutional Autoencoder." ArXiv (2026).
[paper]
[2026.08]
FlexiCrackNet: Xiaoyan Jiang, et al.
"FlexiCrackNet: A Flexible Pipeline for Lightweight Crack Segmentation with Distilled Features from SAM." ArXiv (2026).
[paper]
[code]
[2026.08]
EchoFlow-SAM: Wanting Li, et al.
"EchoFlow-SAM: Motion-guided semi-supervised segmentation for echocardiographic videos." Biomedical Signal Processing and Control (2026).
[paper]
[2026.08]
Geng, Chao, Yajie Wang, Quanming Li, Zhentao Li, Xianfeng Shi, Botao Fu, Wei Li, Cheng Chen, Hong Zhang, Yukai Wang, and et al.
"Automated Detection and Segmentation of Cracks in Urban Underground Structures Based on YOLOv8-SAM2." Buildings (2026).
[paper]
[2026.08]
SAM-CLIP-Thermal: Yiyuan Lin and Chenjiao Tan and Changying Li and Yu Jiang.
"SAM-CLIP-Thermal: Leveraging large multimodal models for reliable and scalable annotation in thermal image segmentation for field plant phenotyping." Plant Phenomics (2026).
[paper]
[code]
[2026.08]
HBF-BCER: Ping, Shengyang, Zhijie Lin, Liliang Lin, Lei Zhao, Lisha Ye, Bangguo Wang, and Tao Wang.
"A Prompt-Preserving SAM ViT-B Adaptation Framework with Historical Branch Fusion and Soft Convolutional Expert Weighting for Medical Image Segmentation." Bioengineering (2026).
[paper]
[2026.08]
Puspitasari, Fachrina Dewi and Zhang, Chaoning and Mandal, Avilasha and Zheng, Sheng and Qin, Caiyan and Kim, Tae-Ho and Lee, Jewon and Wang, Guoqing and Yang, Yang and Shen, Heng Tao.
"Accelerating SAM2 with Efficient Memory Attention Module via Spatiotemporal Token Pruning." TPAMI (2026).
[paper]
[2026.08]
Token-Adaptive LoRA: Xin Chen, Jun Yan, Zhiyu Yan, Jianwen Deng, Jiaqi Wu, Yonghong Gong, Xiaohua Jiang.
"Token-Adaptive LoRA: Enhancing Segment Anything for Remote Sensing Imagery through Parameter-Efficient Fine- Tuning." ArXiv (2026).
[paper]
[2026.08]
Zero-Click-SAM2: Pasierb, Daniel and Wijata, Agata M. and Nalepa, Jakub.
"Zero-Click Brain Tumor Segmentation Using Segment Anything Model 2." ICIP (2026).
[paper]
[code]
[2026.08]
Bui-Tran, Quang-Khai and Nguyen, Thanh-Huy and Le, Bac and Xu, Min.
"Adapting SAM Without Labels: Uncertainty-Aware Source-Free Medical Image Segmentation." ICIP (2026).
[paper]
[2026.08]
SAM2TC: Cocco, Marco and Dunnhofer, Matteo and Micheloni, Christian.
"Representation Compensation of SAM2 for Segmenting Objects under Transformation in Videos." ICIP (2026).
[paper]
[2026.08]
FE-SAM: Gao, Feng and Pan, Zizhe and Wang, Haoting and Hua, Ruzhuang and Cao, Jingchao and Dong, Junyu and Du, Qian.
"Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation." TGRS (2026).
[paper]
[code]
[2026.08]
Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot.
"A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3." ArXiv (2026).
[paper]
[2026.08]
SAM2Dual: JeongRae Kim and Changwon Lim.
"SAM2Dual: Training-Free, Dual Memory for Long-Term Video Object Segmentation." TIP (2026).
[paper]
[2026.08]
SAM2-DPT: Steven Landgraf, Joceline Hinz, Markus Ulrich.
"A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation." ISPRS (2026).
[paper]
[2026.08]
EpigraphNet: Utsav Poudel, Rasik Bhattarai, Siddhartha Pathak, Raghavendra Ramacharna, Gaurav Jaswal.
"Zero-Shot SAM2 Segmentation and Vision Transformer-Based Recognition of Elamite Cuneiform Symbols from Degraded Tablet Images." ArXiv (2026).
[paper]
[code]
[2026.08]
Ce Bian, Xusheng He, Jinrong Zhang, Canyang Wu, Xianjing Han, Jianlong Wu.
"Key-Frame Reasoning with SAM3: Third Place Solution for the MeViS-Text Track of the 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
Cesar Borja, Breck A. McCollum, Jarret E. Byrnes, Kenneth Sebens, Ana C. Murillo.
"Leveraging existing sparse point annotations for benthic imagery dense segmentation." ArXiv (2026).
[paper]
[2026.08]
SSSAM: Ruichao Hou, Boyue Xu, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao.
"S3AM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM-RNO: Wang, F., Liu, Q. & Wu, G.
"Road negative obstacle segmentation via a dual-branch segment anything model with RGB-depth cross-prompting." J Supercomput (2026).
[paper]
[2026.08]
SAM-Med2D-GeoCrop: Wang, Tianqi and Li, Jianuo and Yang, Chenhao and Xu, Jinyi and Zhou, Mian and Dang, Kang and Zhang, Linxue.
"ROI-Focused Geometry-Aware Adaptation for Accurate Small-Structure Segmentation in Medical SAM." ICIP (2026).
[paper]
[2026.08]
RISE: Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang.
"RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning." ArXiv (2026).
[paper]
[2026.08]
DreamX-Phi 1.0: DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang.
"DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation." ArXiv (2026).
[paper]
[code]
[2026.08]
VOS-Agent: Canyang Wu, Jinrong Zhang, Xusheng He, Ce Bian, Xianjing Han, Jianlong Wu.
"VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)." ECCV Workshop(2026).
[paper]
[2026.08]
TCSR-Monito: Hieu D. Pham, Dang P. M. Cao, Thanh Trung Huynh.
"Beyond Uncertainty: Generalizable Failure Monitoring for Surgical Segmentation under Acquisition Degradation." MICCAI Workshop(2026).
[paper]
[code]
[2026.08]
SUGFW+: Xiaochuan Ma, Ning Zhu, Jia Fu, Lanfeng Zhong, Hanyu Jiang, Bin Song, Kang Li, Guotai Wang.
"SUGFW+: An Uncertainty-guided Feature Weighting Framework for Cold Start Active Adaptation of SAM in Medical Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
FE-SAM: Feng Gao, Zizhe Pan, Haoting Wang, Ruzhuang Hua, Jingchao Cao, Junyu Dong, Qian Du.
"Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation." IEEE TGRS (2026).
[paper]
[code]
[2026.08]
Mohammadreza Narimani, Shreyan Mitra, Parastoo Farajpoor.
"From crown candidates to neighborhood screening: integrating optical GeoAI and spatial modeling for urban-canopy assessment in Davis, California." ArXiv (2026).
[paper]
[code]
[dataset]
[2026.08]
Yiwen Ren, Jianing Liu, Yingxin Wang, Kexin Zhang, Licheng Jiao, Lingling Li, Xu Liu.
"Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
RoboSeg: Zhaochen Lan, Mengxiang Lin.
"RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera." ArXiv (2026).
[paper]
[2026.08]
Roni Blushtein-Livnon, Tal Svoray, Osher Rafaeli, Michael Dorman, Itay Fischhendler, Havazelet Yahel, Emir Galilee.
"Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing Imagery." ArXiv (2026).
[paper]
[2026.08]
Seed2GS: Zongjian Ding, Yudong Gao, Jiale Liu, Xinglin Yu, Junxing Ren, Dong Wei, Yajing Chen, Shan Huang, Mingjun Cheng, Min Li.
"Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes via a Single Reference-View Grounding." ArXiv (2026).
[paper]
[2026.08]
Mike Szklarzewski, CJ George, Gavin Smithson, Christopher Stokes, Dakota Fulp, William M. Jones, Benjamin Wynn, Alexander Ur, Agit Yesiloz, Clint Kallenbach, Mark Swartz, Nathan DeBardeleben, Sharmistha Chakrabarti.
"From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection." ICMLA (2026).
[paper]
[2026.08]
BAP-MOS: Satvik Praveen, Shengji Jin, Ahmed Lamidi, Xin Qian, Yi Sheng.
"BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
Shah Imran Ahsan Chowdhury, Kazi Jihadur Rashid, Rajsree Das Tuli, Rahul Saha, Bulbul Ahammad.
"GeoAI-based post-segmentation quality validation of building footprints via spatial feature engineering." ArXiv (2026).
[paper]
[2026.08]
LEGO: Yuning Peng, Haiping Wang, Yuan Liu, Yipeng Lu, Zhen Dong, Bisheng Yang.
"LEGO: Leveled Language Gaussian Splatting." ECCV (2026).
[paper]
[code]
[2026.08]
SSUPER: Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim.
"Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026." ArXiv (2026).
[paper]
[2026.08]
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang.
"Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation." Neurocomputing (2026).
[paper]
[2026.08]
AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel.
"Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations." ArXiv (2026).
[paper]
[2026.08]
VOLA: Yuchen Zhang, Yuan Gao, Sebastian Schmidt, Johannes Betz.
"VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM3tool: Nakul Poudel, Richard Simon, Cristian A. Linte.
"Toward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3)." MIUA (2026).
[paper]
[2026.08]
KD-SAM: Yang Fang and Bingbing Jiang and Haopeng Huo and Uswah Khairuddin and Yeon Lee and Yong Qin.
"KD-SAM: Keep-Awake and Detail-Enhanced Segment Anything Model for Multi-Modality Multi-Organ Medical Image Segmentation." Expert Systems with Applications (2026).
[paper]
[code]
[2026.08]
S3-Diff: Jiaming Liang, QiHui Han, Guangye Ou, Jiawen Liu, Haolin Chen, Xi Zhong, Jiazhou Chen, Xiaoqi Sheng, Hongmin Cai.
"S3-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images." ArXiv (2026).
[paper]
[2026.08]
GeoDistill-Refine: Yonglong Zhang, Zongwu Xie, Yang Liu.
"GeoDistill-Refine: Silhouette-First Geometry Distillation for Annotation-Free Spacecraft Segmentation." ArXiv (2026).
[paper]
[2026.08]
Hao Wang, Yuxuan Zhang, Wei Yang.
"Universal Concept Disruption for SAM3 Image Segmentation." ArXiv (2026).
[paper]
[2026.08]
EgoAfford: Xinyuan Guan, Feifan Chen, Xinyu Zhan, Fu-Cheng Zhang, Cewu Lu, Lixin Yang.
"EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
Jonathan Klingspon, Scott McAvoy, Maurizio Seracini, Falko Kuester.
"Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital Twins." ArXiv (2026).
[paper]
[2026.08]
CROSS: Tingzhang Luo, Ruizhong Liu, Yichao Liu, Cheng Fan, Yu Liu, Jianyuan Guo.
"CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation." ECCV (2026).
[paper]
[code]
[2026.08]
FS-CPL: Rahul Venkataramani, Rachana Sathish.
"Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding." ArXiv (2026).
[paper]
[2026.08]
ACTrack: Wenrui Cai, Yuzhe Li, Qingjie Liu, Yunhong Wang.
"Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking." ArXiv (2026).
[paper]
[2026.08]
GaussianSelector: Baihan Yang, Tiexin Li, Yuheng Liu, Xin Lin, Xinke Li, Xiaohui Xie, Truong Nguyen.
"GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization." ArXiv (2026).
[paper]
[2026.08]
EOVSAM: Haomin Peng, Yongkang Li, Zhaoxiang Liu, Xiaojie Jin, Shiguo Lian, Yunchao Wei, Xinggang Wang.
"EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass." ArXiv (2026).
[paper]
[code]
[2026.08]
PhenoStitch: Xuechen Li.
"PhenoStitch: Training-Free Panoptic Crop Mapping from Satellite Image Time Series." ArXiv (2026).
[paper]
[2026.08]
ELFSS-AR: Xueting Bai, Huan Ni.
"Training-Free Entity-Level Few-Shot Segmentation of Remote Sensing Images with Advection Refinement." ArXiv (2026).
[paper]
[code]
[2026.08]
UltraSAM3: Bo Xu, Quanhao Zhu, Rui Lin, Boling Zhu, Chenyuan Wang, Hongfei Lin, Feng Xia, Chenhua Ji.
"UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM+D: Yu Song, Hao Sun, Shiyu Teng, Ikuko Nishikawa, Yen-wei Chen.
"SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting." ECCV (2026).
[paper]
[code]
[2026.08]
TMMSAM2: Fu, Xiyou and Zhang, Ting and Zhang, Xiaoyu and Lin, Mingying and Lv, Zijun and He, Wangquan and Ren, Qi and Xu, Meng and Jia, Sen.
"TMMSAM2: Tracker-Aided Multitemporal Memory SAM2 for Hyperspectral Object Tracking." TNNLS (2026).
[paper]
[2026.08]
SAM3D-VLA: Zonghe Liu, et al.
"SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models." ArXiv (2026).
[paper]
[2026.07]
Jinghong Liu, Yuchuan Deng, Fanping Liu, Meng Huang, Xirong Li.
"Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.07]
ZMIS-SAM: Dekun Yuan, Zhongwei Li, Zheng Qiao, Jie Zhang.
"ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
Nevio Dubbini, Lisa Yeomans, Marco Pavia, Ramazan Parmaksiz, Ayse Atas Hooglugt, Gabriele Gattiglia, Beatrice Demarchi.
"Multimodal fusion of visual and morphometric features for avian bone classification." ArXiv (2026).
[paper]
[2026.07]
RDVSv2: Tianyu Li, Jiahao He, Keren Fu, Qijun Zhao.
"RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection." ACMMM (2026).
[paper]
[code]
[2026.07]
Sanjay Subramanian, Junwei Yu, Zirui Wang, Rohil Malpani, Maggie Chung, Adam Yala, Dan Klein, Trevor Darrell.
"Open-Ended CT Volume Segmentation with Weak Supervision from Language." ArXiv (2026).
[paper]
[2026.07]
SADe: Hang Xing, Guangjun Liu, Yan Xia, Xueming Ding.
"SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations." ArXiv (2026).
[paper]
[2026.07]
ConFusion: Guo Yurong, He Yufei, Li Yonghao, Chang Dongliang, Zhang Ke, Ma Zhanyu.
"ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion." ArXiv (2026).
[paper]
[code]
[2026.07]
EditCLEVR: Anuraag Gadehothur Karnam, Tarunesh Sathish.
"EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations." ArXiv (2026).
[paper]
[code]
[2026.07]
SurgSAM3: Changjing Liu, Yiming Huang, Beilei Cui, Liangjing Shao, Long Bai, Yanheng Li, Haoxuan Che, Hongliang Ren.
"Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation." MICCAI Workshop (2026).
[paper]
[code]
[2026.07]
Hiu Ching Cheung, Wenchao Yue, Zhengran Han, Mingcong Chen, Guanglin Cao, Hongbin Liu, Hongliang Ren.
"Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots." ICCBS (2026).
[paper]
[2026.07]
LSP-SAM2: Zhaoyuan Wu, Naiyang Guan, Yinghui Gao, Longfei Su, Min Liu.
"A Light-Weight Self-Prompting Foundation Model for Automatous Video Object Segmentation." ICIC (2026).
[paper]
[code]
[2026.07]
FESAM: Chaoyue Wang, Jiang Wang, Lingfang Li, Weijian Hu & Lizhen Cui .
"FESAM: Frequency-Enhanced SAM with Boundary-Aware Decoding for Ultrasound Image Segmentation." ICIC (2026).
[paper]
[2026.07]
Agent-SAM-I2V: Guo Yang, Jiaqi Zhang, Yao Zhu & Longze Fan.
"Agent-SAM-I2V: Self-correcting Promptable Video Segmentation via Agentic Drift Detection and Multi-Prompt Fusion." ICIC (2026).
[paper]
[2026.07]
GB-SAM: Chenlin Xu, Lei Zhang, Lituan Wang, Xinyu Pu, Pengfei Ma, Guangwu Qian.
"GB-SAM: Gaussian-Prior and Boundary-Guided Test-Time Adaptation for Medical Image Segmentation." ICIC (2026).
[paper]
[code]
[2026.07]
CME-SAM: Qiyuan Wang, Jinfu Wang, Chuyu Chen, Pengtao Ren, Shijie Ling, Kejiang Xiao.
"CME-SAM: Contrastive Mask-Enhanced Segment Anything Model for Generalizable Medical Image Segmentation." ICIC (2026).
[paper]
[2026.07]
IBS-EMA: Yizhuo Wang, Shiquan Min, Jiangping Zhu & Pei Zhou.
"IBS-EMA: Mitigating Test-Time Prompt Distribution Shift for Medical Segment Anything Models." ICIC (2026).
[paper]
[2026.07]
HCFNet: Yu, Jiwei, Kecheng Zhou, Ting Wang, Hongxiao Gan, Yu Wang, and Shuzhi Gao.
"HCFNet: A SAM2-Based Hierarchical Cross-Branch Frequency-Aware Network for Industrial Surface Defect Segmentation." Sensors (2026).
[paper]
[2026.07]
ZA-SAM: Jinliang Su, Yun Jiang, Zequn Zhang & Yuhang Li .
"Adaptive Prompted, Zero-Annotation SAM: Weakly Supervised Binary Medical Image Segmentation." ICIC (2026).
[paper]
[2026.07]
SURE-SAM2: Yucan Duan, Chun Wang, Kaiyu Miao & Xiaoyan He.
"SURE-SAM2: Semantic and Uncertainty-aware Refinement SAM2 for Change Detection." ICIC (2026).
[paper]
[2026.07]
PolypSAM-Lite: Hasan, Umar, and Muhammad Ali Nayeem.
"Low-Rank Attention Reparameterization for Parameter-Efficient Adaptation of the Segment Anything Model to Colorectal Polyp Segmentation." Mathematics (2026).
[paper]
[2026.07]
Koki AMANO, Otoha YAMANAKA, Wakana KAWAI, Tatsuya HAYASHI, Nobuo KOCHI, Ippeita DAN.
"Segment-Anything-based AOI Analysis for Eye-tracking: A Gaze Judgment Method Considering the Visual Angle." J-STAGE(2026).
[paper]
[2026.07]
Naka-SAM: Chen, Juan; Wu, Jiajie; Guo, Lei; Ge, Wenping; Ma, Jie.
"Naka-SAM: A Cognition-Inspired Framework with Nakagami Prior for Ultrasound Segmentation." Proceedings of the Annual Meeting of the Cognitive Science Society (2026).
[paper]
[2026.07]
CG-SAM2: Bin He, Zhiwei Chen, Shengmin Zhao, Qinqin Zhou, Aiwen Jiang, Miaohui Zhang.
"CG-SAM2: Confidence-Guided Pseudo-label Refinement for Weakly Supervised Camouflaged Object Detection." ICIC (2026).
[paper]
[2026.07]
Mohammadreza Narimani, Vikram Anand, Parastoo Farajpoor.
"Farmland Extent and Visible Boundary Mapping from 1 m NAIP Imagery Using Residual U-Net and Text-Prompted SAM 3 Refinement." ArXiv (2026).
[paper]
[code]
[dataset]
[2026.07]
FluxGraph: Yihong Sun, Bharath Hariharan.
"Efficient Tracking and Understanding Object Transformations." ArXiv (2026).
[paper]
[code]
[2026.07]
SENSATION-DS: Hakan Calim, Anamaria Dumitrescu, Adarsh Bhandary Panambur, Huzaifa Asif, Andreas Maier.
"Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation." ArXiv (2026).
[paper]
[2026.07]
Lean-SAM2: Xudong Ouyang, Wenlun Zhang, Yimin Xu, Huazhong Liu, Yunshan Zhong.
"Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2." ArXiv (2026).
[paper]
[code]
[2026.07]
Samy Mounir, Mikolaj Cieslak, Najmeddine Dhieb, Hakim Ghazzai, Jonathan Klein, Katja Froehlich, Soeren Pirk, Wojciech Palubicki, Gianluca Setti, Ahmed M. Eltawil, Dominik L. Michels.
"Text-conditioned Segmentation for Tomato Phenotyping via Procedural Synthetic Data." ArXiv (2026).
[paper]
[2026.07]
Scene-SAM3D: Yuqi Zhang, et al.
"Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning." ArXiv (2026).
[paper]
[code]
[2026.07]
OP-HRG: Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Chia-Wei Tang, Zaber Ibn Abdul Hakim, Anuj Karpatne, Chris Thomas.
"Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning." ECCV (2026).
[paper]
[code]
[2026.07]
Silas kwabla Gah, Ebenezer Owusu.
"Training-Free Open-Vocabulary 3D Point-Cloud Segmentation on the Generalized Few-Shot Benchmark." ArXiv (2026).
[paper]
[2026.07]
Jinchang Zhang, Arnold Zumbrun, Jing Lin, and Guoyu Lu.
"Foundation-Assisted Active Learning for Object Detection Annotation." ArXiv (2026).
[paper]
[2026.07]
Lite-Pi: Shivanshu Agnihotri, Snehashis Majhi, Deepak Ranjan Nayak, Dwarikanath Mahapatra, Debesh Jha.
"Induce to Empower: Improving Lightweight Baselines via Foundation Model Induction for Generalized Polyp Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
Joey Páolo Kardolus, Daan Hendriks, Jaap Jansen.
"Direct Clinical Joint Angle Extraction from Parametric Body Model Rotation Matrices." ArXiv (2026).
[paper]
[code]
[2026.07]
Zhe Xin, Hanzhi Chang, Penghui Huang, Yinian Mao, Guoquan Huang.
"Robust Multimodal Dynamic Object Segmentation." ICRA (2026).
[paper]
[2026.07]
Minghui Xu, Chaoyi Zhou, Aaron P. Cecil, Xi Liu, Siyu Huang, Yuhao Xu.
"Digital measurement of droplet flame diameter in microgravity combustion images using Segment Anything Model 2 with automatic prompt selection." ArXiv (2026).
[paper]
[2026.07]
SAMRI-3D: Zhao Wang, Wei Dai, Hongfu Sun, Craig Engstrom, Shekhar S. Chandra.
"SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens." ArXiv (2026).
[paper]
[code]
[2026.07]
IMTrack: Zhiqiang Hou, Chuangye Xu, Sugang Ma, Xiaobao Yang, Lei Pu.
"Robust visual tracking via implicit memory-guided re-detection." EAAI (2026).
[paper]
[2026.07]
ReportMedSAM: Anghong Du, Theodoros N. Arvanitis, Colin Watts, Alejandro F. Frangi, Le Zhang.
"ReportMedSAM: Guiding Segmentation Through Radiology Reports." ArXiv (2026).
[paper]
[2026.07]
ViPSAM: San Lee, Nalee Kim, Jeong Il Yu, Hee Chul Park, Boah Kim.
"ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model." MICCAI (2026).
[paper]
[2026.07]
XCT-SAM: Md Mahedi Hasan, Md Mushfiqur Rahaman, Alan Pachkovskiy, Imtiaz Ahmed, Jeremy Dawson, Srinjoy Das.
"XCT-SAM: Sequential Parameter-Efficient Domain Adaptation of SAM for Industrial XCT Defect Segmentation." ICPR workshop (2026).
[paper]
[code]
[2026.07]
Yuanzhi He.
"Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation." ArXiv (2026).
[paper]
[2026.07]
SARFA: Tyler Ward, Abdullah Imran.
"SARFA: Segment Anything with Radiomic Feature Alignment." ArXiv (2026).
[paper]
[code]
[2026.07]
SAM-PAG: Wenqi Si, Gongyang Li, Shixiang Shi, Weisi Lin.
"Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion." IEEE TMM (2026).
[paper]
[code]
[2026.07]
SERD: Shipeng Liu, Zhanping Song, Liang Zhao, Dengfeng Chen.
"Semantic-Edge Response Decoding of SAM3 for Zero-Shot Crack Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
GFR-SAM: Yilong Yang, Jianxin Tian, Shengchuan Zhang, Liujuan Cao.
"GFR-SAM: Training-Free Referring Camouflaged Object Segmentation via Cross-Image Prompting." ArXiv (2026).
[paper]
[2026.07]
MobileSAM2: Kai Jiang, Jiaxing Huang, Jingyi Zhang, Weiying Xie, Yunsong Li, Yufei Wang, Aoran Xiao, Dacheng Tao.
"MobileSAM2: Lightweight Segment Anything for Spatial Intelligence." ECCV (2026).
[paper]
[2026.07]
REBASE: Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana, Sandesh Hegde, Debarshi Banerjee, Uma Mahesh.
"REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
CtrlVTON: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park.
"CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
SAMPLe: Hossein Rajoli, Fatemeh Lotfi, Niloufar Alipour Talemi, Hossein Kashiani, Xiaolong Ma, Fatemeh Afghah.
"SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs." ECCV (2026).
[paper]
[2026.07]
Mohammad Dabaja, Turgay Celik.
"Promptable Concept Segmentation from Above: Evaluating SAM 3's Zero-Shot and One-Shot Capabilities in Remote Sensing." ArXiv (2026).
[paper]
[2026.07]
SAM-MT: Ruiqi Shen, Chang Liu, Henghui Ding.
"SAM-MT: Real-Time Interactive Multi-Target Video Segmentation." ECCV (2026).
[paper]
[code]
[2026.07]
EP-SAM: Wenhao Li, Fangyi Liu, Bo Du.
"An Edge-aware Prompt-enhanced SAM for Ultrasound Image Segmentation." ICME (2026).
[paper]
[2026.07]
HPR-SAM: Yingzhen Hu, Yiheng Zhong, Keying Zhu, Zimu Zhang, Zihan Ye, Sifan Song, Jionglong Su, Xiaofeng Liu.
"HPR-SAM: Hierarchical Probabilistic Representation Learning for Prompt-free SAM-based Medical Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
DA-SAM3: Ying Chen, Jinyue Li, Kun Wang, Qiankun Li, Yang Liu.
"Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation." MICCAI (2026).
[paper]
[code]
[2026.07]
RVAF: Jin Yang, Ping Wei, Nanning Zheng.
"Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation." IROS (2026).
[paper]
[code]
[2026.07]
GeoSAM-Lite: Yongcong Wang, Jie Zhang, Rui Jiang, Xubing Yang, Ting Yun, Li Zhang.
"GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation." GRSL (2026).
[paper]
[2026.07]
GLLS: Runzhi Deng, Yundi Hu, Yiming Zhong, Zhao Wang, Xixi Liu, Hongsong Wang, Caifeng Shan, Fang Zhao.
"Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection." ECCV (2026).
[paper]
[2026.07]
SharpSplat: Porus Vaid, Shivam Chopra, Vaibhav Kumar.
"SharpSplat: Edge-Regularized 3D Gaussian Splatting for High Fidelity Urban Building Reconstruction from UAV images." IGARSS(2026).
[paper]
[2026.07]
ChatImage: Wencan Jiang, Jiangning Zhang, Yong Liu.
"ChatImage: Navigating Long-Form LLM Answers through Interactive Images." ArXiv (2026).
[paper]
[code]
[project]
[2026.07]
IPS-Seg: Le-Anh Tran.
"Exploring SAM Supervision for Fine-Grained UAV Target Segmentation under Data Scarcity." ArXiv (2026).
[paper]
[2026.07]
Muhammad Aamir, Matthew Wijers, Sangyun Shin, Andrew Loveridge, Andrew Markham.
"A non-invasive video-based method for individual identification of wildlife using gait dynamics." ArXiv (2026).
[paper]
[2026.07]
Improved Iris-SAM: Maduabuchi Kingsley Okorie, et al.
"Improved Iris-SAM -Based Iris Segmentation for Recognition in Biometric Security." EJASET (2026).
[paper]
[2026.07]
GPS: Park, J., Jeong, J.
"GPS: GlobalCLIP-PatchCore-SAM Based Zero-Shot Anomaly Detection and Localization in Smart Manufacturing." ICCSA (2026).
[paper]
[2026.07]
SE-MTDNet: Qing Geng and Kaiqi Ye and Fan Xu and Yu Meng and Miao Huang and Chunyan Yuan and Li Li and Bingbo Gao and Hu Zhou and Jianyu Yang and Ying Li and Jianxi Huang and Xiaochuang Yao.
"A data- and knowledge-driven cropland parcel recognition method based on segment anything model (SAM)." International Journal of Applied Earth Observation and Geoinformation (2026).
[paper]
[2026.07]
SAM2-ICHNet: Wanying Xie, Ronghui Ju, He Li, Wei Guo, Zhaoxuan Gong & Guodong Zhang.
"SAM2-ICHNet: A detection-guided intracranial hemorrhage segmentation framework for tiny lesions and complex backgrounds." SIViP (2026).
[paper]
[2026.07]
G2TAM: Chenming Zhu, Peizhou Cao, Jingli Lin, Wenbo Hu, Yunlong Ran, Jiangmiao Pang, Tai Wang, Xihui Liu.
"G2TAM: Geometry Grounded Track Anything Model." ICML (2026).
[paper]
[2026.07]
HBF-BCER: Shengyang Ping,Zhijie Lin *,Liliang Lin,Lei Zhao,Lisha Ye,Bangguo Wang,Tao Wang.
"A Prompt-Preserving MedSAM Enhancement Framework with Historical Feature Fusion and Soft Convolutional Expert Weighting for Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.07]
SAMURAI: Carlos Perez, Neeru Gupta, Ipek Oruc.
"SAMURAI: A Two-Stage Foundation Model Pipeline for Robust Optic Nerve Head Segmentation in Fundus Images." The 39th Canadian Conference on Artificial Intelligence (2026).
[paper]
[2026.07]
LongEgoRefer: Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai, Nakamasa Inoue, Chenhui Chu.
"LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension." ECCV (2026).
[paper]
[code]
[2026.06]
MMIR-TCM: Lihui Luo, Joongwon Chae, Ziyan Chen, Yang Liu, Siyi Cheng, Weihan Gao, Zelin Zeng, Xiaoming Yin, Samaneh Beheshti Kashi, Dongmei Yu, Lian Zhang, Jing Sui, Zeming Liang, Jiansong Ji, Peter E. Lobie, Peiwu Qin.
"MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support." ArXiv (2026).
[paper]
[code]
[2026.06]
AdaCount: Muhammad Ibraheem Siddiqui, Muhammad Haris Khan.
"AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting." ArXiv (2026).
[paper]
[code]
[2026.06]
Object LeJEPA: Jakob Geusen, Ender Konukoglu.
"Object-centric LeJEPA." ArXiv (2026).
[paper]
[2026.06]
Jian Song, Tian Zi, Shen Guanting.
"From Technical Metrics to User Perception: A User Study of a Multimodal Human–Robot Interaction System for Object Detection and Grasping." ArXiv (2026).
[paper]
[2026.06]
Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado López, Mathias Unberath.
"Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos." ArXiv (2026).
[paper]
[2026.06]
BEE: Zhiqiang Hou, et al.
"Bridging the encoder gap: Stability-aware efficient adaptation of SAM2 for video object segmentation." ArXiv (2026).
[paper]
[2026.06]
EpiSAM: Arnav Sharma, Pratyush Jena, Amal Joseph, Ravi Kiran Sarvadevabhatla.
"EpiSAM: Character Segmentation in Challenging Stone Inscriptions." ICDAR (2026).
[paper]
[code]
[2026.06]
PGE-SAM: Tuan-Duc Nguyen, Anh-Tuan Mai, Duc-Trong Le.
"PGE-SAM: Prompt-Guided Feature Enhancement for Interactive Segmentation under Degradation." ArXiv (2026).
[paper]
[2026.06]
ExACT: Zixiao Zhang, Lingling Li, Pei He, Xu Liu, Licheng Jiao.
"ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images." ArXiv (2026).
[paper]
[2026.06]
SemDynReg: Ruitao Chen, Mozhang Guo, Jinge Li.
"SemDynReg: Semantics-Guided Deformation Regularization for Dynamic 3D Gaussian Splatting." ArXiv (2026).
[paper]
[code]
[2026.06]
Xin Dong, Wenfeng Deng, Yansong Tang.
"Occlusion-Robust Multi-Object Decoupling for Physics-Based Interaction." ArXiv (2026).
[paper]
[2026.06]
SPARK: Bryce Grant, Aryeh Rothenberg, Logan Senning, Zonghe Chua, Zach Patterson, Peng Wang.
"Sequential Planning via Anchored Robotic Keypoints." ArXiv (2026).
[paper]
[code]
[2026.06]
CG-ICS: Zhigang Chen, Xiawu Zheng, Rongrong Ji.
"Toward Robust In-Context Segmentation via Concept Guidance." ECCV (2026).
[paper]
[2026.06]
Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi, Eva Cetinic, Gennaro Vessio, Giovanna Castellano.
"Understanding How MLLMs Describe Artworks Using Token Activation Maps." ArXiv (2026).
[paper]
[code]
[2026.06]
TEP-SAM: Yinghui Xing, Donghao Chu, Shizhou Zhang, Di Xu.
"Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection." ICML (2026).
[paper]
[code]
[2026.06]
Simple-ViLMedSAM: Chengcan Qian, Dong Nie, Geng Chen, Daoqiang Zhang, Xuyun Wen.
"Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation." CVPR (2026).
[paper]
[code]
[2026.06]
ScribSAM: Long Chen, et al.
"ScribSAM: A robust scribble-supervised framework for spatiotemporal segmentation of breast lesions in ultrasound videos." Computerized Medical Imaging and Graphics (2026).
[paper]
[code]
[2026.06]
Stat-SAM: Juan Chen, et al.
"Stat-SAM: Learning Global Echo-Intensity Priors as Prompts for SAM in Ultrasound Image Segmentation." ICMR (2026).
[paper]
[2026.06]
Lv, X. et al.
"Lightweight Shape-Aware Segment Anything for Cardiac Ultrasound Segmentation." SMC-IOT (2026).
[paper]
[2026.06]
Lighted-SAM: Yuhan Jia, Lixin Duan, Wen Li, and Fengmao Lv.
"Lighted-SAM: Lightening Open-World SAM for Low-Light Segmentation." TIP (2026).
[paper]
[code]
[2026.06]
NeuroSeg-MF: Zhehao Xu, Weiyi Liu, Shanshan Liang, Hongbo Jia, Xiaowei Chen, Han Qin, and Xiang Liao.
"NeuroSeg-MF: robust neuron segmentation in two-photon Ca2+ imaging using multi-feature fusion and detection-guided SAM." Biomed. Opt. Express (2026).
[paper]
[2026.06]
CPPS-SAM: Zerong Zhang, Lianghua He.
"SAM foundation model and expert model cross prompting framework for semi-supervised medical image segmentation." Journal of Visual Communication and Image Representation (2026).
[paper]
[code]
[2026.06]
Elakiya Sivakumar.
"Fine-Tuning SAM2 for Coronary Artery Segmentation in X-Ray Fluoroscopy." ArXiv (2026).
[paper]
[2026.06]
Tian, J., Cai, W., Sun, Z. et al.
"Unsupervised Change Detection in Remote Sensing Images Using an Integrated SAM and MAD Method." J Indian Soc Remote Sens (2026).
[paper]
[2026.06]
Hizukuri, A.
"Computerized Classification Method for Glioma Molecular Subtypes on Brain MR Images Using SAM-Med3D with Low-Rank Adaptation." J Digit Imaging. Inform. med.(2026).
[paper]
[2026.06]
M2C: Quan Zhou, Shaoqing Zhai, Qiang Hu Jia Chen, Qiang Li, Zhiwei Wang.
"Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation." MICCAI (2026).
[paper]
[code]
[2026.06]
SAM2Matting: Ruiqi Shen, Guangquan Jie, Chang Liu, Henghui Ding.
"SAM2Matting: Generalized Image and Video Matting." ArXiv (2026).
[paper]
[code]
[website]
[2026.06]
SENTRY: Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi, Muzammal Naseer, Naoufel Werghi, Sajid Javed.
"SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking." ECCV (2026).
[paper]
[code]
[2026.06]
MorVess: Fuyou Mao, Yifei Chen, Beining Wu, Lixin Lin, Jinnan Dai, Zhiling Li, Yilei Chen, Yaqi Wang, Hao Zhang, Yan Tang, Huiyu Zhou, Feiwei Qin.
"MorVess: Morphology-Aware Pulmonary Vessel Segmentation Network." ArXiv (2026).
[paper]
[code]
[2026.06]
Marvin Rüdt, Hao Pang, Constantin Enke, Zäzilia Seibold, Kai Furmans.
"Vision-Language Model Reasoning for Contextual Semantic Mapping in Intralogistics." IEEE ETFA (2026).
[paper]
[2026.06]
FEENet: Yang, Zhiyuan and Xu, Jindong and Ni, Mengying and Su, Menghui and Peng, Jiantao.
"A Fuzzy-Embedded Edge Enhancement Network via Segment Anything Model for VHR Remote Sensing Images Change Detection." TGRS (2026).
[paper]
[2026.06]
DR-MV3D: Jiho Choi, Seonho Lee, Seojeong Park, Hyunjung Shim.
"Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views." ECCV (2026).
[paper]
[code]
[2026.06]
ARTEMIS: Tong Wang, Siwen Wang, Yaolei Qi, Jinxing Zhou, Yuting He, Guanyu Yang, Yutong Xie.
"ARTEMIS: Agent-guided Reliability-aware Temporal Mask Evolution for Imperfectly Supervised Video Polyp Segmentation." IEEE TIP (2026).
[paper]
[code]
[2026.06]
VTOS: Jinchao Ge, Lingqiao Liu, Shuwen Zhao, Lei Wang.
"VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers." ArXiv (2026).
[paper]
[2026.06]
SSL.Prop.: Tatsuya Suzuki, Kazuya Ijuin, Hideki Tomimori, Megumi Chikano, Katsushi Sakai.
"Sparse Point-Guided Fusion of Supervised and Self-Supervised Learning Model for Seaweed Segmentation." ArXiv (2026).
[paper]
[2026.06]
ProC-SAM3: Yanghui Song, Nanqing Liu, Haonan Yin, Yingjie Gao, Chengfu Yang, Qi Ming.
"Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation." GRSL (2026).
[paper]
[code]
[2026.06]
μMatch: Marei Freitag, Olesia Korchevaia, Luca Freckmann, Anwai Archit, Constantin Pape.
"Match: Foundation Models for Semi-supervised Learning and Domain Adaptation in EM." ArXiv (2026).
[paper]
[2026.06]
CM-TTA: Yubo Zhou, Jianghao Wu, Ping Ye, Shaoting Zhang, Guotai Wang.
"Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.06]
Hi-Seg: Hongqiao Dong, Wenhao Chi, Ruobing Liang, Xiaokui Yang, Wenhua Liang, Peng Hou, Wenjun Pu, Yipeng Zhao, Ping Chen, Haiping Liu, Jianxing He, Bo Liu.
"Human and AI collaboration for pulmonary nodule segmentation." ArXiv (2026).
[paper]
[2026.06]
DeformX: Yi Yang, Xiang Fei, Lehong Wang, Chenhao Li, Zilin Dai, Henry Kou, Lu Li, Howie Choset.
"DeformX: A Versatile Co-Simulation Framework for Deformable Linear Objects." IROS (2026).
[paper]
[code]
[2026.06]
SARIF: Dong-Hyun Moon, Ju-Hyeon Nam, Sang-Chul Lee.
"SARIF: Segment Anything for Robust Image Forensics." ECCV (2026).
[paper]
[code]
[2026.06]
Auto-SAM: Yijun Wang and Dongyu Zheng and Mingcai Hou and Hongjun Li and Wei Zeng and Caihua Chen and Sixuan Wu and Yujie Gao and Yifan Bai.
"An auto-prompting Segment Anything Model for dual-modal grain segmentation in rock images." Applied Computing and Geosciences (2026).
[paper]
[2026.06]
SA-VIS: Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain, Ender Konukoglu, Luc Van Gool.
"SA-VIS: Sparse frame Annotations for training Video Instance Segmentation." ArXiv (2026).
[paper]
[2026.06]
SAM3 Self-Distillation for Fine-Grained GOOSE 2D Semantic Segmentation.
"SAM3 Self-Distillation for Fine-Grained GOOSE 2D Semantic Segmentation." ArXiv (2026).
[paper]
[2026.06]
Intrinsic-GS: Hasan Yazar, Mohamed Rayan Barhdadi, Erchin Serpedin, Mehmet Tuncel, Hasan Kurban.
"Intrinsic 4D Gaussian Segmentation from Scene Cues." ArXiv (2026).
[paper]
[code]
[2026.06]
Sonata Simonaitis-Boyd, Soonhong Lee, Lauren N. O'Brien, Brandon T. Turner, Ralph Massarczyk, Steven R. Elliott, Aobo Li, Alexander F. Leder.
"Vision AI Agent for Continuous Material Monitoring of LEGEND-1000 LoFi Reentrant Tube." ArXiv (2026).
[paper]
[2026.06]
PEFT-MedSAM: Asad Channa, Abdullah Khan, Asghar Ali Chandio, Aamir Akbar, Shahzad Memon, Aqib Hussain, Ameer Hamza.
"PEFT-MedSAM: Efficient Fine-Tuning of Medical Foundation Models for Explainable Skin Lesion Segmentation." ArXiv (2026).
[paper]
[2026.06]
Paul Julius Kühn, Saptarshi Neil Sinha, Jakob Hansen, Robin Horst.
"Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows." ArXiv (2026).
[paper]
[2026.06]
Nicholas A. Welsh, Lennon J. Shikhman, Monty Nehru Attazs, Seemanthini K. Putane, Van Minh Nguyen, Ryan T. White.
"Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection." CVPR Workshop (2026).
[paper]
[2026.06]
TVG: Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang.
"Thinking with Visual Grounding." ArXiv (2026).
[paper]
[code]
[dataset]
[2026.06]
Multi-HMR 2: Guénolé Fiche, Philippe Weinzaepfel, Romain Brégier, Fabien Baradel.
"Multi-HMR 2: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking." ArXiv (2026).
[paper]
[2026.06]
DETECTURE: Aviad Cohen Zada, Nadav Orenstein, Shai Avidan, Gal Oren.
"Sub-Semantic Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.06]
MuDuo: Fuyou Mao, Beining Wu, Yanfeng Jiang, Bohan Xu, Lixin Lin, Naye Ji, Hao Zhang, Yan Tang.
"Mutual Distillation of Dual-Foundation Models for Semi-Supervised PET/CT Segmentation." MICCAI (2026).
[paper]
[code]
[2026.06]
Gen-VCoT:: Zhiqiang Zhou, Xu Ling, Junliang Dai.
"Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations." ArXiv (2026).
[paper]
[2026.06]
Nadav Orenstein, Aviad Cohen Zada, Shai Avidan, Gal Oren.
"Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation." ArXiv (2026).
[paper]
[2026.06]
Changwoo Song.
"Parameter-Efficient Adaptation of SAM 3 for Automated ITV Generation from 4DCT Images." ArXiv (2026).
[paper]
[2026.06]
Aniq Ahmad, Heather Bedle, Ahmad Mustafa.
"Domain-Guided Prompting of the Segment Anything Model for Seismic Interpretation: The Role of Attributes, Visualization, and Hybrid Prompts." ArXiv (2026).
[paper]
[2026.06]
Yiping Li, Ronald de Jong, Romy van Jaarsveld, Franco Badaloni, Gino Kuiper, Jelle Ruurda, Josien Pluim, Marcel Breeuwer.
"Object Tokens as a Bridge Between Segmentation and Visual Question Answering in Robotic Surgery." ArXiv (2026).
[paper]
[2026.06]
MaxCode: Qiyue Liang, Steven Ingram, George Vanica, Andi Gavrilescu, Newfel Harrat, Hassan Sipra, Sethuraman Sankaran.
"Agentic Framework for Deep Learning workload migration via In-Context Learning." ArXiv (2026).
[paper]
[code]
[2026.06]
ActiveSAM: Tran Dinh Tien, Zhiqiang Shen.
"ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation." ArXiv (2026).
[paper]
[code]
[2026.06]
MooMIns: Robert Langendörfer, Markus Hillemann, Markus Ulrich.
"MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances." ArXiv (2026).
[paper]
[code]
[2026.06]
Keyi Zhu, Kyle Lammers, Chaaran Arunachalam, Kaixiang Zhang, Renfu Lu, Zhaojian Li.
"A Modular Dual-Arm Apple Harvesting Robot with Enhanced Field Performance." ArXiv (2026).
[paper]
[2026.06]
SAM-Deep-EIoU: Alexander Holmberg.
"SAM-Deep-EIoU: Selective Mask Propagation for Multi-Object Tracking." ArXiv (2026).
[paper]
[2026.06]
Li, Chenying, Xiao Tan, Xinyu Huang, Ling Sa, Nailong Zhang, and Gang Qiu.
"Galloping Target Tracking and Parameter Measurement Method for Overhead Transmission Lines Based on SAM2 Video Segmentation." Electronics (2026).
[paper]
[2026.06]
MSBA-SAM: Tao Guo, Kui Xu, Kailei Chen, Chun Xie, Shi Qiu & Rui Ye.
"MSBA-SAM: a multi-scale and boundary-aware framework for power grid segmentation in aerial image." Energy Informatics(2026).
[paper]
[2026.06]
Suyog Jadhav, Dilip K. Prasad, Krishna Agarwal.
"SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy." CVPRW (2026).
[paper]
[2026.06]
Mix-QSAM3: Navin Ranjan, Andreas Savakis.
"Mix-QSAM3: Mixed-Precision Quantization for the Segment Anything with Concepts Model." CVPRW (2026).
[paper]
[2026.06]
LSG-SAM: Muhammad Imran, Yugyung Lee.
"Latent-Stability Gated SAM: Detecting Hallucinated Segmentations under Domain Shift." CVPRW (2026).
[paper]
[2026.06]
VegSAM: Chenxiang Wu, Chenyu Li, Danfeng Hong.
"VegSAM: Vegetation-aware Adapter for Segment Anything Model in Urban Tree Segmentation." CVPRW (2026).
[paper]
[2026.06]
SAM3Count: Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar.
"SAM3Count for Zero-Shot Open Vocabulary Counting in Images and Videos." CVPRW (2026).
[paper]
[code]
[2026.06]
SAM-OOD: Seher Kanwal, Seung-Ik Lee.
"SAM-OOD: Foundation-Model-Guided Unknown Mining for Object-Level Anomaly Detection." CVPRW (2026).
[paper]
[2026.06]
C-RWHD: Khalil Khazmi, Zied Lachiri.
"Coupled annotation and architecture tailoring for lightweight and robust wheat head detection: SAM-oriented bounding boxes with simplified YOLO variants validated on a Tunisian wheat head dataset." SAT (2026).
[paper]
[2026.06]
Chun Cao, et al.
"A Segment Anything Model adaptation framework for battery visual inspection under complex radiographic imaging conditions." PR (2026).
[paper]
[2026.06]
IFP: Shuqi Xia, Guangze Shi, Jiarui Cao, Aoyuan Shi, Meilin Liu, Xiaoyi Zhang, Yujie Wang, Xueyu Liu, Cai Zhao, Ziyuan He, Yongfei Wu, Mingqiang Wei.
"Instruction-Focus-Prompt:Semantics-Driven Structural Prompts for Universal SAM Segmentation." CVPR Findings (2026).
[paper]
[code]
[2026.06]
UCOD-MKD: Huafeng Chen, Chenguang Zhu, Yueming Lyu, Caifeng Shan.
"Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection." CVPR (2026).
[paper]
[code]
[2026.06]
mSAMUNet: Md. Shariful Alam, et al.
"Cell segmentation in microscopy images using a SAM-based U-Net architecture and a novel dataset." CMPB (2026).
[paper]
[2026.06]
RFD: Ji Xia, et al.
"RFD: A Reducing Feature Discrepancy method for unsupervised cross-modality SAM adaptation." CMIG (2026).
[paper]
[2026.06]
SAMosaic3D: Peng Wang, Yongcai Wang, Wang Chen, Hualong Cao, Kang Yang, Chunxu Li, Jie Wen, Deying Li.
"SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything." CVPR (2026).
[paper]
[code]
[2026.06]
RoSAMDepth: Xuanang Gao, Zhiwei Ning, Gengming Zhang, Jiaxi Cao, Runze Yang, Zhonglong Zheng, Jie Yang, Rong Xiao, Wei Liu.
"RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model." CVPR (2026).
[paper]
[code]
[2026.06]
PASD: Zhang, Yuliang and He, Fang and Peng, Lulu and Yan, Tianyu and Zhang, Pingping and Song, Ting and Du, Lili and Chen, Dunjin.
"3D Segment Anything Model With Visual Mamba for Diagnosing Placenta Accreta Spectrum." TIP (2026).
[paper]
[code]
[2026.06]
CSAM-TUNet: Chao Xu, et al.
"CSAM-TUNet: A SAM-guided contrastive learning framework with topological attention for pericardial adipose tissue segmentation." BSPC (2026).
[paper]
[2026.06]
TransNet–SAM2: Bamwenda, Julius, Mehmet Siraç Özerdem, Orhan Ayyildiz, Veysi Akpolat, and İrem Akpolat.
"TransNet–SAM2: A Transformer–Foundation Model Framework for Prompt-Free Segmentation of White Blood Cells in Microscopic Blood Smear Images." Diagnostics (2026).
[paper]
[2026.06]
SAM-3D-MSF: Yifu Zhang, Chun Shen & Jianbing Li.
"SAM-3D-MSF: Parameter-Efficient Adaptation of Segment Anything Model for 3D Tooth CBCT Segmentation." PAKDD (2026).
[paper]
[2026.06]
Fatih Fehmi Şimşek, Melih Altay, Kaan Kalkan, Mehmet Cengiz Arslanoğlu.
"Assessing the Impact of Spatial Resolution and Hyperparameters on Automatic Agricultural Parcel Delineation Using the Segment Anything Model With Multi-Resolution and Super-Resolved Satellite Imagery." Transactions in GIS (2026).
[paper]
[2026.06]
LSAC: Yuxuan Chen, Haoyuan Xu, Peize He.
"Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models." INTERSPEECH (2026).
[paper]
[2026.06]
ZODS-RS: Zuan Gu, Tianhan Gao, Langxu Zhao.
"ZODS-RS -- Zero-training Oriented Detection & Segmentation for Remote Sensing." ArXiv (2026).
[paper]
[2026.06]
Avi Gupta, Nilotpal Sinha, Vishnu Raj, Sambuddha Saha, Pratik Joshi, Koteswar Rao Jerripothula, Tammam Tillo.
"Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio." ICML Workshop (2026).
[paper]
[2026.06]
IPSM-Bench: Jinglin Xu, Shangyan Zhao, Jiabo Wang, Xinghong Mu, Yulong Lei, Jiacheng Zhang, Hongbo Sun, Yageng Li.
"IPSM-Bench: A New Intermediate Phase Segmentation Benchmark in Microstructure Images of Zinc-Based Absorbable Biomaterials." IJCAI (2026).
[paper]
[2026.06]
Nermeen Abou Baker, Uwe Handmann.
"Don't waste SAM." ESANN (2023).
[paper]
[2026.06]
RTVP: Zekai Zhang, Qinghui Chen, Maomao Xiong, Shijiao Ding, Zhanzhi Su, Xinjie Yao, Yiming Sun, Cong Bai, Jinglin Zhang.
"Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline." AAAI (2025).
[paper]
[code]
[2026.06]
Open-V: Silas Kwabla Gah, Ebenezer Owusu.
"Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration." ArXiv (2026).
[paper]
[2026.06]
SAM-Flow: Haowang Cui, Rui Chen, Tao Luo, Tao Guo, Zheng Qin, Jiaze Wang.
"SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing." ArXiv (2026).
[paper]
[code]
[2026.06]
Pigformer: Mk Bashar, Kuljit Bhatti, Gary Rohrer, Madonna Benjamin, Tami Brown-Brandl, Daniel Morris.
"What's Under the Skin? Estimating Swine Body Condition." ArXiv (2026).
[paper]
[code]
[2026.06]
TopoPult-SSL: Nicolò Savioli, Luca Del Tongo.
"TopoPult-SSL: Gland-Mask-Free Cross-Device Meibomian Gland Segmentation via Self-Distilled Weak Clinical Priors." ArXiv (2026).
[paper]
[2026.06]
MedSAM-BoxPredictor: Amirhossein Movahedisefat, Amirreza Fateh, Mohammad Reza Mohammadi.
"Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.06]
CR-SEG: Yifan Cao, Xiaocui Yang, Faxian Wan, Shi Feng, Daling Wang, Yifei Zhang.
"CR-SEG: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation." ArXiv (2026).
[paper]
[2026.06]
SAMatcher: Xu Pan, Qiyuan Ma, Mingyue Dong, He Chen, Wei Ji, Xianwei Zheng.
"SAMatcher: Co-Visibility Modeling with Segment Anything for Robust Feature Matching." ArXiv (2026).
[paper]
[code]
[2026.06]
Sema Helali, Lina Abu Nadab, Sausan Alqawas, Alaa Abd-Alrazaq, Faleh Tamimi, Rafat Damseh.
"Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models." ArXiv (2026).
[paper]
[2026.06]
FLAME: Yiming Wang, Baiqi Wu, Qingming Li, Jiahao Chen, Tong Zhang, Shouling Ji.
"Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization." ICML (2026).
[paper]
[code]
[2026.06]
PerBite: Ahmad AlMughrabi, Farid Al-Areqi, David Fernández Gómez, Umair Haroon, Marc Bolaños, Ricardo Marques, Petia Radeva.
"PerBite: A Curated Diagnostic Workflow for Bite-Aware Food Volume Estimation." ArXiv (2026).
[paper]
[code]
[2026.06]
LG-SAM: Panav Shah, Geet Sethi, Ashutosh Gandhe.
"Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting." CVPR Workshops (2026).
[paper]
[code]
[2026.06]
Sakib Mohammad, Jarin Ritu, Md Sakhawat Hossain.
"Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models." ArXiv (2026).
[paper]
[2026.06]
GeoSAM-3D: Arun Sharma.
"GeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video." (NeurIPS(2026).
[paper]
[2026.06]
MLAM: Yuliang Zhang, Fang He, Lulu Peng, Tianyu Yan, Pingping Zhang, Ting Song, Lili Du, Dunjin Chen.
"3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum." TIP (2026).
[paper]
[code]
[2026.06]
CLOC: Mengqi Lei, Shuokun Cheng, Wei Bao, Shaoyi Du, Jun-Hai Yong, Siqi Li, Yue Gao.
"Count Anything." ArXiv (2026).
[paper]
[code]
[2026.05]
Suyog Jadhav, Dilip K. Prasad, Krishna Agarwal.
"SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy." CVPR Workshops (2026).
[paper]
[2026.05]
CM-SAM: Jiexin Liang, et al.
"CM-SAM: Chaos-enhanced hybrid encoder for medical image segmentation with segment anything model." Biomedical Signal Processing and Control (2026).
[paper]
[2026.05]
SAM3D-Phys: Xin Dong, Weijian Deng, Lihan Zhang, Tianru Dai, Wenfeng Deng, Yansong Tang.
"SAM3D-Phys: Towards Multi-Object Interactive Simulation in Real World." ArXiv (2026).
[paper]
[code]
[2026.05]
ESAM++: Qin Liu, Lavisha Aggarwal, Saptarashmi Bandyopadhyay, Vikas Bahirwani, Marc Niethammer, Ehsan Adeli, Andrea Colaco.
"ESAM++: Efficient Online 3D Perception on the Edge." CVPR (2026).
[paper]
[code]
[2026.05]
DOST: Bolian Peng, Ying Tang, Xu Liu, Long Sun, Xiaoqiang Lu.
"Turbulence-Robust Dynamic Object Segmentation with Multi-Signal Priors and SAM2 Refinement." ArXiv (2026).
[paper]
[2026.05]
ViTA: Ji-Hoon Hwang, Jisung Bae, Dong-Wook Kim, Yeonkyu Lee, Seung-Woo Seo.
"From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments." ArXiv (2026).
[paper]
[2026.05]
CoP: Sanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim, Hyungseok Seo, Kyungsu Kim.
"One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation." MICCAI (2026).
[paper]
[code]
[2026.05]
Toomas Tahves, Mauro Bellone, Junyi Gu, Raivo Sell.
"SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving." ArXiv (2026).
[paper]
[project]
[dataset]
[2026.05]
Water-AutoSAM: Sun, Yingrui, Yang Hong, Xiaowei Zhou, and Junyu Dong.
"Water-AutoSAM: Dual-Domain Enhanced Auto-Prompting SAM for Underwater Segmentation." Journal of Marine Science and Engineering (2026).
[paper]
[2026.05]
Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla, Yuhao Chen.
"Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion." ArXiv (2026).
[paper]
[2026.05]
PlayClass: Prince Ravi Leow, Neil Scheidwasser, Rebecca Oscarsson, Per Jensen, Samir Bhatt, David Alejandro Duchêne.
"PlayClass: Automated Play Behaviour Classification in Poultry." CVPR Workshop (2026).
[paper]
[code]
[2026.05]
PinPoint: Pouya Sadeghi, Shawn He, Pedro Pablo Guerrero Vela, C. Thomas, Alex Wong, Sirisha Rambhatla.
"PinPoint: Prompting with Informative Interior Points." ArXiv (2026).
[paper]
[2026.05]
Mannat Khurana, Sanyam Jain, Rishav Agarwal.
"Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis." ArXiv (2026).
[paper]
[2026.05]
InstructSAM: Yuqian Yuan, Wentong Li, Zhaocheng Li, Yutong Lin, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang, Wenqiao Zhang.
"InstructSAM: Segment Any Instance with Any Instructions." ArXiv (2026).
[paper]
[code]
[2026.05]
BED-SAM2: Tyler Rust, Dara McNally, Kyle O'Donnell, Colin Kelly, Chandra Kambhamettu.
"BED-SAM2: Boundary-Enhanced-Depth SAM2 via Monocular Geometric Priors." CVPR Workshop (2026).
[paper]
[code]
[2026.05]
ANAUS: Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li.
"Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation." MICCAI (2026).
[paper]
[code]
[2026.05]
RepSAM: Wenhui Chu.
"RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation." IJCAI-ECAI 2026 workshop (2026).
[paper]
[2026.05]
SAM 3-to-YOLOv8: Marcos Vinicius Mendes Faria, Thiago Borges Pereira, Isabella C. F. S. Condotta, Thiago Meireles Paixão, Francisco de Assis Boldt.
"SAM3-Assisted Training of Lightweight YOLO Models for Precision Pig Farming." IEEE SAS (2026).
[paper]
[2026.05]
DeCoDrift: H. M. Shadman Tabib, Md. Shamsuzzoha Bayzid, M Sohel Rahman.
"DeCoDrift: Stabilizing Decoder Coupling in Closed-Loop Foundation Segmentation." ArXiv (2026).
[paper]
[2026.05]
MGNet,: Xia Li, Xinran Liu, Lin Qi, Junyu Dong.
"Weakly Supervised Camouflaged Object Detection Based on the SAM Model and Mask Guidance." ArXiv (2026).
[paper]
[2026.05]
CLIP-Guided SAM: Shayan Jalilian, Abdul Bais.
"CLIP-Guided SAM: Parameter-Efficient Semantic Conditioning for Promptable Segmentation." ArXiv (2026).
[paper]
[2026.05]
ViViD-5K: Xiangzhi Tong, Chengrui Zhang, Mac Flaherty, Andre Matteo Garcia, Dominic Gorman, Jonathan Jaramillo, Justine E. Vanden Heuvel, Yu Jiang.
"ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation." ArXiv (2026).
[paper]
[2026.05]
ConceptSeg-R1: Yuan Zhao, Youwei Pang, Jiaming Zuo, Wei Ji, Kailai Zhou, Bin Fan, Yunkang Cao, Lihe Zhang, Xiaofeng Liu, Huchuan Lu, Weisi Lin, Dacheng Tao, Xiaoqi Zhao.
"ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning." ArXiv (2026).
[paper]
[project]
[code]
[2026.05]
MGGA: Fan Gao, et al.
"MedSAM-guided geometry-aware 2D-3D feature fusion for medical image registration." Neural Networks (2026).
[paper]
[code]
[2026.05]
PromptLessSAM: Mohammed Al-Mustafa Hendo, et al.
"PromptLessSAM: From Foundational Model to Domain Expert via Lightweight Decoder Adaptation for Crack Segmentation." Electrical, Computer and Communications Engineering (2026).
[paper]
[2026.05]
MT-SAM: Zhao, Litao and Zhang, Yuhan and Ji, Libiao and Bao, Jie and Li, Caizi and NG, Chi-Fai and Heng, Pheng-Ann.
"MT-SAM: A Mamba-Transformer Enhanced SAM with Prior-guided Prompting for Multi-modal Prostate Cancer Delineation." TMI (2026).
[paper]
[code]
[2026.05]
STATUARIO-40K: Palmieri, Giorgio and Ranieri, Andrea and Vasarelli, Andrea and Di Angelo, Luca.
"STATUARIO-40K: Fine-Tuning SAM 3 for Instance Segmentation at Scale of Monuments, Statues, and Cracks." ArXiv (2026).
[paper]
[dataset]
[2026.05]
Per-SAM-MCPA: Hu, Chuting, Size Dai, Shifan Wu, Qiaolin Ye, and He Yan.
"Per-SAM-MCPA: A Lightweight Framework for Individual Tree Crown Segmentation from UAV Imagery." Remote Sens (2026).
[paper]
[2026.05]
ColorSAM: Bangcheng Zhan, et al.
"ColorSAM: teaching SAM to segment color medical images via quaternion decoding and prompt generation." ESWA (2026).
[paper]
[code]
[2026.05]
PromptSAMNet: A.R. Revathi, et al.
"PromptSAMNet: A memory-enhanced adaptive prompting and clustering-augmented SAM 2.0 framework for multi-leaf plant disease diagnosis." Applied Soft Computing (2026).
[paper]
[2026.05]
SDDNet: Zhao, Yu and Sun, Jing and Zhang, Guohui and Sun, Fuming and Li, Haojie.
"Enhancing SAM2 for Industrial Defect Detection via Dual-Adapter Fine-Tuning." TIM (2026).
[paper]
[code]
[2026.05]
FS-Grasp: Zhai, Di-Hua and Yu, Sheng and Xia, Yuanqing.
"Fast and Efficient 6-DoF Grasp Estimation With Segment Anything Model in Cluttered Scenes." IEEE/ASME Transactions on Mechatronics (2026).
[paper]
[2026.05]
AISCT-SAM: Kuang, Hulin and Tan, Xianzhen and Li, Shunuo and Kan, Shichao and Liu, Jin and Sun, Jiarui and Zhang, Jingyang and Yang, Chunfeng and Qiu, Wu and Zhang, Jiulou and Chen, Yang and Wang, Jianxin.
"AISCT-SAM: Customized SAM-Med2D with 3D Context Awareness and Self-Prompt Generation for Fully Automatic Acute Ischemic Stroke Lesion Segmentation on Non-Contrast CT Scans." JBHI (2026).
[paper]
[2026.05]
MUP-SAM: Lyuyang Tong, Jingwen Jiang, Bo Du.
"MUP-SAM: Multi-Scale Vision Mamba UNet Prompt Generation for SAM in Multi-Organ Medical Image Segmentation." Neural Networks (2026).
[paper]
[2026.05]
CraterSAM+: Li, Miyu and Li, Junjie and Wang, Yumei and Liu, Yu.
"Self-Improving SAM with Specialist Knowledge via Adaptive Direct Preference Optimization for Crater Segmentation." IEEE Geoscience and Remote Sensing Letters (2026).
[paper]
[2026.05]
MedSAM-COALF: Zhao, Pengyu and Hou, Yonghong and Wu, Jiasai and Yan, Ke and Huo, Shuwei.
"MedSAM-COALF: A Cold-Start One-Shot Active Learning Framework for Medical Image Segmentation via Foundation Model-Guided Proxy Tasks and Uncertainty-Aware Sampling."IEEE Sensors Journal (2026).
[paper]
[2026.05]
SAM-SS: Wang, Yalin and Han, Wei and Peng, Hong and Zheng, Weihao and Li, Xiaoxu and Kang, Zhongfeng and Chan, Sixian.
"SAM-SS: Straightforward and Efficient Designs Based on Segment Anything Model for Semantic Segmentation." TCSS (2026).
[paper]
[2026.05]
S2C-Net: Ning, Hailong and Li, Haojie and Zhang, Wuxia and Lei, Tao and Chen, Yanping and Cao, Xiaopeng and Nandi, Asoke K..
"S2C-Net: SAM2-Based Dual-Domain Feature Reconstruction and Semantic Decoupling for Tiny Remote Sensing Object Counting." TGRS (2026).
[paper]
[code]
[2026.05]
LiteWaveRep-MedSAM: Lieqiang Liu, Chengping Zhao, Tengxiao Xu, Wutao Xiong and Yuxiao Zhang.
"LiteWaveRep-MedSAM: A lightweight medical image segmentation model based on wavelet transform and reparameterization." Biomedical Physics & Engineering Express (2026).
[paper]
[code]
[2026.05]
TorqueSAM: Rahat, Shahzalal Khan, et al.
"TorqueSAM: unsupervised kidney CT analysis with localization and SAM-integrated torque clustering segmentation." ArXiv (2026).
[paper]
[2026.05]
Elgström, Albert, Diaz, Jose, Bosch, Carles.
"Hierarchical Annotation of Mural Paintings Using SAM." ArXiv (2026).
[paper]
[2026.05]
ERSF-AS: Cheng Ju, et al.
"ERSF-AS: Explainable recursive zero-shot anomaly segmentation with spatial-frequency priors via CLIP-SAM collaboration." Neurocomputing (2026).
[paper]
[2026.05]
Tan, L., Xia, Y., Teng, D. et al.
"Comparative evaluation of conventional radiomics and VGG-SAM fusion strategies for MRI-based preoperative prediction of perineural invasion in cervical cancer." Abdom Radiol (2026).
[paper]
[2026.05]
FAST-ME : Kakia Panagidi, Stathes Hadjieftymiadis.
"FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis." ArXiv (2026).
[paper]
[2026.05]
Sebastian Cavada, Francesco Pelosin, Lapo Faggi.
"Training-Free Fine-Grained Semantic Segmentations in Low Data Regimes: A FungiTastic Baseline." CVPRW (2026).
[paper]
[2026.05]
COCOTree: Junhyub Lee, Seunghun Chae, Hyosu Kim.
"COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition." ArXiv (2026).
[paper]
[code]
[2026.05]
SAMOSA: Deyi Zhu, Yuji Wang, Yong Liu, Yansong Tang, Bingyao Yu, Jiwen Lu, Jie Zhou.
"Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking." ArXiv (2026).
[paper]
[code]
[2026.05]
DECA: Lifeng Yang, Linshu Chen, Anxing Hu, et al.
"Underwater image segmentation method based on dual-encoder." ICCIIA (2026).
[paper]
[2026.05]
SGA: Jinjin Zhang, Xiefan Guo, Di Huang.
"Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis." ArXiv (2026).
[paper]
[code]
[2026.05]
HyDAR-Pano3D: Yaoyao Yue, Jérôme Schmid, Xiaoshuang Li, Eduardo Delamare, Jinman Kim.
"HyDAR-Pano3D: A Hybrid Disentangled Anatomical Recovery Framework for Panoramic-to-3D Reconstruction." ArXiv (2026).
[paper]
[2026.05]
SAM-Sode: Wanying Tan, Shuo Yan, Dazhi Huang, Yazheng Liu, Zili Shao, Rufeng Chen, Hechang Chen, Mude Shi, Tianxing Ji, Sihong Xie.
"SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection." ArXiv (2026).
[paper]
[2026.05]
Stream3D: Kaichen Zhou, Zeyang Bai, Xinhai Chang, Mengyu Wang, Paul Liang, Fangneng Zhan.
"Stream3D: Sequential Multi-View 3D Generation via Evidential Memory." ArXiv (2026).
[paper]
[code]
[2026.05]
LCA: Qisai Liu, Alloy Das, Zhanhong Jiang, Joshua R. Waite, Aditya Balu, Adarsh Krishnamurthy, Soumik Sarkar.
"Lighting-aware Unified Model for Instance Segmentation." ArXiv (2026).
[paper]
[2026.05]
VASA: Zilin Wang, Stella X. Yu.
"Vision Harnessing Agent for Open Ad-hoc Segmentation." ArXiv (2026).
[paper]
[2026.05]
DarkLLM: Ye Sun, Xin Wang, Jiaming Zhang, Yifeng Gao, Yixu Wang, Yifan Ding, Qixian Zhang, Henghui Ding, Xingjun Ma, Yu-Gang Jiang.
"DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models." ArXiv (2026).
[paper]
[code]
[2026.05]
Tonghao Zhuang, Shanglong Hu, Yongsheng Luo, Zhiqi Zhang, Yu Li.
"Synergistic Foundation Models for Semi-Supervised Fetal Cardiac Ultrasound Analysis: SAM-Med2D Boundary Refinement and DINOv3 Semantic Enhancement." MICCAI Workshop (2026).
[paper]
[code]
[2026.05]
Ananth Sriram, Neel Mokaria, Rajveer Singh.
"Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification." ArXiv (2026).
[paper]
[code]
[2026.05]
MedFM-Robust: Xiangxiang Cui, Tianjin Huang, Yifang Wang, Lijie Hu, Lu Yin.
"MedFM-Robust: Benchmarking Robustness of Medical Foundation Models." MICCAI (2026).
[paper]
[code]
[2026.05]
OmniVL-Guard Pro: Jinjie Shen, Zheng Huang, Yuchen Zhang, Yujiao Wu, Yaxiong Wang, Lechao Cheng, Shengeng Tang, Tianrui Hui, Nan Pu, Zhun Zhong.
"OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics." ArXiv (2026).
[paper]
[code]
[2026.05]
SegRAG: Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed.
"SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation." ArXiv (2026).
[paper]
[code]
[2026.05]
D3S2: Yachan Guo, JoseLuis Gomez Zurita, Danna Xue, Yi Xiao, AntonioManuel Lopez Pena.
"Metric-Guided Feature Fusion of Visual Foundation Models for Segmentation Tasks." CVPR Findings (2026).
[paper]
[code]
[2026.05]
HyperVision: Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Fengchao Xiong, Jianfeng Lu, Jun Zhou.
"HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone." ArXiv (2026).
[paper]
[code]
[2026.05]
Rad-VLSM: Fengyi Zhang, Xujie Zeng, Mohan Liu, Zengyi Wang, Yalong Jiang.
"Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis." ArXiv (2026).
[paper]
[2026.05]
Divya Joshi, J. D. Peiffer, Colleen Peyton, R. James Cotton.
"Markerless Motion Capture for Biomechanical Whole-Body Kinematic Estimation in Infants." EMBC (2026).
[paper]
[2026.05]
WOW-Seg: Danyang Li, Tianhao Wu, Bin Li, Zhenyuan Chen, Yang Zhang, Yuxuan Li, Ming-Ming Cheng, Xiang Li.
"WOW-Seg: A Word-free Open World Segmentation Model." ICLR (2026).
[paper]
[code]
[2026.05]
TinySAM 2: Zhaoyuan Ding, Yijing Yang, Han Shu, Xinghao Chen.
"TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model." ArXiv (2026).
[paper]
[2026.05]
SparseSAM: Hoai-Chau Tran, Chi H. Nguyen, Duy M. H. Nguyen, Mathias Niepert, Fan Lai, Khoa D. Doan.
"SparseSAM: Structured Sparsification of Activations in Segment Anything Models." ArXiv (2026).
[paper]
[2026.05]
CAR-SAM: Houji Wen, Jiangyong Yu, Jun Li, Dawei Yang.
"CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model." ArXiv (2026).
[paper]
[2026.05]
Eugenia Moris, Alicia Costábile, Sebastián Rey, Irene Ferreiro, Joaquín Hurtado, Lizandra Lissette Luciano, Matías Villagrán, Aisha Espino Vázquez, Jomari Ramos, Isadora Monteiro, María Victoria de Santiago, Pilar Moreno, Gonzalo Moratorio, José Ignacio Orlando.
"End-to-end plaque counting and virus titration from laboratory plate images with deep learning." ArXiv (2026).
[paper]
[2026.05]
Raushan Joshi, Jean-Yves Guillemaut.
"Robust Prior-Guided Segmentation for Editable 3D Gaussian Splatting." ICIP (2026).
[paper]
[2026.05]
MT-SAM: Litao Zhao, Yuhan Zhang, Libiao Ji, Jie Bao, Caizi Li, Chi-Fai NG and Pheng-Ann Heng.
"MT-SAM: A Mamba-Transformer Enhanced SAM with Prior-guided Prompting for Multi-modal Prostate Cancer Delineation." TMI (2026).
[paper]
[code]
[2026.05]
DT-ZSAM: Fan, Zhanpeng, Xinglei Gu, Qiyu Liu, Yangheng Hu, and Liang Yu.
"Fusing Dual-Threshold Prompts with SAM for Shot Peening Coverage Assessment on Aircraft Propeller Blades." Applied Sciences (2026).
[paper]
[2026.05]
DeFakerOne: GuangJian Team, Ant Group.
"Venus-DeFakerOne: Unified Fake Image Detection & Localization." ArXiv (2026).
[paper]
[code]
[2026.05]
PDI-Bench: Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, Xueyan Zou.
"Quantitative Video World Model Evaluation for Geometric-Consistency." ArXiv (2026).
[paper]
[code]
[2026.05]
MedCore: Cenwei Zhang, Suncheng Xiang, Lei You.
"MedCore: Boundary-Preserving Medical Core Pruning for MedSAM." ArXiv (2026).
[paper]
[code]
[2026.05]
Seg-Agent: Chao Hao, Jun Xu, Ji Du, Shuo Ye, Ziyue Qiao, Xiaodong Cun, Guangcong Wang, Xubin Zheng, Zitong Yu.
"Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation." ArXiv (2026).
[paper]
[code]
[2026.05]
PointGS: Yixiao Song, Qingyong Li, Wen Wang, Zhicheng Yan.
"PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting." CVPR (2026).
[paper]
[code]
[2026.05]
FocusDepth: Yuxin Du, Tao Lin, Zile Zhong, Runting Li, Xiyao Chen, Jiting Liu, Chenglin Liu, Ying-Cong Chen, Yuqian Fu, Bo Zhao.
"Focusable Monocular Depth Estimation." ArXiv (2026).
[paper]
[2026.05]
M4-SAM: Jiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong, Deyang Liu, Zhi Liu.
"M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection." CVPR (2026).
[paper]
[code]
[2026.05]
Stefano Colamonaco, Andrei-Bogdan Florea, Jaron Maene.
"Weakly Supervised Segmentation as Semantic-Based Regularization." ArXiv (2026).
[paper]
[2026.05]
RUAC: Hongyou Zhou, Marc Toussaint, Ling Shao, Zihan Ye.
"Segment Anything with Robust Uncertainty-Accuracy Correlation." ICML (2026).
[paper]
[code]
[2026.05]
FSAM: Phuoc-Nguyen Bui, Van-Nguyen Pham, Duc-Tai Le, Junghyun Bum, Hyunseung Choo.
"Frequency Adapter with SAM for Generalized Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.05]
CAFE: Shuang Liang, Zeqing Wang, Yuxian Li, Xihui Liu, Han Wang.
"From Pixels to Concepts: Do Segmentation Models Understand What They Segment?." ArXiv (2026).
[paper]
[code]
[project]
[dataset]
[2026.05]
SAMOFT: Yanchao Wang, Dawei Zhang, Chengzhuan Yang, Wei Liu, Minglu Li, Hua Wang, Zhonglong Zheng, Ming-Hsuan Yang.
"SAMOFT: Robust Multi-Object Tracking via Region and Flow." ArXiv (2026).
[paper]
[2026.05]
R. James Cotton, Pouyan Firouzabadi, Wendy Murray.
"Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models." EMBC (2026).
[paper]
[2026.05]
RCoT-Seg: Junwei Wen, Deshui Miao, Guangming Lu, Xin Li, Wenjie Pei.
"RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation." ArXiv (2026).
[paper]
[code]
[2026.05]
SARA: Jiesong Lian, Zixiang Zhou, Ruizhe Zhong, Yuan Zhou, Qinglin Lu, Rui Wang, Long Hu, Yixue Hao, Baoru Huang.
"SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models." ArXiv (2026).
[paper]
[2026.05]
SAM 3D Animal: Xuyi Hu, Jin Lyu, Jiuming Liu, Yebin Liu, Silvia Zuffi, Liang An, Stefan Goetz.
"SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild." ArXiv (2026).
[paper]
[2026.05]
UniD-Shift: Shuai Zhang, Zhecheng Shi, Zhuxiao Li, Jing Ou, Tengxi Wang, Yuan Liu, Wufan Zhao.
"UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition." ArXiv (2026).
[paper]
[code]
[2026.05]
Qwen3-VL-Seg: Yuan Yao, Qiushi Yang, Humen Zhong, Jiangning Wei, Yifang Men, Shuai Bai, Miaomiao Cui, Zhibo Yang.
"Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding." ArXiv (2026).
[paper]
[2026.05]
AS-SAM2: Lv, Tao and Ding, Shenrun and Wang, Yuji and Li, Yue and Lu, Xiaohuan and Tian, Youliang.
"AS-SAM2: Adaptive Self-correction for Visual Tracking with SAM2." TCSVT (2026).
[paper]
[code]
[2026.05]
GA3T: Siwei Cai, Knut Peterson, Quan Tran, Christian Ricks, Dhanush Parthasarathy, Amir Kaidarov, Neil Deshpande, Sukaina Najm, David Han, Lifeng Zhou.
"GA3T: A Ground-Aerial Terrain Traversability Dataset for Heterogeneous Robot Teams in Unstructured Environments." ArXiv (2026).
[paper]
[code]
[2026.05]
HP-Adapter: Hinako Mitsuoka, Kazuhiro Hotta.
"Prompt-Free and Efficient SAM2 Adaptation for Biomedical Semantic Segmentation via Dual Adapters." ICIP (2026).
[paper]
[2026.05]
Ilov3Splat: Binh Long Nguyen, Kien Nguyen, Sridha Sridharan, Clinton Fookes, Peyman Moghadam.
"Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting." ICPR (2026).
[paper]
[code]
[2026.05]
ZhiXin Sun.
"Example-Based Object Detection." ArXiv (2026).
[paper]
[code]
[2026.05]
X2SAM: Hao Wang, Limeng Qiao, Chi Zhang, Lin Ma, Guanglu Wan, Xiangyuan Lan, Xiaodan Liang.
"X2SAM: Any Segmentation in Images and Videos." ArXiv (2026).
[paper]
[code]
[project]
[2026.05]
GLASSNet: Morteza Moradi, Mohammad Moradi, Simone Palazzo, Ali Borji, Concetto Spampinato.
"Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection." ArXiv (2026).
[paper]
[2026.05]
VL-SAM-v3: Chih-Chung Liu, Zhiwei Lin, Yongtao Wang.
"VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection." ArXiv (2026).
[paper]
[2026.05]
ViewSAM: Jiawei Ge, Xintian Zhang, Jiuxin Cao, Bo Liu, Fabian Deuser, Chang Liu, Gong Wenkang, Siyou Li, Juexi Shao, Wenqing Wu, Chen Feng, Ioannis Patras.
"ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking." ArXiv (2026).
[paper]
[2026.05]
DFUDA: Yerin Cheon, Aruna Balasubramanian, Francois Rameau.
"Dual-Foundation Models for Unsupervised Domain Adaptation." ICPR (2026).
[paper]
[code]
[2026.05]
Chase Cartwright, Gongbo Guo, Sai Teja Pusuluri, Christopher N. Mayhew, Mark Hester, Horacio E. Castillo.
"Approaching human parity in the quality of automated organoid image segmentation." ArXiv (2026).
[paper]
[2026.05]
SAMamba3D: Rui Zhang, Xianzhi Song, Linqi Zhu, Branko Bijeljic, Gensheng Li, Martin J. Blunt.
"SAMamba3D: adapting Segment Anything for generalizable 3D segmentation of multiphase pore-scale images." ArXiv (2026).
[paper]
[code]
[2026.05]
Remote SAMsing: Osmar Luiz Ferreira de Carvalho, Osmar Abílio de Carvalho Júnior, Anesmar Olino de Albuquerque, Daniel Guerreiro e Silva.
"Remote SAMsing: From Segment Anything to Segment Everything." ArXiv (2026).
[paper]
[2026.05]
MmSAM: Qingpeng Wang, Zhou Huang, Ying Chen, Yi Bao.
"MmSAM: multimodal meets SAM2 for efficient remote sensing semantic segmentation." International Journal of Applied Earth Observation and Geoinformation (2026).
[paper]
[code]
[2026.04]
Three-shot SAM2: Zhongyi Zhang, Julie Hides, Enrico De Martino, Gervase Tuxworth.
"Multicenter evaluation of three-shot SAM2 segmentation for group-level quantification of lumbar paraspinal muscles at the L4/L5 level on multi-sequence MRI." European Journal of Radiology (2026).
[paper]
[2026.04]
Lee, H., Jo, Y., Hong, I. et al.
"Segment anything model guided dual-mask framework for anatomically faithful medical image translation." Sci Rep (2026).
[paper]
[2026.04]
SAM-BM: Huanhuan Lv, Songru Jiang, Yuhao Bai, Tuohang Wan, Chang Gou, Lijun Chen.
"SAM-BM: An adversarial benchmark for loss functions and multi-scale objects in segment anything models." CVIU (2026).
[paper]
[code]
[2026.04]
SAM2-PolypNet: Zhaoting Mu, et al.
"SAM2-PolypNet: SAM2 with adaptive context enhancement model for polyp segmentation." BSPC (2026).
[paper]
[code]
[2026.04]
Zhihao Ni, Youwen He, Ziyue Zhou, Hankun Zhang, and Jun Tian.
"Rapid 3D reconstruction of key power plant equipment using SAM foreground segmentation and 3D Gaussian splatting." AMNA(2026).
[paper]
[2026.04]
Haiyu Yang, Miel Hostens.
"Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics." ArXiv (2026).
[paper]
[2026.04]
SAM-FuseNet: Zhu, Chenyang and Wang, Jierui and Zhang, Lanlan and Liang, Jia and Su, Qianxiao and Li, Baihua.
"SAM-FuseNet: Segment Anything Guided Multimodal Fusion for RGB–Thermal Aerial Robotic Perception." TGRS (2026).
[paper]
[2026.04]
MemOVCD: Zuzheng Kuang, Honghao Chang, Boqiang Liang, Haoqian Wang, Lijun He, Fan Li, Haixia Bi.
"MemOVCD: Training-Free Open-Vocabulary Change Detection via Cross-Temporal Memory Reasoning and Global-Local Adaptive Rectification." ArXiv (2026).
[paper]
[code]
[2026.04]
Bridge: Mingbo Hong, Feng Liu, Caroline Gevaert, George Vosselman, Hao Cheng.
"Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalization." CVPR (2026).
[paper]
[code]
[2026.04]
CRC-SAM: Daniel Lao.
"CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images." ISBI (2026).
[paper]
[2026.04]
MAFFNet: Zhiwei Feng and Benyi Yang and Baosong Deng.
"SAM-Assisted Multimodal Collaborative Enhancement for Remote Sensing Image Segmentation." Information Fusion (2026).
[paper]
[2026.04]
Sanghati Basu.
"Robustness Evaluation of a Foundation Segmentation Model Under Simulated Domain Shifts in Abdominal CT: Implications for Health Digital Twin Deployment." ArXiv (2026).
[paper]
[code]
[2026.04]
FastSAM-CD: Zhang, Shuxin and Lei, Tao and Wang, Xingwu and Liu, Tongfei and Lv, Zhiyong and Liu, Daqi and Gong, Maoguo and Nandi, Asoke K.
"FastSAM-CD: Remote Sensing Image Change Detection Using Vision Foundation Models With Stronger Encoder and Decoder." TGRS (2026).
[paper]
[2026.04]
GeoSAM: Wujie Zhou, Jin Xie, Caie Xu, Yuanyuan Liu, Yunchao Wang.
"Adapt, Generate, and Supervise: Geometry-Aware Diffusion-Guided SAM Framework for Remote Sensing Semantic Segmentation." TGRS (2026).
[paper]
[code]
[2026.04]
ATSG: Zhang, Yifan and Jiang, Zhiguo and Zhang, Haopeng.
"ATSG: Adaptive Token Linking With Segment Anything Model Guidance for Weakly Supervised Remote Sensing Image Semantic Segmentation." TGRS (2026).
[paper]
[code]
[2026.04]
SemiSAM-O1: Yichi Zhang, Le Xue, Bichun Xu, Judong Luo, Zhigang Wu, Yu Fu, Zixin Hu, Yuan Cheng, Yuan Qi.
"SemiSAM-O1: How far can we push the boundary of annotation-efficient medical image segmentation?." Medical Image Analysis (2026).
[paper]
[code]
[2026.04]
INSIGHT: Alexander Nikitas Dimopoulos, Joseph Grasso, John Beltz.
"INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public Safety." ArXiv (2026).
[paper]
[2026.04]
AgentRVOS: Deshui Miao, Chao Yang, Chao Tian, Guoqing Zhu, Kai Yang, Zhifan Mo, Xin Li.
"AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method." ArXiv (2026).
[paper]
[2026.04]
OAMVOS: Deshui Miao, Xingsen Huang, Yameng Gu, Xiaogang yu, Xin Li, Ming-Hsuan Yang.
"OAMVOS: 2nd Report for 5th PVUW MOSE Track." ArXiv (2026).
[paper]
[2026.04]
ASR-SaSaSa2VA: Zhiyu Wang, Xudong Kang, Shutao Li.
"2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA." ArXiv (2026).
[paper]
[2026.04]
DiffuSAM: Tal Grossman, Noa Cahan, Lev Ayzenberg, Hayit Greenspan.
"DiffuSAM: Diffusion-Based Prompt-Free SAM2 for Few-Shot and Source-Free Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.04]
SPD: Jingxuan Kang, Ziqi Zhang, Shaoming Zheng, Shuang Li, Uday Bharat Patel, Alexander Harry Fitzhugh, Phillip Lung, Yusuf Kiberu, Nikesh Jathanna, Shahnaz Jamil-Copley, Bernhard Kainz, Chen Qin.
"Learning from Noisy Prompts: Saliency-Guided Prompt Distillation for Robust Segmentation with SAM." CVPR (2026).
[paper]
[2026.04]
SGP-SAM: Zixuan Tang, Shen Zhao.
"SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation." ArXiv (2026).
[paper]
[2026.04]
HFS-TriNet: Xu Lu, Qianhong Peng, Qihao Zhou, Shaopeng Liu, Xiuqin Ye, Chuan Yang, Yuan Yuan.
"HFS-TriNet: A Three-Branch Collaborative Feature Learning Network for Prostate Cancer Classification from TRUS Videos." CVPR (2026).
[paper]
[code]
[2026.04]
HAM-SAM2: Pan, Kanghua and Chen, Guo and Zhu, Wei and Zhao, Danhuai and Lu, Tong.
"HAM-SAM2: Enhancing SAM2 for Visual Object Tracking with Adaptive Motion Modeling and Hierarchical Memory Bank." ICASSP (2026).
[paper]
[2026.04]
Xiao, Haodong and Yu, Wenbo and Fang, Hao and Sun, Shuoyang and Chen, Bin and Wang, Xuan and Xia, Shu-Tao.
"Diffusion-Based Natural Adversarial Perturbations Towards Segment Anything Model." ICASSP (2026).
[paper]
[2026.04]
AnchorDiffusion: Honggang Zhao, et al.
"AnchorDiffusion: High-fidelity local image editing via anchor-SAM masks and dynamic noise fusion." Information Sciences (2026).
[paper]
[code]
[2026.04]
Isaac (Zack) Duitz, et al.
"Accelerating Medical Image Segmentation with EfficientViT-SAM." ArXiv (2026).
[paper]
[2026.04]
SAMDistill: Zhaozhong Wang, Dian Shao, Lei Zhang, Zuowei Zhang & Binglu Wang.
"SAMDistill: SAM-based Spatial-temporal Distillation for Robust 3D Object Detection." MIR (2026).
[paper]
[2026.04]
DBSAM: Zheng, Ziyi and Li, Weixing and Pan, Feng and Wang, Ronghao and Gao, Qi.
"DBSAM: A Dual-Branch Segment Anything Model for Infrared Small Target Detection." TGRS (2026).
[paper]
[2026.04]
Anchor-SAM: Li, Wenhai and Huang, Xiaohui and Yang, Xiaofei and Zhou, Yicong and Peng, Jiangtao and Ban, Yifang and Jiang, Nan.
"Anchor-SAM: Active Mining of Latent Anchors From SAM Encoder for Road Extraction." TGRS (2026).
[paper]
[code]
[2026.04]
PromptReg: Huang, Shiqi and Xu, Tingfa and Li, Jianan and Saeed, Shaheer U. and Shen, Ziyi and Barratt, Dean C. and Hu, Yipeng.
"PromptReg: Interactive Registration by “Corresponding Prompts” for Segment Anything Model (SAM)." TIP (2026).
[paper]
[2026.04]
Echo-SAM: Chang Liu, et al.
"Echo-SAM: fully exploits the performance of SAM for echocardiography segmentation." Biomedical Signal Processing and Control (2026).
[paper]
[2026.04]
Med-JSCC: Yang, Fan and Sun, Shuo and Jin, Chanyuan and Gao, Zhen and Niyato, Dusit.
"MedSAM-2 Large Model-Driven Medical Image Semantic Communication for Telemedicine." IEEE Internet of Things Journal (2026).
[paper]
[2026.04]
FMTW-SAM: Wenjie Cai, et al.
"FMTW-SAM: Foreground mixing and temporally weighted SAM feature fusion for cross-domain semi-supervised segmentation of type-B aortic dissection in computed tomography angiography." Neurocomputing (2026).
[paper]
[2026.04]
LAES-UNet: Tingru Liu, Yantong Zhan, Yan Wang and Delong Shao.
"An EfficientSAM-based Integrated Network for Ore Image Segmentation." Engineering Research Express (2026).
[paper]
[2026.04]
DualSplat: Xu Wang, Zhiru Wang, Shiyun Xie, Chengwei Pan, Yisong Chen.
"DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failures." CVPR (2026).
[paper]
[code]
[2026.04]
Amodal SAM: Bo Zhang, Zhuotao Tian, Xin Tao, Songlin Tang, Jun Yu, Wenjie Pei.
"Amodal SAM: A Unified Amodal Segmentation Framework with Generalization." ArXiv (2026).
[paper]
[2026.04]
DualGaze-VLM: Zehong Ke, Yanbo Jiang, Jinhao Li, Zhiyuan Liu, Yiqian Tu, Qingwen Meng, Heye Huang, Jianqiang Wang.
"From Scene to Object: Text-Guided Dual-Gaze Prediction." ArXiv (2026).
[paper]
[2026.04]
Semantic-Fast-SAM: Byunghyun Kim.
"Semantic-Fast-SAM: Efficient Semantic Segmenter." APSIPA ASC (2026).
[paper]
[code]
[2026.04]
SHP-SAM: Xiao, Fen and Huang, Ruozhuo and Wu, Zhenwei and Gao, Xieping.
"Scribble-guided Hierarchical Prompt for SAM-Based Weakly Supervised Salient Object Detection." TCSVT (2026).
[paper]
[code]
[2026.04]
YOLOv10–SAM: Verma, Pooja and Paul, Ayan and Machavaram, Rajendra and Bhattacharya, Mahua.
"Toward Grounded YOLO-SAM: Unified Detection–Segmentation Framework for Agricultural Intelligence." ACDSA (2026).
[paper]
[2026.04]
CoCo-SAM3: Yanhui Chen, Baoyao Yang, Siqi Liu, Jingchao Wang.
"CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation." ArXiv (2026).
[paper]
[2026.04]
LiteBounD: Shivanshu Agnihotri, Snehashis Majhi, Deepak Ranjan Nayak.
"Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models." ArXiv (2026).
[paper]
[code]
[2026.04]
ASTM Grain Size Estimator: Abdul Mueez, Shruti Vyas.
"Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images." CVPR Workshops (2026).
[paper]
[code]
[2026.04]
RefAtt-SAM: Wei, Xingji and Liu, Nanqing and Lei, Sen and Li, Heng-Chao.
"Reference and Attention Guided Few-Shot Adaptation of Segment Anything Model for Remote Sensing Images." TGRS (2026).
[paper]
[code]
[2026.04]
SEM-SAM: Linlin Wei, Yongfeng Xiao, Yifei Wang, Shan Ye, Tao Xu, Congyun Liu, Shuai Shao, Linge Ma, Jihong Cheng, Haowei Pei, Shuping Yin, Zhihua Han, Fuguo Jiang.
"From Optical Inversion to AI Vision: an AI-driven SEM Workflow for Empowering Precise Granulometric Analysis." Clean Energy (2026).
[paper]
[2026.04]
SGCT-Net: Jiang, Yubo and Yuan, Zheming and Zhou, Tairan and Chen, Jing and Xie, Fengying and Jiang, Zhiguo and Zhang, Haopeng.
"SGCT-Net: SAM-Guided Cross-Teaching Network for Weakly Supervised Semantic Segmentation for Generating High-Quality CAMs in High-Resolution Remote Sensing Imagery." JSTARS (2026).
[paper]
[2026.04]
PLS: Zhang, Aoran and Ling, Zhigang and Tan, Haoran and Wang, Yaonan.
"A Part-aware Learning Network for Weakly Supervised Semantic Segmentation." TMM (2026).
[paper]
[2026.04]
DiffuSAM: Geet Sethi, Panav Shah, Ashutosh Gandhe, Soumitra Darshan Nayak.
"DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery." ICLR Workshop (2026).
[paper]
[2026.04]
ViperSAM: Dawar Jyoti Deka.
"Inference-Time Temporal Probability Smoothing for Stable Video Segmentation with SAM2 under Weak Prompts." ArXiv (2026).
[paper]
[2026.04]
Qiuyu Kong, Shakiba Sharifi, Zanxi Ruan, Yiming Wang, Marco Cristani.
"Is SAM3 ready for pathology segmentation?." ArXiv (2026).
[paper]
[2026.04]
Yifei Yan, Yankai Liao, Linqi Ye.
"A Rapid Deployment Pipeline for Autonomous Humanoid Grasping Based on Foundation Models." ArXiv (2026).
[paper]
[2026.04]
Dual-Anchoring: Kangyi Wu, Pengna Li, Kailin Lyu, Lin Zhao, Qingrong He, Jinjun Wang, Jianyi Liu.
"Dual-Anchoring: Addressing State Drift in Vision-Language Navigation." ArXiv (2026).
[paper]
[2026.04]
Islam Mansour, Francescopaolo Sica, Michael Schmitt.
"Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery." ArXiv (2026).
[paper]
[2026.04]
Petro-SAM: Yili Ren, Shiqi Wen, Li Hou, Dingwen Xiao, Weiming Zhang, Caleb Chen Cao, Lin Wang, Zilu Zheng, Qianxiao Su, Mingjun Zhao, Lei Chen.
"From Boundaries to Semantics: Prompt-Guided Multi-Task Learning for Petrographic Thin-section Segmentation." ArXiv (2026).
[paper]
[2026.04]
WILD-SAM: Yucheng Pan, Heping Li, Zhangle Liu, Sajid Hussain, Bin Pan.
"WILD-SAM: Phase-Aware Expert Adaptation of SAM for Landslide Detection in Wrapped InSAR Interferograms." ArXiv (2026).
[paper]
[2026.04]
TF-SMOT: Laurence Bonat, Francesco Tonini, Elisa Ricci, Lorenzo Vaquero.
"Training-Free Semantic Multi-Object Tracking with Vision-Language Models." FG (2026).
[paper]
[2026.04]
Hayato Inoue, Shota Harada, Shumpei Takezaki, Ryoma Bise.
"Cell Instance Segmentation via Multi-Task Image-to-Image Schrödinger Bridge." ArXiv (2026).
[paper]
[2026.04]
Pi-HOC: Sravan Chittupalli, Ayush Jain, Dong Huang.
"Pi-HOC: Pairwise 3D Human-Object Contact Estimation." ArXiv (2026).
[paper]
[code]
[2026.04]
Caiwen Jiang, Lei Zeng, Wei Liu.
"A 3D SAM-Based Progressive Prompting Framework for Multi-Task Segmentation of Radiotherapy-induced Normal Tissue Injuries in Limited-Data Settings." Medical Image Analysis (2026).
[paper]
[2026.04]
Hao Wang, Jiqing Zhang, Xin Yang, Baocai Yin, Lu Jiang, Zetian Mi, Huibing Wang.
"Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection." ArXiv (2026).
[paper]
[2026.04]
PR-MaGIC: Minjae Lee, Sungwoo Hur, Soojin Hwang, Won Hwa Kim.
"PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation." ArXiv (2026).
[paper]
[code]
[2026.04]
OccSAM-Bench: Nhan Ho, Luu Le, Thanh-Huy Nguyen, Thien Nguyen, Xiaofeng Liu, Ulas Bagci.
"Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models." CVPR workshop (2026).
[paper]
[2026.04]
VLMaterial: Jiangyou Zhu, He Chen.
"VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification." ArXiv (2026).
[paper]
[2026.04]
H-SPAM: Julien Walther, Rémi Giraud, Michaël Clément.
"H-SPAM: Hierarchical Superpixel Anything Model." ArXiv (2026).
[paper]
[2026.04]
SeSAM : Anurag Das, Anna Kukleva, Xinting Hu, Yuki M. Asano, Bernt Schiele.
"Do Instance Priors Help Weakly Supervised Semantic Segmentation?." ArXiv (2026).
[paper]
[2026.04]
Boxes2Pixels: Camile Lendering, Erkut Akdag, Egor Bondarev.
"Boxes2Pixels: Learning Defect Segmentation from Noisy SAM Masks." CVPR workshop (2026).
[paper]
[code]
[2026.04]
Kaden Stillwagon, Alexandra Dunnum VandeLoo, Benjamin Magondu, Craig R. Forest.
"Self-supervised Pretraining of Cell Segmentation Models." ArXiv (2026).
[paper]
[2026.04]
RobustMedSAM: Jieru Li, Matthew Chen, Micky C. Nnamdi, J. Ben Tamo, Benoit L. Marteau, May D. Wang.
"RobustMedSAM: Degradation-Resilient Medical Image Segmentation via Robust Foundation Model Adaptation." ArXiv (2026).
[paper]
[2026.04]
PASTA: Melanie Neubauer, Elmar Rueckert, Christian Rauch.
"PASTA: Vision Transformer Patch Aggregation for Weakly Supervised Target and Anomaly Segmentation." ArXiv (2026).
[paper]
[2026.04]
MV3DIS: Yibo Zhao, Yigong Zhang, Jin Xie.
"MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.04]
Adrian Manchado, Tanner Cellio, Jonathan Keane, Yiyang Wang.
"AI Driven Soccer Analysis Using Computer Vision." ArXiv (2026).
[paper]
[2026.04]
Lars Lundqvist, Earl Ranario, Hamid Kamangir, Heesup Yun, Christine Diepenbrock, Brian N. Bailey, J. Mason Earles.
"Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection." ArXiv (2026).
[paper]
[2026.04]
DiTTA: Jihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin Yoon.
"Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation." CVPR (2026).
[paper]
[code]
[2026.04]
OVS-DINO: Haoxi Zeng, Qiankun Liu, Yi Bin, Haiyue Zhang, Yujuan Ding, Guoqing Wang, Deqiang Ouyang, Heng Tao Shen.
"OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance." ArXiv (2026).
[paper]
[2026.04]
Tarot-SAM3: Weiming Zhang, Dingwen Xiao, Songyue Guo, Guangyu Xiang, Shiqi Wen, Minwei Zhao, Lei Chen, Lin Wang.
"Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation." ArXiv (2026).
[paper]
[2026.04]
PanoSAM2: Dingwen Xiao, Weiming Zhang, Shiqi Wen, Lin Wang.
"PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation." ArXiv (2026).
[paper]
[2026.04]
UniSurgSAM: Haofeng Liu, Ziyue Wang, Alex Y. W. Kong, Guanyi Qin, Yunqiu Xu, Chang Han Low, Mingqi Gao, Lap Yan Lennon Chan, Yueming Jin.
"UniSurgSAM: A Unified Promptable Model for Reliable Surgical Video Segmentation." ArXiv (2026).
[paper]
[code]
[2026.04]
Boxer: Daniel DeTone, Tianwei Shen, Fan Zhang, Lingni Ma, Julian Straub, Richard Newcombe, Jakob Engel.
"Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D." ArXiv (2026).
[paper]
[code]
[2026.04]
Ye Bi, Bimala Acharya, David Rosero, Juan Steibel.
"Automated Segmentation and Tracking of Group Housed Pigs Using Foundation Models." ArXiv (2026).
[paper]
[2026.04]
Abdelmoamen Nasser, Yousef Baba'a, Murad Mebrahtu, Nadya Abdel Madjid, Jorge Dias, Majid Khonji.
"Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs." ArXiv (2026).
[paper]
[2026.04]
SGPer: Shijie Wang, Zijian Wang, Yadan Luo, Scott Chapman, Xin Yu, Zi Huang.
"Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation." ArXiv (2026).
[paper]
[2026.04]
Pickalo: Alessandro Tarsi, Matteo Mastrogiuseppe, Saverio Taliani, Simone Cortinovis, Ugo Pattacini.
"Pickalo: Leveraging 6D Pose Estimation for Low-Cost Industrial Bin Picking." ArXiv (2026).
[paper]
[code]
[2026.04]
CardioSAM: Ujjwal Jain.
"CardioSAM: Topology-Aware Decoder Design for High-Precision Cardiac MRI Segmentation." ArXiv (2026).
[paper]
[2026.04]
FSS-SAM3: Yi-Jen Tsai, Yen-Yu Lin, Chien-Yao Wang.
"Few-Shot Semantic Segmentation Meets SAM3." ArXiv (2026).
[paper]
[code]
[2026.04]
XSeg: Hongxia Gao, Litao Li, Yixin Chen, Jiali Wen, Kaijie Zhang, Qianyun Liu.
"XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening." CVPR (2026).
[paper]
[code]
[2026.04]
IndoorCrowd: Sebastian-Ion Nae, Radu Moldoveanu, Alexandra Stefania Ghita, Adina Magda Florea.
"IndoorCrowd: A Multi-Scene Dataset for Human Detection, Segmentation, and Tracking with an Automated Annotation Pipeline." CVPR Workshop (2026).
[paper]
[code]
[2026.04]
GRAZE: Syed Ahsan Masud Zaidi, Lior Shamir, William Hsu, Scott Dietrich, Talha Zaidi.
"GRAZE: Grounded Refinement and Motion-Aware Zero-Shot Event Localization ." CVPR Workshop (2026).
[paper]
[code]
[2026.04]
DPMO: Hongru Chen, Jiyang Huang, Jia Wan, Antoni B. Chan.
"Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation." ArXiv (2026).
[paper]
[2026.04]
Derek Austin.
"Better Rigs, Not Bigger Networks: A Body Model Ablation for Gaussian Avatars." ArXiv (2026).
[paper]
[2026.04]
Xusheng He, Canyang Wu, Jinrong Zhang, Weili Guan, Jianlong Wu, Liqiang Nie.
"The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation." CVPR Workshop (2026).
[paper]
[code]
[2026.04]
TEP: Jinrong Zhang, Canyang Wu, Xusheng He, Weili Guan, Jianlong Wu, Liqiang Nie.
"Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge." CVPR Workshop (2026).
[paper]
[2026.04]
APRVOS: Deshui Miao, Yameng Gu, Chao Yang, Xin Li, Haijun Zhang, Ming-Hsuan Yang.
"APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track." ArXiv (2026).
[paper]
[2026.04]
AdaLoRA-QAT: Prantik Deb, Srimanth Dhondy, N. Ramakrishna, Anu Kapoor, Raju S. Bapi, Tapabrata Chakraborti.
"AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation." ISBI (2026).
[paper]
[code]
[2026.04]
TF-SSD: Zhijin He, Shuo Jin, Siyue Yu, Shuwei Wu, Bingfeng Zhang, Li Yu, Jimin Xiao.
"TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection." CVPR (2026).
[paper]
[code]
[2026.04]
PC-SAM: Chengcheng Lv, Rushi Li, Mincheng Wu, Xiufang Shi, Zhenyu Wen, Shibo He.
"PC-SAM: Patch-Constrained Fine-Grained Interactive Road Segmentation in High-Resolution Remote Sensing Images." ArXiv (2026).
[paper]
[code]
[2026.04]
LunarRockSAM: Wang, Yinan and Ye, Hongxia and Fa, Wenzhe.
"LunarRockSAM: A Domain-Adapted SAM with Bright-Spots Prompting and Conditional Screening for Lunar Rock Extraction." IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (2026).
[paper]
[2026.03]
DBM-SAM: Wei Gao, Teng Li, Cunang Jiang, Sicheng Wang, Yu Dai.
"DBM-SAM: Dual-branch multiscale adaptation of SAM for medical ultrasound segmentation." Displays (2026).
[paper]
[2026.03]
SAM2-RoadNet: Feng, Ruyue, Ziyou Guo, Xiao Du, and Tieru Wu.
"SAM2-RoadNet: Topology-Aware Multi-Scale Road Extraction from High-Resolution Remote Sensing Images." Remote Sensing (2026).
[paper]
[2026.03]
IDRG-mSAM: Wang, Leiquan and Meng, Yu and Luo, Chunbo and Xu, Mingming and Wu, Chunlei and Li, Zhongwei.
"SAM-Based Multi-Scale Fine-Tuning with Inter-layer Difference Guidance for Remote Sensing Change Detection." IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (2026).
[paper]
[2026.03]
SAM-ColonPolypGen: Shasha Zhang and Yuang Cai and Yijun Chen and Xiang Cai and Peng Li.
"SAM-ColonPolypGen: Enhancing automated colon polyp report generation via reinforcement learning and prompt chaining." Biomedical Signal Processing and Control (2026).
[paper]
[2026.03]
SemiBUVS: Long Chen and Qingqing Zheng and Yingying Chen and Faqin Lv and Qiong Wang.
"SAM-Guided Semi-Supervised Breast Lesion Segmentation in Ultrasound Videos with A New Dataset." Expert Systems with Applications (2026).
[paper]
[code]
[2026.03]
Mask-CDKD: Daoyu Shu and Zhan Zhang and Xiao Huang and Ru Wang and Nan Jia and Xinzhe Fu and Bingnan Yang and Fang Wan and Jianzhong Lu and Jianya Gong.
"Mask-CDKD: A source-free and label-free cross-domain knowledge distillation framework from SAM for satellite onboard VHR land-cover mapping." ISPRS Journal of Photogrammetry and Remote Sensing (2026).
[paper]
[2026.03]
SAM2-WaveUNet: Shuzhou Lv and Shubin Zhang and Xiaoshuang Huang and Dong An and Jincun Liu and Yan Meng and Yaoguang Wei.
"SAM2-WaveUNet: A Frequency-Enhanced Segmentation Network for Fine-Grained Marine Organism Delineation." Expert Systems with Applications (2026).
[paper]
[2026.03]
VLP-SAM: Sakurai, Kosuke, Ryotaro Shimizu, and Masayuki Goto.
"Vision and Language Reference for a Segment Anything Model for Few-Shot Segmentation." Journal of Imaging(2026).
[paper]
[2026.03]
AutoPrompt-SAM3D: Cheng, W., Tang, J., Wang, T. et al.
"AutoPrompt-SAM3D: integrated generation and selection for SAM2-based 3D medical segmentation." BMC Bioinformatics (2026).
[paper]
[2026.03]
Shata, Dina, Simon Denman, Sara Omrani, Robin Drogemuller, Hend Ali, and Ayman Wagdy.
"Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation." Buildings (2026).
[paper]
[2026.03]
HATSAM: Tang, T., Rao, Z., Wang, Y. et al.
"HATSAM: hierarchical adaptation strategy for segment anything model in medical imaging." SIViP (2026).
[paper]
[2026.03]
SaSaSaSa2VA: Dengxian Gong, Quanzhu Niu, Shihao Chen, Yuanzheng Wu, Yikang Zhou, Tao Zhang, Haobo Yuan, Lu Qi, Shunping Ji.
"SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track." ArXiv (2026).
[paper]
[2026.03]
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou, Hakim Hacid, Ngoc Dung Huynh, Phuc H. Le Khac, Sanath Narayan, Wamiq Reyaz Para, Ankit Singh.
"Falcon Perception." ArXiv (2026).
[paper]
[code]
[2026.03]
FT-FSOD: Xuanlong Yu, Youyang Sha, Longfei Liu, Xi Shen, Di Yang.
"A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps." CVPR (2026).
[paper]
[code]
[2026.03]
LIT: Xinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan, Jennifer J. Sun.
"Live Interactive Training for Video Segmentation." CVPR (2026).
[paper]
[code]
[2026.03]
Xinyao Zhang, Chang Liu, Xiao Liang, Minghui Zheng, Sara Behdad.
"Evaluating Large and Lightweight Vision Models for Irregular Component Segmentation in E-Waste Disassembly." MSEC (2026).
[paper]
[2026.03]
Syn4Seg: Guohuan Xie, Xin He, Dingying Fan, Le Zhang, Ming-Ming Cheng, Yun Liu.
"Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation." ArXiv (2026).
[paper]
[2026.03]
IP-SAM: Huiyao Zhang, Jin Bai, Rui Guo, JianWen Tan, HongFei Wang, Ye Li.
"IP-SAM: Prompt-Space Conditioning for Prompt-Absent Camouflaged Object Detection." ECCV (2026).
[paper]
[2026.03]
OpenDPR: Qi Guo, Jue Wang, Yinhe Liu, Yanfei Zhong.
"OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery." CVPR (2026).
[paper]
[code]
[2026.03]
Industrial3D: Chao Yin, Hongzhe Yue, Qing Han, Difeng Hu, Zhenyu Liang, Fangzhou Lin, Bing Sun, Boyu Wang, Mingkai Li, Wei Yao, Jack C. P. Cheng.
"Industrial3D: A Terrestrial LiDAR Point Cloud Dataset and CrossParadigm Benchmark for Industrial Infrastructure." ArXiv (2026).
[paper]
[code]
[2026.03]
CFR-SAM: Jingze Su, Tianle Zhu, Jiaxin Cai, Zhiyi Wang, Qi Li, Xiao Zhang, Tong Tong, Shu Wang, Wenxi Liu.
"Adapting SAM to Nuclei Instance Segmentation and Classification via Cooperative Fine-Grained Refinement." ArXiv (2026).
[paper]
[2026.03]
RAP: Zhihao Mao, Bangpu Chen.
"RAP: Retrieve, Adapt, and Prompt-Fit for Training-Free Few-Shot Medical Image Segmentation." IJCNN (2026).
[paper]
[2026.03]
Samik Some, Vinay P. Namboodiri.
"Can Unsupervised Segmentation Reduce Annotation Costs for Video Semantic Segmentation?." ICVGIP (2026).
[paper]
[2026.03]
M. Fazri Nizar.
"Domain-Guided YOLO26 with Composite BCE-Dice-Lovász Loss for Multi-Class Fetal Head Ultrasound Segmentation." ArXiv (2026).
[paper]
[2026.03]
Mask-CDKD: Daoyu Shu and Zhan Zhang and Xiao Huang and Ru Wang and Nan Jia and Xinzhe Fu and Bingnan Yang and Fang Wan and Jianzhong Lu and Jianya Gong.
"Mask-CDKD: A source-free and label-free cross-domain knowledge distillation framework from SAM for satellite onboard VHR land-cover mapping." ISPRS Journal of Photogrammetry and Remote Sensing (2026).
[paper]
[code]
[2026.03]
Colon-Bench: Abdullah Hamdi, Changchun Yang, Xin Gao.
"Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos." ArXiv (2026).
[paper]
[code]
[2026.03]
Nitin Kulkarni, Akhil Devarashetti, Charlie Cluss, Livio Forte, Philip Schneider, Chunming Qiao, Alina Vereshchaka.
"Drive-Through 3D Vehicle Exterior Reconstruction via Dynamic-Scene SfM and Distortion-Aware Gaussian Splatting." ArXiv (2026).
[paper]
[2026.03]
Guoping Xu, Jayaram K. Udupa, Yubing Tong, Xin Long, Ying Zhang, Jie Deng, Weiguo Lu, You Zhang.
"Adapting Segment Anything Model 3 for Concept-Driven Lesion Segmentation inMedical Images: An Experimental Study." ArXiv (2026).
[paper]
[code]
[2026.03]
Sieradzki, Alexander, Kamil Koszela, Szymon Koszykowski, Jakub Bednarek, and Jarosław Kurek.
"Zero-Shot Vertebral Instance Segmentation on DICOM Spine Radiographs Using Promptable Segment Anything Models." Journal of Clinical Medicine (2026).
[paper]
[2026.03]
SemiBUVS: Long Chen and Qingqing Zheng and Yingying Chen and Faqin Lv and Qiong Wang.
"SAM-Guided Semi-Supervised Breast Lesion Segmentation in Ultrasound Videos with A New Dataset." ESWA (2026).
[paper]
[code]
[2026.03]
GridVAD: Mohamed Eltahir, Ahmed O. Ibrahim, Obada Siralkhatim, Tabarak Abdallah, Sondos Mohamed.
"GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids." ArXiv (2026).
[paper]
[code]
[2026.03]
XAI-SAM: Abu Noman Md Sakib, Merjulah Roby, Zijie Zhang, Satish Muluk, Mark K. Eskandari, Ender A. Finol.
"Dissecting Model Failures in Abdominal Aortic Aneurysm Segmentation through Explainability-Driven Analysis." CVPR (2026).
[paper]
[2026.03]
ET-SAM: Xike Zhang, Maoyuan Ye, Juhua Liu, Bo Du.
"ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis." ECCV (2026).
[paper]
[code]
[2026.03]
UW-VOS: Hongshen Zhao, Jingkang Tai, Yuhang Wu, Wenkang Zhang, Xi Lan, Shangyan Wang, Tianyu Zhang, Wankou Yang.
"UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation." ArXiv (2026).
[paper]
[2026.03]
Mingqi Gao, Sijie Li, Jungong Han.
"Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track." ArXiv (2026).
[paper]
[2026.03]
AgentRVOS: Woojeong Jin, Jaeho Lee, Heeseong Shin, Seungho Jang, Junhwan Heo, Seungryong Kim.
"AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation." ArXiv (2026).
[paper]
[code]
[2026.03]
FCL-COD: Jingchen Ni, Quan Zhang, Dan Jiang, Keyu Lv, Ke Zhang, Chun Yuan.
"FCL-COD: Weakly Supervised Camouflaged Object Detection with Frequency-aware and Contrastive Learning." CVPR (2026).
[paper]
[2026.03]
Miquel Lopez Escoriza, Pau Amargant Alvarez.
"Automatic Segmentation of 3D CT scans with SAM2 using a zero-shot approach." ArXiv (2026).
[paper]
[2026.03]
VIRST-Audio: Jihwan Hong, Jaeyoung Do.
"3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio." CVPR workshop (2026).
[paper]
[code]
[2026.03]
FoB: Yuntian Bo, Yazhou Zhu, Piotr Koniusz, Haofeng Zhang.
"Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting." CVPR (2026).
[paper]
[code]
[2026.03]
CataractSAM-2: Mohammad Eslami, Dhanvinkumar Ganeshkumar, Saber Kazeminasab, Michael G. Morley, Michael V. Boland, Michael M. Lin, John B. Miller, David S. Friedman, Nazlee Zebardast, Lucia Sobrin, Tobias Elze.
"CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation." ArXiv (2026).
[paper]
[2026.03]
Lei Huang, Kai-Li Wang, Zhang Chen, Zhen-Huang, Saidjafar Murodzoda, Xin Chen, Jing Chen, Chun-Hao Chen, Yu Xia, Yu-Tong Yang, Jia-Cheng Li, Dilshod Nematov, Ilhan Yavuz, Zhao-Kui Wang.
"SAM Molecular Stacking with Heterogeneous Orientationfor High-Performance Perovskite Photovoltaics." ArXiv (2026).
[paper]
[2026.03]
Thomas Mendelson, Joshua Francois, Galit Lahav, Tammy Riklin-Raviv.
"Boundary-Aware Instance Segmentation in Microscopy Imaging." ISBI (2026).
[paper]
[2026.03]
Muhammad Hassan Maqsood, Yanming Zhu, Alfred Lam, Getamesay Dagnaw, Xuefei Yin, Alan Wee-Chung Liew.
"Prompt-Free Lightweight SAM Adaptation for Histopathology Nuclei Segmentation with Strong Cross-Dataset Generalization." ISBI (2026).
[paper]
[2026.03]
Carolin Teuber, Anwai Archit, Tobias Boothe, Peter Ditte, Jochen Rink, Constantin Pape.
"Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy." ArXiv (2026).
[paper]
[2026.03]
Distillation-SAM: Tang, Jiyang and Han, Hu and Shan, Shiguang and Chen, Xilin.
"Distillation-SAM: Knowledge Distillation Based Auto-prompt Embedding Learning for Surgical Image Segmentation." TMI (2026).
[paper]
[code]
[2026.03]
EventVCOD: Zhang, H., Lyu, Y., Liu, H., Song, J., Yuan, D., & Yang, Y.
"Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired Data." AAAI (2026).
[paper]
[code]
[2026.03]
GoalVLM: MoniJesu James, Amir Atef Habel, Aleksey Fedoseev, Dzmitry Tsetserokou.
"GoalVLM: VLM-driven Object Goal Navigation for Multi-Agent System." ArXiv (2026).
[paper]
[2026.03]
Perceptio: Yuchen Li, Amanmeet Garg, Shalini Chaudhuri, Rui Zhao, Garin Kessler.
"Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
Truncated — view the full README on GitHub.
This repository is for the first comprehensive survey on Meta AI's Segment Anything Model (SAM).
1,222
1,367 commits
updated Sep 23, 2026
The First Comprehensive SAM Survey: A Comprehensive Survey on Segment Anything Model for Vision and Beyond. Chunhui Zhang, Li Liu, Yawen Cui, Guanjie Huang, Weilin Lin, Yiqian Yang, Yuehong Hu. [paper] [homepage][中文解读]
Abstract: Artificial intelligence (AI) is evolving towards artificial general intelligence, which refers to the ability of an AI system to perform a wide range of tasks and exhibit a level of intelligence similar to that of a human being. This is in contrast to narrow or specialized AI, which is designed to perform specific tasks with a high degree of efficiency. Therefore, it is urgent to design a general class of models, which we term foundation models, trained on broad data that can be adapted to various downstream tasks. The recently proposed segment anything model (SAM) has made significant progress in breaking the boundaries of segmentation, greatly promoting the development of foundation models for computer vision. To fully comprehend SAM, we conduct a survey study. As the first to comprehensively review the progress of segmenting anything task for vision and beyond based on the foundation model of SAM, this work focuses on its applications to various tasks and data types by discussing its historical development, recent progress, and profound impact on broad applications. We first introduce the background and terminology for foundation models including SAM, as well as state-of-the-art methods contemporaneous with SAM that are significant for segmenting anything task. Then, we analyze and summarize the advantages and limitations of SAM across various image processing applications, including software scenes, real-world scenes, and complex scenes. Importantly, many insights are drawn to guide future research to develop more versatile foundation models and improve the architecture of SAM. We also summarize massive other amazing applications of SAM in vision and beyond. Finally, we maintain a continuously updated paper list and an open-source project summary for foundation model SAM at here.
Awesome Segment Anything Models: A curated list of awesome segment anything models in computer vision and beyond. This repository supplements our survey paper. We intend to continuously update it.
:boom:SAM 3.1: ''SAM 3.1 Object Multiplex'' was released.
:boom:SAM Audio: ''SAM Audio: Segment Anything in Audio'' was released.
:boom:SAM 3D: ''SAM 3D: 3Dfy Anything in Images'' was released.
:boom:SAM 3: ''SAM 3: Segment Anything with Concepts'' was released.
:boom:SAM 2: ''Segment Anything in Images and Videos'' was released.
:boom:SAM: ''Segment Anything'' was released.
:boom:SAM & SAM2 for videos: The first survey on Segment Anything for Videos: A Systematic Survey was online.
- 2026.06.05: SAM 3D won the CVPR 2026 Best Paper Honorable Mention.
- 2026.03.27: SAM 3.1 Object Multiplex was released.
- 2025.12.15: SAM Audio was released.
- 2025.11.19: SAM 3 and SAM 3D were released.
- 2025.10.11: SAM 3 arrives! Officially announced and set to launch.
- 2025.04.22: SAM 2 won the ICLR 2025 Best Paper Honorable Mention.
- 2024.07.31: The first survey on SAM & SAM2 for Videos was online.
- 2024.07.29: The SAM 2 was released.
- 2023.07.14: "Segment Anything" was accepted by ICCV 2023 (Best Paper Honorable Mention).
- 2023.05.16: An initial version of this Awesome-Segment-Anything project.
- 2023.05.14: The first comprehensive SAM survey was online.
- 2023.04.05: The paper of "Segment Anything" was online.
If you find our work useful in your research, please consider citing:
@article{zhang2023comprehensive,
title={A Comprehensive Survey on Segment Anything Model for Vision and Beyond},
author={Zhang, Chunhui and Liu, Li and Cui, Yawen and Huang, Guanjie and Lin, Weilin and Yang, Yiqian and Hu, Yuehong},
journal={arXiv preprint arXiv:2305.08196},
year={2023}
}
@article{zhang2024segment,
title={Segment Anything for Videos: A Systematic Survey},
author={Zhang, Chunhui and Cui, Yawen and Lin, Weilin and Huang, Guanjie and Rong, Yan and Liu, Li and Shan, Shiguang},
journal={arXiv preprint arXiv:2408.08315},
year={2024}
}
The First Comprehensive SAM Survey: Chunhui Zhang, Li Liu, Yawen Cui, Guanjie Huang, Weilin Lin, Yiqian Yang, Yuehong Hu.
"A Comprehensive Survey on Segment Anything Model for Vision and Beyond." ArXiv (2024).
[paper]
[homepage]
[中文解读]
[2023.05]
The First Survey on SAM & SAM2 for Videos: Chunhui Zhang, Yawen Cui, Weilin Lin, Guanjie Huang, Yan Rong, Li Liu, Shiguang Shan.
"Segment Anything for Videos: A Systematic Survey." ArXiv (2024).
[ArXiv]
[ChinaXiv]
[ResearchGate]
[Project]
[中文解读]
[2024.07]
SAM4MIS: Yichi Zhang, Rushi Jiao.
"Towards Segment Anything Model (SAM) for Medical Image Segmentation: A Survey." CBM (2024).
[paper]
[project]
[2023.05]
Yichi Zhang, Zhenrong Shen.
"Unleashing the Potential of SAM2 for Biomedical Images and Videos: A Survey." ArXiv (2024).
[paper]
[code]
[2024.08]
Tianfei Zhou, Fei Zhang, Boyu Chang, Wenguan Wang, Ye Yuan, Ender Konukoglu, Daniel Cremers.
"Image Segmentation in Foundation Model Era: A Survey." ArXiv (2024).
[paper]
[2024.08]
Chaoning Zhang, Fachrina Dewi Puspitasari, Sheng Zheng, Chenghao Li, Yu Qiao, Taegoo Kang, Xinru Shan, Chenshuang Zhang, Caiyan Qin, Francois Rameau, Lik-Hang Lee, Sung-Ho Bae, Choong Seon Hong.
"A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering." ArXiv (2024).
[paper]
[2023.05]
Xiaorui Sun, Jun Liu, Heng Tao Shen, Xiaofeng Zhu, Ping Hu.
"On Efficient Variants of Segment Anything Model: A Survey." IJCV (2025).
[paper]
[2024.10]
Mudassar Ali and Tong Wu and Haoji Hu and Qiong Luo and Dong Xu and Weizeng Zheng and Neng Jin and Chen Yang and Jincao Yao.
"A review of the Segment Anything Model (SAM) for medical image analysis: Accomplishments and perspectives." Computerized Medical Imaging and Graphics (2024).
[paper]
[2024.12]
Zhang Jiaxing, Tang Hao.
"SAM2 for Image and Video Segmentation: A Comprehensive Survey." ArXiv (2025).
[paper]
[2025.03]
Kang Wang.
"A survey on SAM-based methods for medical image segmentation." IS-AII (2025).
[paper]
[2025.07]
Guoping Xu, Jayaram K. Udupa, Yajun Yu, Hua-Chieh Shao, Songlin Zhao, Wei Liu, You Zhang.
"Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future." ArXiv (2025).
[paper]
[2025.07]
WanSAM4RS-Tracker: Zhipeng Wan and Sheng Wang and Wei Han and Yuewei Wang and Xiaohui Huang and Xiaohan Zhang and Xiaodao Chen and Yunliang Chen.
"A systematic survey and meta-analysis of the segment anything model in remote sensing image processing: Challenges, advances, applications, and opportunities." ISPRS Journal of Photogrammetry and Remote Sensing (2025).
[paper]
[project]
[2025.09]
Yang, Yizai and Cheng, Lechao and Wang, Yaxiong and Hui, Tianrui and Li, Wenjing and Zhong, Zhun.
"A Survey for Point Prompt of Segment Anything Model." MMAsia Workshops (2025).
[paper]
[2025.12]
SAM: Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, Ross Girshick.
"Segment Anything." ICCV (2023) Best Paper Honorable Mention.
[paper]
[homepage]
[code]
[Zhihu]
[Reddit]
[2023.04]
SAM 2: Nikhila Ravi∗,†, Valentin Gabeur∗, Yuan-Ting Hu∗, Ronghang Hu∗, Chaitanya Ryali∗, Tengyu Ma∗, Haitham Khedr∗, Roman Rädle∗ Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár†, Christoph Feichtenhofer∗,†.
"SAM 2: Segment Anything in Images and Videos." ICLR (2025) Best Paper Honorable Mention.
[paper]
[demo]]
[code]
[project]]
[dataset]
[blog]
[2024.07]
SAM 3: Nicolas Carion*, Laura Gustafson*, Yuan-Ting Hu*, Shoubhik Debnath*, Ronghang Hu*, Didac Suris*, Chaitanya Ryali*, Kalyan Vasudev Alwala*, Haitham Khedr*, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman Rädle, Triantafyllos Afouras, Effrosyni Mavroudi, Katherine Xu°, Tsung-Han Wu°, Yu Zhou°, Liliane Momeni°, Rishi Hazra°, Shuangrui Ding°, Sagar Vaze°, Francois Porcher°, Feng Li°, Siyuan Li°, Aishwarya Kamath°, Ho Kei Cheng°, Piotr Dollar†, Nikhila Ravi†, Kate Saenko†, Pengchuan Zhang†, Christoph Feichtenhofer†.
"SAM 3: Segment Anything with Concepts." ICLR (2026).
[paper]
[arXiv]
[code]
[homepage]
[中文解读]
[2025.10]
SAM 3D: SAM 3D Team, Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jiawei Liu, Ziqi Ma, Anushka Sagar, Bowen Song, Xiaodong Wang, Jianing Yang, Bowen Zhang, Piotr Dollár, Georgia Gkioxari, MattFeiszli, Jitendra Malik.
"SAM 3D: 3Dfy Anything in Images." CVPR (2026). CVPR (2026) Best Paper Honorable Mention.
[paper]
[code]
[project]
[demo]
[blog]
[中文解读]
[2025.11]
SAM 3D Body: Xitong Yang⋆, Devansh Kukreja⋆, Don Pinkus⋆, Anushka Sagar, Taosha Fan, Jinhyung Park◦, Soyong Shin◦, Jinkun Cao, Jiawei Liu, Nicolas Ugrinovic, Matt Feiszli†, Jitendra Malik†, Piotr Dollar†, Kris Kitani†.
"SAM 3D Body: Robust Full-Body Human Mesh Recovery." ArXiv (2025).
[paper]
[code]
[project]
[2025.11]
SAM Audio: Bowen Shi∗, Andros Tjandra∗, John Hoffman∗, Helin Wang∗, Yi-Chiao Wu∗, Luya Gao∗, Julius Richter†,Matt Le†, Apoorv Vyas†, Sanyuan Chen†, Christoph Feichtenhofer‡, Piotr Dollár‡, Wei-Ning Hsu‡, Ann Lee‡.
"SAM Audio: Segment Anything in Audio." ArXiv (2025).
[paper]
[code]
[project]
[demo]
[2025.12]
GPT-4V: OpenAI.
"GPT-4V(ision) System Card." ArXiv (2023).
[paper]
[homepage]
[2023.09]
Gemini: Gemini Team, Google.
"Gemini: A Family of Highly Capable Multimodal Models." ArXiv (2023).
[paper]
[homepage]
[blog]
[2023.12]
SEEM: Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, Yong Jae Lee.
"Segment Everything Everywhere All at Once." NeurIPS (2023).
[paper]
[code]
[2023.04]
SegGPT: Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, Tiejun Huang.
"SegGPT: Segmenting Everything In Context." ICCV (2023).
[paper]
[code]
[2023.04]
Grounding DINO: Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang.
"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection." ArXiv (2023).
[paper]
[code]
[2023.04]
ImageBind: Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, Ishan Misra.
"ImageBind: One Embedding Space To Bind Them All." CVPR (2023).
[paper]
[homepage]
[code]
[2023.05]
LanguageBind: Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Wancai Zhang, Zhifeng Li, Wei Liu, Li Yuan.
"LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment." ArXiv (2023).
[paper]
[code]
Meta-Transformer: Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, Xiangyu Yue.
"Meta-Transformer: A Unified Framework for Multimodal Learning." ArXiv (2023).
[paper]
[homepage]
[code]
[中文解读]
[2023.07]
OpenSeeD: Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianfeng Gao, Jianwei Yang, Lei Zhang.
"A Simple Framework for Open-Vocabulary Segmentation and Detection." ICCV (2023).
[paper]
[code]
[2023.03]
RAM: Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, Lei Zhang.
"Recognize Anything: A Strong Image Tagging Model." ArXiv (2023).
[paper]
[homepage]
[code]
[2023.06]
PACGen: Yuheng Li, Haotian Liu, Yangming Wen, Yong Jae Lee.
"Generate Anything Anywhere in Any Scene." ArXiv (2023).
[paper]
[homepage]
[code]
[2023.06]
ASM: Weiyun Wang, Min Shi, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen, Hao Li, Xizhou Zhu, Zhiguo Cao, Yushi Chen, Tong Lu, Jifeng Dai, Yu Qiao.
"The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World." ArXiv (2023).
[paper]
[homepage]
[demo]
[2023.08]
OneFormer: Jitesh Jain, Jiachen Li, MangTik Chiu, Ali Hassani, Nikita Orlov, Humphrey Shi.
"OneFormer: One Transformer to Rule Universal Image Segmentation." CVPR (2023).
[paper]
[homepage]
[code]
[2022.11]
OVSeg: Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, Diana Marculescu.
"Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP." CVPR (2023).
[paper]
[homepage]
[code]
[2022.10]
WAM: Tom Sander, Pierre Fernandez, Alain Durmus, Teddy Furon, Matthijs Douze.
"Watermark Anything with Localized Messages." ArXiv (2024).
[paper]
[code]
[2024.11]
Sa2VA: Haobo Yuan, Xiangtai Li, Tao Zhang, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang.
"Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos." ArXiv (2025).
[paper]
[code]
[project]
[hugging face]
[2025.01]
SAMTok: Yikang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu, Ye Tian, Haochen Wang, Haobo Yuan, Jiacong Wang, Lu Qi, Hao Fei, Anran Wang, Zhuochen Wang, Yujing Wang, Cheng Chen, Shunping Ji, Xiangtai Li.
"SAMTok: Representing Any Mask with Two Words." ArXiv (2026).
[paper]
[code]
[project]
[hugging face]
[demo]
[2026.01]
DAM: Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui.
"Describe Anything: Detailed Localized Image and Video Captioning." ArXiv (2025).
[paper]
[code]
[project]
[huggingface]
[2025.04]
DINOv2: Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, Piotr Bojanowski.
"DINOv2: Learning Robust Visual Features without Supervision." TMLR (2024).
[paper]
[code]
[project]
[2023.04]
DINOv3: Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana, Claire Roberts, Andrea Vedaldi, Jamie Tolan, John Brandt, Camille Couprie, Julien Mairal, Hervé Jégou, Patrick Labatut, Piotr Bojanowski.
"DINOv3." ArXiv (2025).
[paper]
[code]
[2025.08]
Rex-Omni: Qing Jiang, Junan Huo, Xingyu Chen, Yuda Xiong, Zhaoyang Zeng, Yihao Chen, Tianhe Ren, Junzhi Yu, Lei Zhang.
"Detect Anything via Next Point Prediction." ArXiv (2025).
[paper]
[project]
[code]
[2025.10]
Mamba-3: Anonymous authors.
"Mamba-3: Improved Sequence Modeling using State Space Principles." ICLR (2026).
[paper]
[2025.11]
Depth Anything 3: Haotong Lin, Sili Chen, Junhao Liew, Donny Y. Chen, Zhenyu Li, Guang Shi, Jiashi Feng, Bingyi Kang.
"Depth Anything 3: Recovering the Visual Space from Any Views." ICLR (2026).
[paper]
[code]
[2025.11]
Vision Banana: Valentin Gabeur, Shangbang Long, Songyou Peng, Paul Voigtlaender, Shuyang Sun, Yanan Bao, Karen Truong, Zhicheng Wang, Wenlei Zhou, Jonathan T. Barron, Kyle Genova, Nithish Kannen, Sherry Ben, Yandong Li, Mandy Guo, Suhas Yogin, Yiming Gu, Huizhong Chen, Oliver Wang, Saining Xie, Howard Zhou, Kaiming He, Thomas Funkhouser, Jean-Baptiste Alayrac, Radu Soricut.
"Image Generators are Generalist Vision Learners." ArXiv (2026).
[paper]
[code]
[2026.04]
RelateAnything: Maëlic Neau.
"RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs." ArXiv (2026).
[paper]
[code]
[models]
[dataset]
[2026.09]
:boom:AgentDSM: Wentao Sun, Zhengsen Xu, Yiping Chen, John S. Zelek, Jonathan Li.
"Agentic Building-Aware Satellite Gaussian Splatting for Auditable Urban DSM Reconstruction." ArXiv (2026).
[paper]
[code]
[2026.09]
:boom:SAM-V: Jiangshan Gong, Yuqun Wu, Qiqian Fu, Yao Xiao, Chuhang Zou, Shenlong Wang, Derek Hoiem.
"SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
:boom:Sepideh Gohari, Goodarz Mehr, Azim Eskandarian.
"Real-World Perception for Autonomous Driving in Adverse Weather: Enhancing Standard Detectors via Foundation-Guided Auto-Annotation." TITS (2026).
[paper]
[2026.09]
:boom:SAFe: Anja Delić, Jurica Runtas, Marin Oršić, Ivan Marković, Ivan Petrović.
"SAFe: Segment-guided Aggregation of Feature Densities for Anomaly-aware Segmentation." ArXiv (2026).
[paper]
[2026.09]
:boom:PhysReflect: Shuheng Ge, Hongwei Ren, Li Zhang, Xiangqian Wu.
"PhysReflect: Geometry and Perception Guided Diffusion for Physically-Plausible Mirror Reflections." ArXiv (2026).
[paper]
[2026.09]
:boom:Dhruv Gamdha, James Afful, Shambhavi Joshi, Ulrike Passe, Adarsh Krishnamurthy, Baskar Ganapathysubramanian.
"Semi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms." ArXiv (2026).
[paper]
[2026.09]
:boom:SRPR-Net: Lufei Liu, Guojie Li, Suncheng Xiang, Fan Zhang.
"SRPR-Net: Semantic and Relational Prompt Refinement for Automated SAM-based Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
:boom:RoboDistill: Ziying Song, Lin Liu, Hongyu Pan, Shaoqing Xu, Lei Yang, Mingzhe Guo, Caiyan Jia.
"Towards robust multimodal 3D object detection via visual foundation models." ArXiv (2026).
[paper]
[2026.09]
:boom:CODY-SAM3: Laura Cif, Zohra Souei, Diane Demailly, Mayte Castro-Jimenez, Juan Dario Ortigoza-Escobar, Muhammad Mushhood Ur Rehman, Morgan Dornadic, Sophie Huby, Gun-Marie Hariz, Cecile Hubsch, Nathalie Dorison, Eduardo M. Moraud, Jocelyne Bloch, Gabriella Horvath, Olivier Oullier, Xavier Vasques.
"Foundation-model-based multi-label phenotyping of combined hyperkinetic movement disorders." ArXiv (2026).
[paper]
[code]
[data]
[2026.09]
:boom:FOM-SAM3: Haolong Meng, Fangbo Qin, Mengchen Bai, Houwu Wang, Cirong Liu, Shan Yu.
"Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning." ICRA (2026).
[paper]
[code]
[2026.09]
:boom:AgenticSwarm: Muhammad Ahsan Mustafa, Yasheerah Yaqoot, Faryal Batool, Roohan Ahmed Khan, Valerii Serpiva, Dzmitry Tsetserukou.
"AgenticSwarm: Semantic Perception and Adaptive Task Allocation for Heterogeneous Multi-UAV Missions." ArXiv (2026).
[paper]
[2026.09]
:boom:P3-SAM: Qian Xu, Hang Xiong, Anpeng Wang, Sam Kwong, Cong Zhang, Runmin Cong.
"P3-SAM: SAM with Perceptual Parallel Prompt for Few-Shot Strip Steel Surface Defect Segmentation." ICME (2026).
[paper]
[2026.09]
:boom:FedSAM-3D: Xinran Wu, Rencheng Zheng, Yuxiang Dai, Hui Zhang, Xueqin Xia, Yu Cheng, Chengyan Wang, He Wang.
"FedSAM-3D: Adapter-Constrained Federated Adaptation for Transferable Medical Segmentation Foundation Models." TBME (2026).
[paper]
[code]
[2026.09]
:boom:Shujun Lv, Bo Fang, Yongfei Wu, Kun Li, Qiankun Li, Junxin Chen.
"From SAM 1 to SAM 3: Benchmarking Zero-Shot Cross-Domain Medical Image Segmentation." Expert Systems (2026).
[paper]
[2026.09]
:boom:PVPSAM: Han, Fangzhou and Li, Xiaoci and Gu, Li and Li, Li and Mi, Ke and Gu, Shenming and Zhang, Hailong.
"PVPSAM: Method and Benchmark for Weakly Supervised Object-level Photovoltaic Panels Extraction in Remote Sensing Imagery." JSTARS(2026).
[paper]
[2026.09]
:boom:PBP-SAM: Liangchao Chen, Guanying Huo, Weifeng Kong, Ziheng Cao, and Jiaying Chen.
"PBP-SAM: polarization-driven boundary prompt SAM for camouflaged object detection." ArXiv (2026).
[paper]
[2026.09]
:boom:GeoFuse-SAM: Pengtao Ren, et al.
"GeoFuse-SAM: A Multimodal Data Fusion Frameworkfor Boundary-Aware Foundation Model Adaptation inMedical Image Segmentation." ArXiv (2026).
[paper]
[2026.09]
:boom:DPSAM2: Wenbo Lei, Long Yu & Shengwei Tian.
"DPSAM2: memory-guided dual-path adaptation of SAM2 for boundary-aware low-contrast segmentation." The Visual Computer (2026).
[paper]
[code]
[2026.09]
:boom:Marcel Hudcovič, et al.
"From Prompt to Plot: Proof-of-Concept forSegmentation of Agricultural Landscapes in AerialImagery Using SAM 3 Agent." IGARSS (2026).
[paper]
[2026.09]
:boom:Sampaio, Filipe A. and Astudillo, Carlos A. and Souza, Alan and Miranda, Daniel and Borin, Edson.
"Improving SAM-Based Seismic Facies Segmentation With Logits Feedback." LGRS (2026).
[paper]
[2026.09]
:boom:PE-MedSAM2: Yuan, Xuejia and Yang, Zongjian and Guo, Yu and Kong, Fanhui and Ma, Jiquan.
"PE-MedSAM2: Parameter-Efficient Adaptation of MedSAM2 for 2D Medical Image Segmentation." TBME (2026).
[paper]
[code]
[2026.09]
:boom:ReliefSAM: Yihang Chen, Xiang Lyu, Rui Xu, Jiao Pan, Fadjar Ibnu Thufail, Brahmantara, Jiaqing Liu, Satoshi Tanaka & Liang Li.
"ReliefSAM: A Geometry-Augmented Multi-prior Adapter for Bas-Relief Segmentation." ECCV (2026).
[paper]
[code]
[2026.09]
:boom:ODG-SAM2-Morph: Yang, Dongxu, Xirui Xu, Shengmao Zhang, Zuli Wu, Tianfei Cheng, Jianglong Que, Siyao Wu, and Fei Wang.
"Morphometric Information for Yangtze Finless Porpoises Using Detection-Guided SAM2 Segmentation with UAV Imagery." Fishes (2026).
[paper]
[2026.09]
:boom:SnakeSAM: Jingwen Li, et al.
"SnakeSAM: A topology-preserving foundation model for medical curvilinear segmentation." Array(2026).
[paper]
[2026.09]
:boom:Chen, Xuan, and Shaolong Chen.
"Dynamic Consistency-Aware Multi-View Learning with SAM3 for 3D Medical Image Segmentation." Sensors (2026).
[paper]
[2026.09]
:boom:ASAM2-UNet: Xie, Caiyun, Linfeng Zhang, Zhaokun Chen, and Junyun Wu.
"ASAM2-UNet: An Attention-Enhanced SAM2 U-Net for Polyp Segmentation." Electronics (2026).
[paper]
[2026.09]
:boom:Shaghayegh Chavoshian, Ali Barzegar Khanghah & Atena Roshan Fekr.
"Transfer Learning on Segment Anything Model for Footwear Outsole Segmentation to Predict Footwear Slip Resistance." Annals of Biomedical Engineering (2026).
[paper]
[2026.09]
:boom:FST-SAM3: Guanhao Wu, Guilian Chen, Huisi Wu, and Jin Qin.
"FST-SAM3: Taming SAM 3 with Frequency-Spatio-Temporal Refinement for Video Polyp Segmentation." ECCV (2026).
[paper]
[code]
[2026.09]
:boom:FWSAM-Net: Shuchi Chen, Shengbing Chen, Qian Chen.
"FWSAM-Net: Wavelet-enhanced SAM2-based framework with frequency-aware adapter for Infrared Small Target Detection." Infrared Physics & Technology (2026).
[paper]
[2026.09]
:boom:SAM3_Remote_Sensing_LoRA: Nermeen Abou Baker.
"Parameter-Efficient Adaptation of SAM3 for Remote Sensing Segmentation Beyond Single-Domain Prompting." ICANN (2026).
[paper]
[code]
[2026.09]
:boom:MTGF-SAM: Zhang, Liangdong and Liu, Xiaohui and Zhang, Junxiao and Shao, Qinglong and Xing, Huaqiao and Zhu, Qing.
"MTGF-SAM: Multi-Level Terrain-Gated Fusion of Segment Anything Model for Landslide Detection in Remote Sensing Imagery." JSTARS (2026).
[paper]
[2026.09]
:boom:YLSAM2: Jinghui Yang and Liang Wang and Shuyin Hu and Bohao Zhang and Huiyuan Pang and Longqin Xu and Meng Cui and Shuangyin Liu.
"YLSAM2: Attention-guided LoRA enhanced underwater multi-scene fish segmentation and counting based on YOLO11 prompting SAM2." Aquacultural Engineering (2026).
[paper]
[2026.09]
:boom:Busra Aslan.
"YOLO–SAM-Guided ROI-Based Deep Learning for Non-Invasive Neonatal Jaundice Detection." BALKAN JOURNAL OF ELECTRICAL & COMPUTER ENGINEERING(2026).
[paper]
[2026.09]
:boom:ES-SAM: Xudong Yang, Xinnan Fan, Peiyu Zhao, Qi Sun, Pengfei Shi.
"ES-SAM: An Enhanced Semantic-SAM for semantic segmentation." PR (2026).
[paper]
[2026.09]
:boom:WOFT-SAM: Jonáš Šerých ⋅ Jiri Matas.
"Segmentation-Guided Homography Estimation for Long-Term Planar Tracking." ECCV (2026).
[paper]
[code]
[2026.09]
Ömer Faruk Deniz, Mustafa Taha Koçyiğit.
"Open-vocabulary 3D object detection with promptable segmentation." ArXiv (2026).
[paper]
[2026.09]
Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan.
"Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings." ArXiv (2026).
[paper]
[2026.09]
HYDRA: Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Damith Ranasinghe.
"Queries Knew More Than We Thought: Uncovering Latent Knowledge in Segmentation Models." ArXiv (2026).
[paper]
[2026.09]
Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins.
"Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions." ArXiv (2026).
[paper]
[code]
[2026.09]
SetPlanner: Dawei Yan, Yuezhe Yang, Menglan Ruan, Chunfeng Yang, Yudong Zhang.
"SetPlanner: A Lightweight Plug-in Point-Set Planner for Frozen SAM." ArXiv (2026).
[paper]
[code]
[2026.09]
PSMP-CLIP: Xuezhi Xiang, Guanghao Wu, Heqi Xiang, Jiayao Liu, Xiaoheng Li, Yiming Chen, Shanjun Zhang.
"PSMP-CLIP: Patch-Prompt SAM and Multi-Semantic Prompting for CLIP-Based Zero-Shot Anomaly Detection." ArXiv (2026).
[paper]
[2026.09]
VPRef: Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang.
"VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
ReliSAM: Shen Jiang and Xiaoyan Kui and Zhipeng Hu and Yangyang Shi and Ziwei Zou and Zexin Ji and Zeru Hai and Qinsong Li and Zuheng Ming and Haodong Xu and Beiji Zou.
"Reliability-guided dual-view learning with SAM distillation for semi-supervised medical image segmentation." Expert Systems with Applications (2026).
[paper]
[2026.09]
ViCo-SAM3: Qiangqiang Zhou, Wenjun Tang, Yong Chen, Dandan Zhu, Jiawei Xu.
"ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation." ArXiv (2026).
[paper]
[2026.09]
PEFT-SAM-Liver-CT: Ramtin Mojtahedi, Mohammad Hamghalam, Jacob J. Peoples, Richard K. G. Do, Amber L. Simpson.
"Parameter-Efficient Fine-Tuning of Foundation Models for Liver Tumor Segmentation in CT." SPIE Medical Imaging(2026).
[paper]
[code]
[2026.09]
PRR: Weihong Qi, Chen Ling.
"Perceive, Refine, Reason: A Calibrated Pipeline for Measuring Indicators in Strategic Visual Communication on Social Media." ArXiv (2026).
[paper]
[2026.09]
WorldMem: Aditi Tiwari, Akshit Bhalla, Darshan Prasad, Heng Ji.
"Does Video Memory Use What It Retrieves? A Causal Audit of Memory Specificity." ArXiv (2026).
[paper]
[2026.09]
ResoSeg: Chunkai Li, Junhao Yin, Ke Li, Jingde Chen.
"ResoSeg: Resonance Tagger using Transformer and Segment Model." ArXiv (2026).
[paper]
[code]
[2026.09]
MIE-SAM: Ze Li, Ying Ying Zhang, Shuai Zhang, Zhi Peng Wang.
"Multi-modal interaction enhanced segment anything model (MIE-SAM) for RGB-T salient object detection." Neural Networks (2026).
[paper]
[code]
[2026.09]
Silas Kwabla Gah, Ebenezer Owusu.
"Beyond Argmax: A Mechanistic Study of Semantic Retention in Frozen Foundation-Model Composition for Generalized Few-Shot 3D Segmentation." ArXiv (2026).
[paper]
[2026.09]
SAMV-DUSt3R: Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao.
"SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views." ArXiv (2026).
[paper]
[2026.09]
DINO-Med: Boya Wang, Ruizhe Li, Chao Chen, Xin Chen.
"DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging." ArXiv (2026).
[paper]
[2026.09]
BruNet: Qiming Wang, Richard J. Motley, Ebube E. Obi, Xianfang Sun, Paul L. Rosin.
"BruNet: A Cross-Domain Transfer Framework for Bruise Segmentation." ArXiv (2026).
[paper]
[2026.09]
DiSECT: Ramtin Mojtahedi, Mohammad Hamghalam, Jacob J. Peoples, Natalie Gangai, Mithat Gonen, Yun Shin Chun, HyunSeon Christine Kang, Richard K. G. Do, Amber L. Simpson.
"Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed Tomography." ArXiv (2026).
[paper]
[2026.09]
LeCor: Yi Luo, Yike Guo, Wenxuan Li, Zongwei Zhou, Rui Zhang, Kai Ding.
"LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour Segmentation." ArXiv (2026).
[paper]
[2026.09]
Guided SAM 3D: Jerred Chen, Simon Weber, Ronald Clark.
"Guiding Image-to-3D Generation with Test-Time Partial Observations." ArXiv (2026).
[paper]
[2026.09]
PBP-SAM: Liangchao Chen, Guanying Huo, Weifeng Kong, Ziheng Cao, and Jiaying Chen.
"PBP-SAM: polarization-driven boundary prompt SAM for camouflaged object detection." Applied Optics (2026).
[paper]
[2026.09]
LSVOS: Chang Liu, Henghui Ding, Lingyi Hong, Ning Xu, Linjie Yang, Yuchen Fan, Canyang Wu, Jinrong Zhang, Xusheng He, Ce Bian, Xianjing Han, Jianlong Wu, Mingqi Gao, Sijie Li, Jungong Han, JeongRae Kim, Chaehyun Kim, Changwon Lim, Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim, Pranjal Aggarwal, Sean Welleck, Yiwen Ren, Jianing Liu, Yingxin Wang, Kexin Zhang, Licheng Jiao, Lingling Li, Xu Liu, Jinxing Zhou, Suiyi Zhao, Yanghao Zhou, Ruohao Guo, Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu.
"Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation." ECCVW (2026).
[paper]
[code]
[2026.09]
MorphoOrgaAgent: Hanyi Zhang, Maximilian Hoermann, Lion J. Gleiter, Yiling Xu, Bettina Katalin Budai, Hans-Ulrich Kauczor, Carsten Marr, Tingying Peng.
"MorphoOrgaAgent: A Foundation-Model-Based Multi-Agent System for Autonomous Organoid Analysis." MICCAI workshop (2026).
[paper]
[code]
[2026.09]
SAM3-O2D2: Lucas Görnhardt, Timo Bartels, Tim Fingscheidt.
"SAM3-O2D2: Zero-Shot Object Out-of-Distribution Detection by Object Class Prompting of the SAM3-Image Model." ArXiv (2026).
[paper]
[2026.09]
MR-RS-SDFR: Quanxin Zheng, Shuai Zhao.
"MRI-Guided Reslice-Refined Cross-Slice SDF Reconstruction of the Left Ventricle from Cardiac MRI with Sparse Axial Supervision." ArXiv (2026).
[paper]
[2026.09]
Andreas Gilson, Laura Hennig, Peter Pietrzyk.
"Zero-Shot 3D Plant Organ Segmentation with SAM3 and Semantic NeRFs." ECCVW (2026).
[paper]
[2026.09]
CoRe-SAM3: Shipeng Liu, Liang Zhao, Dengfeng Chen.
"CoRe-SAM3: Conditional Semantic--Visual Reconciliation for SAM3 Crack Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
DriveZero: Hao He, Chengcheng Hu, Zirun Su, Heng Zhang, Haisong Liu, Jinke Li, Haochen Tian, Zhenwei Shen, Hongyang Li, Zhichao Li, Yunchen Yang, Bochao Huang, Siyu Zhang, Kuangye Chen, Xiongjie Zhang, Wentao Dai, Hengchen Dai, Siyuan Liu, Zehao Huang, Naiyan Wang.
"DriveZero: End-to-End Driving Beyond Human Demonstrations." ArXiv (2026).
[paper]
[code]
[2026.09]
CrACK: Feifei Liu, Jintao Cheng, Chi Man Vong, Xiaoyu Tang.
"CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models." ArXiv (2026).
[paper]
[2026.09]
SAM-Radar: Jue Wang, Xuan Wang, Hao Zhou, Ruixiang Zhou, Yixuan Zhou, Tianshuo Yuan, Jieming Ma, Jie Zhang, Fei Luo.
"Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking." ArXiv (2026).
[paper]
[2026.09]
SeGDeP: Linnan Zhao, Xu Liu, Lingling Li, Licheng Jiao, Fang Liu, Wenping Ma.
"SeGDeP: Semantic- and Geometric-Aware Decoupled Prompts for Reasoning Segmentation." ArXiv (2026).
[paper]
[2026.09]
Diffuse2Seg: Christoph Hümmer, Joachim Sicking, Fabian Hüger, Hanno Gottschalk.
"Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision." ArXiv (2026).
[paper]
[2026.09]
Hesham, S.A.S., Liu, Y., Sun, G. et al.
"Evaluating SAM2 for Video Semantic Segmentation." Mach. Intell. Res.(2026).
[paper]
[2026.09]
Carballo Pérez, Áurea Genoveva, Pablo Magariños-Docampo, Pedro Orgeira-Crespo, and Fernando Aguado-Agelet.
"Weakly Supervised Segmentation of Macroalgae Through Gradient Analysis in Convolutional Neural Networks and Segment Anything Model." Applied Sciences(2026).
[paper]
[2026.09]
STCM: Dong, Wuzhou, Yulin Chen, Zhipan Wang, and Qingling Zhang.
"Semantic–Texture Complementation and Prediction-Guided SAM Fusion for Plastic Mulch Segmentation in GF-7 Imagery." Remote Sensing (2026).
[paper]
[2026.09]
SarSAM: Fatih Fehmi Ş İMŞEK, Melih ALTAY, Saygin ABDIKAN.
"SarSAM: An Integrated Framework for Agricultural Field Boundary Delineation and Mapping Using High-Resolution SAR (PAZ) Imagery and the Segment Anything Model." Advances in Space Research (2026).
[paper]
[2026.09]
SMART: Gao, Fang and Shi, Lei and Jin, Yan and Zheng, Hanbo and Huang, Qingbao and Yu, Jun.
"Say, Move, and Remember: Enhancing SAM2 for Referring Video Segmentation via Motion Modeling and Key Action-Aware Memory." TMM (2026).
[paper]
[code]
[2026.09]
TeF-SAM: Jinxin Liang, Xiaoming Liu, Zhiyuan Zang, Xin Wang, Yicheng Qi, Xiang Li.
"TeF-SAM: Prototype memory for medical lesion segmentation with text-free inference." Computerized Medical Imaging and Graphics (2026).
[paper]
[code]
[2026.09]
FastSAM-GAL: Zhu, Zhifu and Yuan, Xiping and Gan, Shu and Luo, Weidong and Chen, Cheng and Li, Xuan.
"FastSAM Guided Adversarial Learning for Unsupervised Multimodal Remote Sensing Change Detection." TGRS (2026).
[paper]
[2026.09]
A2SAM: Yong Chen and Renyi Chen and Qingbo Kang and Rui Wang and He Lyu and Zekun Jiang and Hongqiu Wang and Huan Song and Kang Li.
"Informativeness-driven active adaptation of SAM: Structural prompts and contrastive parameter selection for medical tubular segmentation." MIA (2026).
[paper]
[code]
[2026.09]
Yuxiang Huang, et al.
"Automated Zero-Shot Video Segmentation of Indoor Walls via SAM2 with Pretrained Semantic Prompts." Construction Research Congress(2026).
[paper]
[2026.09]
SAM3TGNet: Zhang, Jiayin, Nan Mo, Gege Ma, and Bangyan Tang.
"SAM3TGNet: A SAM3 Feature Encoding and Global Context Spatiotemporal Attention-Enhanced Change Detection Method for Optical Remote Sensing Images." Sensors (2026).
[paper]
[2026.09]
SAM-AUT: Amir-M. Naddaf-Sh, Vinay S Baburao, Hassan Zargarzadeh.
"SAM for Weld Defect Detection in Ultrasonic B-Scans." ArXiv (2026).
[paper]
[code]
[2026.09]
SWIFT: Lucie Bracq, Romain Guiet, Sandra Offner, Béatrice Kunz, Gisou van der Goot & Nathalie Brandenberg.
"A Single-organoid Workflow for quantitative Imaging classiFication and Tracking (SWIFT)." Communications Biology (2026).
[paper]
[2026.09]
SarSAM: Fatih Fehmi ŞİMŞEK and Melih ALTAY and Saygin ABDIKAN.
"SarSAM: An Integrated Framework for Agricultural Field Boundary Delineation and Mapping Using High-Resolution SAR (PAZ) Imagery and the Segment Anything Model." Advances in Space Research (2026).
[paper]
[2026.09]
ReliSAM: Shen Jiang and Xiaoyan Kui and Zhipeng Hu and Yangyang Shi and Ziwei Zou and Zexin Ji and Zeru Hai and Qinsong Li and Zuheng Ming and Haodong Xu and Beiji Zou.
"Reliability-guided dual-view learning with SAM distillation for semi-supervised medical image segmentation." Expert Systems with Applications (2026).
[paper]
[2026.09]
SAMLoRA: Wang, Xuewu, Wenlu Zhao, Cai Wang, Xu Chen, Yan Xu, Zuoman Zhang, Xirui Qiao, Bing Cao, Huifang Wang, and Hao Liu.
"High-Resolution Mapping and Spatial Pattern Analysis of Areca Palm Plantations in Sanya, China, Using SAMLoRA." Forests (2026).
[paper]
[2026.09]
Kai Zhao and Chenchen Kang and Suzy Rogiers and Oula Ghannoum and Yi Guo.
"Adapting SAM3 for 3D fruit counting with cross-view contrastive learning and Hough voting." Computers and Electronics in Agriculture (2026).
[paper]
[2026.09]
SAM-FSYOLO: Zhongyi Wang, Luohua Zhang, Changning Wei, Richu Jin, Dongjun Zhang, Tijun Bie, and Yonghui Yang.
"SAM-FSYOLO: An Integrated Framework for Miniature Covert Imaging Device Detection in Hotel Environments." ArXiv (2026).
[paper]
[2026.09]
CLON: Seojin Ji, Yoojin Kwon, Hyung-Sin Kim.
"CLON: Cue-Calibrated Linguistic Object Onboarding for Zero-Shot 6D Pose Front-Ends." ArXiv (2026).
[paper]
[2026.09]
WireSeg-32K: Zilin Dai, Lehong Wang, Yi Yang, Xiang Fei.
"WireSeg-32K: A Physics-Grounded Synthetic Dataset for Wire Instance Segmentation." CVPRW (2026).
[paper]
[code]
[2026.09]
Hailong Ning, Hao Wang, Yimeng Wang, Tao Lei, Renwei Dian, Asoke K. Nandi.
"Progressive Pseudo-Label Optimization for Point-Supervised Change Detection." ArXiv (2026).
[paper]
[2026.09]
FreNet: Yinan Liu, Jiankang Hong, Zhen Gao, Ye Lu.
"Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation." ArXiv (2026).
[paper]
[2026.09]
ENEAS: Javier del Pino, Salvador Rodríguez, Alejandro Garabito, Javier Álvarez, Chema Garabito.
"ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
SAM3-LoRA: P. Malaisree, S. Youwai, S. Janrungautai, D. Amorndechaphon, P. Rojanavasu, W. Songkitti.
"SAM3-LoRA: Parameter-Efficient Adaptation of a Concept-Promptable Foundation Model for Multi-Class Structural Defect Segmentation." ArXiv (2026).
[paper]
[code]
[2026.09]
Udo Schlegel, Shubhangi, Gabriel Dax, Sai Rahul Kaminwar, Florian Karl, Thomas Seidl.
"Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting." ECML-PKDD(2026).
[paper]
[2026.09]
NeuSOGA: Qingde Li, Qingqi Hong, Jie Tian.
"Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations." ArXiv (2026).
[paper]
[code]
[2026.09]
AcrossVAM1.0: Yafei Zhang, Nan Wu.
"AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction." ArXiv (2026).
[paper]
[2026.08]
OPUS: Xiaoyan Wei, Zhimin Yao, Ruilin Yang, Wei Zhang, Yong Dai, Yi Zhang, Wei Ge.
"OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection." ArXiv (2026).
[paper]
[2026.08]
Marin Maletic, Goran Vasiljevic.
"Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring." ICASSP (2026).
[paper]
[2026.08]
MemMTL: Yangyang Xu, Haobo Yuan, Yuzhu Wang, Duo Su, Xi Ye, Yibo Yang, Jun Zhu.
"Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction." ArXiv (2026).
[paper]
[2026.08]
MariSat: Amir Abbes, Ines Harrabi, Lucas Justin Yirepoa Kinda, Rim Trabelsi, Adnane Cabani, Fatma Abdelkefi.
"MariSat: A Maritime Dataset for Instance Segmentation of Objects in Satellite and Aerial Images." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM-3D-Body: R. James Cotton, J. D. Peiffer, Lucinda Williamson, John Leske, Georgios Pavlakos.
"Biomechanical 3D Body: Self-Supervised Distillation of Biomechanical Pose from a 3D Body Foundation Model." ECCVW (2026).
[paper]
[2026.08]
FoundYou: Gabriele Trivigno, Marcos Alfaro, Claudia Cuttano, Gabriele Berton, Luis Payá, Carlo Masone.
"FoundYou: A Unified Model for Personalized Segmentation and Retrieval." ECCV (2026).
[paper]
[code]
[2026.08]
SAM-STIR: Hu, Xiao and Zhou, Yun and Lv, Jian and Lai, Wenjie and Jiang, Yadong.
"SAM-STIR: Detection-Guided Memory Update for SAM 3-Based Infrared Small and Tiny Object Tracking." TGRS (2026).
[paper]
[2026.08]
RASP-SAM: Peng Zhang and Chen Liu and Zeyu Liu and Guanglei Zhang and Hongming Shan and Wenjian Wang.
"Adapting foundation models to weakly-supervised few-shot medical image segmentation via retrieval-augmented semantic prompting." Pattern Recognition (2026).
[paper]
[2026.08]
AffectOmni: Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua.
"AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes." ArXiv (2026).
[paper]
[code]
[2026.08]
SD-SAM 3: Ali Lesani, Chul Min Yeum, Su-Min Kang.
"Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery." ArXiv (2026).
[paper]
[2026.08]
T2S: Kumju Jo, Heesun Jung, Sungyong Baik.
"Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator." KBS (2026).
[paper]
[2026.08]
FAN-LoRA: Ziquan Liu, Zhewei Zhu, Xuyang Shi.
"FAN-LoRA: A Fourier-Adaptive Nonlinear Low-Rank Adaptor for Medical Foundation Model Domain Adaptation." ArXiv (2026).
[paper]
[2026.08]
LiDAR-SAM2: Jihun Kim, Hyun-Kurl Jang, Hyemin Yang, Jinnyeong Yang, Hyeokjun Kweon, Kuk-Jin Yoon.
"Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models." ECCV Workshop(2026).
[paper]
[2026.08]
HPMA: Xinning Yao, Jingjing Wang, Jinghua Yue, Xiaoyan Luo, Fugen Zhou, Bo Liu.
"Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation." ArXiv (2026).
[paper]
[2026.08]
ReGround-Surg: Jiaxin Wen, Ming Yin, Lu Liu, Zeyu Fu.
"ReGround-Surg: Reliability-Guided Anchor Grounding for Referring Surgical Video Segmentation." PRCV (2026).
[paper]
[code]
[2026.08]
OptiSight: Alperen Avan, Jordi Sanchez-Riera.
"OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation." ArXiv (2026).
[paper]
[code]
[2026.08]
Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu.
"MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
SAM3Dual: JeongRae Kim, Chaehyun Kim, Changwon Lim.
"SAM3Dual: A 3rd Place Solution to the MOSEv2 Track, 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
Mingqi Gao, Sijie Li, Jungong Han.
"Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
RAVP: Salvatore Calcagno, Marco Finocchiaro, Giovanni Bellitto, Daniela Giordano, Concetto Spampinato, Federica Proietto Salanitri.
"Retrieval-Augmented Visual Prompting: Guiding Foundation Models in Two-Photon Imaging." ArXiv (2026).
[paper]
[2026.08]
SPARK-SAM: Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu.
"SPARK-SAM: Self-Prompt Adaptation with Response Knowledge for SAM in Infrared Small Target Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
GAP-SAM: Haozhen Yan, Siyuan Shan, Zijian Yu, Youqi Wang, Yan Hong, Jun Lan, Jianfu Zhang.
"GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization." ArXiv (2026).
[paper]
[2026.08]
Manani, Hiren.
"From Detection to Segmentation: A Foundation Model Approach to Organoid Brightfield Image Analysis Using SAM 3." ArXiv (2026).
[paper]
[2026.08]
DPD-SAM: Ma, Yize and Liang, Zhenyang and Yin, Feiyu and Song, Pengfei and Song, Junyi and Yang, Shengxiang and Wu, Guoqing and Ren, Yan and Yu, Jinhua.
"Latent Domain-Specific Prompt-Driven SAM via Conditional Diffusion Refinement for Postoperative Glioma Segmentation." TII (2026).
[paper]
[2026.08]
Hilla Fred and Mogens Agerbo Krogh and Britt {Bang Jensen} and Laura Ruotsalainen and Jouni Vielma and Matti Pastell.
"Automatic visual detection of fish in Recirculated Aquaculture Systems using the Segment Anything Model." Aquaculture (2026).
[paper]
[2026.08]
SNFusion: Yang Fang and Jingjing Chen and Dahang Wan and Xianli Lang and Shuangbao Shu and Rongsheng Lu and Qianqian Wu.
"SNFusion: A SAMv2-guided boundary-aware fusion framework for visible-infrared maritime vessel detection." Ocean Engineering (2026).
[paper]
[code]
[2026.08]
ACRIS-SAM2: Jiaxiang Luo, Weiwen Chen.
"ACRIS-SAM2: Attribute-driven cross-modal interaction and dual semantic prompting for few-shot segmentation." Neurocomputing (2026).
[paper]
[2026.08]
SAM2 R-CNN: Mehdi Gharbage, Céline Teulière, Pierre Bouges & Thierry Chateau .
"SAM2 R-CNN: Transferring SAM 2 Knowledge for Data Efficient Instance Segmentation." ICPR (2026).
[paper]
[code]
[2026.08]
Yang, Chenzheng, Shenhua Yang, Pu Wang, Weijun Wang, and Zeyang Huang.
"Constrained Boundary Enhancement for SAM 2-Based Ship Segmentation in UAV Berthing and Unberthing Videos." Applied Sciences (2026).
[paper]
[2026.08]
Swapnil Biswas.
"Enhancing Skin Lesion Classification in Teledermatology via SAM-Based Segmentation and Wavelet-Based Convolutional Autoencoder." ArXiv (2026).
[paper]
[2026.08]
FlexiCrackNet: Xiaoyan Jiang, et al.
"FlexiCrackNet: A Flexible Pipeline for Lightweight Crack Segmentation with Distilled Features from SAM." ArXiv (2026).
[paper]
[code]
[2026.08]
EchoFlow-SAM: Wanting Li, et al.
"EchoFlow-SAM: Motion-guided semi-supervised segmentation for echocardiographic videos." Biomedical Signal Processing and Control (2026).
[paper]
[2026.08]
Geng, Chao, Yajie Wang, Quanming Li, Zhentao Li, Xianfeng Shi, Botao Fu, Wei Li, Cheng Chen, Hong Zhang, Yukai Wang, and et al.
"Automated Detection and Segmentation of Cracks in Urban Underground Structures Based on YOLOv8-SAM2." Buildings (2026).
[paper]
[2026.08]
SAM-CLIP-Thermal: Yiyuan Lin and Chenjiao Tan and Changying Li and Yu Jiang.
"SAM-CLIP-Thermal: Leveraging large multimodal models for reliable and scalable annotation in thermal image segmentation for field plant phenotyping." Plant Phenomics (2026).
[paper]
[code]
[2026.08]
HBF-BCER: Ping, Shengyang, Zhijie Lin, Liliang Lin, Lei Zhao, Lisha Ye, Bangguo Wang, and Tao Wang.
"A Prompt-Preserving SAM ViT-B Adaptation Framework with Historical Branch Fusion and Soft Convolutional Expert Weighting for Medical Image Segmentation." Bioengineering (2026).
[paper]
[2026.08]
Puspitasari, Fachrina Dewi and Zhang, Chaoning and Mandal, Avilasha and Zheng, Sheng and Qin, Caiyan and Kim, Tae-Ho and Lee, Jewon and Wang, Guoqing and Yang, Yang and Shen, Heng Tao.
"Accelerating SAM2 with Efficient Memory Attention Module via Spatiotemporal Token Pruning." TPAMI (2026).
[paper]
[2026.08]
Token-Adaptive LoRA: Xin Chen, Jun Yan, Zhiyu Yan, Jianwen Deng, Jiaqi Wu, Yonghong Gong, Xiaohua Jiang.
"Token-Adaptive LoRA: Enhancing Segment Anything for Remote Sensing Imagery through Parameter-Efficient Fine- Tuning." ArXiv (2026).
[paper]
[2026.08]
Zero-Click-SAM2: Pasierb, Daniel and Wijata, Agata M. and Nalepa, Jakub.
"Zero-Click Brain Tumor Segmentation Using Segment Anything Model 2." ICIP (2026).
[paper]
[code]
[2026.08]
Bui-Tran, Quang-Khai and Nguyen, Thanh-Huy and Le, Bac and Xu, Min.
"Adapting SAM Without Labels: Uncertainty-Aware Source-Free Medical Image Segmentation." ICIP (2026).
[paper]
[2026.08]
SAM2TC: Cocco, Marco and Dunnhofer, Matteo and Micheloni, Christian.
"Representation Compensation of SAM2 for Segmenting Objects under Transformation in Videos." ICIP (2026).
[paper]
[2026.08]
FE-SAM: Gao, Feng and Pan, Zizhe and Wang, Haoting and Hua, Ruzhuang and Cao, Jingchao and Dong, Junyu and Du, Qian.
"Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation." TGRS (2026).
[paper]
[code]
[2026.08]
Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot.
"A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3." ArXiv (2026).
[paper]
[2026.08]
SAM2Dual: JeongRae Kim and Changwon Lim.
"SAM2Dual: Training-Free, Dual Memory for Long-Term Video Object Segmentation." TIP (2026).
[paper]
[2026.08]
SAM2-DPT: Steven Landgraf, Joceline Hinz, Markus Ulrich.
"A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation." ISPRS (2026).
[paper]
[2026.08]
EpigraphNet: Utsav Poudel, Rasik Bhattarai, Siddhartha Pathak, Raghavendra Ramacharna, Gaurav Jaswal.
"Zero-Shot SAM2 Segmentation and Vision Transformer-Based Recognition of Elamite Cuneiform Symbols from Degraded Tablet Images." ArXiv (2026).
[paper]
[code]
[2026.08]
Ce Bian, Xusheng He, Jinrong Zhang, Canyang Wu, Xianjing Han, Jianlong Wu.
"Key-Frame Reasoning with SAM3: Third Place Solution for the MeViS-Text Track of the 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
Cesar Borja, Breck A. McCollum, Jarret E. Byrnes, Kenneth Sebens, Ana C. Murillo.
"Leveraging existing sparse point annotations for benthic imagery dense segmentation." ArXiv (2026).
[paper]
[2026.08]
SSSAM: Ruichao Hou, Boyue Xu, Tongwei Ren, Dongming Zhou, Gangshan Wu, Jinde Cao.
"S3AM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM-RNO: Wang, F., Liu, Q. & Wu, G.
"Road negative obstacle segmentation via a dual-branch segment anything model with RGB-depth cross-prompting." J Supercomput (2026).
[paper]
[2026.08]
SAM-Med2D-GeoCrop: Wang, Tianqi and Li, Jianuo and Yang, Chenhao and Xu, Jinyi and Zhou, Mian and Dang, Kang and Zhang, Linxue.
"ROI-Focused Geometry-Aware Adaptation for Accurate Small-Structure Segmentation in Medical SAM." ICIP (2026).
[paper]
[2026.08]
RISE: Yanbo Jiang, Haotian Zheng, Jiahao Wang, Hanxiao Ren, Yitao Xu, Yining Xing, Zehong Ke, Hao Cheng, Yiqian Tu, Jinhao Li, Zhiyuan Xuan, Fang Zhang, Jianqiang Wang.
"RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning." ArXiv (2026).
[paper]
[2026.08]
DreamX-Phi 1.0: DreamX Team, Rui Chen, Xiangxiang Chu, Geng Li, Jifan Li, Qingfeng Shi, Datao Tang, Jing Tang, Jun Wang, Pengfei Zhang.
"DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation." ArXiv (2026).
[paper]
[code]
[2026.08]
VOS-Agent: Canyang Wu, Jinrong Zhang, Xusheng He, Ce Bian, Xianjing Han, Jianlong Wu.
"VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)." ECCV Workshop(2026).
[paper]
[2026.08]
TCSR-Monito: Hieu D. Pham, Dang P. M. Cao, Thanh Trung Huynh.
"Beyond Uncertainty: Generalizable Failure Monitoring for Surgical Segmentation under Acquisition Degradation." MICCAI Workshop(2026).
[paper]
[code]
[2026.08]
SUGFW+: Xiaochuan Ma, Ning Zhu, Jia Fu, Lanfeng Zhong, Hanyu Jiang, Bin Song, Kang Li, Guotai Wang.
"SUGFW+: An Uncertainty-guided Feature Weighting Framework for Cold Start Active Adaptation of SAM in Medical Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
FE-SAM: Feng Gao, Zizhe Pan, Haoting Wang, Ruzhuang Hua, Jingchao Cao, Junyu Dong, Qian Du.
"Frequency and Edge-Guided Segment Anything Model for Remote Sensing Image Semantic Segmentation." IEEE TGRS (2026).
[paper]
[code]
[2026.08]
Mohammadreza Narimani, Shreyan Mitra, Parastoo Farajpoor.
"From crown candidates to neighborhood screening: integrating optical GeoAI and spatial modeling for urban-canopy assessment in Davis, California." ArXiv (2026).
[paper]
[code]
[dataset]
[2026.08]
Yiwen Ren, Jianing Liu, Yingxin Wang, Kexin Zhang, Licheng Jiao, Lingling Li, Xu Liu.
"Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge." ArXiv (2026).
[paper]
[2026.08]
RoboSeg: Zhaochen Lan, Mengxiang Lin.
"RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera." ArXiv (2026).
[paper]
[2026.08]
Roni Blushtein-Livnon, Tal Svoray, Osher Rafaeli, Michael Dorman, Itay Fischhendler, Havazelet Yahel, Emir Galilee.
"Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing Imagery." ArXiv (2026).
[paper]
[2026.08]
Seed2GS: Zongjian Ding, Yudong Gao, Jiale Liu, Xinglin Yu, Junxing Ren, Dong Wei, Yajing Chen, Shan Huang, Mingjun Cheng, Min Li.
"Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes via a Single Reference-View Grounding." ArXiv (2026).
[paper]
[2026.08]
Mike Szklarzewski, CJ George, Gavin Smithson, Christopher Stokes, Dakota Fulp, William M. Jones, Benjamin Wynn, Alexander Ur, Agit Yesiloz, Clint Kallenbach, Mark Swartz, Nathan DeBardeleben, Sharmistha Chakrabarti.
"From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection." ICMLA (2026).
[paper]
[2026.08]
BAP-MOS: Satvik Praveen, Shengji Jin, Ahmed Lamidi, Xin Qian, Yi Sheng.
"BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
Shah Imran Ahsan Chowdhury, Kazi Jihadur Rashid, Rajsree Das Tuli, Rahul Saha, Bulbul Ahammad.
"GeoAI-based post-segmentation quality validation of building footprints via spatial feature engineering." ArXiv (2026).
[paper]
[2026.08]
LEGO: Yuning Peng, Haiping Wang, Yuan Liu, Yipeng Lu, Zhen Dong, Bisheng Yang.
"LEGO: Leveled Language Gaussian Splatting." ECCV (2026).
[paper]
[code]
[2026.08]
SSUPER: Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim.
"Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026." ArXiv (2026).
[paper]
[2026.08]
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang.
"Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation." Neurocomputing (2026).
[paper]
[2026.08]
AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel.
"Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations." ArXiv (2026).
[paper]
[2026.08]
VOLA: Yuchen Zhang, Yuan Gao, Sebastian Schmidt, Johannes Betz.
"VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM3tool: Nakul Poudel, Richard Simon, Cristian A. Linte.
"Toward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3)." MIUA (2026).
[paper]
[2026.08]
KD-SAM: Yang Fang and Bingbing Jiang and Haopeng Huo and Uswah Khairuddin and Yeon Lee and Yong Qin.
"KD-SAM: Keep-Awake and Detail-Enhanced Segment Anything Model for Multi-Modality Multi-Organ Medical Image Segmentation." Expert Systems with Applications (2026).
[paper]
[code]
[2026.08]
S3-Diff: Jiaming Liang, QiHui Han, Guangye Ou, Jiawen Liu, Haolin Chen, Xi Zhong, Jiazhou Chen, Xiaoqi Sheng, Hongmin Cai.
"S3-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images." ArXiv (2026).
[paper]
[2026.08]
GeoDistill-Refine: Yonglong Zhang, Zongwu Xie, Yang Liu.
"GeoDistill-Refine: Silhouette-First Geometry Distillation for Annotation-Free Spacecraft Segmentation." ArXiv (2026).
[paper]
[2026.08]
Hao Wang, Yuxuan Zhang, Wei Yang.
"Universal Concept Disruption for SAM3 Image Segmentation." ArXiv (2026).
[paper]
[2026.08]
EgoAfford: Xinyuan Guan, Feifan Chen, Xinyu Zhan, Fu-Cheng Zhang, Cewu Lu, Lixin Yang.
"EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
Jonathan Klingspon, Scott McAvoy, Maurizio Seracini, Falko Kuester.
"Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital Twins." ArXiv (2026).
[paper]
[2026.08]
CROSS: Tingzhang Luo, Ruizhong Liu, Yichao Liu, Cheng Fan, Yu Liu, Jianyuan Guo.
"CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation." ECCV (2026).
[paper]
[code]
[2026.08]
FS-CPL: Rahul Venkataramani, Rachana Sathish.
"Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding." ArXiv (2026).
[paper]
[2026.08]
ACTrack: Wenrui Cai, Yuzhe Li, Qingjie Liu, Yunhong Wang.
"Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking." ArXiv (2026).
[paper]
[2026.08]
GaussianSelector: Baihan Yang, Tiexin Li, Yuheng Liu, Xin Lin, Xinke Li, Xiaohui Xie, Truong Nguyen.
"GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization." ArXiv (2026).
[paper]
[2026.08]
EOVSAM: Haomin Peng, Yongkang Li, Zhaoxiang Liu, Xiaojie Jin, Shiguo Lian, Yunchao Wei, Xinggang Wang.
"EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass." ArXiv (2026).
[paper]
[code]
[2026.08]
PhenoStitch: Xuechen Li.
"PhenoStitch: Training-Free Panoptic Crop Mapping from Satellite Image Time Series." ArXiv (2026).
[paper]
[2026.08]
ELFSS-AR: Xueting Bai, Huan Ni.
"Training-Free Entity-Level Few-Shot Segmentation of Remote Sensing Images with Advection Refinement." ArXiv (2026).
[paper]
[code]
[2026.08]
UltraSAM3: Bo Xu, Quanhao Zhu, Rui Lin, Boling Zhu, Chenyuan Wang, Hongfei Lin, Feng Xia, Chenhua Ji.
"UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.08]
SAM+D: Yu Song, Hao Sun, Shiyu Teng, Ikuko Nishikawa, Yen-wei Chen.
"SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting." ECCV (2026).
[paper]
[code]
[2026.08]
TMMSAM2: Fu, Xiyou and Zhang, Ting and Zhang, Xiaoyu and Lin, Mingying and Lv, Zijun and He, Wangquan and Ren, Qi and Xu, Meng and Jia, Sen.
"TMMSAM2: Tracker-Aided Multitemporal Memory SAM2 for Hyperspectral Object Tracking." TNNLS (2026).
[paper]
[2026.08]
SAM3D-VLA: Zonghe Liu, et al.
"SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models." ArXiv (2026).
[paper]
[2026.07]
Jinghong Liu, Yuchuan Deng, Fanping Liu, Meng Huang, Xirong Li.
"Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.07]
ZMIS-SAM: Dekun Yuan, Zhongwei Li, Zheng Qiao, Jie Zhang.
"ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
Nevio Dubbini, Lisa Yeomans, Marco Pavia, Ramazan Parmaksiz, Ayse Atas Hooglugt, Gabriele Gattiglia, Beatrice Demarchi.
"Multimodal fusion of visual and morphometric features for avian bone classification." ArXiv (2026).
[paper]
[2026.07]
RDVSv2: Tianyu Li, Jiahao He, Keren Fu, Qijun Zhao.
"RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection." ACMMM (2026).
[paper]
[code]
[2026.07]
Sanjay Subramanian, Junwei Yu, Zirui Wang, Rohil Malpani, Maggie Chung, Adam Yala, Dan Klein, Trevor Darrell.
"Open-Ended CT Volume Segmentation with Weak Supervision from Language." ArXiv (2026).
[paper]
[2026.07]
SADe: Hang Xing, Guangjun Liu, Yan Xia, Xueming Ding.
"SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations." ArXiv (2026).
[paper]
[2026.07]
ConFusion: Guo Yurong, He Yufei, Li Yonghao, Chang Dongliang, Zhang Ke, Ma Zhanyu.
"ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion." ArXiv (2026).
[paper]
[code]
[2026.07]
EditCLEVR: Anuraag Gadehothur Karnam, Tarunesh Sathish.
"EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations." ArXiv (2026).
[paper]
[code]
[2026.07]
SurgSAM3: Changjing Liu, Yiming Huang, Beilei Cui, Liangjing Shao, Long Bai, Yanheng Li, Haoxuan Che, Hongliang Ren.
"Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation." MICCAI Workshop (2026).
[paper]
[code]
[2026.07]
Hiu Ching Cheung, Wenchao Yue, Zhengran Han, Mingcong Chen, Guanglin Cao, Hongbin Liu, Hongliang Ren.
"Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots." ICCBS (2026).
[paper]
[2026.07]
LSP-SAM2: Zhaoyuan Wu, Naiyang Guan, Yinghui Gao, Longfei Su, Min Liu.
"A Light-Weight Self-Prompting Foundation Model for Automatous Video Object Segmentation." ICIC (2026).
[paper]
[code]
[2026.07]
FESAM: Chaoyue Wang, Jiang Wang, Lingfang Li, Weijian Hu & Lizhen Cui .
"FESAM: Frequency-Enhanced SAM with Boundary-Aware Decoding for Ultrasound Image Segmentation." ICIC (2026).
[paper]
[2026.07]
Agent-SAM-I2V: Guo Yang, Jiaqi Zhang, Yao Zhu & Longze Fan.
"Agent-SAM-I2V: Self-correcting Promptable Video Segmentation via Agentic Drift Detection and Multi-Prompt Fusion." ICIC (2026).
[paper]
[2026.07]
GB-SAM: Chenlin Xu, Lei Zhang, Lituan Wang, Xinyu Pu, Pengfei Ma, Guangwu Qian.
"GB-SAM: Gaussian-Prior and Boundary-Guided Test-Time Adaptation for Medical Image Segmentation." ICIC (2026).
[paper]
[code]
[2026.07]
CME-SAM: Qiyuan Wang, Jinfu Wang, Chuyu Chen, Pengtao Ren, Shijie Ling, Kejiang Xiao.
"CME-SAM: Contrastive Mask-Enhanced Segment Anything Model for Generalizable Medical Image Segmentation." ICIC (2026).
[paper]
[2026.07]
IBS-EMA: Yizhuo Wang, Shiquan Min, Jiangping Zhu & Pei Zhou.
"IBS-EMA: Mitigating Test-Time Prompt Distribution Shift for Medical Segment Anything Models." ICIC (2026).
[paper]
[2026.07]
HCFNet: Yu, Jiwei, Kecheng Zhou, Ting Wang, Hongxiao Gan, Yu Wang, and Shuzhi Gao.
"HCFNet: A SAM2-Based Hierarchical Cross-Branch Frequency-Aware Network for Industrial Surface Defect Segmentation." Sensors (2026).
[paper]
[2026.07]
ZA-SAM: Jinliang Su, Yun Jiang, Zequn Zhang & Yuhang Li .
"Adaptive Prompted, Zero-Annotation SAM: Weakly Supervised Binary Medical Image Segmentation." ICIC (2026).
[paper]
[2026.07]
SURE-SAM2: Yucan Duan, Chun Wang, Kaiyu Miao & Xiaoyan He.
"SURE-SAM2: Semantic and Uncertainty-aware Refinement SAM2 for Change Detection." ICIC (2026).
[paper]
[2026.07]
PolypSAM-Lite: Hasan, Umar, and Muhammad Ali Nayeem.
"Low-Rank Attention Reparameterization for Parameter-Efficient Adaptation of the Segment Anything Model to Colorectal Polyp Segmentation." Mathematics (2026).
[paper]
[2026.07]
Koki AMANO, Otoha YAMANAKA, Wakana KAWAI, Tatsuya HAYASHI, Nobuo KOCHI, Ippeita DAN.
"Segment-Anything-based AOI Analysis for Eye-tracking: A Gaze Judgment Method Considering the Visual Angle." J-STAGE(2026).
[paper]
[2026.07]
Naka-SAM: Chen, Juan; Wu, Jiajie; Guo, Lei; Ge, Wenping; Ma, Jie.
"Naka-SAM: A Cognition-Inspired Framework with Nakagami Prior for Ultrasound Segmentation." Proceedings of the Annual Meeting of the Cognitive Science Society (2026).
[paper]
[2026.07]
CG-SAM2: Bin He, Zhiwei Chen, Shengmin Zhao, Qinqin Zhou, Aiwen Jiang, Miaohui Zhang.
"CG-SAM2: Confidence-Guided Pseudo-label Refinement for Weakly Supervised Camouflaged Object Detection." ICIC (2026).
[paper]
[2026.07]
Mohammadreza Narimani, Vikram Anand, Parastoo Farajpoor.
"Farmland Extent and Visible Boundary Mapping from 1 m NAIP Imagery Using Residual U-Net and Text-Prompted SAM 3 Refinement." ArXiv (2026).
[paper]
[code]
[dataset]
[2026.07]
FluxGraph: Yihong Sun, Bharath Hariharan.
"Efficient Tracking and Understanding Object Transformations." ArXiv (2026).
[paper]
[code]
[2026.07]
SENSATION-DS: Hakan Calim, Anamaria Dumitrescu, Adarsh Bhandary Panambur, Huzaifa Asif, Andreas Maier.
"Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation." ArXiv (2026).
[paper]
[2026.07]
Lean-SAM2: Xudong Ouyang, Wenlun Zhang, Yimin Xu, Huazhong Liu, Yunshan Zhong.
"Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2." ArXiv (2026).
[paper]
[code]
[2026.07]
Samy Mounir, Mikolaj Cieslak, Najmeddine Dhieb, Hakim Ghazzai, Jonathan Klein, Katja Froehlich, Soeren Pirk, Wojciech Palubicki, Gianluca Setti, Ahmed M. Eltawil, Dominik L. Michels.
"Text-conditioned Segmentation for Tomato Phenotyping via Procedural Synthetic Data." ArXiv (2026).
[paper]
[2026.07]
Scene-SAM3D: Yuqi Zhang, et al.
"Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning." ArXiv (2026).
[paper]
[code]
[2026.07]
OP-HRG: Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Chia-Wei Tang, Zaber Ibn Abdul Hakim, Anuj Karpatne, Chris Thomas.
"Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning." ECCV (2026).
[paper]
[code]
[2026.07]
Silas kwabla Gah, Ebenezer Owusu.
"Training-Free Open-Vocabulary 3D Point-Cloud Segmentation on the Generalized Few-Shot Benchmark." ArXiv (2026).
[paper]
[2026.07]
Jinchang Zhang, Arnold Zumbrun, Jing Lin, and Guoyu Lu.
"Foundation-Assisted Active Learning for Object Detection Annotation." ArXiv (2026).
[paper]
[2026.07]
Lite-Pi: Shivanshu Agnihotri, Snehashis Majhi, Deepak Ranjan Nayak, Dwarikanath Mahapatra, Debesh Jha.
"Induce to Empower: Improving Lightweight Baselines via Foundation Model Induction for Generalized Polyp Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
Joey Páolo Kardolus, Daan Hendriks, Jaap Jansen.
"Direct Clinical Joint Angle Extraction from Parametric Body Model Rotation Matrices." ArXiv (2026).
[paper]
[code]
[2026.07]
Zhe Xin, Hanzhi Chang, Penghui Huang, Yinian Mao, Guoquan Huang.
"Robust Multimodal Dynamic Object Segmentation." ICRA (2026).
[paper]
[2026.07]
Minghui Xu, Chaoyi Zhou, Aaron P. Cecil, Xi Liu, Siyu Huang, Yuhao Xu.
"Digital measurement of droplet flame diameter in microgravity combustion images using Segment Anything Model 2 with automatic prompt selection." ArXiv (2026).
[paper]
[2026.07]
SAMRI-3D: Zhao Wang, Wei Dai, Hongfu Sun, Craig Engstrom, Shekhar S. Chandra.
"SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens." ArXiv (2026).
[paper]
[code]
[2026.07]
IMTrack: Zhiqiang Hou, Chuangye Xu, Sugang Ma, Xiaobao Yang, Lei Pu.
"Robust visual tracking via implicit memory-guided re-detection." EAAI (2026).
[paper]
[2026.07]
ReportMedSAM: Anghong Du, Theodoros N. Arvanitis, Colin Watts, Alejandro F. Frangi, Le Zhang.
"ReportMedSAM: Guiding Segmentation Through Radiology Reports." ArXiv (2026).
[paper]
[2026.07]
ViPSAM: San Lee, Nalee Kim, Jeong Il Yu, Hee Chul Park, Boah Kim.
"ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model." MICCAI (2026).
[paper]
[2026.07]
XCT-SAM: Md Mahedi Hasan, Md Mushfiqur Rahaman, Alan Pachkovskiy, Imtiaz Ahmed, Jeremy Dawson, Srinjoy Das.
"XCT-SAM: Sequential Parameter-Efficient Domain Adaptation of SAM for Industrial XCT Defect Segmentation." ICPR workshop (2026).
[paper]
[code]
[2026.07]
Yuanzhi He.
"Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation." ArXiv (2026).
[paper]
[2026.07]
SARFA: Tyler Ward, Abdullah Imran.
"SARFA: Segment Anything with Radiomic Feature Alignment." ArXiv (2026).
[paper]
[code]
[2026.07]
SAM-PAG: Wenqi Si, Gongyang Li, Shixiang Shi, Weisi Lin.
"Weakly-Supervised RGB-D Salient Object Detection via SAM-driven Pseudo Annotation and State Space Interaction-based Diffusion." IEEE TMM (2026).
[paper]
[code]
[2026.07]
SERD: Shipeng Liu, Zhanping Song, Liang Zhao, Dengfeng Chen.
"Semantic-Edge Response Decoding of SAM3 for Zero-Shot Crack Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
GFR-SAM: Yilong Yang, Jianxin Tian, Shengchuan Zhang, Liujuan Cao.
"GFR-SAM: Training-Free Referring Camouflaged Object Segmentation via Cross-Image Prompting." ArXiv (2026).
[paper]
[2026.07]
MobileSAM2: Kai Jiang, Jiaxing Huang, Jingyi Zhang, Weiying Xie, Yunsong Li, Yufei Wang, Aoran Xiao, Dacheng Tao.
"MobileSAM2: Lightweight Segment Anything for Spatial Intelligence." ECCV (2026).
[paper]
[2026.07]
REBASE: Mantha Sai Gopal, Jaison Saji Chacko, Harsh Nandwana, Sandesh Hegde, Debarshi Banerjee, Uma Mahesh.
"REBASE: Reference-Background Subspace Elimination for Training-Free In-Context Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
CtrlVTON: Seungyong Lee, Hyun Jun Jang, Sangoh Kim, Sungjoon Park.
"CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
SAMPLe: Hossein Rajoli, Fatemeh Lotfi, Niloufar Alipour Talemi, Hossein Kashiani, Xiaolong Ma, Fatemeh Afghah.
"SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs." ECCV (2026).
[paper]
[2026.07]
Mohammad Dabaja, Turgay Celik.
"Promptable Concept Segmentation from Above: Evaluating SAM 3's Zero-Shot and One-Shot Capabilities in Remote Sensing." ArXiv (2026).
[paper]
[2026.07]
SAM-MT: Ruiqi Shen, Chang Liu, Henghui Ding.
"SAM-MT: Real-Time Interactive Multi-Target Video Segmentation." ECCV (2026).
[paper]
[code]
[2026.07]
EP-SAM: Wenhao Li, Fangyi Liu, Bo Du.
"An Edge-aware Prompt-enhanced SAM for Ultrasound Image Segmentation." ICME (2026).
[paper]
[2026.07]
HPR-SAM: Yingzhen Hu, Yiheng Zhong, Keying Zhu, Zimu Zhang, Zihan Ye, Sifan Song, Jionglong Su, Xiaofeng Liu.
"HPR-SAM: Hierarchical Probabilistic Representation Learning for Prompt-free SAM-based Medical Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.07]
DA-SAM3: Ying Chen, Jinyue Li, Kun Wang, Qiankun Li, Yang Liu.
"Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation." MICCAI (2026).
[paper]
[code]
[2026.07]
RVAF: Jin Yang, Ping Wei, Nanning Zheng.
"Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation." IROS (2026).
[paper]
[code]
[2026.07]
GeoSAM-Lite: Yongcong Wang, Jie Zhang, Rui Jiang, Xubing Yang, Ting Yun, Li Zhang.
"GeoSAM-Lite: A Lightweight Foundation Model for Onboard Remote Sensing Segmentation." GRSL (2026).
[paper]
[2026.07]
GLLS: Runzhi Deng, Yundi Hu, Yiming Zhong, Zhao Wang, Xixi Liu, Hongsong Wang, Caifeng Shan, Fang Zhao.
"Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection." ECCV (2026).
[paper]
[2026.07]
SharpSplat: Porus Vaid, Shivam Chopra, Vaibhav Kumar.
"SharpSplat: Edge-Regularized 3D Gaussian Splatting for High Fidelity Urban Building Reconstruction from UAV images." IGARSS(2026).
[paper]
[2026.07]
ChatImage: Wencan Jiang, Jiangning Zhang, Yong Liu.
"ChatImage: Navigating Long-Form LLM Answers through Interactive Images." ArXiv (2026).
[paper]
[code]
[project]
[2026.07]
IPS-Seg: Le-Anh Tran.
"Exploring SAM Supervision for Fine-Grained UAV Target Segmentation under Data Scarcity." ArXiv (2026).
[paper]
[2026.07]
Muhammad Aamir, Matthew Wijers, Sangyun Shin, Andrew Loveridge, Andrew Markham.
"A non-invasive video-based method for individual identification of wildlife using gait dynamics." ArXiv (2026).
[paper]
[2026.07]
Improved Iris-SAM: Maduabuchi Kingsley Okorie, et al.
"Improved Iris-SAM -Based Iris Segmentation for Recognition in Biometric Security." EJASET (2026).
[paper]
[2026.07]
GPS: Park, J., Jeong, J.
"GPS: GlobalCLIP-PatchCore-SAM Based Zero-Shot Anomaly Detection and Localization in Smart Manufacturing." ICCSA (2026).
[paper]
[2026.07]
SE-MTDNet: Qing Geng and Kaiqi Ye and Fan Xu and Yu Meng and Miao Huang and Chunyan Yuan and Li Li and Bingbo Gao and Hu Zhou and Jianyu Yang and Ying Li and Jianxi Huang and Xiaochuang Yao.
"A data- and knowledge-driven cropland parcel recognition method based on segment anything model (SAM)." International Journal of Applied Earth Observation and Geoinformation (2026).
[paper]
[2026.07]
SAM2-ICHNet: Wanying Xie, Ronghui Ju, He Li, Wei Guo, Zhaoxuan Gong & Guodong Zhang.
"SAM2-ICHNet: A detection-guided intracranial hemorrhage segmentation framework for tiny lesions and complex backgrounds." SIViP (2026).
[paper]
[2026.07]
G2TAM: Chenming Zhu, Peizhou Cao, Jingli Lin, Wenbo Hu, Yunlong Ran, Jiangmiao Pang, Tai Wang, Xihui Liu.
"G2TAM: Geometry Grounded Track Anything Model." ICML (2026).
[paper]
[2026.07]
HBF-BCER: Shengyang Ping,Zhijie Lin *,Liliang Lin,Lei Zhao,Lisha Ye,Bangguo Wang,Tao Wang.
"A Prompt-Preserving MedSAM Enhancement Framework with Historical Feature Fusion and Soft Convolutional Expert Weighting for Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.07]
SAMURAI: Carlos Perez, Neeru Gupta, Ipek Oruc.
"SAMURAI: A Two-Stage Foundation Model Pipeline for Robust Optic Nerve Head Segmentation in Fundus Images." The 39th Canadian Conference on Artificial Intelligence (2026).
[paper]
[2026.07]
LongEgoRefer: Shunya Kato, Taiki Miyanishi, Shuhei Kurita, Mahiro Ukai, Nakamasa Inoue, Chenhui Chu.
"LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension." ECCV (2026).
[paper]
[code]
[2026.06]
MMIR-TCM: Lihui Luo, Joongwon Chae, Ziyan Chen, Yang Liu, Siyi Cheng, Weihan Gao, Zelin Zeng, Xiaoming Yin, Samaneh Beheshti Kashi, Dongmei Yu, Lian Zhang, Jing Sui, Zeming Liang, Jiansong Ji, Peter E. Lobie, Peiwu Qin.
"MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support." ArXiv (2026).
[paper]
[code]
[2026.06]
AdaCount: Muhammad Ibraheem Siddiqui, Muhammad Haris Khan.
"AdaCount: Training-Free Similarity-Guided Spatial and Feature Adaptation for Zero-Shot Object Counting." ArXiv (2026).
[paper]
[code]
[2026.06]
Object LeJEPA: Jakob Geusen, Ender Konukoglu.
"Object-centric LeJEPA." ArXiv (2026).
[paper]
[2026.06]
Jian Song, Tian Zi, Shen Guanting.
"From Technical Metrics to User Perception: A User Study of a Multimodal Human–Robot Interaction System for Object Detection and Grasping." ArXiv (2026).
[paper]
[2026.06]
Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado López, Mathias Unberath.
"Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos." ArXiv (2026).
[paper]
[2026.06]
BEE: Zhiqiang Hou, et al.
"Bridging the encoder gap: Stability-aware efficient adaptation of SAM2 for video object segmentation." ArXiv (2026).
[paper]
[2026.06]
EpiSAM: Arnav Sharma, Pratyush Jena, Amal Joseph, Ravi Kiran Sarvadevabhatla.
"EpiSAM: Character Segmentation in Challenging Stone Inscriptions." ICDAR (2026).
[paper]
[code]
[2026.06]
PGE-SAM: Tuan-Duc Nguyen, Anh-Tuan Mai, Duc-Trong Le.
"PGE-SAM: Prompt-Guided Feature Enhancement for Interactive Segmentation under Degradation." ArXiv (2026).
[paper]
[2026.06]
ExACT: Zixiao Zhang, Lingling Li, Pei He, Xu Liu, Licheng Jiao.
"ExACT: Exemplar-Driven Calibrated Refinement for Training-Free Visual Grounding in Remote Sensing Images." ArXiv (2026).
[paper]
[2026.06]
SemDynReg: Ruitao Chen, Mozhang Guo, Jinge Li.
"SemDynReg: Semantics-Guided Deformation Regularization for Dynamic 3D Gaussian Splatting." ArXiv (2026).
[paper]
[code]
[2026.06]
Xin Dong, Wenfeng Deng, Yansong Tang.
"Occlusion-Robust Multi-Object Decoupling for Physics-Based Interaction." ArXiv (2026).
[paper]
[2026.06]
SPARK: Bryce Grant, Aryeh Rothenberg, Logan Senning, Zonghe Chua, Zach Patterson, Peng Wang.
"Sequential Planning via Anchored Robotic Keypoints." ArXiv (2026).
[paper]
[code]
[2026.06]
CG-ICS: Zhigang Chen, Xiawu Zheng, Rongrong Ji.
"Toward Robust In-Context Segmentation via Concept Guidance." ECCV (2026).
[paper]
[2026.06]
Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi, Eva Cetinic, Gennaro Vessio, Giovanna Castellano.
"Understanding How MLLMs Describe Artworks Using Token Activation Maps." ArXiv (2026).
[paper]
[code]
[2026.06]
TEP-SAM: Yinghui Xing, Donghao Chu, Shizhou Zhang, Di Xu.
"Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection." ICML (2026).
[paper]
[code]
[2026.06]
Simple-ViLMedSAM: Chengcan Qian, Dong Nie, Geng Chen, Daoqiang Zhang, Xuyun Wen.
"Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image Segmentation." CVPR (2026).
[paper]
[code]
[2026.06]
ScribSAM: Long Chen, et al.
"ScribSAM: A robust scribble-supervised framework for spatiotemporal segmentation of breast lesions in ultrasound videos." Computerized Medical Imaging and Graphics (2026).
[paper]
[code]
[2026.06]
Stat-SAM: Juan Chen, et al.
"Stat-SAM: Learning Global Echo-Intensity Priors as Prompts for SAM in Ultrasound Image Segmentation." ICMR (2026).
[paper]
[2026.06]
Lv, X. et al.
"Lightweight Shape-Aware Segment Anything for Cardiac Ultrasound Segmentation." SMC-IOT (2026).
[paper]
[2026.06]
Lighted-SAM: Yuhan Jia, Lixin Duan, Wen Li, and Fengmao Lv.
"Lighted-SAM: Lightening Open-World SAM for Low-Light Segmentation." TIP (2026).
[paper]
[code]
[2026.06]
NeuroSeg-MF: Zhehao Xu, Weiyi Liu, Shanshan Liang, Hongbo Jia, Xiaowei Chen, Han Qin, and Xiang Liao.
"NeuroSeg-MF: robust neuron segmentation in two-photon Ca2+ imaging using multi-feature fusion and detection-guided SAM." Biomed. Opt. Express (2026).
[paper]
[2026.06]
CPPS-SAM: Zerong Zhang, Lianghua He.
"SAM foundation model and expert model cross prompting framework for semi-supervised medical image segmentation." Journal of Visual Communication and Image Representation (2026).
[paper]
[code]
[2026.06]
Elakiya Sivakumar.
"Fine-Tuning SAM2 for Coronary Artery Segmentation in X-Ray Fluoroscopy." ArXiv (2026).
[paper]
[2026.06]
Tian, J., Cai, W., Sun, Z. et al.
"Unsupervised Change Detection in Remote Sensing Images Using an Integrated SAM and MAD Method." J Indian Soc Remote Sens (2026).
[paper]
[2026.06]
Hizukuri, A.
"Computerized Classification Method for Glioma Molecular Subtypes on Brain MR Images Using SAM-Med3D with Low-Rank Adaptation." J Digit Imaging. Inform. med.(2026).
[paper]
[2026.06]
M2C: Quan Zhou, Shaoqing Zhai, Qiang Hu Jia Chen, Qiang Li, Zhiwei Wang.
"Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation." MICCAI (2026).
[paper]
[code]
[2026.06]
SAM2Matting: Ruiqi Shen, Guangquan Jie, Chang Liu, Henghui Ding.
"SAM2Matting: Generalized Image and Video Matting." ArXiv (2026).
[paper]
[code]
[website]
[2026.06]
SENTRY: Mohamad Alansari, Yonathan Michael, Hasan AlMarzouqi, Muzammal Naseer, Naoufel Werghi, Sajid Javed.
"SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking." ECCV (2026).
[paper]
[code]
[2026.06]
MorVess: Fuyou Mao, Yifei Chen, Beining Wu, Lixin Lin, Jinnan Dai, Zhiling Li, Yilei Chen, Yaqi Wang, Hao Zhang, Yan Tang, Huiyu Zhou, Feiwei Qin.
"MorVess: Morphology-Aware Pulmonary Vessel Segmentation Network." ArXiv (2026).
[paper]
[code]
[2026.06]
Marvin Rüdt, Hao Pang, Constantin Enke, Zäzilia Seibold, Kai Furmans.
"Vision-Language Model Reasoning for Contextual Semantic Mapping in Intralogistics." IEEE ETFA (2026).
[paper]
[2026.06]
FEENet: Yang, Zhiyuan and Xu, Jindong and Ni, Mengying and Su, Menghui and Peng, Jiantao.
"A Fuzzy-Embedded Edge Enhancement Network via Segment Anything Model for VHR Remote Sensing Images Change Detection." TGRS (2026).
[paper]
[2026.06]
DR-MV3D: Jiho Choi, Seonho Lee, Seojeong Park, Hyunjung Shim.
"Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views." ECCV (2026).
[paper]
[code]
[2026.06]
ARTEMIS: Tong Wang, Siwen Wang, Yaolei Qi, Jinxing Zhou, Yuting He, Guanyu Yang, Yutong Xie.
"ARTEMIS: Agent-guided Reliability-aware Temporal Mask Evolution for Imperfectly Supervised Video Polyp Segmentation." IEEE TIP (2026).
[paper]
[code]
[2026.06]
VTOS: Jinchao Ge, Lingqiao Liu, Shuwen Zhao, Lei Wang.
"VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers." ArXiv (2026).
[paper]
[2026.06]
SSL.Prop.: Tatsuya Suzuki, Kazuya Ijuin, Hideki Tomimori, Megumi Chikano, Katsushi Sakai.
"Sparse Point-Guided Fusion of Supervised and Self-Supervised Learning Model for Seaweed Segmentation." ArXiv (2026).
[paper]
[2026.06]
ProC-SAM3: Yanghui Song, Nanqing Liu, Haonan Yin, Yingjie Gao, Chengfu Yang, Qi Ming.
"Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation." GRSL (2026).
[paper]
[code]
[2026.06]
μMatch: Marei Freitag, Olesia Korchevaia, Luca Freckmann, Anwai Archit, Constantin Pape.
"Match: Foundation Models for Semi-supervised Learning and Domain Adaptation in EM." ArXiv (2026).
[paper]
[2026.06]
CM-TTA: Yubo Zhou, Jianghao Wu, Ping Ye, Shaoting Zhang, Guotai Wang.
"Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.06]
Hi-Seg: Hongqiao Dong, Wenhao Chi, Ruobing Liang, Xiaokui Yang, Wenhua Liang, Peng Hou, Wenjun Pu, Yipeng Zhao, Ping Chen, Haiping Liu, Jianxing He, Bo Liu.
"Human and AI collaboration for pulmonary nodule segmentation." ArXiv (2026).
[paper]
[2026.06]
DeformX: Yi Yang, Xiang Fei, Lehong Wang, Chenhao Li, Zilin Dai, Henry Kou, Lu Li, Howie Choset.
"DeformX: A Versatile Co-Simulation Framework for Deformable Linear Objects." IROS (2026).
[paper]
[code]
[2026.06]
SARIF: Dong-Hyun Moon, Ju-Hyeon Nam, Sang-Chul Lee.
"SARIF: Segment Anything for Robust Image Forensics." ECCV (2026).
[paper]
[code]
[2026.06]
Auto-SAM: Yijun Wang and Dongyu Zheng and Mingcai Hou and Hongjun Li and Wei Zeng and Caihua Chen and Sixuan Wu and Yujie Gao and Yifan Bai.
"An auto-prompting Segment Anything Model for dual-modal grain segmentation in rock images." Applied Computing and Geosciences (2026).
[paper]
[2026.06]
SA-VIS: Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain, Ender Konukoglu, Luc Van Gool.
"SA-VIS: Sparse frame Annotations for training Video Instance Segmentation." ArXiv (2026).
[paper]
[2026.06]
SAM3 Self-Distillation for Fine-Grained GOOSE 2D Semantic Segmentation.
"SAM3 Self-Distillation for Fine-Grained GOOSE 2D Semantic Segmentation." ArXiv (2026).
[paper]
[2026.06]
Intrinsic-GS: Hasan Yazar, Mohamed Rayan Barhdadi, Erchin Serpedin, Mehmet Tuncel, Hasan Kurban.
"Intrinsic 4D Gaussian Segmentation from Scene Cues." ArXiv (2026).
[paper]
[code]
[2026.06]
Sonata Simonaitis-Boyd, Soonhong Lee, Lauren N. O'Brien, Brandon T. Turner, Ralph Massarczyk, Steven R. Elliott, Aobo Li, Alexander F. Leder.
"Vision AI Agent for Continuous Material Monitoring of LEGEND-1000 LoFi Reentrant Tube." ArXiv (2026).
[paper]
[2026.06]
PEFT-MedSAM: Asad Channa, Abdullah Khan, Asghar Ali Chandio, Aamir Akbar, Shahzad Memon, Aqib Hussain, Ameer Hamza.
"PEFT-MedSAM: Efficient Fine-Tuning of Medical Foundation Models for Explainable Skin Lesion Segmentation." ArXiv (2026).
[paper]
[2026.06]
Paul Julius Kühn, Saptarshi Neil Sinha, Jakob Hansen, Robin Horst.
"Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows." ArXiv (2026).
[paper]
[2026.06]
Nicholas A. Welsh, Lennon J. Shikhman, Monty Nehru Attazs, Seemanthini K. Putane, Van Minh Nguyen, Ryan T. White.
"Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection." CVPR Workshop (2026).
[paper]
[2026.06]
TVG: Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang.
"Thinking with Visual Grounding." ArXiv (2026).
[paper]
[code]
[dataset]
[2026.06]
Multi-HMR 2: Guénolé Fiche, Philippe Weinzaepfel, Romain Brégier, Fabien Baradel.
"Multi-HMR 2: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking." ArXiv (2026).
[paper]
[2026.06]
DETECTURE: Aviad Cohen Zada, Nadav Orenstein, Shai Avidan, Gal Oren.
"Sub-Semantic Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.06]
MuDuo: Fuyou Mao, Beining Wu, Yanfeng Jiang, Bohan Xu, Lixin Lin, Naye Ji, Hao Zhang, Yan Tang.
"Mutual Distillation of Dual-Foundation Models for Semi-Supervised PET/CT Segmentation." MICCAI (2026).
[paper]
[code]
[2026.06]
Gen-VCoT:: Zhiqiang Zhou, Xu Ling, Junliang Dai.
"Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations." ArXiv (2026).
[paper]
[2026.06]
Nadav Orenstein, Aviad Cohen Zada, Shai Avidan, Gal Oren.
"Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation." ArXiv (2026).
[paper]
[2026.06]
Changwoo Song.
"Parameter-Efficient Adaptation of SAM 3 for Automated ITV Generation from 4DCT Images." ArXiv (2026).
[paper]
[2026.06]
Aniq Ahmad, Heather Bedle, Ahmad Mustafa.
"Domain-Guided Prompting of the Segment Anything Model for Seismic Interpretation: The Role of Attributes, Visualization, and Hybrid Prompts." ArXiv (2026).
[paper]
[2026.06]
Yiping Li, Ronald de Jong, Romy van Jaarsveld, Franco Badaloni, Gino Kuiper, Jelle Ruurda, Josien Pluim, Marcel Breeuwer.
"Object Tokens as a Bridge Between Segmentation and Visual Question Answering in Robotic Surgery." ArXiv (2026).
[paper]
[2026.06]
MaxCode: Qiyue Liang, Steven Ingram, George Vanica, Andi Gavrilescu, Newfel Harrat, Hassan Sipra, Sethuraman Sankaran.
"Agentic Framework for Deep Learning workload migration via In-Context Learning." ArXiv (2026).
[paper]
[code]
[2026.06]
ActiveSAM: Tran Dinh Tien, Zhiqiang Shen.
"ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation." ArXiv (2026).
[paper]
[code]
[2026.06]
MooMIns: Robert Langendörfer, Markus Hillemann, Markus Ulrich.
"MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances." ArXiv (2026).
[paper]
[code]
[2026.06]
Keyi Zhu, Kyle Lammers, Chaaran Arunachalam, Kaixiang Zhang, Renfu Lu, Zhaojian Li.
"A Modular Dual-Arm Apple Harvesting Robot with Enhanced Field Performance." ArXiv (2026).
[paper]
[2026.06]
SAM-Deep-EIoU: Alexander Holmberg.
"SAM-Deep-EIoU: Selective Mask Propagation for Multi-Object Tracking." ArXiv (2026).
[paper]
[2026.06]
Li, Chenying, Xiao Tan, Xinyu Huang, Ling Sa, Nailong Zhang, and Gang Qiu.
"Galloping Target Tracking and Parameter Measurement Method for Overhead Transmission Lines Based on SAM2 Video Segmentation." Electronics (2026).
[paper]
[2026.06]
MSBA-SAM: Tao Guo, Kui Xu, Kailei Chen, Chun Xie, Shi Qiu & Rui Ye.
"MSBA-SAM: a multi-scale and boundary-aware framework for power grid segmentation in aerial image." Energy Informatics(2026).
[paper]
[2026.06]
Suyog Jadhav, Dilip K. Prasad, Krishna Agarwal.
"SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy." CVPRW (2026).
[paper]
[2026.06]
Mix-QSAM3: Navin Ranjan, Andreas Savakis.
"Mix-QSAM3: Mixed-Precision Quantization for the Segment Anything with Concepts Model." CVPRW (2026).
[paper]
[2026.06]
LSG-SAM: Muhammad Imran, Yugyung Lee.
"Latent-Stability Gated SAM: Detecting Hallucinated Segmentations under Domain Shift." CVPRW (2026).
[paper]
[2026.06]
VegSAM: Chenxiang Wu, Chenyu Li, Danfeng Hong.
"VegSAM: Vegetation-aware Adapter for Segment Anything Model in Urban Tree Segmentation." CVPRW (2026).
[paper]
[2026.06]
SAM3Count: Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar.
"SAM3Count for Zero-Shot Open Vocabulary Counting in Images and Videos." CVPRW (2026).
[paper]
[code]
[2026.06]
SAM-OOD: Seher Kanwal, Seung-Ik Lee.
"SAM-OOD: Foundation-Model-Guided Unknown Mining for Object-Level Anomaly Detection." CVPRW (2026).
[paper]
[2026.06]
C-RWHD: Khalil Khazmi, Zied Lachiri.
"Coupled annotation and architecture tailoring for lightweight and robust wheat head detection: SAM-oriented bounding boxes with simplified YOLO variants validated on a Tunisian wheat head dataset." SAT (2026).
[paper]
[2026.06]
Chun Cao, et al.
"A Segment Anything Model adaptation framework for battery visual inspection under complex radiographic imaging conditions." PR (2026).
[paper]
[2026.06]
IFP: Shuqi Xia, Guangze Shi, Jiarui Cao, Aoyuan Shi, Meilin Liu, Xiaoyi Zhang, Yujie Wang, Xueyu Liu, Cai Zhao, Ziyuan He, Yongfei Wu, Mingqiang Wei.
"Instruction-Focus-Prompt:Semantics-Driven Structural Prompts for Universal SAM Segmentation." CVPR Findings (2026).
[paper]
[code]
[2026.06]
UCOD-MKD: Huafeng Chen, Chenguang Zhu, Yueming Lyu, Caifeng Shan.
"Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection." CVPR (2026).
[paper]
[code]
[2026.06]
mSAMUNet: Md. Shariful Alam, et al.
"Cell segmentation in microscopy images using a SAM-based U-Net architecture and a novel dataset." CMPB (2026).
[paper]
[2026.06]
RFD: Ji Xia, et al.
"RFD: A Reducing Feature Discrepancy method for unsupervised cross-modality SAM adaptation." CMIG (2026).
[paper]
[2026.06]
SAMosaic3D: Peng Wang, Yongcai Wang, Wang Chen, Hualong Cao, Kang Yang, Chunxu Li, Jie Wen, Deying Li.
"SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything." CVPR (2026).
[paper]
[code]
[2026.06]
RoSAMDepth: Xuanang Gao, Zhiwei Ning, Gengming Zhang, Jiaxi Cao, Runze Yang, Zhonglong Zheng, Jie Yang, Rong Xiao, Wei Liu.
"RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model." CVPR (2026).
[paper]
[code]
[2026.06]
PASD: Zhang, Yuliang and He, Fang and Peng, Lulu and Yan, Tianyu and Zhang, Pingping and Song, Ting and Du, Lili and Chen, Dunjin.
"3D Segment Anything Model With Visual Mamba for Diagnosing Placenta Accreta Spectrum." TIP (2026).
[paper]
[code]
[2026.06]
CSAM-TUNet: Chao Xu, et al.
"CSAM-TUNet: A SAM-guided contrastive learning framework with topological attention for pericardial adipose tissue segmentation." BSPC (2026).
[paper]
[2026.06]
TransNet–SAM2: Bamwenda, Julius, Mehmet Siraç Özerdem, Orhan Ayyildiz, Veysi Akpolat, and İrem Akpolat.
"TransNet–SAM2: A Transformer–Foundation Model Framework for Prompt-Free Segmentation of White Blood Cells in Microscopic Blood Smear Images." Diagnostics (2026).
[paper]
[2026.06]
SAM-3D-MSF: Yifu Zhang, Chun Shen & Jianbing Li.
"SAM-3D-MSF: Parameter-Efficient Adaptation of Segment Anything Model for 3D Tooth CBCT Segmentation." PAKDD (2026).
[paper]
[2026.06]
Fatih Fehmi Şimşek, Melih Altay, Kaan Kalkan, Mehmet Cengiz Arslanoğlu.
"Assessing the Impact of Spatial Resolution and Hyperparameters on Automatic Agricultural Parcel Delineation Using the Segment Anything Model With Multi-Resolution and Super-Resolved Satellite Imagery." Transactions in GIS (2026).
[paper]
[2026.06]
LSAC: Yuxuan Chen, Haoyuan Xu, Peize He.
"Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models." INTERSPEECH (2026).
[paper]
[2026.06]
ZODS-RS: Zuan Gu, Tianhan Gao, Langxu Zhao.
"ZODS-RS -- Zero-training Oriented Detection & Segmentation for Remote Sensing." ArXiv (2026).
[paper]
[2026.06]
Avi Gupta, Nilotpal Sinha, Vishnu Raj, Sambuddha Saha, Pratik Joshi, Koteswar Rao Jerripothula, Tammam Tillo.
"Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio." ICML Workshop (2026).
[paper]
[2026.06]
IPSM-Bench: Jinglin Xu, Shangyan Zhao, Jiabo Wang, Xinghong Mu, Yulong Lei, Jiacheng Zhang, Hongbo Sun, Yageng Li.
"IPSM-Bench: A New Intermediate Phase Segmentation Benchmark in Microstructure Images of Zinc-Based Absorbable Biomaterials." IJCAI (2026).
[paper]
[2026.06]
Nermeen Abou Baker, Uwe Handmann.
"Don't waste SAM." ESANN (2023).
[paper]
[2026.06]
RTVP: Zekai Zhang, Qinghui Chen, Maomao Xiong, Shijiao Ding, Zhanzhi Su, Xinjie Yao, Yiming Sun, Cong Bai, Jinglin Zhang.
"Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline." AAAI (2025).
[paper]
[code]
[2026.06]
Open-V: Silas Kwabla Gah, Ebenezer Owusu.
"Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration." ArXiv (2026).
[paper]
[2026.06]
SAM-Flow: Haowang Cui, Rui Chen, Tao Luo, Tao Guo, Zheng Qin, Jiaze Wang.
"SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing." ArXiv (2026).
[paper]
[code]
[2026.06]
Pigformer: Mk Bashar, Kuljit Bhatti, Gary Rohrer, Madonna Benjamin, Tami Brown-Brandl, Daniel Morris.
"What's Under the Skin? Estimating Swine Body Condition." ArXiv (2026).
[paper]
[code]
[2026.06]
TopoPult-SSL: Nicolò Savioli, Luca Del Tongo.
"TopoPult-SSL: Gland-Mask-Free Cross-Device Meibomian Gland Segmentation via Self-Distilled Weak Clinical Priors." ArXiv (2026).
[paper]
[2026.06]
MedSAM-BoxPredictor: Amirhossein Movahedisefat, Amirreza Fateh, Mohammad Reza Mohammadi.
"Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation." ArXiv (2026).
[paper]
[code]
[2026.06]
CR-SEG: Yifan Cao, Xiaocui Yang, Faxian Wan, Shi Feng, Daling Wang, Yifei Zhang.
"CR-SEG: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation." ArXiv (2026).
[paper]
[2026.06]
SAMatcher: Xu Pan, Qiyuan Ma, Mingyue Dong, He Chen, Wei Ji, Xianwei Zheng.
"SAMatcher: Co-Visibility Modeling with Segment Anything for Robust Feature Matching." ArXiv (2026).
[paper]
[code]
[2026.06]
Sema Helali, Lina Abu Nadab, Sausan Alqawas, Alaa Abd-Alrazaq, Faleh Tamimi, Rafat Damseh.
"Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models." ArXiv (2026).
[paper]
[2026.06]
FLAME: Yiming Wang, Baiqi Wu, Qingming Li, Jiahao Chen, Tong Zhang, Shouling Ji.
"Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization." ICML (2026).
[paper]
[code]
[2026.06]
PerBite: Ahmad AlMughrabi, Farid Al-Areqi, David Fernández Gómez, Umair Haroon, Marc Bolaños, Ricardo Marques, Petia Radeva.
"PerBite: A Curated Diagnostic Workflow for Bite-Aware Food Volume Estimation." ArXiv (2026).
[paper]
[code]
[2026.06]
LG-SAM: Panav Shah, Geet Sethi, Ashutosh Gandhe.
"Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting." CVPR Workshops (2026).
[paper]
[code]
[2026.06]
Sakib Mohammad, Jarin Ritu, Md Sakhawat Hossain.
"Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models." ArXiv (2026).
[paper]
[2026.06]
GeoSAM-3D: Arun Sharma.
"GeoSAM-3D: Geodesic Prompt Propagation for Open-Vocabulary 3D Scene Segmentation from Monocular Video." (NeurIPS(2026).
[paper]
[2026.06]
MLAM: Yuliang Zhang, Fang He, Lulu Peng, Tianyu Yan, Pingping Zhang, Ting Song, Lili Du, Dunjin Chen.
"3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum." TIP (2026).
[paper]
[code]
[2026.06]
CLOC: Mengqi Lei, Shuokun Cheng, Wei Bao, Shaoyi Du, Jun-Hai Yong, Siqi Li, Yue Gao.
"Count Anything." ArXiv (2026).
[paper]
[code]
[2026.05]
Suyog Jadhav, Dilip K. Prasad, Krishna Agarwal.
"SAM for Robust Mitochondria Instance Segmentation in Fluorescence Microscopy." CVPR Workshops (2026).
[paper]
[2026.05]
CM-SAM: Jiexin Liang, et al.
"CM-SAM: Chaos-enhanced hybrid encoder for medical image segmentation with segment anything model." Biomedical Signal Processing and Control (2026).
[paper]
[2026.05]
SAM3D-Phys: Xin Dong, Weijian Deng, Lihan Zhang, Tianru Dai, Wenfeng Deng, Yansong Tang.
"SAM3D-Phys: Towards Multi-Object Interactive Simulation in Real World." ArXiv (2026).
[paper]
[code]
[2026.05]
ESAM++: Qin Liu, Lavisha Aggarwal, Saptarashmi Bandyopadhyay, Vikas Bahirwani, Marc Niethammer, Ehsan Adeli, Andrea Colaco.
"ESAM++: Efficient Online 3D Perception on the Edge." CVPR (2026).
[paper]
[code]
[2026.05]
DOST: Bolian Peng, Ying Tang, Xu Liu, Long Sun, Xiaoqiang Lu.
"Turbulence-Robust Dynamic Object Segmentation with Multi-Signal Priors and SAM2 Refinement." ArXiv (2026).
[paper]
[2026.05]
ViTA: Ji-Hoon Hwang, Jisung Bae, Dong-Wook Kim, Yeonkyu Lee, Seung-Woo Seo.
"From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments." ArXiv (2026).
[paper]
[2026.05]
CoP: Sanghyun Jo, Seo Jin Lee, Seohyung Hong, Yoorim Gang, Hyeongsub Kim, Hyungseok Seo, Kyungsu Kim.
"One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation." MICCAI (2026).
[paper]
[code]
[2026.05]
Toomas Tahves, Mauro Bellone, Junyi Gu, Raivo Sell.
"SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving." ArXiv (2026).
[paper]
[project]
[dataset]
[2026.05]
Water-AutoSAM: Sun, Yingrui, Yang Hong, Xiaowei Zhou, and Junyu Dong.
"Water-AutoSAM: Dual-Domain Enhanced Auto-Prompting SAM for Underwater Segmentation." Journal of Marine Science and Engineering (2026).
[paper]
[2026.05]
Dmytro Klepachevskyi, Alexander Wong, Sirisha Rambhatla, Yuhao Chen.
"Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion." ArXiv (2026).
[paper]
[2026.05]
PlayClass: Prince Ravi Leow, Neil Scheidwasser, Rebecca Oscarsson, Per Jensen, Samir Bhatt, David Alejandro Duchêne.
"PlayClass: Automated Play Behaviour Classification in Poultry." CVPR Workshop (2026).
[paper]
[code]
[2026.05]
PinPoint: Pouya Sadeghi, Shawn He, Pedro Pablo Guerrero Vela, C. Thomas, Alex Wong, Sirisha Rambhatla.
"PinPoint: Prompting with Informative Interior Points." ArXiv (2026).
[paper]
[2026.05]
Mannat Khurana, Sanyam Jain, Rishav Agarwal.
"Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis." ArXiv (2026).
[paper]
[2026.05]
InstructSAM: Yuqian Yuan, Wentong Li, Zhaocheng Li, Yutong Lin, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang, Wenqiao Zhang.
"InstructSAM: Segment Any Instance with Any Instructions." ArXiv (2026).
[paper]
[code]
[2026.05]
BED-SAM2: Tyler Rust, Dara McNally, Kyle O'Donnell, Colin Kelly, Chandra Kambhamettu.
"BED-SAM2: Boundary-Enhanced-Depth SAM2 via Monocular Geometric Priors." CVPR Workshop (2026).
[paper]
[code]
[2026.05]
ANAUS: Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li.
"Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation." MICCAI (2026).
[paper]
[code]
[2026.05]
RepSAM: Wenhui Chu.
"RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation." IJCAI-ECAI 2026 workshop (2026).
[paper]
[2026.05]
SAM 3-to-YOLOv8: Marcos Vinicius Mendes Faria, Thiago Borges Pereira, Isabella C. F. S. Condotta, Thiago Meireles Paixão, Francisco de Assis Boldt.
"SAM3-Assisted Training of Lightweight YOLO Models for Precision Pig Farming." IEEE SAS (2026).
[paper]
[2026.05]
DeCoDrift: H. M. Shadman Tabib, Md. Shamsuzzoha Bayzid, M Sohel Rahman.
"DeCoDrift: Stabilizing Decoder Coupling in Closed-Loop Foundation Segmentation." ArXiv (2026).
[paper]
[2026.05]
MGNet,: Xia Li, Xinran Liu, Lin Qi, Junyu Dong.
"Weakly Supervised Camouflaged Object Detection Based on the SAM Model and Mask Guidance." ArXiv (2026).
[paper]
[2026.05]
CLIP-Guided SAM: Shayan Jalilian, Abdul Bais.
"CLIP-Guided SAM: Parameter-Efficient Semantic Conditioning for Promptable Segmentation." ArXiv (2026).
[paper]
[2026.05]
ViViD-5K: Xiangzhi Tong, Chengrui Zhang, Mac Flaherty, Andre Matteo Garcia, Dominic Gorman, Jonathan Jaramillo, Justine E. Vanden Heuvel, Yu Jiang.
"ViViD-5K: Vineyard vision dataset for field-based berry detection and segmentation and grape cluster closure estimation." ArXiv (2026).
[paper]
[2026.05]
ConceptSeg-R1: Yuan Zhao, Youwei Pang, Jiaming Zuo, Wei Ji, Kailai Zhou, Bin Fan, Yunkang Cao, Lihe Zhang, Xiaofeng Liu, Huchuan Lu, Weisi Lin, Dacheng Tao, Xiaoqi Zhao.
"ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning." ArXiv (2026).
[paper]
[project]
[code]
[2026.05]
MGGA: Fan Gao, et al.
"MedSAM-guided geometry-aware 2D-3D feature fusion for medical image registration." Neural Networks (2026).
[paper]
[code]
[2026.05]
PromptLessSAM: Mohammed Al-Mustafa Hendo, et al.
"PromptLessSAM: From Foundational Model to Domain Expert via Lightweight Decoder Adaptation for Crack Segmentation." Electrical, Computer and Communications Engineering (2026).
[paper]
[2026.05]
MT-SAM: Zhao, Litao and Zhang, Yuhan and Ji, Libiao and Bao, Jie and Li, Caizi and NG, Chi-Fai and Heng, Pheng-Ann.
"MT-SAM: A Mamba-Transformer Enhanced SAM with Prior-guided Prompting for Multi-modal Prostate Cancer Delineation." TMI (2026).
[paper]
[code]
[2026.05]
STATUARIO-40K: Palmieri, Giorgio and Ranieri, Andrea and Vasarelli, Andrea and Di Angelo, Luca.
"STATUARIO-40K: Fine-Tuning SAM 3 for Instance Segmentation at Scale of Monuments, Statues, and Cracks." ArXiv (2026).
[paper]
[dataset]
[2026.05]
Per-SAM-MCPA: Hu, Chuting, Size Dai, Shifan Wu, Qiaolin Ye, and He Yan.
"Per-SAM-MCPA: A Lightweight Framework for Individual Tree Crown Segmentation from UAV Imagery." Remote Sens (2026).
[paper]
[2026.05]
ColorSAM: Bangcheng Zhan, et al.
"ColorSAM: teaching SAM to segment color medical images via quaternion decoding and prompt generation." ESWA (2026).
[paper]
[code]
[2026.05]
PromptSAMNet: A.R. Revathi, et al.
"PromptSAMNet: A memory-enhanced adaptive prompting and clustering-augmented SAM 2.0 framework for multi-leaf plant disease diagnosis." Applied Soft Computing (2026).
[paper]
[2026.05]
SDDNet: Zhao, Yu and Sun, Jing and Zhang, Guohui and Sun, Fuming and Li, Haojie.
"Enhancing SAM2 for Industrial Defect Detection via Dual-Adapter Fine-Tuning." TIM (2026).
[paper]
[code]
[2026.05]
FS-Grasp: Zhai, Di-Hua and Yu, Sheng and Xia, Yuanqing.
"Fast and Efficient 6-DoF Grasp Estimation With Segment Anything Model in Cluttered Scenes." IEEE/ASME Transactions on Mechatronics (2026).
[paper]
[2026.05]
AISCT-SAM: Kuang, Hulin and Tan, Xianzhen and Li, Shunuo and Kan, Shichao and Liu, Jin and Sun, Jiarui and Zhang, Jingyang and Yang, Chunfeng and Qiu, Wu and Zhang, Jiulou and Chen, Yang and Wang, Jianxin.
"AISCT-SAM: Customized SAM-Med2D with 3D Context Awareness and Self-Prompt Generation for Fully Automatic Acute Ischemic Stroke Lesion Segmentation on Non-Contrast CT Scans." JBHI (2026).
[paper]
[2026.05]
MUP-SAM: Lyuyang Tong, Jingwen Jiang, Bo Du.
"MUP-SAM: Multi-Scale Vision Mamba UNet Prompt Generation for SAM in Multi-Organ Medical Image Segmentation." Neural Networks (2026).
[paper]
[2026.05]
CraterSAM+: Li, Miyu and Li, Junjie and Wang, Yumei and Liu, Yu.
"Self-Improving SAM with Specialist Knowledge via Adaptive Direct Preference Optimization for Crater Segmentation." IEEE Geoscience and Remote Sensing Letters (2026).
[paper]
[2026.05]
MedSAM-COALF: Zhao, Pengyu and Hou, Yonghong and Wu, Jiasai and Yan, Ke and Huo, Shuwei.
"MedSAM-COALF: A Cold-Start One-Shot Active Learning Framework for Medical Image Segmentation via Foundation Model-Guided Proxy Tasks and Uncertainty-Aware Sampling."IEEE Sensors Journal (2026).
[paper]
[2026.05]
SAM-SS: Wang, Yalin and Han, Wei and Peng, Hong and Zheng, Weihao and Li, Xiaoxu and Kang, Zhongfeng and Chan, Sixian.
"SAM-SS: Straightforward and Efficient Designs Based on Segment Anything Model for Semantic Segmentation." TCSS (2026).
[paper]
[2026.05]
S2C-Net: Ning, Hailong and Li, Haojie and Zhang, Wuxia and Lei, Tao and Chen, Yanping and Cao, Xiaopeng and Nandi, Asoke K..
"S2C-Net: SAM2-Based Dual-Domain Feature Reconstruction and Semantic Decoupling for Tiny Remote Sensing Object Counting." TGRS (2026).
[paper]
[code]
[2026.05]
LiteWaveRep-MedSAM: Lieqiang Liu, Chengping Zhao, Tengxiao Xu, Wutao Xiong and Yuxiao Zhang.
"LiteWaveRep-MedSAM: A lightweight medical image segmentation model based on wavelet transform and reparameterization." Biomedical Physics & Engineering Express (2026).
[paper]
[code]
[2026.05]
TorqueSAM: Rahat, Shahzalal Khan, et al.
"TorqueSAM: unsupervised kidney CT analysis with localization and SAM-integrated torque clustering segmentation." ArXiv (2026).
[paper]
[2026.05]
Elgström, Albert, Diaz, Jose, Bosch, Carles.
"Hierarchical Annotation of Mural Paintings Using SAM." ArXiv (2026).
[paper]
[2026.05]
ERSF-AS: Cheng Ju, et al.
"ERSF-AS: Explainable recursive zero-shot anomaly segmentation with spatial-frequency priors via CLIP-SAM collaboration." Neurocomputing (2026).
[paper]
[2026.05]
Tan, L., Xia, Y., Teng, D. et al.
"Comparative evaluation of conventional radiomics and VGG-SAM fusion strategies for MRI-based preoperative prediction of perineural invasion in cervical cancer." Abdom Radiol (2026).
[paper]
[2026.05]
FAST-ME : Kakia Panagidi, Stathes Hadjieftymiadis.
"FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis." ArXiv (2026).
[paper]
[2026.05]
Sebastian Cavada, Francesco Pelosin, Lapo Faggi.
"Training-Free Fine-Grained Semantic Segmentations in Low Data Regimes: A FungiTastic Baseline." CVPRW (2026).
[paper]
[2026.05]
COCOTree: Junhyub Lee, Seunghun Chae, Hyosu Kim.
"COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition." ArXiv (2026).
[paper]
[code]
[2026.05]
SAMOSA: Deyi Zhu, Yuji Wang, Yong Liu, Yansong Tang, Bingyao Yu, Jiwen Lu, Jie Zhou.
"Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking." ArXiv (2026).
[paper]
[code]
[2026.05]
DECA: Lifeng Yang, Linshu Chen, Anxing Hu, et al.
"Underwater image segmentation method based on dual-encoder." ICCIIA (2026).
[paper]
[2026.05]
SGA: Jinjin Zhang, Xiefan Guo, Di Huang.
"Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis." ArXiv (2026).
[paper]
[code]
[2026.05]
HyDAR-Pano3D: Yaoyao Yue, Jérôme Schmid, Xiaoshuang Li, Eduardo Delamare, Jinman Kim.
"HyDAR-Pano3D: A Hybrid Disentangled Anatomical Recovery Framework for Panoramic-to-3D Reconstruction." ArXiv (2026).
[paper]
[2026.05]
SAM-Sode: Wanying Tan, Shuo Yan, Dazhi Huang, Yazheng Liu, Zili Shao, Rufeng Chen, Hechang Chen, Mude Shi, Tianxing Ji, Sihong Xie.
"SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection." ArXiv (2026).
[paper]
[2026.05]
Stream3D: Kaichen Zhou, Zeyang Bai, Xinhai Chang, Mengyu Wang, Paul Liang, Fangneng Zhan.
"Stream3D: Sequential Multi-View 3D Generation via Evidential Memory." ArXiv (2026).
[paper]
[code]
[2026.05]
LCA: Qisai Liu, Alloy Das, Zhanhong Jiang, Joshua R. Waite, Aditya Balu, Adarsh Krishnamurthy, Soumik Sarkar.
"Lighting-aware Unified Model for Instance Segmentation." ArXiv (2026).
[paper]
[2026.05]
VASA: Zilin Wang, Stella X. Yu.
"Vision Harnessing Agent for Open Ad-hoc Segmentation." ArXiv (2026).
[paper]
[2026.05]
DarkLLM: Ye Sun, Xin Wang, Jiaming Zhang, Yifeng Gao, Yixu Wang, Yifan Ding, Qixian Zhang, Henghui Ding, Xingjun Ma, Yu-Gang Jiang.
"DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models." ArXiv (2026).
[paper]
[code]
[2026.05]
Tonghao Zhuang, Shanglong Hu, Yongsheng Luo, Zhiqi Zhang, Yu Li.
"Synergistic Foundation Models for Semi-Supervised Fetal Cardiac Ultrasound Analysis: SAM-Med2D Boundary Refinement and DINOv3 Semantic Enhancement." MICCAI Workshop (2026).
[paper]
[code]
[2026.05]
Ananth Sriram, Neel Mokaria, Rajveer Singh.
"Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification." ArXiv (2026).
[paper]
[code]
[2026.05]
MedFM-Robust: Xiangxiang Cui, Tianjin Huang, Yifang Wang, Lijie Hu, Lu Yin.
"MedFM-Robust: Benchmarking Robustness of Medical Foundation Models." MICCAI (2026).
[paper]
[code]
[2026.05]
OmniVL-Guard Pro: Jinjie Shen, Zheng Huang, Yuchen Zhang, Yujiao Wu, Yaxiong Wang, Lechao Cheng, Shengeng Tang, Tianrui Hui, Nan Pu, Zhun Zhong.
"OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics." ArXiv (2026).
[paper]
[code]
[2026.05]
SegRAG: Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed.
"SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation." ArXiv (2026).
[paper]
[code]
[2026.05]
D3S2: Yachan Guo, JoseLuis Gomez Zurita, Danna Xue, Yi Xiao, AntonioManuel Lopez Pena.
"Metric-Guided Feature Fusion of Visual Foundation Models for Segmentation Tasks." CVPR Findings (2026).
[paper]
[code]
[2026.05]
HyperVision: Guanyiman Fu, Jingtao Li, Zihang Cheng, Zhuanfeng Li, Diqi Chen, Yan Xu, Fengchao Xiong, Jianfeng Lu, Jun Zhou.
"HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone." ArXiv (2026).
[paper]
[code]
[2026.05]
Rad-VLSM: Fengyi Zhang, Xujie Zeng, Mohan Liu, Zengyi Wang, Yalong Jiang.
"Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis." ArXiv (2026).
[paper]
[2026.05]
Divya Joshi, J. D. Peiffer, Colleen Peyton, R. James Cotton.
"Markerless Motion Capture for Biomechanical Whole-Body Kinematic Estimation in Infants." EMBC (2026).
[paper]
[2026.05]
WOW-Seg: Danyang Li, Tianhao Wu, Bin Li, Zhenyuan Chen, Yang Zhang, Yuxuan Li, Ming-Ming Cheng, Xiang Li.
"WOW-Seg: A Word-free Open World Segmentation Model." ICLR (2026).
[paper]
[code]
[2026.05]
TinySAM 2: Zhaoyuan Ding, Yijing Yang, Han Shu, Xinghao Chen.
"TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model." ArXiv (2026).
[paper]
[2026.05]
SparseSAM: Hoai-Chau Tran, Chi H. Nguyen, Duy M. H. Nguyen, Mathias Niepert, Fan Lai, Khoa D. Doan.
"SparseSAM: Structured Sparsification of Activations in Segment Anything Models." ArXiv (2026).
[paper]
[2026.05]
CAR-SAM: Houji Wen, Jiangyong Yu, Jun Li, Dawei Yang.
"CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model." ArXiv (2026).
[paper]
[2026.05]
Eugenia Moris, Alicia Costábile, Sebastián Rey, Irene Ferreiro, Joaquín Hurtado, Lizandra Lissette Luciano, Matías Villagrán, Aisha Espino Vázquez, Jomari Ramos, Isadora Monteiro, María Victoria de Santiago, Pilar Moreno, Gonzalo Moratorio, José Ignacio Orlando.
"End-to-end plaque counting and virus titration from laboratory plate images with deep learning." ArXiv (2026).
[paper]
[2026.05]
Raushan Joshi, Jean-Yves Guillemaut.
"Robust Prior-Guided Segmentation for Editable 3D Gaussian Splatting." ICIP (2026).
[paper]
[2026.05]
MT-SAM: Litao Zhao, Yuhan Zhang, Libiao Ji, Jie Bao, Caizi Li, Chi-Fai NG and Pheng-Ann Heng.
"MT-SAM: A Mamba-Transformer Enhanced SAM with Prior-guided Prompting for Multi-modal Prostate Cancer Delineation." TMI (2026).
[paper]
[code]
[2026.05]
DT-ZSAM: Fan, Zhanpeng, Xinglei Gu, Qiyu Liu, Yangheng Hu, and Liang Yu.
"Fusing Dual-Threshold Prompts with SAM for Shot Peening Coverage Assessment on Aircraft Propeller Blades." Applied Sciences (2026).
[paper]
[2026.05]
DeFakerOne: GuangJian Team, Ant Group.
"Venus-DeFakerOne: Unified Fake Image Detection & Localization." ArXiv (2026).
[paper]
[code]
[2026.05]
PDI-Bench: Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, Xueyan Zou.
"Quantitative Video World Model Evaluation for Geometric-Consistency." ArXiv (2026).
[paper]
[code]
[2026.05]
MedCore: Cenwei Zhang, Suncheng Xiang, Lei You.
"MedCore: Boundary-Preserving Medical Core Pruning for MedSAM." ArXiv (2026).
[paper]
[code]
[2026.05]
Seg-Agent: Chao Hao, Jun Xu, Ji Du, Shuo Ye, Ziyue Qiao, Xiaodong Cun, Guangcong Wang, Xubin Zheng, Zitong Yu.
"Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation." ArXiv (2026).
[paper]
[code]
[2026.05]
PointGS: Yixiao Song, Qingyong Li, Wen Wang, Zhicheng Yan.
"PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting." CVPR (2026).
[paper]
[code]
[2026.05]
FocusDepth: Yuxin Du, Tao Lin, Zile Zhong, Runting Li, Xiyao Chen, Jiting Liu, Chenglin Liu, Ying-Cong Chen, Yuqian Fu, Bo Zhao.
"Focusable Monocular Depth Estimation." ArXiv (2026).
[paper]
[2026.05]
M4-SAM: Jiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong, Deyang Liu, Zhi Liu.
"M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection." CVPR (2026).
[paper]
[code]
[2026.05]
Stefano Colamonaco, Andrei-Bogdan Florea, Jaron Maene.
"Weakly Supervised Segmentation as Semantic-Based Regularization." ArXiv (2026).
[paper]
[2026.05]
RUAC: Hongyou Zhou, Marc Toussaint, Ling Shao, Zihan Ye.
"Segment Anything with Robust Uncertainty-Accuracy Correlation." ICML (2026).
[paper]
[code]
[2026.05]
FSAM: Phuoc-Nguyen Bui, Van-Nguyen Pham, Duc-Tai Le, Junghyun Bum, Hyunseung Choo.
"Frequency Adapter with SAM for Generalized Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.05]
CAFE: Shuang Liang, Zeqing Wang, Yuxian Li, Xihui Liu, Han Wang.
"From Pixels to Concepts: Do Segmentation Models Understand What They Segment?." ArXiv (2026).
[paper]
[code]
[project]
[dataset]
[2026.05]
SAMOFT: Yanchao Wang, Dawei Zhang, Chengzhuan Yang, Wei Liu, Minglu Li, Hua Wang, Zhonglong Zheng, Ming-Hsuan Yang.
"SAMOFT: Robust Multi-Object Tracking via Region and Flow." ArXiv (2026).
[paper]
[2026.05]
R. James Cotton, Pouyan Firouzabadi, Wendy Murray.
"Monocular Biomechanical Tracking of Fingers with Inverse Kinematics to Foundation Models." EMBC (2026).
[paper]
[2026.05]
RCoT-Seg: Junwei Wen, Deshui Miao, Guangming Lu, Xin Li, Wenjie Pei.
"RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation." ArXiv (2026).
[paper]
[code]
[2026.05]
SARA: Jiesong Lian, Zixiang Zhou, Ruizhe Zhong, Yuan Zhou, Qinglin Lu, Rui Wang, Long Hu, Yixue Hao, Baoru Huang.
"SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models." ArXiv (2026).
[paper]
[2026.05]
SAM 3D Animal: Xuyi Hu, Jin Lyu, Jiuming Liu, Yebin Liu, Silvia Zuffi, Liang An, Stefan Goetz.
"SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild." ArXiv (2026).
[paper]
[2026.05]
UniD-Shift: Shuai Zhang, Zhecheng Shi, Zhuxiao Li, Jing Ou, Tengxi Wang, Yuan Liu, Wufan Zhao.
"UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition." ArXiv (2026).
[paper]
[code]
[2026.05]
Qwen3-VL-Seg: Yuan Yao, Qiushi Yang, Humen Zhong, Jiangning Wei, Yifang Men, Shuai Bai, Miaomiao Cui, Zhibo Yang.
"Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding." ArXiv (2026).
[paper]
[2026.05]
AS-SAM2: Lv, Tao and Ding, Shenrun and Wang, Yuji and Li, Yue and Lu, Xiaohuan and Tian, Youliang.
"AS-SAM2: Adaptive Self-correction for Visual Tracking with SAM2." TCSVT (2026).
[paper]
[code]
[2026.05]
GA3T: Siwei Cai, Knut Peterson, Quan Tran, Christian Ricks, Dhanush Parthasarathy, Amir Kaidarov, Neil Deshpande, Sukaina Najm, David Han, Lifeng Zhou.
"GA3T: A Ground-Aerial Terrain Traversability Dataset for Heterogeneous Robot Teams in Unstructured Environments." ArXiv (2026).
[paper]
[code]
[2026.05]
HP-Adapter: Hinako Mitsuoka, Kazuhiro Hotta.
"Prompt-Free and Efficient SAM2 Adaptation for Biomedical Semantic Segmentation via Dual Adapters." ICIP (2026).
[paper]
[2026.05]
Ilov3Splat: Binh Long Nguyen, Kien Nguyen, Sridha Sridharan, Clinton Fookes, Peyman Moghadam.
"Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting." ICPR (2026).
[paper]
[code]
[2026.05]
ZhiXin Sun.
"Example-Based Object Detection." ArXiv (2026).
[paper]
[code]
[2026.05]
X2SAM: Hao Wang, Limeng Qiao, Chi Zhang, Lin Ma, Guanglu Wan, Xiangyuan Lan, Xiaodan Liang.
"X2SAM: Any Segmentation in Images and Videos." ArXiv (2026).
[paper]
[code]
[project]
[2026.05]
GLASSNet: Morteza Moradi, Mohammad Moradi, Simone Palazzo, Ali Borji, Concetto Spampinato.
"Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection." ArXiv (2026).
[paper]
[2026.05]
VL-SAM-v3: Chih-Chung Liu, Zhiwei Lin, Yongtao Wang.
"VL-SAM-v3: Memory-Guided Visual Priors for Open-World Object Detection." ArXiv (2026).
[paper]
[2026.05]
ViewSAM: Jiawei Ge, Xintian Zhang, Jiuxin Cao, Bo Liu, Fabian Deuser, Chang Liu, Gong Wenkang, Siyou Li, Juexi Shao, Wenqing Wu, Chen Feng, Ioannis Patras.
"ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking." ArXiv (2026).
[paper]
[2026.05]
DFUDA: Yerin Cheon, Aruna Balasubramanian, Francois Rameau.
"Dual-Foundation Models for Unsupervised Domain Adaptation." ICPR (2026).
[paper]
[code]
[2026.05]
Chase Cartwright, Gongbo Guo, Sai Teja Pusuluri, Christopher N. Mayhew, Mark Hester, Horacio E. Castillo.
"Approaching human parity in the quality of automated organoid image segmentation." ArXiv (2026).
[paper]
[2026.05]
SAMamba3D: Rui Zhang, Xianzhi Song, Linqi Zhu, Branko Bijeljic, Gensheng Li, Martin J. Blunt.
"SAMamba3D: adapting Segment Anything for generalizable 3D segmentation of multiphase pore-scale images." ArXiv (2026).
[paper]
[code]
[2026.05]
Remote SAMsing: Osmar Luiz Ferreira de Carvalho, Osmar Abílio de Carvalho Júnior, Anesmar Olino de Albuquerque, Daniel Guerreiro e Silva.
"Remote SAMsing: From Segment Anything to Segment Everything." ArXiv (2026).
[paper]
[2026.05]
MmSAM: Qingpeng Wang, Zhou Huang, Ying Chen, Yi Bao.
"MmSAM: multimodal meets SAM2 for efficient remote sensing semantic segmentation." International Journal of Applied Earth Observation and Geoinformation (2026).
[paper]
[code]
[2026.04]
Three-shot SAM2: Zhongyi Zhang, Julie Hides, Enrico De Martino, Gervase Tuxworth.
"Multicenter evaluation of three-shot SAM2 segmentation for group-level quantification of lumbar paraspinal muscles at the L4/L5 level on multi-sequence MRI." European Journal of Radiology (2026).
[paper]
[2026.04]
Lee, H., Jo, Y., Hong, I. et al.
"Segment anything model guided dual-mask framework for anatomically faithful medical image translation." Sci Rep (2026).
[paper]
[2026.04]
SAM-BM: Huanhuan Lv, Songru Jiang, Yuhao Bai, Tuohang Wan, Chang Gou, Lijun Chen.
"SAM-BM: An adversarial benchmark for loss functions and multi-scale objects in segment anything models." CVIU (2026).
[paper]
[code]
[2026.04]
SAM2-PolypNet: Zhaoting Mu, et al.
"SAM2-PolypNet: SAM2 with adaptive context enhancement model for polyp segmentation." BSPC (2026).
[paper]
[code]
[2026.04]
Zhihao Ni, Youwen He, Ziyue Zhou, Hankun Zhang, and Jun Tian.
"Rapid 3D reconstruction of key power plant equipment using SAM foreground segmentation and 3D Gaussian splatting." AMNA(2026).
[paper]
[2026.04]
Haiyu Yang, Miel Hostens.
"Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics." ArXiv (2026).
[paper]
[2026.04]
SAM-FuseNet: Zhu, Chenyang and Wang, Jierui and Zhang, Lanlan and Liang, Jia and Su, Qianxiao and Li, Baihua.
"SAM-FuseNet: Segment Anything Guided Multimodal Fusion for RGB–Thermal Aerial Robotic Perception." TGRS (2026).
[paper]
[2026.04]
MemOVCD: Zuzheng Kuang, Honghao Chang, Boqiang Liang, Haoqian Wang, Lijun He, Fan Li, Haixia Bi.
"MemOVCD: Training-Free Open-Vocabulary Change Detection via Cross-Temporal Memory Reasoning and Global-Local Adaptive Rectification." ArXiv (2026).
[paper]
[code]
[2026.04]
Bridge: Mingbo Hong, Feng Liu, Caroline Gevaert, George Vosselman, Hao Cheng.
"Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalization." CVPR (2026).
[paper]
[code]
[2026.04]
CRC-SAM: Daniel Lao.
"CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images." ISBI (2026).
[paper]
[2026.04]
MAFFNet: Zhiwei Feng and Benyi Yang and Baosong Deng.
"SAM-Assisted Multimodal Collaborative Enhancement for Remote Sensing Image Segmentation." Information Fusion (2026).
[paper]
[2026.04]
Sanghati Basu.
"Robustness Evaluation of a Foundation Segmentation Model Under Simulated Domain Shifts in Abdominal CT: Implications for Health Digital Twin Deployment." ArXiv (2026).
[paper]
[code]
[2026.04]
FastSAM-CD: Zhang, Shuxin and Lei, Tao and Wang, Xingwu and Liu, Tongfei and Lv, Zhiyong and Liu, Daqi and Gong, Maoguo and Nandi, Asoke K.
"FastSAM-CD: Remote Sensing Image Change Detection Using Vision Foundation Models With Stronger Encoder and Decoder." TGRS (2026).
[paper]
[2026.04]
GeoSAM: Wujie Zhou, Jin Xie, Caie Xu, Yuanyuan Liu, Yunchao Wang.
"Adapt, Generate, and Supervise: Geometry-Aware Diffusion-Guided SAM Framework for Remote Sensing Semantic Segmentation." TGRS (2026).
[paper]
[code]
[2026.04]
ATSG: Zhang, Yifan and Jiang, Zhiguo and Zhang, Haopeng.
"ATSG: Adaptive Token Linking With Segment Anything Model Guidance for Weakly Supervised Remote Sensing Image Semantic Segmentation." TGRS (2026).
[paper]
[code]
[2026.04]
SemiSAM-O1: Yichi Zhang, Le Xue, Bichun Xu, Judong Luo, Zhigang Wu, Yu Fu, Zixin Hu, Yuan Cheng, Yuan Qi.
"SemiSAM-O1: How far can we push the boundary of annotation-efficient medical image segmentation?." Medical Image Analysis (2026).
[paper]
[code]
[2026.04]
INSIGHT: Alexander Nikitas Dimopoulos, Joseph Grasso, John Beltz.
"INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public Safety." ArXiv (2026).
[paper]
[2026.04]
AgentRVOS: Deshui Miao, Chao Yang, Chao Tian, Guoqing Zhu, Kai Yang, Zhifan Mo, Xin Li.
"AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method." ArXiv (2026).
[paper]
[2026.04]
OAMVOS: Deshui Miao, Xingsen Huang, Yameng Gu, Xiaogang yu, Xin Li, Ming-Hsuan Yang.
"OAMVOS: 2nd Report for 5th PVUW MOSE Track." ArXiv (2026).
[paper]
[2026.04]
ASR-SaSaSa2VA: Zhiyu Wang, Xudong Kang, Shutao Li.
"2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA." ArXiv (2026).
[paper]
[2026.04]
DiffuSAM: Tal Grossman, Noa Cahan, Lev Ayzenberg, Hayit Greenspan.
"DiffuSAM: Diffusion-Based Prompt-Free SAM2 for Few-Shot and Source-Free Medical Image Segmentation." ArXiv (2026).
[paper]
[2026.04]
SPD: Jingxuan Kang, Ziqi Zhang, Shaoming Zheng, Shuang Li, Uday Bharat Patel, Alexander Harry Fitzhugh, Phillip Lung, Yusuf Kiberu, Nikesh Jathanna, Shahnaz Jamil-Copley, Bernhard Kainz, Chen Qin.
"Learning from Noisy Prompts: Saliency-Guided Prompt Distillation for Robust Segmentation with SAM." CVPR (2026).
[paper]
[2026.04]
SGP-SAM: Zixuan Tang, Shen Zhao.
"SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation." ArXiv (2026).
[paper]
[2026.04]
HFS-TriNet: Xu Lu, Qianhong Peng, Qihao Zhou, Shaopeng Liu, Xiuqin Ye, Chuan Yang, Yuan Yuan.
"HFS-TriNet: A Three-Branch Collaborative Feature Learning Network for Prostate Cancer Classification from TRUS Videos." CVPR (2026).
[paper]
[code]
[2026.04]
HAM-SAM2: Pan, Kanghua and Chen, Guo and Zhu, Wei and Zhao, Danhuai and Lu, Tong.
"HAM-SAM2: Enhancing SAM2 for Visual Object Tracking with Adaptive Motion Modeling and Hierarchical Memory Bank." ICASSP (2026).
[paper]
[2026.04]
Xiao, Haodong and Yu, Wenbo and Fang, Hao and Sun, Shuoyang and Chen, Bin and Wang, Xuan and Xia, Shu-Tao.
"Diffusion-Based Natural Adversarial Perturbations Towards Segment Anything Model." ICASSP (2026).
[paper]
[2026.04]
AnchorDiffusion: Honggang Zhao, et al.
"AnchorDiffusion: High-fidelity local image editing via anchor-SAM masks and dynamic noise fusion." Information Sciences (2026).
[paper]
[code]
[2026.04]
Isaac (Zack) Duitz, et al.
"Accelerating Medical Image Segmentation with EfficientViT-SAM." ArXiv (2026).
[paper]
[2026.04]
SAMDistill: Zhaozhong Wang, Dian Shao, Lei Zhang, Zuowei Zhang & Binglu Wang.
"SAMDistill: SAM-based Spatial-temporal Distillation for Robust 3D Object Detection." MIR (2026).
[paper]
[2026.04]
DBSAM: Zheng, Ziyi and Li, Weixing and Pan, Feng and Wang, Ronghao and Gao, Qi.
"DBSAM: A Dual-Branch Segment Anything Model for Infrared Small Target Detection." TGRS (2026).
[paper]
[2026.04]
Anchor-SAM: Li, Wenhai and Huang, Xiaohui and Yang, Xiaofei and Zhou, Yicong and Peng, Jiangtao and Ban, Yifang and Jiang, Nan.
"Anchor-SAM: Active Mining of Latent Anchors From SAM Encoder for Road Extraction." TGRS (2026).
[paper]
[code]
[2026.04]
PromptReg: Huang, Shiqi and Xu, Tingfa and Li, Jianan and Saeed, Shaheer U. and Shen, Ziyi and Barratt, Dean C. and Hu, Yipeng.
"PromptReg: Interactive Registration by “Corresponding Prompts” for Segment Anything Model (SAM)." TIP (2026).
[paper]
[2026.04]
Echo-SAM: Chang Liu, et al.
"Echo-SAM: fully exploits the performance of SAM for echocardiography segmentation." Biomedical Signal Processing and Control (2026).
[paper]
[2026.04]
Med-JSCC: Yang, Fan and Sun, Shuo and Jin, Chanyuan and Gao, Zhen and Niyato, Dusit.
"MedSAM-2 Large Model-Driven Medical Image Semantic Communication for Telemedicine." IEEE Internet of Things Journal (2026).
[paper]
[2026.04]
FMTW-SAM: Wenjie Cai, et al.
"FMTW-SAM: Foreground mixing and temporally weighted SAM feature fusion for cross-domain semi-supervised segmentation of type-B aortic dissection in computed tomography angiography." Neurocomputing (2026).
[paper]
[2026.04]
LAES-UNet: Tingru Liu, Yantong Zhan, Yan Wang and Delong Shao.
"An EfficientSAM-based Integrated Network for Ore Image Segmentation." Engineering Research Express (2026).
[paper]
[2026.04]
DualSplat: Xu Wang, Zhiru Wang, Shiyun Xie, Chengwei Pan, Yisong Chen.
"DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failures." CVPR (2026).
[paper]
[code]
[2026.04]
Amodal SAM: Bo Zhang, Zhuotao Tian, Xin Tao, Songlin Tang, Jun Yu, Wenjie Pei.
"Amodal SAM: A Unified Amodal Segmentation Framework with Generalization." ArXiv (2026).
[paper]
[2026.04]
DualGaze-VLM: Zehong Ke, Yanbo Jiang, Jinhao Li, Zhiyuan Liu, Yiqian Tu, Qingwen Meng, Heye Huang, Jianqiang Wang.
"From Scene to Object: Text-Guided Dual-Gaze Prediction." ArXiv (2026).
[paper]
[2026.04]
Semantic-Fast-SAM: Byunghyun Kim.
"Semantic-Fast-SAM: Efficient Semantic Segmenter." APSIPA ASC (2026).
[paper]
[code]
[2026.04]
SHP-SAM: Xiao, Fen and Huang, Ruozhuo and Wu, Zhenwei and Gao, Xieping.
"Scribble-guided Hierarchical Prompt for SAM-Based Weakly Supervised Salient Object Detection." TCSVT (2026).
[paper]
[code]
[2026.04]
YOLOv10–SAM: Verma, Pooja and Paul, Ayan and Machavaram, Rajendra and Bhattacharya, Mahua.
"Toward Grounded YOLO-SAM: Unified Detection–Segmentation Framework for Agricultural Intelligence." ACDSA (2026).
[paper]
[2026.04]
CoCo-SAM3: Yanhui Chen, Baoyao Yang, Siqi Liu, Jingchao Wang.
"CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation." ArXiv (2026).
[paper]
[2026.04]
LiteBounD: Shivanshu Agnihotri, Snehashis Majhi, Deepak Ranjan Nayak.
"Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models." ArXiv (2026).
[paper]
[code]
[2026.04]
ASTM Grain Size Estimator: Abdul Mueez, Shruti Vyas.
"Bridging Foundation Models and ASTM Metallurgical Standards for Automated Grain Size Estimation from Microscopy Images." CVPR Workshops (2026).
[paper]
[code]
[2026.04]
RefAtt-SAM: Wei, Xingji and Liu, Nanqing and Lei, Sen and Li, Heng-Chao.
"Reference and Attention Guided Few-Shot Adaptation of Segment Anything Model for Remote Sensing Images." TGRS (2026).
[paper]
[code]
[2026.04]
SEM-SAM: Linlin Wei, Yongfeng Xiao, Yifei Wang, Shan Ye, Tao Xu, Congyun Liu, Shuai Shao, Linge Ma, Jihong Cheng, Haowei Pei, Shuping Yin, Zhihua Han, Fuguo Jiang.
"From Optical Inversion to AI Vision: an AI-driven SEM Workflow for Empowering Precise Granulometric Analysis." Clean Energy (2026).
[paper]
[2026.04]
SGCT-Net: Jiang, Yubo and Yuan, Zheming and Zhou, Tairan and Chen, Jing and Xie, Fengying and Jiang, Zhiguo and Zhang, Haopeng.
"SGCT-Net: SAM-Guided Cross-Teaching Network for Weakly Supervised Semantic Segmentation for Generating High-Quality CAMs in High-Resolution Remote Sensing Imagery." JSTARS (2026).
[paper]
[2026.04]
PLS: Zhang, Aoran and Ling, Zhigang and Tan, Haoran and Wang, Yaonan.
"A Part-aware Learning Network for Weakly Supervised Semantic Segmentation." TMM (2026).
[paper]
[2026.04]
DiffuSAM: Geet Sethi, Panav Shah, Ashutosh Gandhe, Soumitra Darshan Nayak.
"DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery." ICLR Workshop (2026).
[paper]
[2026.04]
ViperSAM: Dawar Jyoti Deka.
"Inference-Time Temporal Probability Smoothing for Stable Video Segmentation with SAM2 under Weak Prompts." ArXiv (2026).
[paper]
[2026.04]
Qiuyu Kong, Shakiba Sharifi, Zanxi Ruan, Yiming Wang, Marco Cristani.
"Is SAM3 ready for pathology segmentation?." ArXiv (2026).
[paper]
[2026.04]
Yifei Yan, Yankai Liao, Linqi Ye.
"A Rapid Deployment Pipeline for Autonomous Humanoid Grasping Based on Foundation Models." ArXiv (2026).
[paper]
[2026.04]
Dual-Anchoring: Kangyi Wu, Pengna Li, Kailin Lyu, Lin Zhao, Qingrong He, Jinjun Wang, Jianyi Liu.
"Dual-Anchoring: Addressing State Drift in Vision-Language Navigation." ArXiv (2026).
[paper]
[2026.04]
Islam Mansour, Francescopaolo Sica, Michael Schmitt.
"Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery." ArXiv (2026).
[paper]
[2026.04]
Petro-SAM: Yili Ren, Shiqi Wen, Li Hou, Dingwen Xiao, Weiming Zhang, Caleb Chen Cao, Lin Wang, Zilu Zheng, Qianxiao Su, Mingjun Zhao, Lei Chen.
"From Boundaries to Semantics: Prompt-Guided Multi-Task Learning for Petrographic Thin-section Segmentation." ArXiv (2026).
[paper]
[2026.04]
WILD-SAM: Yucheng Pan, Heping Li, Zhangle Liu, Sajid Hussain, Bin Pan.
"WILD-SAM: Phase-Aware Expert Adaptation of SAM for Landslide Detection in Wrapped InSAR Interferograms." ArXiv (2026).
[paper]
[2026.04]
TF-SMOT: Laurence Bonat, Francesco Tonini, Elisa Ricci, Lorenzo Vaquero.
"Training-Free Semantic Multi-Object Tracking with Vision-Language Models." FG (2026).
[paper]
[2026.04]
Hayato Inoue, Shota Harada, Shumpei Takezaki, Ryoma Bise.
"Cell Instance Segmentation via Multi-Task Image-to-Image Schrödinger Bridge." ArXiv (2026).
[paper]
[2026.04]
Pi-HOC: Sravan Chittupalli, Ayush Jain, Dong Huang.
"Pi-HOC: Pairwise 3D Human-Object Contact Estimation." ArXiv (2026).
[paper]
[code]
[2026.04]
Caiwen Jiang, Lei Zeng, Wei Liu.
"A 3D SAM-Based Progressive Prompting Framework for Multi-Task Segmentation of Radiotherapy-induced Normal Tissue Injuries in Limited-Data Settings." Medical Image Analysis (2026).
[paper]
[2026.04]
Hao Wang, Jiqing Zhang, Xin Yang, Baocai Yin, Lu Jiang, Zetian Mi, Huibing Wang.
"Modality-Agnostic Prompt Learning for Multi-Modal Camouflaged Object Detection." ArXiv (2026).
[paper]
[2026.04]
PR-MaGIC: Minjae Lee, Sungwoo Hur, Soojin Hwang, Won Hwa Kim.
"PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation." ArXiv (2026).
[paper]
[code]
[2026.04]
OccSAM-Bench: Nhan Ho, Luu Le, Thanh-Huy Nguyen, Thien Nguyen, Xiaofeng Liu, Ulas Bagci.
"Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models." CVPR workshop (2026).
[paper]
[2026.04]
VLMaterial: Jiangyou Zhu, He Chen.
"VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification." ArXiv (2026).
[paper]
[2026.04]
H-SPAM: Julien Walther, Rémi Giraud, Michaël Clément.
"H-SPAM: Hierarchical Superpixel Anything Model." ArXiv (2026).
[paper]
[2026.04]
SeSAM : Anurag Das, Anna Kukleva, Xinting Hu, Yuki M. Asano, Bernt Schiele.
"Do Instance Priors Help Weakly Supervised Semantic Segmentation?." ArXiv (2026).
[paper]
[2026.04]
Boxes2Pixels: Camile Lendering, Erkut Akdag, Egor Bondarev.
"Boxes2Pixels: Learning Defect Segmentation from Noisy SAM Masks." CVPR workshop (2026).
[paper]
[code]
[2026.04]
Kaden Stillwagon, Alexandra Dunnum VandeLoo, Benjamin Magondu, Craig R. Forest.
"Self-supervised Pretraining of Cell Segmentation Models." ArXiv (2026).
[paper]
[2026.04]
RobustMedSAM: Jieru Li, Matthew Chen, Micky C. Nnamdi, J. Ben Tamo, Benoit L. Marteau, May D. Wang.
"RobustMedSAM: Degradation-Resilient Medical Image Segmentation via Robust Foundation Model Adaptation." ArXiv (2026).
[paper]
[2026.04]
PASTA: Melanie Neubauer, Elmar Rueckert, Christian Rauch.
"PASTA: Vision Transformer Patch Aggregation for Weakly Supervised Target and Anomaly Segmentation." ArXiv (2026).
[paper]
[2026.04]
MV3DIS: Yibo Zhao, Yigong Zhang, Jin Xie.
"MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation." ArXiv (2026).
[paper]
[code]
[2026.04]
Adrian Manchado, Tanner Cellio, Jonathan Keane, Yiyang Wang.
"AI Driven Soccer Analysis Using Computer Vision." ArXiv (2026).
[paper]
[2026.04]
Lars Lundqvist, Earl Ranario, Hamid Kamangir, Heesup Yun, Christine Diepenbrock, Brian N. Bailey, J. Mason Earles.
"Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection." ArXiv (2026).
[paper]
[2026.04]
DiTTA: Jihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin Yoon.
"Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time Adaptation." CVPR (2026).
[paper]
[code]
[2026.04]
OVS-DINO: Haoxi Zeng, Qiankun Liu, Yi Bin, Haiyue Zhang, Yujuan Ding, Guoqing Wang, Deqiang Ouyang, Heng Tao Shen.
"OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance." ArXiv (2026).
[paper]
[2026.04]
Tarot-SAM3: Weiming Zhang, Dingwen Xiao, Songyue Guo, Guangyu Xiang, Shiqi Wen, Minwei Zhao, Lei Chen, Lin Wang.
"Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation." ArXiv (2026).
[paper]
[2026.04]
PanoSAM2: Dingwen Xiao, Weiming Zhang, Shiqi Wen, Lin Wang.
"PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation." ArXiv (2026).
[paper]
[2026.04]
UniSurgSAM: Haofeng Liu, Ziyue Wang, Alex Y. W. Kong, Guanyi Qin, Yunqiu Xu, Chang Han Low, Mingqi Gao, Lap Yan Lennon Chan, Yueming Jin.
"UniSurgSAM: A Unified Promptable Model for Reliable Surgical Video Segmentation." ArXiv (2026).
[paper]
[code]
[2026.04]
Boxer: Daniel DeTone, Tianwei Shen, Fan Zhang, Lingni Ma, Julian Straub, Richard Newcombe, Jakob Engel.
"Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D." ArXiv (2026).
[paper]
[code]
[2026.04]
Ye Bi, Bimala Acharya, David Rosero, Juan Steibel.
"Automated Segmentation and Tracking of Group Housed Pigs Using Foundation Models." ArXiv (2026).
[paper]
[2026.04]
Abdelmoamen Nasser, Yousef Baba'a, Murad Mebrahtu, Nadya Abdel Madjid, Jorge Dias, Majid Khonji.
"Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs." ArXiv (2026).
[paper]
[2026.04]
SGPer: Shijie Wang, Zijian Wang, Yadan Luo, Scott Chapman, Xin Yu, Zi Huang.
"Learning to Synergize Semantic and Geometric Priors for Limited-Data Wheat Disease Segmentation." ArXiv (2026).
[paper]
[2026.04]
Pickalo: Alessandro Tarsi, Matteo Mastrogiuseppe, Saverio Taliani, Simone Cortinovis, Ugo Pattacini.
"Pickalo: Leveraging 6D Pose Estimation for Low-Cost Industrial Bin Picking." ArXiv (2026).
[paper]
[code]
[2026.04]
CardioSAM: Ujjwal Jain.
"CardioSAM: Topology-Aware Decoder Design for High-Precision Cardiac MRI Segmentation." ArXiv (2026).
[paper]
[2026.04]
FSS-SAM3: Yi-Jen Tsai, Yen-Yu Lin, Chien-Yao Wang.
"Few-Shot Semantic Segmentation Meets SAM3." ArXiv (2026).
[paper]
[code]
[2026.04]
XSeg: Hongxia Gao, Litao Li, Yixin Chen, Jiali Wen, Kaijie Zhang, Qianyun Liu.
"XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security Screening." CVPR (2026).
[paper]
[code]
[2026.04]
IndoorCrowd: Sebastian-Ion Nae, Radu Moldoveanu, Alexandra Stefania Ghita, Adina Magda Florea.
"IndoorCrowd: A Multi-Scene Dataset for Human Detection, Segmentation, and Tracking with an Automated Annotation Pipeline." CVPR Workshop (2026).
[paper]
[code]
[2026.04]
GRAZE: Syed Ahsan Masud Zaidi, Lior Shamir, William Hsu, Scott Dietrich, Talha Zaidi.
"GRAZE: Grounded Refinement and Motion-Aware Zero-Shot Event Localization ." CVPR Workshop (2026).
[paper]
[code]
[2026.04]
DPMO: Hongru Chen, Jiyang Huang, Jia Wan, Antoni B. Chan.
"Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation." ArXiv (2026).
[paper]
[2026.04]
Derek Austin.
"Better Rigs, Not Bigger Networks: A Body Model Ablation for Gaussian Avatars." ArXiv (2026).
[paper]
[2026.04]
Xusheng He, Canyang Wu, Jinrong Zhang, Weili Guan, Jianlong Wu, Liqiang Nie.
"The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation." CVPR Workshop (2026).
[paper]
[code]
[2026.04]
TEP: Jinrong Zhang, Canyang Wu, Xusheng He, Weili Guan, Jianlong Wu, Liqiang Nie.
"Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge." CVPR Workshop (2026).
[paper]
[2026.04]
APRVOS: Deshui Miao, Yameng Gu, Chao Yang, Xin Li, Haijun Zhang, Ming-Hsuan Yang.
"APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track." ArXiv (2026).
[paper]
[2026.04]
AdaLoRA-QAT: Prantik Deb, Srimanth Dhondy, N. Ramakrishna, Anu Kapoor, Raju S. Bapi, Tapabrata Chakraborti.
"AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation." ISBI (2026).
[paper]
[code]
[2026.04]
TF-SSD: Zhijin He, Shuo Jin, Siyue Yu, Shuwei Wu, Bingfeng Zhang, Li Yu, Jimin Xiao.
"TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection." CVPR (2026).
[paper]
[code]
[2026.04]
PC-SAM: Chengcheng Lv, Rushi Li, Mincheng Wu, Xiufang Shi, Zhenyu Wen, Shibo He.
"PC-SAM: Patch-Constrained Fine-Grained Interactive Road Segmentation in High-Resolution Remote Sensing Images." ArXiv (2026).
[paper]
[code]
[2026.04]
LunarRockSAM: Wang, Yinan and Ye, Hongxia and Fa, Wenzhe.
"LunarRockSAM: A Domain-Adapted SAM with Bright-Spots Prompting and Conditional Screening for Lunar Rock Extraction." IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (2026).
[paper]
[2026.03]
DBM-SAM: Wei Gao, Teng Li, Cunang Jiang, Sicheng Wang, Yu Dai.
"DBM-SAM: Dual-branch multiscale adaptation of SAM for medical ultrasound segmentation." Displays (2026).
[paper]
[2026.03]
SAM2-RoadNet: Feng, Ruyue, Ziyou Guo, Xiao Du, and Tieru Wu.
"SAM2-RoadNet: Topology-Aware Multi-Scale Road Extraction from High-Resolution Remote Sensing Images." Remote Sensing (2026).
[paper]
[2026.03]
IDRG-mSAM: Wang, Leiquan and Meng, Yu and Luo, Chunbo and Xu, Mingming and Wu, Chunlei and Li, Zhongwei.
"SAM-Based Multi-Scale Fine-Tuning with Inter-layer Difference Guidance for Remote Sensing Change Detection." IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (2026).
[paper]
[2026.03]
SAM-ColonPolypGen: Shasha Zhang and Yuang Cai and Yijun Chen and Xiang Cai and Peng Li.
"SAM-ColonPolypGen: Enhancing automated colon polyp report generation via reinforcement learning and prompt chaining." Biomedical Signal Processing and Control (2026).
[paper]
[2026.03]
SemiBUVS: Long Chen and Qingqing Zheng and Yingying Chen and Faqin Lv and Qiong Wang.
"SAM-Guided Semi-Supervised Breast Lesion Segmentation in Ultrasound Videos with A New Dataset." Expert Systems with Applications (2026).
[paper]
[code]
[2026.03]
Mask-CDKD: Daoyu Shu and Zhan Zhang and Xiao Huang and Ru Wang and Nan Jia and Xinzhe Fu and Bingnan Yang and Fang Wan and Jianzhong Lu and Jianya Gong.
"Mask-CDKD: A source-free and label-free cross-domain knowledge distillation framework from SAM for satellite onboard VHR land-cover mapping." ISPRS Journal of Photogrammetry and Remote Sensing (2026).
[paper]
[2026.03]
SAM2-WaveUNet: Shuzhou Lv and Shubin Zhang and Xiaoshuang Huang and Dong An and Jincun Liu and Yan Meng and Yaoguang Wei.
"SAM2-WaveUNet: A Frequency-Enhanced Segmentation Network for Fine-Grained Marine Organism Delineation." Expert Systems with Applications (2026).
[paper]
[2026.03]
VLP-SAM: Sakurai, Kosuke, Ryotaro Shimizu, and Masayuki Goto.
"Vision and Language Reference for a Segment Anything Model for Few-Shot Segmentation." Journal of Imaging(2026).
[paper]
[2026.03]
AutoPrompt-SAM3D: Cheng, W., Tang, J., Wang, T. et al.
"AutoPrompt-SAM3D: integrated generation and selection for SAM2-based 3D medical segmentation." BMC Bioinformatics (2026).
[paper]
[2026.03]
Shata, Dina, Simon Denman, Sara Omrani, Robin Drogemuller, Hend Ali, and Ayman Wagdy.
"Parameter-Efficient Adaptation of Generative-Foundation (Flux, Qwen) vs. Zero-Shot (Gemini, SAM3) Models for Aerial Image Segmentation." Buildings (2026).
[paper]
[2026.03]
HATSAM: Tang, T., Rao, Z., Wang, Y. et al.
"HATSAM: hierarchical adaptation strategy for segment anything model in medical imaging." SIViP (2026).
[paper]
[2026.03]
SaSaSaSa2VA: Dengxian Gong, Quanzhu Niu, Shihao Chen, Yuanzheng Wu, Yikang Zhou, Tao Zhang, Haobo Yuan, Lu Qi, Shunping Ji.
"SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track." ArXiv (2026).
[paper]
[2026.03]
Aviraj Bevli, Sofian Chaybouti, Yasser Dahou, Hakim Hacid, Ngoc Dung Huynh, Phuc H. Le Khac, Sanath Narayan, Wamiq Reyaz Para, Ankit Singh.
"Falcon Perception." ArXiv (2026).
[paper]
[code]
[2026.03]
FT-FSOD: Xuanlong Yu, Youyang Sha, Longfei Liu, Xi Shen, Di Yang.
"A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps." CVPR (2026).
[paper]
[code]
[2026.03]
LIT: Xinyu Yang, Haozheng Yu, Yihong Sun, Bharath Hariharan, Jennifer J. Sun.
"Live Interactive Training for Video Segmentation." CVPR (2026).
[paper]
[code]
[2026.03]
Xinyao Zhang, Chang Liu, Xiao Liang, Minghui Zheng, Sara Behdad.
"Evaluating Large and Lightweight Vision Models for Irregular Component Segmentation in E-Waste Disassembly." MSEC (2026).
[paper]
[2026.03]
Syn4Seg: Guohuan Xie, Xin He, Dingying Fan, Le Zhang, Ming-Ming Cheng, Yun Liu.
"Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation." ArXiv (2026).
[paper]
[2026.03]
IP-SAM: Huiyao Zhang, Jin Bai, Rui Guo, JianWen Tan, HongFei Wang, Ye Li.
"IP-SAM: Prompt-Space Conditioning for Prompt-Absent Camouflaged Object Detection." ECCV (2026).
[paper]
[2026.03]
OpenDPR: Qi Guo, Jue Wang, Yinhe Liu, Yanfei Zhong.
"OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery." CVPR (2026).
[paper]
[code]
[2026.03]
Industrial3D: Chao Yin, Hongzhe Yue, Qing Han, Difeng Hu, Zhenyu Liang, Fangzhou Lin, Bing Sun, Boyu Wang, Mingkai Li, Wei Yao, Jack C. P. Cheng.
"Industrial3D: A Terrestrial LiDAR Point Cloud Dataset and CrossParadigm Benchmark for Industrial Infrastructure." ArXiv (2026).
[paper]
[code]
[2026.03]
CFR-SAM: Jingze Su, Tianle Zhu, Jiaxin Cai, Zhiyi Wang, Qi Li, Xiao Zhang, Tong Tong, Shu Wang, Wenxi Liu.
"Adapting SAM to Nuclei Instance Segmentation and Classification via Cooperative Fine-Grained Refinement." ArXiv (2026).
[paper]
[2026.03]
RAP: Zhihao Mao, Bangpu Chen.
"RAP: Retrieve, Adapt, and Prompt-Fit for Training-Free Few-Shot Medical Image Segmentation." IJCNN (2026).
[paper]
[2026.03]
Samik Some, Vinay P. Namboodiri.
"Can Unsupervised Segmentation Reduce Annotation Costs for Video Semantic Segmentation?." ICVGIP (2026).
[paper]
[2026.03]
M. Fazri Nizar.
"Domain-Guided YOLO26 with Composite BCE-Dice-Lovász Loss for Multi-Class Fetal Head Ultrasound Segmentation." ArXiv (2026).
[paper]
[2026.03]
Mask-CDKD: Daoyu Shu and Zhan Zhang and Xiao Huang and Ru Wang and Nan Jia and Xinzhe Fu and Bingnan Yang and Fang Wan and Jianzhong Lu and Jianya Gong.
"Mask-CDKD: A source-free and label-free cross-domain knowledge distillation framework from SAM for satellite onboard VHR land-cover mapping." ISPRS Journal of Photogrammetry and Remote Sensing (2026).
[paper]
[code]
[2026.03]
Colon-Bench: Abdullah Hamdi, Changchun Yang, Xin Gao.
"Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos." ArXiv (2026).
[paper]
[code]
[2026.03]
Nitin Kulkarni, Akhil Devarashetti, Charlie Cluss, Livio Forte, Philip Schneider, Chunming Qiao, Alina Vereshchaka.
"Drive-Through 3D Vehicle Exterior Reconstruction via Dynamic-Scene SfM and Distortion-Aware Gaussian Splatting." ArXiv (2026).
[paper]
[2026.03]
Guoping Xu, Jayaram K. Udupa, Yubing Tong, Xin Long, Ying Zhang, Jie Deng, Weiguo Lu, You Zhang.
"Adapting Segment Anything Model 3 for Concept-Driven Lesion Segmentation inMedical Images: An Experimental Study." ArXiv (2026).
[paper]
[code]
[2026.03]
Sieradzki, Alexander, Kamil Koszela, Szymon Koszykowski, Jakub Bednarek, and Jarosław Kurek.
"Zero-Shot Vertebral Instance Segmentation on DICOM Spine Radiographs Using Promptable Segment Anything Models." Journal of Clinical Medicine (2026).
[paper]
[2026.03]
SemiBUVS: Long Chen and Qingqing Zheng and Yingying Chen and Faqin Lv and Qiong Wang.
"SAM-Guided Semi-Supervised Breast Lesion Segmentation in Ultrasound Videos with A New Dataset." ESWA (2026).
[paper]
[code]
[2026.03]
GridVAD: Mohamed Eltahir, Ahmed O. Ibrahim, Obada Siralkhatim, Tabarak Abdallah, Sondos Mohamed.
"GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids." ArXiv (2026).
[paper]
[code]
[2026.03]
XAI-SAM: Abu Noman Md Sakib, Merjulah Roby, Zijie Zhang, Satish Muluk, Mark K. Eskandari, Ender A. Finol.
"Dissecting Model Failures in Abdominal Aortic Aneurysm Segmentation through Explainability-Driven Analysis." CVPR (2026).
[paper]
[2026.03]
ET-SAM: Xike Zhang, Maoyuan Ye, Juhua Liu, Bo Du.
"ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis." ECCV (2026).
[paper]
[code]
[2026.03]
UW-VOS: Hongshen Zhao, Jingkang Tai, Yuhang Wu, Wenkang Zhang, Xi Lan, Shangyan Wang, Tianyu Zhang, Wankou Yang.
"UW-VOS: A Large-Scale Dataset for Underwater Video Object Segmentation." ArXiv (2026).
[paper]
[2026.03]
Mingqi Gao, Sijie Li, Jungong Han.
"Re-Prompting SAM 3 via Object Retrieval: 3rd of the 5th PVUW MOSE Track." ArXiv (2026).
[paper]
[2026.03]
AgentRVOS: Woojeong Jin, Jaeho Lee, Heeseong Shin, Seungho Jang, Junhwan Heo, Seungryong Kim.
"AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation." ArXiv (2026).
[paper]
[code]
[2026.03]
FCL-COD: Jingchen Ni, Quan Zhang, Dan Jiang, Keyu Lv, Ke Zhang, Chun Yuan.
"FCL-COD: Weakly Supervised Camouflaged Object Detection with Frequency-aware and Contrastive Learning." CVPR (2026).
[paper]
[2026.03]
Miquel Lopez Escoriza, Pau Amargant Alvarez.
"Automatic Segmentation of 3D CT scans with SAM2 using a zero-shot approach." ArXiv (2026).
[paper]
[2026.03]
VIRST-Audio: Jihwan Hong, Jaeyoung Do.
"3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio." CVPR workshop (2026).
[paper]
[code]
[2026.03]
FoB: Yuntian Bo, Yazhou Zhu, Piotr Koniusz, Haofeng Zhang.
"Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric Prompting." CVPR (2026).
[paper]
[code]
[2026.03]
CataractSAM-2: Mohammad Eslami, Dhanvinkumar Ganeshkumar, Saber Kazeminasab, Michael G. Morley, Michael V. Boland, Michael M. Lin, John B. Miller, David S. Friedman, Nazlee Zebardast, Lucia Sobrin, Tobias Elze.
"CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation." ArXiv (2026).
[paper]
[2026.03]
Lei Huang, Kai-Li Wang, Zhang Chen, Zhen-Huang, Saidjafar Murodzoda, Xin Chen, Jing Chen, Chun-Hao Chen, Yu Xia, Yu-Tong Yang, Jia-Cheng Li, Dilshod Nematov, Ilhan Yavuz, Zhao-Kui Wang.
"SAM Molecular Stacking with Heterogeneous Orientationfor High-Performance Perovskite Photovoltaics." ArXiv (2026).
[paper]
[2026.03]
Thomas Mendelson, Joshua Francois, Galit Lahav, Tammy Riklin-Raviv.
"Boundary-Aware Instance Segmentation in Microscopy Imaging." ISBI (2026).
[paper]
[2026.03]
Muhammad Hassan Maqsood, Yanming Zhu, Alfred Lam, Getamesay Dagnaw, Xuefei Yin, Alan Wee-Chung Liew.
"Prompt-Free Lightweight SAM Adaptation for Histopathology Nuclei Segmentation with Strong Cross-Dataset Generalization." ISBI (2026).
[paper]
[2026.03]
Carolin Teuber, Anwai Archit, Tobias Boothe, Peter Ditte, Jochen Rink, Constantin Pape.
"Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy." ArXiv (2026).
[paper]
[2026.03]
Distillation-SAM: Tang, Jiyang and Han, Hu and Shan, Shiguang and Chen, Xilin.
"Distillation-SAM: Knowledge Distillation Based Auto-prompt Embedding Learning for Surgical Image Segmentation." TMI (2026).
[paper]
[code]
[2026.03]
EventVCOD: Zhang, H., Lyu, Y., Liu, H., Song, J., Yuan, D., & Yang, Y.
"Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired Data." AAAI (2026).
[paper]
[code]
[2026.03]
GoalVLM: MoniJesu James, Amir Atef Habel, Aleksey Fedoseev, Dzmitry Tsetserokou.
"GoalVLM: VLM-driven Object Goal Navigation for Multi-Agent System." ArXiv (2026).
[paper]
[2026.03]
Perceptio: Yuchen Li, Amanmeet Garg, Shalini Chaudhuri, Rui Zhao, Garin Kessler.
"Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
Truncated — view the full README on GitHub.