GuoleiSun/Awesome-SAM2

This repo aims to include materials (papers, codes, slides) about SAM2 (segment anything in images and videos). We are continuously improving the project. Welcome to PR the works (papers, repos) that are missed.

156

21 commits

updated Oct 1, 2025

See the code

README

Awesome SAM2 (Segment Anything in Images and Videos)

Awesome SAM2 Stars Forks Last Updated

📖 About This Repository

This repository aims to be the most comprehensive collection of materials (papers, codes, datasets, demos) about SAM2 (Segment Anything in Images and Videos), Meta AI's groundbreaking vision foundation model.

SAM2 represents a significant advancement in computer vision, extending the capabilities of the original SAM to handle both images and videos with unprecedented accuracy and efficiency. This curated list covers the rapidly expanding ecosystem of SAM2 applications across diverse domains - from medical imaging to robotics, from 3D reconstruction to video generation.

🔥 Why SAM2?

SAM2 has revolutionized segmentation tasks by:

  • Providing unified image and video segmentation capabilities
  • Enabling zero-shot generalization across domains
  • Offering efficient real-time processing
  • Supporting diverse prompting mechanisms (points, boxes, masks, text)

📈 Repository Stats: Currently tracking 500+ papers and projects across 15+ domains

🤝 Contributing: We continuously improve this collection. Feel free to submit PRs for missed works, corrections, or new categories!


This repo aims to include materials (papers, codes, slides) about SAM2 (segment anything in images and videos), a vision foundation model released by Meta AI . We are continuously improving the project. Welcome to PR the works (papers, repos) that are missed.

SAM2

📊 Repository Statistics

DomainPapersKey Highlights
🏥 Medical80+Surgery, 3D medical imaging, pathology
🎬 Video Processing70+Object tracking, temporal consistency
🤖 Robotics50+Manipulation, navigation, human-robot interaction
🛰️ Remote Sensing20+Satellite imagery, environmental monitoring
🎨 Generation/Editing35+Video synthesis, image editing, creative tools
🧊 3D Processing45+Point clouds, mesh processing, reconstruction
🎯 Core Segmentation60+Novel applications and improvements
Total500+Across 15+ domains

💡 Quick Navigation: Use Ctrl+F to search for specific keywords, or click on the Table of Contents links above for domain-specific papers.

Contents

Papers/Projects

📚 Surveys & Reviews

Comprehensive overviews and systematic analyses of SAM2 applications across various domains

🎯 Traditional Segmentation

Core image and video segmentation applications, including novel architectures and domain-specific adaptations

🖼️ Image Segmentation

Segmentation Applications
Other Image Tasks

🎬 Video Segmentation

Temporal segmentation, object tracking, and video understanding applications

Referring/Reasoning Video Object Segmentation

Other Video Tasks

ReleaseTitleCode
2025.03MMCD: Memory-Based Multimodal Change DetectionNA
2025.03EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting🕒Soon
2025.03Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene🌐Project page
2025.03EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining🔗 Code
2025.03High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flight🔗 Code
2025.03FusionSegReID: Advancing Person Re-Identification with Multimodal Retrieval and Precise SegmentationNA
2025.04Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting🔗 Code
2025.04How Can Objects Help Video-Language Understanding?🕒Soon
2025.04SAMJAM:Zero-Shot Video Scene Graph Generation for Egocentric Kitchen VideosNA
2025.05Research on a traffic flow statistical algorithm based on YBOVDT and SAM2📊 Data
2025.05One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object TrajectoryNA
2025.06ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer🔗 Code
2025.06Track Any Object:A Granular Video Anomaly Detection Pipeline🌐Project page
2025.06Open-World Object Counting in Videos🔗 Code
2025.06Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D AnnotationsNA
2025.06SAM2RL:Towards Reinforcement Learning Memory Control in Segment Anything Model 2NA
2025.07Visual tracking by matching points using diffusion model🔗 Code
2025.07Intelligent and quantitative ligament breakup event analysis in 65 kHz off-axis holographic video of swirl spray🔗 Code
2025.07Towards Blind Bitstream-corrupted Video Recovery: AVisual Foundation Model-driven FrameworkNA
2025.07SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking🔗 Code
2025.08Towards automated video-based human behavior analysis: leveraging AI capabilities for spatial behavior detectionNA

🔊 Audio-visual segmentation (AVS)

Multi-modal approaches combining audio and visual information for segmentation

📊 Graph Learning

Scene graph generation and graph-based reasoning with SAM2

🏥 Medical Domain

Healthcare applications including surgery, diagnostics, and biomedical research

Medical Video & 3D Segmentation

ReleaseTitleCode
2024.08Segment anything in medical images and videos: Benchmark and deployment🔗 Code
2024.08SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video SegmentationNA
2024.08Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video SegmentationNA
2024.08Novel adaptation of video segmentation to 3D MRI: efficient zero-shot knee segmentation with SAM2NA
2024.08Biomedical SAM 2: Segment anything in biomedical images and videos🔗 Code
2024.08Polyp SAM 2: Advancing Zero-shot Polyp Segmentation in Colorectal Cancer Detection🔗 Code
2024.08Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning🔗 Code
2024.08Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video SegmentationNA
2024.09SAM-OCTA2: Layer Sequence OCTA Segmentation with Fine-tuned Segment Anything Model 2🔗 Code
2024.09Self-Prompting Polyp Segmentation in Colonoscopy using Hybrid Yolo-SAM 2 Model🔗 Code
2024.10A-MFST: Adaptive Multi-Flow Sparse Tracker for Real-Time Tissue Tracking Under OcclusionNA
2024.10ECHOPulse: ECG controlled echocardio-grams video generation🔗 Code
2024.11Phase-Informed Tool Segmentation for Manual Small-Incision Cataract SurgeryNA
2025.02SASVi - Segment Any Surgical Video🔗 Code
2025.02Text-Promtable propagation for referring medical image sequence segmentationNA
2025.02Less is More? Revisiting the Importance of Frame Rate in Real-Time Zero-Shot Surgical Video Segmentation
2025.03SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection🔗 Code(& dataset)
2025.03Surgical Gaussian Surfels: Highly Accurate Real-time Surgical Scene Rendering🔗 Code
2025.03Rethinking Few-Shot Medical Image Segmentation by SAM2: A Training-Free Framework with Augmentative Prompting and Dynamic MatchingNA
2025.03Self-Prompting Driven SAM2 for 3D Medical Image SegmentationNA
2025.04RP-SAM2: Refining Point Prompts for Stable Surgical Instrument Segmentation🔗 Code
2025.04Agglomerating Large Vision Encoders via Distillation for VFSS SegmentationNA
2025.04MedSAM2: Segment Anything in 3D Medical Images and Videos🔗 Code
2025.05Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos🕒Soon
2025.05Adapting Segment Anything 2 for Diabetic Retinopathy Lesion SegmentationNA
2025.07Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical AssistanceNA
2025.07Towards Affordable Tumor Segmentation and Visualization for 3D Breast MRI Using SAM2NA
2025.08Edge2Prompt: Modality-Agnostic Model for Out-of-Distribution Liver SegmentationNA
2025.08F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery🔗 Code
2025.08TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios🔗 Code
2025.08SAM2Med3D: Leveraging video foundation models for 3D breast MRI segmentationData Upon Request

Medical Image Segmentation

ReleaseTitleCode
2024.08SAM & SAM 2 in 3D Slicer: SegmentWithSAM Extension for Annotating Medical Images🔗 Code
2024.08SAM2-PATH: A better segment anything model for semantic segmentation in digital pathology🔗 Code
2024.08Is SAM 2 Better than SAM in Medical Image Segmentation?NA
2024.08A Short Review and Evaluation of SAM2's Performance in 3D CT Image Segmentation🔗 Code
2024.08Interactive 3D Medical Image Segmentation🔗 Code
2024.08SAM2-UNet: Segment Anything 2 Makes Strong Encoder for Natural and Medical Image Segmentation🔗 Code
2024.08SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More🔗 Code
2024.08Retrieval-augmented Few-shot Medical Image Segmentation with Foundation Models
2024.10SAM-Swin: SAM-Driven Dual-Swin Transformers with Adaptive Lesion Enhancement for Laryngo-Pharyngeal Tumor Detection🔗 Code
2024.11A multi-task learning model for clinically interpretable sesamoiditis gradingNA
2024.11Zero-shot capability of SAM-family models for bone segmentation in CT scansNA
2024.11SAM-I2I: Unleash the Power of Segment Anything Model for Medical Image TranslationNA
2024.12Medical SAM 2: Segment Medical Images As Video Via Segment Anything Model 2🔗 Code
2025.03Self-Prompting Driven SAM2 for 3D Medical Image SegmentationNA
2025.03WeakMedSAM: Weakly-Supervised Medical Image Segmentation via SAM with Sub-Class Exploration and Prompt Affinity Mining🔗 Code
2025.03Research on recognition of diabetic retinopathy hemorrhage lesions based on fine tuning of segment anything modelNA
2025.04HRMedSeg: Unlocking High-resolution Medical Image segmentation via Memory-efficient Attention Modeling🔗 Code
2025.04Prompt Once, Segment Everything: Leveraging SAM 2 Potential for Infinite Medical Image Segmentation with a Single Prompt🔗 Code
2025.04Semi-automated segmentation of magnitude images in 4D flow MR scans using segment anything model 2 (SAM 2)NA
2025.05ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking🔗 Code
2025.06MorphSAM: Learning the Morphological Prompts from Atlases for Spine Image SegmentationNA
2025.06SynPo: Boosting Training-Free Few-Shot Medical Segmentation via High-Quality Negative Prompts🌐Project page
2025.06Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning🔗 Code
2025.07Speckle2Self: Self-Supervised Ultrasound Speckle Reduction Without Clean Data🕒Soon
2025.08 Training-Free Breast Ultrasound Image Segmentation with Retrieval-based SAM2NA
2025.08CXR-ODDet: An Omni-Decoupled Multi-ClassLesion Localization Framework for Automatic ChestX-Ray AnalysisNA
2025.08LGFFM: A Localized and Globalized Frequency Fusion Model for Ultrasound Image Segmentation🔗 Code
2025.04Semi-automated segmentation of magnitude images in 4D flow MR scans using segment anything model 2 (SAM 2)NA
2025.09Co-Seg: Mutual Prompt-Guided Collaborative Learning for Tissue and Nuclei Segmentation🔗 Code

Other Medical Applications

ReleaseTitleCode
2025.03Flip Learning: Weakly supervised erase to segment nodules in breast ultrasoundNA
2025.03From Monocular Vision to Autonomous Action:Guiding Tumor Resection via 3D ReconstructionNA
2025.03Operating Room Workflow Analysis via Reasoning Segmentation over Digital TwinsNA
2025.03Early Detection and Classification of Lung Cancer using Segment Anything Model 2 and Dense NetNA
2025.04Zero-Shot 4D Lidar Panoptic SegmentationNA
2025.04SYNTHFM: Training Modality-Agnostic Foundation Models for Medical Image Segmentation Without Real Medical DataNA
2025.04VoxelFeat: Voxel-wise foundation model featuresNA
2025.06Leadership Assessment in Pediatric Intensive Care Unit Team TrainingNA
2025.08GM-ABS: Promptable Generalist Model Drives Active Barely Supervised Training in Specialist Model for 3D Medical Image Segmentation🔗 Code
2025.08Transgene-free generation of mouse post-gastrulation whole embryo models solely from naive ESCs and iPSCsNA
2025.09A Machine Learning Assisted Tool and Numerical Model for Analyzing Lipid Nanoparticles🔗 Code
2025.09Evaluating the Efficacy of Mebendazole Repurposing for Ovarian Cancer Therapy Using Optical Coherence TomographyNA
2025.09Evaluation of Radio Frequency Ablation in Human Left Atrial Tissues for Atrial Fibrillation Using Optical Coherence TomographyNA
2025.09Abundance of Maternal Mitochondrial Genome Is Dispensable up to the Mitochondrial Genome Activation in Post-Implantation EmbryosNA
2025.09The Evaluation of a Deep Learning Approach to Automatic Segmentation of Teeth and Shade Guides for Tooth Shade Matching Using the SAM2 AlgorithmNA
2025.09Neck-focused Remote Photoplethysmography (rPPG): A comparative study using clinical data and the PyVHR frameworkNA

🎭 Camouflaged Object Detection (COD)

Detecting and segmenting objects that blend with their surroundings

Video COD

Image COD

🛰️ Remote Sensing

Satellite imagery analysis, environmental monitoring, and geospatial applications

ReleaseTitleCode
2024.11DED-SAM: Adapting Segment Anything Model 2 for Dual Encoder-Decoder Change DetectionNA
2025.01Prompt-Based Segmentation at Multiple Resolutions and Lighting Conditions using Segment Anything Model 2NA
2025.03Customized SAM 2 for Referring Remote Sensing Image SegmentationNA
2025.05InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition🔗 Code
2025.06Baltimore Atlas: FreqWeaver Adapter for Semi-supervised Ultra-high Spatial Resolution Land Cover ClassificationNA
2025.06Bundle adjustment for multi-source Mars orbiter imagery with generalized control constraintsNA
2025.07Leveraging SAM 2 and LiDAR for Automated Individual Tree Crown Delineation: A Comparative Evaluation of Prompting Methods🔗 Code
2025.07Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment🔗 Code
2025.07CSW-SAM: a cross-scale algorithm for very-high-resolution water body segmentation based on segment anything model 2NA
2025.07A Fine Agricultural Flood Segmentation Model For HJ-2E S-band SAR DataNA
2025.08DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction🔗 Code
2025.09PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection🔗 Code
2025.09An Automatic Sample Augmentation Method for Paddy Rice Mapping Based on Segment Anything Model and Phenological Features—A Case Study in Southwest ChinaNA
2025.09SOPSeg: Prompt-based Small Object Instance Segmentation in Remote SensingImageryNA
2025.09BiSAM-CD: Zero-Shot Remote Sensing Change Detection via Bidirectional Temporal Memory in SAM2🔗 Code
2025.10SinkSAM-Net: Knowledge-driven self-supervised sinkhole segmentation using topographic priors and Segment Anything Model🌐Project page

🧊 3D Processing & Point Clouds

Three-dimensional data processing, reconstruction, and analysis

3D Segmentation

3D Reconstruction

Other 3D Applications

ReleaseTitleCode
2025.03LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View Scenes🌐Project page
2025.03DecoupledGaussian: Object-Scene Decoupling for Physics-Based Interaction🌐Project page
2025.03Free Your Hands: Lightweight Relightable Turntable Capture PipelineNA
2025.03WildSeg3D: Segment Any 3D Objects in the Wild from 2D Images🕒Soon
2025.03Pseudo-LiDAR With Two-Dimensional Instance for Monocular Three-Dimensional Object TrackingNA
2025.03SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion ConstraintNA
2025.03SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining🔗 Code
2025.03COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian Splitting🔗 Code
2025.03Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields🌐Project page
2025.03Semantic Consistent Language Gaussian Splatting for Point-Level Open-vocabulary Querying🌐Project page
2025.04Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting🌐Project page
2025.04 FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive AgentsNA
2025.06 GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation🌐Project page
2025.05Constructing a 3D Town from a Single Image🌐Project page
2025.06GenMOJO: Robust Multi-Object 4D Generation for In-the-wild Videos🌐Project page
2025.06CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image🌐Project page
2025.06BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing🌐Project page
2025.07LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion🌐Project page
2025.07 Consistent Bokeh for Multi-View Images With 3D Gaussian Splatting🔗 Code
2025.07Defect segmentation and 3D reconstruction in concrete structures using SAM 2 and 3D Gaussian splattingUpon Request
2025.07Image-Guided Shape-from-Template Using Mesh Inextensibility Constraints🔗 Code
2025.07MG-Mono: A Lightweight Multi-Granularity Method for Self-Supervised Monocular Depth Estimation🔗 Code
2025.08Training-free automatic instance segmentation of girder bridge point cloud via large model fusion with reverse entity modelling verificationNA
2025.08SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass🌐Project page

🎨 Image or Video Generation & Editing

Creative applications including content generation, editing, and artistic tools

ReleaseTitleCode
2024.10AdaptiveDrag: Semantic-Driven Dragging on Diffusion-Based Image Editing🔗 Code
2024.11VideoDirector: Precise Video Editing via Text-to-Video Models🌐Project page
2024.11VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing🌐Project page
2024.11Generative Omnimatte: Learning to Decompose Video into Layers🌐Project page
2024.12InterDyn: Controllable Interactive Dynamics with Video Diffusion Models🌐Project page
2025.01MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis🌐Project page
2025.01BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations🌐Project page
2025.03TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video SynthesisNA
2025.03Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias🔗 Code
2025.03Unified Dense Prediction of Video DiffusionNA
2025.03DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image🕒Soon
2025.03FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing🌐Project page
2025.03MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance🌐Project page
2025.03Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting🌐Project page
2025.03Multi-Subject and Motion Customization of Text-to-Video Diffusion Models🌐Project page
2025.04DreamFuse: Adaptive Image Fusion with Diffusion Transformer🌐Project page
2025.04Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution🔗 Code
2025.06 Keyframe-Guided Creative Video Inpainting🌐Project page
2025.06OmniGen2: Exploration to Advanced Multimodal Generation🌐Project page
2025.07Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning🔗 Code
2025.07Enhanced Velocity Field Modeling for Gaussian Video ReconstructionNA
2025.08NEP: Autoregressive Image Editing via Next Editing Token Prediction🌐Project page

🗺️ Simultaneous Localization and Mapping (SLAM / VO)

Navigation, mapping, and localization applications

💡 Light Field Segmentation

Advanced imaging techniques and multi-dimensional visual processing

🤖 Robotics

Autonomous systems, manipulation, navigation, and human-robot interaction

ReleaseTitleCode
2024.10A Pipeline for Segmenting and Structuring RGB-D Data for Robotics ApplicationsNA
2024.10VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model🌐Project page
2025.02Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos🌐Project page
2025.02Map Space Belief Prediction for Manipulation-Enhanced Mapping(To be released)
2025.02Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids🌐Project page
2025.03DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping🌐Project page
2025.03Autonomous Dissection in Robotic CholecystectomyNA
2025.03MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model🌐Project page
2025.03LuciBot: Automated Robot Policy Learning from Generated Videos🌐Project page
2025.03IMPACT : Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models🌐Project page
2025.03VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and InvisibilityNA
2025.03ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis🌐Project page
2025.04Slot-Level Robotic Placement via Visual Imitation from Single Human Video🌐Project page
2025.04Entangled chip removal utilizing mass-spring model with mobile manipulatorNA
2025.05Symbolically-Guided Visual Plan Inference from Uncurated Video DataNA
2025.05Geometry-Consistent Video Diffusion for Robotic Visual Policy Transfer🌐Project page
2025.05Grasp the Invisibility by Vision-Language guided Active View PlanningNA
2025.07Geometry-aware 4D Video Generation for Robot Manipulation🌐Project page
2025.07Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation LearningNA
2025.07GraspGen: A Diffusion-based Framework for 6-DOF Grasping🌐Project page
2025.07RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping🔗 Code
2025.08Train Once, Deploy Anywhere: Realizing Data-Efficient Dynamic Object Manipulation🔗 Code
2025.09ObjectReact: Learning Object-Relative Control for Visual Navigation🌐Project page

⚡ Adaptation, Compression & Edge Applications

Efficiency optimizations, model compression, and deployment on resource-constrained devices

📖 Training

Resources for model training, datasets, and learning frameworks

Datasets

Curated datasets and benchmarks for SAM2 training and evaluation

ReleaseTitleCode
2025.02SurgPose: a Dataset for Articulated Robotic Surgical Tool Pose Estimation and Tracking🌐Project page
2025.02The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionNA
2025.02Picking the Cream of the Crop:Visual-Centric Data Selection with Collaborative Agents🔗 Code
2025.03Phantom: Training Robots Without Robots Using Only Human Videos🌐Project page
2025.03Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups🌐Project page
2025.03What Are You Doing? A Closer Look at Controllable Human Video Generation🔗 Code
2025.03Instrument-Splatting: Controllable Photorealistic Reconstruction of Surgical Instruments Using Gaussian Splatting🕒Soon
2025.03Referring to Any Person🌐Project page
2025.03AUTV: Creating Underwater Video Datasets with Pixel-wise AnnotationsNA
2025.03DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios🌐Project page
2025.05A fusion network for multi-modality medical image registration with progressive feature alignment🔗 Code
2025.04InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable GaussiansNA
2025.04VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing🌐Project page
2025.04UrbanWaste: In-the-Bin Dataset for Waste Disposal Inspection with Multi-Granularity Hierarchical Labels🌐Project page
2025.06HD-EPIC: A Highly-Detailed Egocentric Video Dataset🌐Project page
2025.06 GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities🌐Project page
2025.06INTERNSPATIAL: A Comprehensive Dataset for Spatial Reasoning in Vision-Language ModelsNA
2025.06BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos🌐Project page
2025.06SAM4D:Segment Anything in Camera and LiDAR Streams🌐Project page
2025.06XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation🌐Project page
2025.07A New Dataset and Performance Benchmark for Real-time Spacecraft Segmentation in Onboard Flight Computers🔗 Code
2025.07Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation🌐Project page
2025.08MOSE: Complex Video Object Segmentation Dataset🌐Project page
2025.08DreamVE: Unified Instruction-based Image and Video Editing🌐Project page
2025.09SPATIALVID: A Large-scale Video Dataset with Spatial Annotations🌐Project page

🔄 Used for Data Augmentation (/Tool)

Tools and methods for data synthesis, augmentation, and preprocessing

ReleaseTitleCode
2025.02Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach🌐Project page
2025.03A Taxonomy for Evaluating Generalist Robot Policies🌐Project page
2025.03CRESTE: Scalable Mapless Navigation with Internet Scale Priors and Counterfactual Guidance🌐Project page
2025.03Shaken, Not Stirred: A Novel Dataset for Visual Understanding of Glasses in Human-Robot Bartending Tasks🌐Project page
2025.03YOLOE: Real-Time Seeing Anything🔗 Code
2025.03VACE: All-in-One Video Creation and Editing🌐Project page
2025.03V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video🌐Project page
2025.03Better Together: Unified Motion Capture and 3D Avatar ReconstructionNA
2025.03CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based GuidanceNA
2025.03RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance 🌐Project page
2025.03Evaluating the FLUX.1 Synthetic Data on YOLOv9 for AI-Powered Poultry FarmingNA
2025.03Any2Caption : Interpreting Any Condition to Caption for Controllable Video Generation🌐Project page
2025.05LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration🔗 Code
2025.04VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning🌐Project page
2025.04M2Flow: A Motion Information Fusion Framework for Enhanced Unsupervised Optical Flow Estimation in Autonomous DrivingNA
2025.05Interspatial Attention for Efficient 4D Human Video Generation🌐Project page
2025.06Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting GarmentsNA
2025.06Impact of Synthetic Data from Diffusion Models on Weed Detection PerformanceNA
2025.06VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models🌐Project page
2025.06Building Software for Analyzing Muck Piles After Blasting in Laboratory Conditions with Integrated Artificial IntelligenceNA
2025.06WeedSwin hierarchical vision transformer with SAM-2 for multi-stage weed detection and classificationOn Request
2025.07Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data🔗 Code
2025.07 RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding🕒Soon

🛠️ Training Helper

Supporting tools and frameworks for model training and fine-tuning

📊 Performance Evaluations

Benchmarking, evaluation metrics, and comparative studies

Post Processing

🛡️ Robustness

Security, adversarial robustness, and reliability studies

🌟 Unique Applications/Usage

Novel and creative applications that don't fit traditional categories

ReleaseTitleCode
2024.09Towards Robust Automation of Surgical Systems via Digital Twin-based Scene Representations from Foundation ModelsNA
2024.09Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models🔗 Code
2024.09Point of Interest Recognition and Tracking in Aerial Video during Live Cycling BroadcastsNA
2024.10ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting🔗 Code
2024.10GRS: Generating Robotic Simulation Tasks from Real-World ImagesNA
2024.10Iterative Optimization Annotation Pipeline and ALSS-YOLO-Seg for Efficient Banana Plantation Segmentation in UAV ImageryNA
2024.10Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting🔗 Code
2025.01Zero-Shot Pupil Segmentation with SAM 2: A Case Study of Over 14 Million Images📊Data
2025.02Best Foot Forward: Robust Foot Reconstruction in-the-wild
2025.03ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment🌐Project page
2025.03JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse🌐Project page
2025.04MORPHEUS: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments🌐Project page
2025.05Air-Ground Collaboration for Language-Specified Missions in Unknown EnvironmentsNA
2025.06In Situ Detection and Measurement of Broccoli Heads Under Different Lighting Conditions Using Proximal Remote SensingNA
2025.07 Zero-Shot Recognition of Test Tube Types by Automatically Collecting and Labeling RGB DataNA
2025.07Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse AutomationNA
2025.07Phys2Real: Physically-Informed Gaussian Splatting for Adaptive Sim-to-Real Transfer in Robotic ManipulationNA

🤝 Contributing

We welcome contributions from the community! Here's how you can help improve this repository:

How to Contribute

  1. Submit a Pull Request with new papers/projects
  2. Report Issues for broken links or incorrect information
  3. Suggest New Categories for emerging SAM2 applications
  4. Improve Organization by suggesting better categorization

Contribution Guidelines

  • Format: Follow the existing table format with Release Date | Title | Code links
  • Quality: Include only peer-reviewed papers or significant projects
  • Recency: Focus on 2024+ publications (SAM2 era)
  • Completeness: Provide accurate metadata and working links when available
  • Categories: Place papers in the most appropriate domain category

Adding New Papers

| YYYY.MM | [Paper Title](paper_link) | [🔗 Code](code_link) / [🌐Project page](project_link) / NA |
  • 🔗 [🔗 Code] - Source code repositories
  • 🌐 [🌐Project page] - Official project websites
  • 📊 [📊Data] - Datasets and benchmarks
  • 🖥️ [🖥️Demo] - Interactive demonstrations
  • 📖 [📖Repo] - Documentation repositories
  • 🕒 🕒Soon - Coming soon
  • NA - Not available

📈 Repository Metrics

  • Total Papers: 500+ and growing
  • Domains Covered: 15+ major application areas
  • Time Range: 2024-2025 (SAM2 era)
  • Update Frequency: Weekly additions
  • Community: Open for contributions

🎉 Acknowledgments

Special thanks to:

  • Meta AI for releasing SAM2 and advancing the field
  • Research Community for the incredible pace of SAM2 adoption and innovation
  • Contributors who help maintain and improve this collection
  • Reviewers who ensure quality and accuracy

📄 License

This collection is maintained under the MIT License. Individual papers and projects retain their original licenses.


⭐ Star this repository if you find it helpful! ⭐
Last updated: August 2025 | Maintained by the Community

GuoleiSun/Awesome-SAM2

This repo aims to include materials (papers, codes, slides) about SAM2 (segment anything in images and videos). We are continuously improving the project. Welcome to PR the works (papers, repos) that are missed.

156

21 commits

updated Oct 1, 2025

See the code

README

Awesome SAM2 (Segment Anything in Images and Videos)

Awesome SAM2 Stars Forks Last Updated

📖 About This Repository

This repository aims to be the most comprehensive collection of materials (papers, codes, datasets, demos) about SAM2 (Segment Anything in Images and Videos), Meta AI's groundbreaking vision foundation model.

SAM2 represents a significant advancement in computer vision, extending the capabilities of the original SAM to handle both images and videos with unprecedented accuracy and efficiency. This curated list covers the rapidly expanding ecosystem of SAM2 applications across diverse domains - from medical imaging to robotics, from 3D reconstruction to video generation.

🔥 Why SAM2?

SAM2 has revolutionized segmentation tasks by:

  • Providing unified image and video segmentation capabilities
  • Enabling zero-shot generalization across domains
  • Offering efficient real-time processing
  • Supporting diverse prompting mechanisms (points, boxes, masks, text)

📈 Repository Stats: Currently tracking 500+ papers and projects across 15+ domains

🤝 Contributing: We continuously improve this collection. Feel free to submit PRs for missed works, corrections, or new categories!


This repo aims to include materials (papers, codes, slides) about SAM2 (segment anything in images and videos), a vision foundation model released by Meta AI . We are continuously improving the project. Welcome to PR the works (papers, repos) that are missed.

SAM2

📊 Repository Statistics

DomainPapersKey Highlights
🏥 Medical80+Surgery, 3D medical imaging, pathology
🎬 Video Processing70+Object tracking, temporal consistency
🤖 Robotics50+Manipulation, navigation, human-robot interaction
🛰️ Remote Sensing20+Satellite imagery, environmental monitoring
🎨 Generation/Editing35+Video synthesis, image editing, creative tools
🧊 3D Processing45+Point clouds, mesh processing, reconstruction
🎯 Core Segmentation60+Novel applications and improvements
Total500+Across 15+ domains

💡 Quick Navigation: Use Ctrl+F to search for specific keywords, or click on the Table of Contents links above for domain-specific papers.

Contents

Papers/Projects

📚 Surveys & Reviews

Comprehensive overviews and systematic analyses of SAM2 applications across various domains

🎯 Traditional Segmentation

Core image and video segmentation applications, including novel architectures and domain-specific adaptations

🖼️ Image Segmentation

Segmentation Applications
Other Image Tasks

🎬 Video Segmentation

Temporal segmentation, object tracking, and video understanding applications

Referring/Reasoning Video Object Segmentation

Other Video Tasks

ReleaseTitleCode
2025.03MMCD: Memory-Based Multimodal Change DetectionNA
2025.03EgoSplat: Open-Vocabulary Egocentric Scene Understanding with Language Embedded 3D Gaussian Splatting🕒Soon
2025.03Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene🌐Project page
2025.03EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining🔗 Code
2025.03High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flight🔗 Code
2025.03FusionSegReID: Advancing Person Re-Identification with Multimodal Retrieval and Precise SegmentationNA
2025.04Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting🔗 Code
2025.04How Can Objects Help Video-Language Understanding?🕒Soon
2025.04SAMJAM:Zero-Shot Video Scene Graph Generation for Egocentric Kitchen VideosNA
2025.05Research on a traffic flow statistical algorithm based on YBOVDT and SAM2📊 Data
2025.05One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object TrajectoryNA
2025.06ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer🔗 Code
2025.06Track Any Object:A Granular Video Anomaly Detection Pipeline🌐Project page
2025.06Open-World Object Counting in Videos🔗 Code
2025.06Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D AnnotationsNA
2025.06SAM2RL:Towards Reinforcement Learning Memory Control in Segment Anything Model 2NA
2025.07Visual tracking by matching points using diffusion model🔗 Code
2025.07Intelligent and quantitative ligament breakup event analysis in 65 kHz off-axis holographic video of swirl spray🔗 Code
2025.07Towards Blind Bitstream-corrupted Video Recovery: AVisual Foundation Model-driven FrameworkNA
2025.07SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking🔗 Code
2025.08Towards automated video-based human behavior analysis: leveraging AI capabilities for spatial behavior detectionNA

🔊 Audio-visual segmentation (AVS)

Multi-modal approaches combining audio and visual information for segmentation

📊 Graph Learning

Scene graph generation and graph-based reasoning with SAM2

🏥 Medical Domain

Healthcare applications including surgery, diagnostics, and biomedical research

Medical Video & 3D Segmentation

ReleaseTitleCode
2024.08Segment anything in medical images and videos: Benchmark and deployment🔗 Code
2024.08SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video SegmentationNA
2024.08Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video SegmentationNA
2024.08Novel adaptation of video segmentation to 3D MRI: efficient zero-shot knee segmentation with SAM2NA
2024.08Biomedical SAM 2: Segment anything in biomedical images and videos🔗 Code
2024.08Polyp SAM 2: Advancing Zero-shot Polyp Segmentation in Colorectal Cancer Detection🔗 Code
2024.08Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning🔗 Code
2024.08Performance and Non-adversarial Robustness of the Segment Anything Model 2 in Surgical Video SegmentationNA
2024.09SAM-OCTA2: Layer Sequence OCTA Segmentation with Fine-tuned Segment Anything Model 2🔗 Code
2024.09Self-Prompting Polyp Segmentation in Colonoscopy using Hybrid Yolo-SAM 2 Model🔗 Code
2024.10A-MFST: Adaptive Multi-Flow Sparse Tracker for Real-Time Tissue Tracking Under OcclusionNA
2024.10ECHOPulse: ECG controlled echocardio-grams video generation🔗 Code
2024.11Phase-Informed Tool Segmentation for Manual Small-Incision Cataract SurgeryNA
2025.02SASVi - Segment Any Surgical Video🔗 Code
2025.02Text-Promtable propagation for referring medical image sequence segmentationNA
2025.02Less is More? Revisiting the Importance of Frame Rate in Real-Time Zero-Shot Surgical Video Segmentation
2025.03SurgiSAM2: Fine-tuning a foundational model for surgical video anatomy segmentation and detection🔗 Code(& dataset)
2025.03Surgical Gaussian Surfels: Highly Accurate Real-time Surgical Scene Rendering🔗 Code
2025.03Rethinking Few-Shot Medical Image Segmentation by SAM2: A Training-Free Framework with Augmentative Prompting and Dynamic MatchingNA
2025.03Self-Prompting Driven SAM2 for 3D Medical Image SegmentationNA
2025.04RP-SAM2: Refining Point Prompts for Stable Surgical Instrument Segmentation🔗 Code
2025.04Agglomerating Large Vision Encoders via Distillation for VFSS SegmentationNA
2025.04MedSAM2: Segment Anything in 3D Medical Images and Videos🔗 Code
2025.05Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos🕒Soon
2025.05Adapting Segment Anything 2 for Diabetic Retinopathy Lesion SegmentationNA
2025.07Beyond Rigid AI: Towards Natural Human-Machine Symbiosis for Interoperative Surgical AssistanceNA
2025.07Towards Affordable Tumor Segmentation and Visualization for 3D Breast MRI Using SAM2NA
2025.08Edge2Prompt: Modality-Agnostic Model for Out-of-Distribution Liver SegmentationNA
2025.08F2PASeg: Feature Fusion for Pituitary Anatomy Segmentation in Endoscopic Surgery🔗 Code
2025.08TSMS-SAM2: Multi-scale Temporal Sampling Augmentation and Memory-Splitting Pruning for Promptable Video Object Segmentation and Tracking in Surgical Scenarios🔗 Code
2025.08SAM2Med3D: Leveraging video foundation models for 3D breast MRI segmentationData Upon Request

Medical Image Segmentation

ReleaseTitleCode
2024.08SAM & SAM 2 in 3D Slicer: SegmentWithSAM Extension for Annotating Medical Images🔗 Code
2024.08SAM2-PATH: A better segment anything model for semantic segmentation in digital pathology🔗 Code
2024.08Is SAM 2 Better than SAM in Medical Image Segmentation?NA
2024.08A Short Review and Evaluation of SAM2's Performance in 3D CT Image Segmentation🔗 Code
2024.08Interactive 3D Medical Image Segmentation🔗 Code
2024.08SAM2-UNet: Segment Anything 2 Makes Strong Encoder for Natural and Medical Image Segmentation🔗 Code
2024.08SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More🔗 Code
2024.08Retrieval-augmented Few-shot Medical Image Segmentation with Foundation Models
2024.10SAM-Swin: SAM-Driven Dual-Swin Transformers with Adaptive Lesion Enhancement for Laryngo-Pharyngeal Tumor Detection🔗 Code
2024.11A multi-task learning model for clinically interpretable sesamoiditis gradingNA
2024.11Zero-shot capability of SAM-family models for bone segmentation in CT scansNA
2024.11SAM-I2I: Unleash the Power of Segment Anything Model for Medical Image TranslationNA
2024.12Medical SAM 2: Segment Medical Images As Video Via Segment Anything Model 2🔗 Code
2025.03Self-Prompting Driven SAM2 for 3D Medical Image SegmentationNA
2025.03WeakMedSAM: Weakly-Supervised Medical Image Segmentation via SAM with Sub-Class Exploration and Prompt Affinity Mining🔗 Code
2025.03Research on recognition of diabetic retinopathy hemorrhage lesions based on fine tuning of segment anything modelNA
2025.04HRMedSeg: Unlocking High-resolution Medical Image segmentation via Memory-efficient Attention Modeling🔗 Code
2025.04Prompt Once, Segment Everything: Leveraging SAM 2 Potential for Infinite Medical Image Segmentation with a Single Prompt🔗 Code
2025.04Semi-automated segmentation of magnitude images in 4D flow MR scans using segment anything model 2 (SAM 2)NA
2025.05ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-term Tracking🔗 Code
2025.06MorphSAM: Learning the Morphological Prompts from Atlases for Spine Image SegmentationNA
2025.06SynPo: Boosting Training-Free Few-Shot Medical Segmentation via High-Quality Negative Prompts🌐Project page
2025.06Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning🔗 Code
2025.07Speckle2Self: Self-Supervised Ultrasound Speckle Reduction Without Clean Data🕒Soon
2025.08 Training-Free Breast Ultrasound Image Segmentation with Retrieval-based SAM2NA
2025.08CXR-ODDet: An Omni-Decoupled Multi-ClassLesion Localization Framework for Automatic ChestX-Ray AnalysisNA
2025.08LGFFM: A Localized and Globalized Frequency Fusion Model for Ultrasound Image Segmentation🔗 Code
2025.04Semi-automated segmentation of magnitude images in 4D flow MR scans using segment anything model 2 (SAM 2)NA
2025.09Co-Seg: Mutual Prompt-Guided Collaborative Learning for Tissue and Nuclei Segmentation🔗 Code

Other Medical Applications

ReleaseTitleCode
2025.03Flip Learning: Weakly supervised erase to segment nodules in breast ultrasoundNA
2025.03From Monocular Vision to Autonomous Action:Guiding Tumor Resection via 3D ReconstructionNA
2025.03Operating Room Workflow Analysis via Reasoning Segmentation over Digital TwinsNA
2025.03Early Detection and Classification of Lung Cancer using Segment Anything Model 2 and Dense NetNA
2025.04Zero-Shot 4D Lidar Panoptic SegmentationNA
2025.04SYNTHFM: Training Modality-Agnostic Foundation Models for Medical Image Segmentation Without Real Medical DataNA
2025.04VoxelFeat: Voxel-wise foundation model featuresNA
2025.06Leadership Assessment in Pediatric Intensive Care Unit Team TrainingNA
2025.08GM-ABS: Promptable Generalist Model Drives Active Barely Supervised Training in Specialist Model for 3D Medical Image Segmentation🔗 Code
2025.08Transgene-free generation of mouse post-gastrulation whole embryo models solely from naive ESCs and iPSCsNA
2025.09A Machine Learning Assisted Tool and Numerical Model for Analyzing Lipid Nanoparticles🔗 Code
2025.09Evaluating the Efficacy of Mebendazole Repurposing for Ovarian Cancer Therapy Using Optical Coherence TomographyNA
2025.09Evaluation of Radio Frequency Ablation in Human Left Atrial Tissues for Atrial Fibrillation Using Optical Coherence TomographyNA
2025.09Abundance of Maternal Mitochondrial Genome Is Dispensable up to the Mitochondrial Genome Activation in Post-Implantation EmbryosNA
2025.09The Evaluation of a Deep Learning Approach to Automatic Segmentation of Teeth and Shade Guides for Tooth Shade Matching Using the SAM2 AlgorithmNA
2025.09Neck-focused Remote Photoplethysmography (rPPG): A comparative study using clinical data and the PyVHR frameworkNA

🎭 Camouflaged Object Detection (COD)

Detecting and segmenting objects that blend with their surroundings

Video COD

Image COD

🛰️ Remote Sensing

Satellite imagery analysis, environmental monitoring, and geospatial applications

ReleaseTitleCode
2024.11DED-SAM: Adapting Segment Anything Model 2 for Dual Encoder-Decoder Change DetectionNA
2025.01Prompt-Based Segmentation at Multiple Resolutions and Lighting Conditions using Segment Anything Model 2NA
2025.03Customized SAM 2 for Referring Remote Sensing Image SegmentationNA
2025.05InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition🔗 Code
2025.06Baltimore Atlas: FreqWeaver Adapter for Semi-supervised Ultra-high Spatial Resolution Land Cover ClassificationNA
2025.06Bundle adjustment for multi-source Mars orbiter imagery with generalized control constraintsNA
2025.07Leveraging SAM 2 and LiDAR for Automated Individual Tree Crown Delineation: A Comparative Evaluation of Prompting Methods🔗 Code
2025.07Aerial Visual Localization over Low Level-of-Detail City Models using Explicit Silhouette Alignment🔗 Code
2025.07CSW-SAM: a cross-scale algorithm for very-high-resolution water body segmentation based on segment anything model 2NA
2025.07A Fine Agricultural Flood Segmentation Model For HJ-2E S-band SAR DataNA
2025.08DeH4R: A Decoupled and Hybrid Method for Road Network Graph Extraction🔗 Code
2025.09PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection🔗 Code
2025.09An Automatic Sample Augmentation Method for Paddy Rice Mapping Based on Segment Anything Model and Phenological Features—A Case Study in Southwest ChinaNA
2025.09SOPSeg: Prompt-based Small Object Instance Segmentation in Remote SensingImageryNA
2025.09BiSAM-CD: Zero-Shot Remote Sensing Change Detection via Bidirectional Temporal Memory in SAM2🔗 Code
2025.10SinkSAM-Net: Knowledge-driven self-supervised sinkhole segmentation using topographic priors and Segment Anything Model🌐Project page

🧊 3D Processing & Point Clouds

Three-dimensional data processing, reconstruction, and analysis

3D Segmentation

3D Reconstruction

Other 3D Applications

ReleaseTitleCode
2025.03LP-Gaussians: Learnable Parametric Gaussian Splatting for Efficient Dynamic Reconstruction of Single-View Scenes🌐Project page
2025.03DecoupledGaussian: Object-Scene Decoupling for Physics-Based Interaction🌐Project page
2025.03Free Your Hands: Lightweight Relightable Turntable Capture PipelineNA
2025.03WildSeg3D: Segment Any 3D Objects in the Wild from 2D Images🕒Soon
2025.03Pseudo-LiDAR With Two-Dimensional Instance for Monocular Three-Dimensional Object TrackingNA
2025.03SED-MVS: Segmentation-Driven and Edge-Aligned Deformation Multi-View Stereo with Depth Restoration and Occlusion ConstraintNA
2025.03SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining🔗 Code
2025.03COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian Splitting🔗 Code
2025.03Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields🌐Project page
2025.03Semantic Consistent Language Gaussian Splatting for Point-Level Open-vocabulary Querying🌐Project page
2025.04Augmented Unseen Region Alignment for Reference-based 360° Unbounded Scene Inpainting🌐Project page
2025.04 FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive AgentsNA
2025.06 GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation🌐Project page
2025.05Constructing a 3D Town from a Single Image🌐Project page
2025.06GenMOJO: Robust Multi-Object 4D Generation for In-the-wild Videos🌐Project page
2025.06CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image🌐Project page
2025.06BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing🌐Project page
2025.07LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion🌐Project page
2025.07 Consistent Bokeh for Multi-View Images With 3D Gaussian Splatting🔗 Code
2025.07Defect segmentation and 3D reconstruction in concrete structures using SAM 2 and 3D Gaussian splattingUpon Request
2025.07Image-Guided Shape-from-Template Using Mesh Inextensibility Constraints🔗 Code
2025.07MG-Mono: A Lightweight Multi-Granularity Method for Self-Supervised Monocular Depth Estimation🔗 Code
2025.08Training-free automatic instance segmentation of girder bridge point cloud via large model fusion with reverse entity modelling verificationNA
2025.08SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass🌐Project page

🎨 Image or Video Generation & Editing

Creative applications including content generation, editing, and artistic tools

ReleaseTitleCode
2024.10AdaptiveDrag: Semantic-Driven Dragging on Diffusion-Based Image Editing🔗 Code
2024.11VideoDirector: Precise Video Editing via Text-to-Video Models🌐Project page
2024.11VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing🌐Project page
2024.11Generative Omnimatte: Learning to Decompose Video into Layers🌐Project page
2024.12InterDyn: Controllable Interactive Dynamics with Video Diffusion Models🌐Project page
2025.01MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis🌐Project page
2025.01BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations🌐Project page
2025.03TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video SynthesisNA
2025.03Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias🔗 Code
2025.03Unified Dense Prediction of Video DiffusionNA
2025.03DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image🕒Soon
2025.03FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing🌐Project page
2025.03MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance🌐Project page
2025.03Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting🌐Project page
2025.03Multi-Subject and Motion Customization of Text-to-Video Diffusion Models🌐Project page
2025.04DreamFuse: Adaptive Image Fusion with Diffusion Transformer🌐Project page
2025.04Enhanced Semantic Extraction and Guidance for UGC Image Super Resolution🔗 Code
2025.06 Keyframe-Guided Creative Video Inpainting🌐Project page
2025.06OmniGen2: Exploration to Advanced Multimodal Generation🌐Project page
2025.07Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning🔗 Code
2025.07Enhanced Velocity Field Modeling for Gaussian Video ReconstructionNA
2025.08NEP: Autoregressive Image Editing via Next Editing Token Prediction🌐Project page

🗺️ Simultaneous Localization and Mapping (SLAM / VO)

Navigation, mapping, and localization applications

💡 Light Field Segmentation

Advanced imaging techniques and multi-dimensional visual processing

🤖 Robotics

Autonomous systems, manipulation, navigation, and human-robot interaction

ReleaseTitleCode
2024.10A Pipeline for Segmenting and Structuring RGB-D Data for Robotics ApplicationsNA
2024.10VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model🌐Project page
2025.02Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos🌐Project page
2025.02Map Space Belief Prediction for Manipulation-Enhanced Mapping(To be released)
2025.02Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids🌐Project page
2025.03DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping🌐Project page
2025.03Autonomous Dissection in Robotic CholecystectomyNA
2025.03MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model🌐Project page
2025.03LuciBot: Automated Robot Policy Learning from Generated Videos🌐Project page
2025.03IMPACT : Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models🌐Project page
2025.03VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and InvisibilityNA
2025.03ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis🌐Project page
2025.04Slot-Level Robotic Placement via Visual Imitation from Single Human Video🌐Project page
2025.04Entangled chip removal utilizing mass-spring model with mobile manipulatorNA
2025.05Symbolically-Guided Visual Plan Inference from Uncurated Video DataNA
2025.05Geometry-Consistent Video Diffusion for Robotic Visual Policy Transfer🌐Project page
2025.05Grasp the Invisibility by Vision-Language guided Active View PlanningNA
2025.07Geometry-aware 4D Video Generation for Robot Manipulation🌐Project page
2025.07Object-Centric Mobile Manipulation through SAM2-Guided Perception and Imitation LearningNA
2025.07GraspGen: A Diffusion-based Framework for 6-DOF Grasping🌐Project page
2025.07RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping🔗 Code
2025.08Train Once, Deploy Anywhere: Realizing Data-Efficient Dynamic Object Manipulation🔗 Code
2025.09ObjectReact: Learning Object-Relative Control for Visual Navigation🌐Project page

⚡ Adaptation, Compression & Edge Applications

Efficiency optimizations, model compression, and deployment on resource-constrained devices

📖 Training

Resources for model training, datasets, and learning frameworks

Datasets

Curated datasets and benchmarks for SAM2 training and evaluation

ReleaseTitleCode
2025.02SurgPose: a Dataset for Articulated Robotic Surgical Tool Pose Estimation and Tracking🌐Project page
2025.02The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionNA
2025.02Picking the Cream of the Crop:Visual-Centric Data Selection with Collaborative Agents🔗 Code
2025.03Phantom: Training Robots Without Robots Using Only Human Videos🌐Project page
2025.03Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups🌐Project page
2025.03What Are You Doing? A Closer Look at Controllable Human Video Generation🔗 Code
2025.03Instrument-Splatting: Controllable Photorealistic Reconstruction of Surgical Instruments Using Gaussian Splatting🕒Soon
2025.03Referring to Any Person🌐Project page
2025.03AUTV: Creating Underwater Video Datasets with Pixel-wise AnnotationsNA
2025.03DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios🌐Project page
2025.05A fusion network for multi-modality medical image registration with progressive feature alignment🔗 Code
2025.04InteractAvatar: Modeling Hand-Face Interaction in Photorealistic Avatars with Deformable GaussiansNA
2025.04VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing🌐Project page
2025.04UrbanWaste: In-the-Bin Dataset for Waste Disposal Inspection with Multi-Granularity Hierarchical Labels🌐Project page
2025.06HD-EPIC: A Highly-Detailed Egocentric Video Dataset🌐Project page
2025.06 GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities🌐Project page
2025.06INTERNSPATIAL: A Comprehensive Dataset for Spatial Reasoning in Vision-Language ModelsNA
2025.06BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos🌐Project page
2025.06SAM4D:Segment Anything in Camera and LiDAR Streams🌐Project page
2025.06XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation🌐Project page
2025.07A New Dataset and Performance Benchmark for Real-time Spacecraft Segmentation in Onboard Flight Computers🔗 Code
2025.07Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation🌐Project page
2025.08MOSE: Complex Video Object Segmentation Dataset🌐Project page
2025.08DreamVE: Unified Instruction-based Image and Video Editing🌐Project page
2025.09SPATIALVID: A Large-scale Video Dataset with Spatial Annotations🌐Project page

🔄 Used for Data Augmentation (/Tool)

Tools and methods for data synthesis, augmentation, and preprocessing

ReleaseTitleCode
2025.02Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach🌐Project page
2025.03A Taxonomy for Evaluating Generalist Robot Policies🌐Project page
2025.03CRESTE: Scalable Mapless Navigation with Internet Scale Priors and Counterfactual Guidance🌐Project page
2025.03Shaken, Not Stirred: A Novel Dataset for Visual Understanding of Glasses in Human-Robot Bartending Tasks🌐Project page
2025.03YOLOE: Real-Time Seeing Anything🔗 Code
2025.03VACE: All-in-One Video Creation and Editing🌐Project page
2025.03V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video🌐Project page
2025.03Better Together: Unified Motion Capture and 3D Avatar ReconstructionNA
2025.03CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based GuidanceNA
2025.03RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance 🌐Project page
2025.03Evaluating the FLUX.1 Synthetic Data on YOLOv9 for AI-Powered Poultry FarmingNA
2025.03Any2Caption : Interpreting Any Condition to Caption for Controllable Video Generation🌐Project page
2025.05LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration🔗 Code
2025.04VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning🌐Project page
2025.04M2Flow: A Motion Information Fusion Framework for Enhanced Unsupervised Optical Flow Estimation in Autonomous DrivingNA
2025.05Interspatial Attention for Efficient 4D Human Video Generation🌐Project page
2025.06Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting GarmentsNA
2025.06Impact of Synthetic Data from Diffusion Models on Weed Detection PerformanceNA
2025.06VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models🌐Project page
2025.06Building Software for Analyzing Muck Piles After Blasting in Laboratory Conditions with Integrated Artificial IntelligenceNA
2025.06WeedSwin hierarchical vision transformer with SAM-2 for multi-stage weed detection and classificationOn Request
2025.07Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data🔗 Code
2025.07 RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding🕒Soon

🛠️ Training Helper

Supporting tools and frameworks for model training and fine-tuning

📊 Performance Evaluations

Benchmarking, evaluation metrics, and comparative studies

Post Processing

🛡️ Robustness

Security, adversarial robustness, and reliability studies

🌟 Unique Applications/Usage

Novel and creative applications that don't fit traditional categories

ReleaseTitleCode
2024.09Towards Robust Automation of Surgical Systems via Digital Twin-based Scene Representations from Foundation ModelsNA
2024.09Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models🔗 Code
2024.09Point of Interest Recognition and Tracking in Aerial Video during Live Cycling BroadcastsNA
2024.10ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting🔗 Code
2024.10GRS: Generating Robotic Simulation Tasks from Real-World ImagesNA
2024.10Iterative Optimization Annotation Pipeline and ALSS-YOLO-Seg for Efficient Banana Plantation Segmentation in UAV ImageryNA
2024.10Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting🔗 Code
2025.01Zero-Shot Pupil Segmentation with SAM 2: A Case Study of Over 14 Million Images📊Data
2025.02Best Foot Forward: Robust Foot Reconstruction in-the-wild
2025.03ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment🌐Project page
2025.03JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse🌐Project page
2025.04MORPHEUS: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments🌐Project page
2025.05Air-Ground Collaboration for Language-Specified Missions in Unknown EnvironmentsNA
2025.06In Situ Detection and Measurement of Broccoli Heads Under Different Lighting Conditions Using Proximal Remote SensingNA
2025.07 Zero-Shot Recognition of Test Tube Types by Automatically Collecting and Labeling RGB DataNA
2025.07Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse AutomationNA
2025.07Phys2Real: Physically-Informed Gaussian Splatting for Adaptive Sim-to-Real Transfer in Robotic ManipulationNA

🤝 Contributing

We welcome contributions from the community! Here's how you can help improve this repository:

How to Contribute

  1. Submit a Pull Request with new papers/projects
  2. Report Issues for broken links or incorrect information
  3. Suggest New Categories for emerging SAM2 applications
  4. Improve Organization by suggesting better categorization

Contribution Guidelines

  • Format: Follow the existing table format with Release Date | Title | Code links
  • Quality: Include only peer-reviewed papers or significant projects
  • Recency: Focus on 2024+ publications (SAM2 era)
  • Completeness: Provide accurate metadata and working links when available
  • Categories: Place papers in the most appropriate domain category

Adding New Papers

| YYYY.MM | [Paper Title](paper_link) | [🔗 Code](code_link) / [🌐Project page](project_link) / NA |
  • 🔗 [🔗 Code] - Source code repositories
  • 🌐 [🌐Project page] - Official project websites
  • 📊 [📊Data] - Datasets and benchmarks
  • 🖥️ [🖥️Demo] - Interactive demonstrations
  • 📖 [📖Repo] - Documentation repositories
  • 🕒 🕒Soon - Coming soon
  • NA - Not available

📈 Repository Metrics

  • Total Papers: 500+ and growing
  • Domains Covered: 15+ major application areas
  • Time Range: 2024-2025 (SAM2 era)
  • Update Frequency: Weekly additions
  • Community: Open for contributions

🎉 Acknowledgments

Special thanks to:

  • Meta AI for releasing SAM2 and advancing the field
  • Research Community for the incredible pace of SAM2 adoption and innovation
  • Contributors who help maintain and improve this collection
  • Reviewers who ensure quality and accuracy

📄 License

This collection is maintained under the MIT License. Individual papers and projects retain their original licenses.


⭐ Star this repository if you find it helpful! ⭐
Last updated: August 2025 | Maintained by the Community