KyanChen/Awesome-Referring-Remote-Sensing-Image-Segmentation

Awesome Referring Remote Sensing Image Segmentation

38

0 commits

updated Sep 25, 2025

See the code

README

Awesome Referring Remote Sensing Image Segmentation

Contributions Welcome

:loudspeaker: Call for Contribution

We actively welcome:

  • :page_facing_up: New papers (CVPR/ICCV/ECCV/JSTARS/TGRS/etc.)
  • :computer: Open-source implementations
  • :bar_chart: Dataset annotations
  • :bookmark_tabs: Technical summaries

:book: Table of Contents

  1. Technical Background
  2. Latest Papers
  3. Datasets
  4. Codebases
  5. Evaluation

:rocket: Technical Background

Referring Remote Sensing Image Segmentation (RRSIS) combines natural language descriptions with pixel-level understanding of aerial or satellite imagery. Key challenges include:

  • Multiscale Objects: Buildings (1-100m) vs Vehicles (2-5m)
  • Sensor Variations: 0.3m (WorldView) to 10m (Sentinel-2)
  • Linguistic Ambiguity: "The red-roofed building northeast of the river"
  • Occlusion Complexity: Partial visibility due to clouds/shadows (avg. 15-30% coverage in satellite imagery)
  • Temporal Dynamics: Seasonal changes (e.g. snow cover) affecting object appearance
  • Geometric Distortion: Parallax errors up to 5-10 pixels in oblique aerial views

:newspaper: Latest Papers

Segmentation Models

MethodYearVenueTitleCode
SAARN2025ArxivRIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation:computer: Code
RSRefSeg 22025ArxivRSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models:computer: Code
RRSECS2025GRSMRRSECS: Referring remote sensing expression comprehension and segmentation:computer: Code
MRSNet2025ArxivA Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark:computer: Code
LSCF2025TGRSLSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image Segmentation-
MPBF2025RSMultimodal Prompt-Guided Bidirectional Fusion for Referring Remote Sensing Image Segmentation-
DiffRIS2025ArxivDiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models-
CADFormer2025JSTARSCADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation:computer: Code
RS2-SAM 22025ArxivCustomized SAM 2 for Referring Remote Sensing Image Segmentation-
PSLGSAM2025ArxivSemantic Localization Guiding Segment Anything Model For Reference Remote Sensing Image Segmentation-
SegEarth-R12025ArxivSegEarth-R1: Geospatial Pixel Reasoning via Large Language Model:computer: Code
MAFN2025GRSLMultimodal-Aware Fusion Network For Referring Remote Sensing Image Segmentation:computer: Code
RS2-SAM 22025ArxivCustomized SAM 2 for Referring Remote Sensing Image Segmentation-
BTDNet2025ArxivReferring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction:computer: Code
AeroReformer2025ArxivAerial Referring Transformer for UAV-based Referring Image Segmentation:computer: Code
RSRefSeg2025IGARSSReferring Remote Sensing Image Segmentation with Foundation Models:computer: Code
SBANet2025ArxivScale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation-
RSSep2024ACCVWRSSep: Sequence-to-Sequence Model for Simultaneous Referring Remote Sensing Segmentation and Detection-
CroBIM2024ArxivCross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation:computer: Code
DANet2024ACMMMRethinking the Implicit Optimization Paradigm with Dual Alignments for Referring Remote Sensing Image Segmentation-
-2024IGARSSReferring Image Segmentation for Remote Sensing Data-
RMSIN2024CVPRRotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation:computer: Code
FIANet2024TGRSExploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation:computer: Code
LGCE2024TGRSRRSIS: Referring Remote Sensing Image Segmentation:computer: Code

:floppy_disk: Datasets

YearDatasetSizeDownload Links
2025RIS-LAD13,871 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2025EarthReason30,000 image-question-mask-answer quadruples:page_facing_up: Paper | :floppy_disk: Data
2025RefDIOR38,320 image-caption-box-mask quadruples:page_facing_up: Paper | :floppy_disk: Data
2025NWPU-Refer49,745 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2024RISBench52,472 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2024RRSIS-D17,402 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2024RefSegRS4,420 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data

:computer: Codebases

FrameworkLanguageStarsFeaturesReference Code
MMSegmentationPythonGitHub StarsMulti-model support
RSIS pipelines
Pre-trained models
:rocket: Demo

:chart_with_upwards_trend: Evaluation

Core Metrics

Basic IoU (Intersection over Union)

$$\text{IoU} = \frac{|P \cap G|}{|P \cup G|}$$

  • Where ( P ) = Predicted segmentation, ( G ) = Ground truth
  • Fundamental measure for pixel-wise segmentation accuracy

Generalized IoU (gIoU)

$$\text{gIoU} = \frac{1}{N} \sum_{i=1}^{N} \text{IoU}_i$$

  • Calculated by averaging IoU scores across all test samples
  • Reflects average performance on individual instances
  • Primary metric (more robust to outliers)

Cumulative IoU (cIoU)

$$\text{cIoU} = \frac{\sum_{i=1}^{N} |P_i \cap G_i|}{\sum_{i=1}^{N} |P_i \cup G_i|}$$

  • Computes ratio of cumulative intersections to unions
  • Sensitive to large target areas (higher variance)

Precision@X (Pr@X)

ThresholdDefinitionTypical Value
Pr@0.5% samples with IoU > 50%65.2%
Pr@0.6% samples with IoU > 60%48.7%
Pr@0.7% samples with IoU > 70%32.1%
Pr@0.8% samples with IoU > 80%15.6%
Pr@0.9% samples with IoU > 90%5.3%

Implementation

Standard metric implementation reference:
iou_metrics.py based on TorchMetrics.


:construction: Project under active development - Contribution Guidelines

KyanChen/Awesome-Referring-Remote-Sensing-Image-Segmentation

Awesome Referring Remote Sensing Image Segmentation

38

0 commits

updated Sep 25, 2025

See the code

README

Awesome Referring Remote Sensing Image Segmentation

Contributions Welcome

:loudspeaker: Call for Contribution

We actively welcome:

  • :page_facing_up: New papers (CVPR/ICCV/ECCV/JSTARS/TGRS/etc.)
  • :computer: Open-source implementations
  • :bar_chart: Dataset annotations
  • :bookmark_tabs: Technical summaries

:book: Table of Contents

  1. Technical Background
  2. Latest Papers
  3. Datasets
  4. Codebases
  5. Evaluation

:rocket: Technical Background

Referring Remote Sensing Image Segmentation (RRSIS) combines natural language descriptions with pixel-level understanding of aerial or satellite imagery. Key challenges include:

  • Multiscale Objects: Buildings (1-100m) vs Vehicles (2-5m)
  • Sensor Variations: 0.3m (WorldView) to 10m (Sentinel-2)
  • Linguistic Ambiguity: "The red-roofed building northeast of the river"
  • Occlusion Complexity: Partial visibility due to clouds/shadows (avg. 15-30% coverage in satellite imagery)
  • Temporal Dynamics: Seasonal changes (e.g. snow cover) affecting object appearance
  • Geometric Distortion: Parallax errors up to 5-10 pixels in oblique aerial views

:newspaper: Latest Papers

Segmentation Models

MethodYearVenueTitleCode
SAARN2025ArxivRIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation:computer: Code
RSRefSeg 22025ArxivRSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models:computer: Code
RRSECS2025GRSMRRSECS: Referring remote sensing expression comprehension and segmentation:computer: Code
MRSNet2025ArxivA Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark:computer: Code
LSCF2025TGRSLSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image Segmentation-
MPBF2025RSMultimodal Prompt-Guided Bidirectional Fusion for Referring Remote Sensing Image Segmentation-
DiffRIS2025ArxivDiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models-
CADFormer2025JSTARSCADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation:computer: Code
RS2-SAM 22025ArxivCustomized SAM 2 for Referring Remote Sensing Image Segmentation-
PSLGSAM2025ArxivSemantic Localization Guiding Segment Anything Model For Reference Remote Sensing Image Segmentation-
SegEarth-R12025ArxivSegEarth-R1: Geospatial Pixel Reasoning via Large Language Model:computer: Code
MAFN2025GRSLMultimodal-Aware Fusion Network For Referring Remote Sensing Image Segmentation:computer: Code
RS2-SAM 22025ArxivCustomized SAM 2 for Referring Remote Sensing Image Segmentation-
BTDNet2025ArxivReferring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction:computer: Code
AeroReformer2025ArxivAerial Referring Transformer for UAV-based Referring Image Segmentation:computer: Code
RSRefSeg2025IGARSSReferring Remote Sensing Image Segmentation with Foundation Models:computer: Code
SBANet2025ArxivScale-wise Bidirectional Alignment Network for Referring Remote Sensing Image Segmentation-
RSSep2024ACCVWRSSep: Sequence-to-Sequence Model for Simultaneous Referring Remote Sensing Segmentation and Detection-
CroBIM2024ArxivCross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation:computer: Code
DANet2024ACMMMRethinking the Implicit Optimization Paradigm with Dual Alignments for Referring Remote Sensing Image Segmentation-
-2024IGARSSReferring Image Segmentation for Remote Sensing Data-
RMSIN2024CVPRRotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation:computer: Code
FIANet2024TGRSExploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation:computer: Code
LGCE2024TGRSRRSIS: Referring Remote Sensing Image Segmentation:computer: Code

:floppy_disk: Datasets

YearDatasetSizeDownload Links
2025RIS-LAD13,871 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2025EarthReason30,000 image-question-mask-answer quadruples:page_facing_up: Paper | :floppy_disk: Data
2025RefDIOR38,320 image-caption-box-mask quadruples:page_facing_up: Paper | :floppy_disk: Data
2025NWPU-Refer49,745 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2024RISBench52,472 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2024RRSIS-D17,402 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data
2024RefSegRS4,420 image-caption-mask triplets:page_facing_up: Paper | :floppy_disk: Data

:computer: Codebases

FrameworkLanguageStarsFeaturesReference Code
MMSegmentationPythonGitHub StarsMulti-model support
RSIS pipelines
Pre-trained models
:rocket: Demo

:chart_with_upwards_trend: Evaluation

Core Metrics

Basic IoU (Intersection over Union)

$$\text{IoU} = \frac{|P \cap G|}{|P \cup G|}$$

  • Where ( P ) = Predicted segmentation, ( G ) = Ground truth
  • Fundamental measure for pixel-wise segmentation accuracy

Generalized IoU (gIoU)

$$\text{gIoU} = \frac{1}{N} \sum_{i=1}^{N} \text{IoU}_i$$

  • Calculated by averaging IoU scores across all test samples
  • Reflects average performance on individual instances
  • Primary metric (more robust to outliers)

Cumulative IoU (cIoU)

$$\text{cIoU} = \frac{\sum_{i=1}^{N} |P_i \cap G_i|}{\sum_{i=1}^{N} |P_i \cup G_i|}$$

  • Computes ratio of cumulative intersections to unions
  • Sensitive to large target areas (higher variance)

Precision@X (Pr@X)

ThresholdDefinitionTypical Value
Pr@0.5% samples with IoU > 50%65.2%
Pr@0.6% samples with IoU > 60%48.7%
Pr@0.7% samples with IoU > 70%32.1%
Pr@0.8% samples with IoU > 80%15.6%
Pr@0.9% samples with IoU > 90%5.3%

Implementation

Standard metric implementation reference:
iou_metrics.py based on TorchMetrics.


:construction: Project under active development - Contribution Guidelines