A Benchmark and Awesome Collection of Methods for Remote Sensing Image-Text Retrieval (RSITR)๏ฝ Remote Sensing Cross-Model Retrieval (RSCMR) | Remote Sensing Vision-Lanuage Models (RSVLMs)
70
28 commits
updated Mar 10, 2025
A Benchmark and Awesome Collection of Methods for Remote Sensing Image-Text Retrieval (RSITR) ๏ฝ Remote Sensing Cross-model Retrieval (RSCMR) from the Internet, if there are any omissions, please contact me jiancheng.pan.plus@gmail.com.
Record the major news of RSVLMs community.
Collect the more popular image-text pairs datasets on remote sensing, and welcome contact for additions if there are more.
| Dataset Name | No. of images | Image Resolution | Vision-Lanuage Model |
|---|---|---|---|
| UCM-Captions | 613 | 256โรโ256 | - |
| Sydney-Captions | 2,100 | 500โรโ500 | - |
| RSICD | 10,921 | 224 ร 224 | - |
| RSITMD | 4,743 | 256 ร 256 | - |
| NWPU-Captions | 31,500 | 256 ร 256 | - |
| RET-3, SEG-4, DET-10 | - | All Resolutions | RemoteCLIP |
| RS5M | 5 million+ | All Resolutions | GeoRSCLIP |
| SkyScript | 5.2 million+ | All Resolutions | SkyCLIP |
Welcome to add more RSITR | RSCMR methods.
๐ Cross-Modal Retrieval on RSICD:
๐ [Papers with Code]
๐ Cross-Modal Retrieval on RSITMD:
๐ [Papers with Code]
๐ Closed-Domain Method: Training and testing on a single dataset.
๐ Open-Domain Method: Using extra datasets for pre-training to gain more inter-domain knowledge.
โก๏ธ Hashing Method: Efficient retrieval on large-scale datasets becomes feasible.
[AAAI 2024] | SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing | [๐ Paper] [๐ GitHub]
[TGRS 2023] | RemoteCLIP: A Vision Language Foundation Model for Remote Sensing | [๐ Paper] [๐ GitHub]
[TGRS 2023] | RS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation Model | [๐ Paper] [๐ GitHub]
[ArXiv 2023] | RSGPT: A Remote Sensing Vision Language Model and Benchmark | [๐ Paper]
[TGRS 2023] | Parameter-Efficient Transfer Learning for Remote Sensing ImageโText Retrieval | [๐ Paper]
[ACMMM 2023] | A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval | [๐ Paper] [๐ GitHub]
[TGRS 2023] | Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval | [๐ Paper]
[Sensors 2023] | A Fine-Grained Semantic Alignment Method Specific to Aggregate Multi-Scale Information for Cross-Modal Remote Sensing Image Retrieval | [๐ Paper]
[Remote Sensing 2023] | A Fusion Encoder with Multi-Task Guidance for Cross-Modal TextโImage Retrieval in Remote Sensing | [๐ Paper]
[IGARSS 2023] | A Texture and Saliency Enhanced Image Learning Method For Cross-Modal Remote Sensing Image-Text Retrieval | [๐ Paper]
[IGARSS 2023] | A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge Distillation | [๐ Paper]
[TGRS 2023] | Knowledge-Aided Momentum Contrastive Learning for Remote-Sensing Image Text Retrieval | [๐ Paper]
[Mathematics 2023] | An End-to-End Framework Based on Vision-Language Fusion for Remote Sensing Cross-Modal Text-Image Retrieval | [๐ Paper]
[TGRS 2023] Hypersphere-based Remote Sensing Cross-Modal Text-Image Retrieval via Curriculum Learning | [๐ Paper]
[TGRS 2023] | Interacting-Enhancing Feature Transformer for Cross-Modal Remote-Sensing Image and Text Retrieval | [๐ Paper]
[ICMR 2023] | Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval | [๐ Paper] [๐ GitHub]
[CDCEO 2022] | Knowledge-Aware Cross-Modal Text-Image Retrieval for Remote Sensing Images | [๐ Paper]
[IGARSS 2022] | A transformer-based cross-modal image-text retrieval method using feature decoupling and reconstruction | [๐ Paper]
[INT J APPL EARTH OBS 2022] | MCRN: A Multi-source Cross-modal Retrieval Network for remote sensing | [๐ Paper]
[JSTARS 2022] | Multilanguage Transformer for Improved Text to Remote Sensing Image Retrieval | [๐ Paper]
[Applied Sciences 2022] | Contrasting Dual Transformer Architectures for Multi-Modal Remote Sensing Image Retrieval | [๐ Paper]
[TGRS 2022] | Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information | [๐ Paper] [๐ GitHub]
[TGRS 2021] | A Lightweight Multi-Scale Crossmodal Text-Image Retrieval Method in Remote Sensing | [๐ Paper]
[TGRS 2021] | Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval | [๐ Paper] [๐ GitHub]
[JSTARS 2021] | A Deep Semantic Alignment Network for the Cross-Modal Image-Text Retrieval in Remote Sensing | [๐ Paper]
[LGRS 2021] | Fusion-Based Correlation Learning Model for Cross-Modal Remote Sensing Image Retrieval | [๐ Paper]
[Remote Sensing 2020] | TextRS: Deep Bidirectional Triplet Network for Matching Text to Remote Sensing Images | [๐ Paper]
[JSTARS 2022] | Remote Sensing Cross-Modal Retrieval by Deep Image-Voice Hashing | [๐ Paper]
[ArXiv 2022] | Deep Unsupervised Contrastive Hashing for Large-Scale Cross-Modal Text-Image Retrieval in Remote Sensing | [๐ Paper]
[ICIP 2022] | An Unsupervised Cross-Modal Hashing Method Robust to Noisy Training Image-Text Correspondences in Remote Sensing | [๐ Paper]
A Benchmark and Awesome Collection of Methods for Remote Sensing Image-Text Retrieval (RSITR)๏ฝ Remote Sensing Cross-Model Retrieval (RSCMR) | Remote Sensing Vision-Lanuage Models (RSVLMs)
70
28 commits
updated Mar 10, 2025
A Benchmark and Awesome Collection of Methods for Remote Sensing Image-Text Retrieval (RSITR) ๏ฝ Remote Sensing Cross-model Retrieval (RSCMR) from the Internet, if there are any omissions, please contact me jiancheng.pan.plus@gmail.com.
Record the major news of RSVLMs community.
Collect the more popular image-text pairs datasets on remote sensing, and welcome contact for additions if there are more.
| Dataset Name | No. of images | Image Resolution | Vision-Lanuage Model |
|---|---|---|---|
| UCM-Captions | 613 | 256โรโ256 | - |
| Sydney-Captions | 2,100 | 500โรโ500 | - |
| RSICD | 10,921 | 224 ร 224 | - |
| RSITMD | 4,743 | 256 ร 256 | - |
| NWPU-Captions | 31,500 | 256 ร 256 | - |
| RET-3, SEG-4, DET-10 | - | All Resolutions | RemoteCLIP |
| RS5M | 5 million+ | All Resolutions | GeoRSCLIP |
| SkyScript | 5.2 million+ | All Resolutions | SkyCLIP |
Welcome to add more RSITR | RSCMR methods.
๐ Cross-Modal Retrieval on RSICD:
๐ [Papers with Code]
๐ Cross-Modal Retrieval on RSITMD:
๐ [Papers with Code]
๐ Closed-Domain Method: Training and testing on a single dataset.
๐ Open-Domain Method: Using extra datasets for pre-training to gain more inter-domain knowledge.
โก๏ธ Hashing Method: Efficient retrieval on large-scale datasets becomes feasible.
[AAAI 2024] | SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing | [๐ Paper] [๐ GitHub]
[TGRS 2023] | RemoteCLIP: A Vision Language Foundation Model for Remote Sensing | [๐ Paper] [๐ GitHub]
[TGRS 2023] | RS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation Model | [๐ Paper] [๐ GitHub]
[ArXiv 2023] | RSGPT: A Remote Sensing Vision Language Model and Benchmark | [๐ Paper]
[TGRS 2023] | Parameter-Efficient Transfer Learning for Remote Sensing ImageโText Retrieval | [๐ Paper]
[ACMMM 2023] | A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval | [๐ Paper] [๐ GitHub]
[TGRS 2023] | Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval | [๐ Paper]
[Sensors 2023] | A Fine-Grained Semantic Alignment Method Specific to Aggregate Multi-Scale Information for Cross-Modal Remote Sensing Image Retrieval | [๐ Paper]
[Remote Sensing 2023] | A Fusion Encoder with Multi-Task Guidance for Cross-Modal TextโImage Retrieval in Remote Sensing | [๐ Paper]
[IGARSS 2023] | A Texture and Saliency Enhanced Image Learning Method For Cross-Modal Remote Sensing Image-Text Retrieval | [๐ Paper]
[IGARSS 2023] | A Fast and Accurate Method for Remote Sensing Image-Text Retrieval Based On Large Model Knowledge Distillation | [๐ Paper]
[TGRS 2023] | Knowledge-Aided Momentum Contrastive Learning for Remote-Sensing Image Text Retrieval | [๐ Paper]
[Mathematics 2023] | An End-to-End Framework Based on Vision-Language Fusion for Remote Sensing Cross-Modal Text-Image Retrieval | [๐ Paper]
[TGRS 2023] Hypersphere-based Remote Sensing Cross-Modal Text-Image Retrieval via Curriculum Learning | [๐ Paper]
[TGRS 2023] | Interacting-Enhancing Feature Transformer for Cross-Modal Remote-Sensing Image and Text Retrieval | [๐ Paper]
[ICMR 2023] | Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval | [๐ Paper] [๐ GitHub]
[CDCEO 2022] | Knowledge-Aware Cross-Modal Text-Image Retrieval for Remote Sensing Images | [๐ Paper]
[IGARSS 2022] | A transformer-based cross-modal image-text retrieval method using feature decoupling and reconstruction | [๐ Paper]
[INT J APPL EARTH OBS 2022] | MCRN: A Multi-source Cross-modal Retrieval Network for remote sensing | [๐ Paper]
[JSTARS 2022] | Multilanguage Transformer for Improved Text to Remote Sensing Image Retrieval | [๐ Paper]
[Applied Sciences 2022] | Contrasting Dual Transformer Architectures for Multi-Modal Remote Sensing Image Retrieval | [๐ Paper]
[TGRS 2022] | Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information | [๐ Paper] [๐ GitHub]
[TGRS 2021] | A Lightweight Multi-Scale Crossmodal Text-Image Retrieval Method in Remote Sensing | [๐ Paper]
[TGRS 2021] | Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval | [๐ Paper] [๐ GitHub]
[JSTARS 2021] | A Deep Semantic Alignment Network for the Cross-Modal Image-Text Retrieval in Remote Sensing | [๐ Paper]
[LGRS 2021] | Fusion-Based Correlation Learning Model for Cross-Modal Remote Sensing Image Retrieval | [๐ Paper]
[Remote Sensing 2020] | TextRS: Deep Bidirectional Triplet Network for Matching Text to Remote Sensing Images | [๐ Paper]
[JSTARS 2022] | Remote Sensing Cross-Modal Retrieval by Deep Image-Voice Hashing | [๐ Paper]
[ArXiv 2022] | Deep Unsupervised Contrastive Hashing for Large-Scale Cross-Modal Text-Image Retrieval in Remote Sensing | [๐ Paper]
[ICIP 2022] | An Unsupervised Cross-Modal Hashing Method Robust to Noisy Training Image-Text Correspondences in Remote Sensing | [๐ Paper]