lzw-lzw/awesome-remote-sensing-vision-language-models

Awesome-Remote-Sensing-Vision-Language-Models

195

68 commits

updated Apr 27, 2024

See the code

README

Awesome remote sensing vision language models

This is a repository for visual language models in remote sensing, including advanced methods and commonly used datasets in different applications, such as image-text retrieval, visual question answering, pretraining, etc.

If you find any relevant papers that are not included here, please feel free to pull requests at any time.

PRs Welcome

Table of Contents

Surveys

Remote Sensing Vision Language Model

Applications

Pretraining

Image Captioning

PaperPublished inCode/Project
Deep Semantic Understanding of High Resolution Remote Sensing ImageCITS 2016-
Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?TGRS 2017-
Exploring models and data for remote sensing image caption generationTGRS 2017code
Natural language escription of remote sensing images based on deep learningIGARSS 2017-
Description Generation for Remote Sensing Images Using Attribute Attention MechanismRemote Sensing 2019-
Vaa:Visual aligning attention model for remote sensing image captioningIEEE Access 2019-
Exploring Multi-Level Attention and Semantic Relationship for Remote Sensing Image CaptioningIEEE Access 2019-
A multi-level attention model for remote sensing image captionsRemote Sensing 2020-
Remote sensing image captioning via variational autoencoder and reinforcement learningKnowledge-Based Systems 2020-
Truncation cross entropy loss for remote sensing image captioninTGRS 2020-
Word–Sentence Framework for Remote Sensing Image CaptioningTGRS 2020code
A novel SVM-based decoder for remote sensing image captioningTGRS 2021-
High-resolution remote sensing image captioning based on structured attentionTGRS 2021code
Exploring transformer and multilabel classification for remote sensing image captioningGRSL 2022-
NWPU-captions dataset and mlca-net for remote sensing image captioningTGRS 2022-
Remote Sensing Image Change Captioning With Dual-Branch Transformers: A New Method and a Large Scale DatasetTGRS 2022code
Transforming remote sensing images to textual descriptionsINT J APPL EARTH OBS 2022-
Remote-sensing image captioning based on multilayer aggregated transformerGRSL 2022-
Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioningJ SYST ENG ELECTRON 2023-
Multi-source interactive stair attention for remote sensing image captioningRemote Sensing 2023-
Changes to Captions: An Attentive Network for Remote Sensing Change Captioningarxiv 2023code
Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioningarxiv 2023code

Text-based Image Generation

Image-text Retrieval

Visual Question Answering

PaperPublished inCode/Project
RSVQA: Visual question answering for remote sensing dataTGRS 2020code
Mutual Attention Inception Network for Remote Sensing Visual Question AnsweringTGRS 2021code
How to find a good image-text embedding for remote sensing visual question answering?ECML-PKDD 2021-
Cross-Modal Visual Question Answering for Remote Sensing Data: The International Conference on Digital Image Computing: Techniques and ApplicationsDICTA 2021-
RSVQA meets bigearthnet: a new,large-scale, visual question answering dataset for remote sensingIGARSS 2021code
Self-Paced Curriculum Learning for Visual Question Answering on Remote Sensing DataIGARSS 2021-
From easy to hard: Learning language-guided curriculum for visual question answering on remote sensing dataTGRS 2022code
Language transformers for remote sensing visual question answeringIGARSS 2022-
Open-ended remote sensing visual question answering with transformersIJRS 2022-
Bi-modal transformer-based approach for visual question answering in remote sensing imageryTGRS 2022-
Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answeringCVPRW 2022-
Change detection meets visual question answeringTGRS 2022code
A spatial hierarchical reasoning network for remote sensing visual question answeringTGRS 2023-
Multilingual Augmentation for Robust Visual Question Answering in Remote Sensing ImagesJURSE 2023-
LiT-4-RSVQA: Lightweight Transformer-based Visual Question Answering in Remote SensingIGARSS 2023code
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMsarXiv 2023code

Visual Grounding

Scene Classification

Object Detection

Semantic Segmentation

Others

Dataset

Image Captioning Dataset

Text-based Image Generation Dataset

Text-based Image Retrieval Dataset

DatasetHome/ProjectDownload link
RSITMDGithub[BaiduYun] [Google Drive]

Visual Question Answering Dataset

Visual Grounding Dataset

DatasetHome/ProjectDownload link
DIOR-RSVGGithub[Google Drive]

Scene Classification Dataset

Object Detection Dataset

Semantic Segmentation Dataset

DatasetHome/ProjectDownload link
VaihingenHome[BaiduYun]
PotsdamHome[BaiduYun]
TorontoHome-
GIDHome[BaiduYun code:GID5] [OneDrive]

Significant stargazers

Cristina

189 followers · starred Jan 2024

lzw-lzw/awesome-remote-sensing-vision-language-models

Awesome-Remote-Sensing-Vision-Language-Models

195

68 commits

updated Apr 27, 2024

See the code

README

Awesome remote sensing vision language models

This is a repository for visual language models in remote sensing, including advanced methods and commonly used datasets in different applications, such as image-text retrieval, visual question answering, pretraining, etc.

If you find any relevant papers that are not included here, please feel free to pull requests at any time.

PRs Welcome

Table of Contents

Surveys

Remote Sensing Vision Language Model

Applications

Pretraining

Image Captioning

PaperPublished inCode/Project
Deep Semantic Understanding of High Resolution Remote Sensing ImageCITS 2016-
Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?TGRS 2017-
Exploring models and data for remote sensing image caption generationTGRS 2017code
Natural language escription of remote sensing images based on deep learningIGARSS 2017-
Description Generation for Remote Sensing Images Using Attribute Attention MechanismRemote Sensing 2019-
Vaa:Visual aligning attention model for remote sensing image captioningIEEE Access 2019-
Exploring Multi-Level Attention and Semantic Relationship for Remote Sensing Image CaptioningIEEE Access 2019-
A multi-level attention model for remote sensing image captionsRemote Sensing 2020-
Remote sensing image captioning via variational autoencoder and reinforcement learningKnowledge-Based Systems 2020-
Truncation cross entropy loss for remote sensing image captioninTGRS 2020-
Word–Sentence Framework for Remote Sensing Image CaptioningTGRS 2020code
A novel SVM-based decoder for remote sensing image captioningTGRS 2021-
High-resolution remote sensing image captioning based on structured attentionTGRS 2021code
Exploring transformer and multilabel classification for remote sensing image captioningGRSL 2022-
NWPU-captions dataset and mlca-net for remote sensing image captioningTGRS 2022-
Remote Sensing Image Change Captioning With Dual-Branch Transformers: A New Method and a Large Scale DatasetTGRS 2022code
Transforming remote sensing images to textual descriptionsINT J APPL EARTH OBS 2022-
Remote-sensing image captioning based on multilayer aggregated transformerGRSL 2022-
Vlca: vision-language aligning model with cross-modal attention for bilingual remote sensing image captioningJ SYST ENG ELECTRON 2023-
Multi-source interactive stair attention for remote sensing image captioningRemote Sensing 2023-
Changes to Captions: An Attentive Network for Remote Sensing Change Captioningarxiv 2023code
Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioningarxiv 2023code

Text-based Image Generation

Image-text Retrieval

Visual Question Answering

PaperPublished inCode/Project
RSVQA: Visual question answering for remote sensing dataTGRS 2020code
Mutual Attention Inception Network for Remote Sensing Visual Question AnsweringTGRS 2021code
How to find a good image-text embedding for remote sensing visual question answering?ECML-PKDD 2021-
Cross-Modal Visual Question Answering for Remote Sensing Data: The International Conference on Digital Image Computing: Techniques and ApplicationsDICTA 2021-
RSVQA meets bigearthnet: a new,large-scale, visual question answering dataset for remote sensingIGARSS 2021code
Self-Paced Curriculum Learning for Visual Question Answering on Remote Sensing DataIGARSS 2021-
From easy to hard: Learning language-guided curriculum for visual question answering on remote sensing dataTGRS 2022code
Language transformers for remote sensing visual question answeringIGARSS 2022-
Open-ended remote sensing visual question answering with transformersIJRS 2022-
Bi-modal transformer-based approach for visual question answering in remote sensing imageryTGRS 2022-
Prompt-RSVQA: Prompting visual context to a language model for remote sensing visual question answeringCVPRW 2022-
Change detection meets visual question answeringTGRS 2022code
A spatial hierarchical reasoning network for remote sensing visual question answeringTGRS 2023-
Multilingual Augmentation for Robust Visual Question Answering in Remote Sensing ImagesJURSE 2023-
LiT-4-RSVQA: Lightweight Transformer-based Visual Question Answering in Remote SensingIGARSS 2023code
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMsarXiv 2023code

Visual Grounding

Scene Classification

Object Detection

Semantic Segmentation

Others

Dataset

Image Captioning Dataset

Text-based Image Generation Dataset

Text-based Image Retrieval Dataset

DatasetHome/ProjectDownload link
RSITMDGithub[BaiduYun] [Google Drive]

Visual Question Answering Dataset

Visual Grounding Dataset

DatasetHome/ProjectDownload link
DIOR-RSVGGithub[Google Drive]

Scene Classification Dataset

Object Detection Dataset

Semantic Segmentation Dataset

DatasetHome/ProjectDownload link
VaihingenHome[BaiduYun]
PotsdamHome[BaiduYun]
TorontoHome-
GIDHome[BaiduYun code:GID5] [OneDrive]

Significant stargazers

Cristina

189 followers · starred Jan 2024