Collection of Remote Sensing Vision-Language models and papers
To add your work to this repo, feel free to submit the request or contact me at zilun.zhang@zju.edu.cn
EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answering (2023.12) [pdf]
A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval (2023.10) [pdf]
A Fine-Grained Semantic Alignment Method Specific to Aggregate Multi-Scale Information for Cross-Modal Remote Sensing Image Retrieval (2023.10) [pdf]
Multilanguage Transformer for Improved Text to Remote Sensing Image Retrieval (2023.10) [pdf]
A Fusion Encoder with Multi-Task Guidance for Cross-Modal Text–Image Retrieval in Remote Sensing (2023.09) [pdf]
Parameter-Efficient Transfer Learning for Remote Sensing Image-Text Retrieval (2023.09) [pdf]
Hypersphere-based remote sensing cross-modal text–image retrieval via curriculum learning (2023.09) [pdf]
RS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation Model (2023.06) [pdf]
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing (2023.06) [pdf]
Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval (2023.06) [pdf]
Vision-Language Models in Remote Sensing: Current Progress and Future Trends (2023.05) [pdf]
MCRN: A Multi-source Cross-modal Retrieval Network for remote sensing (2022.12) [pdf]
RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data (2022.10) [pdf]
Learning to Evaluate Performance of Multi-modal Semantic Localization (2022.09) [pdf]
Knowledge-Aware Cross-Modal Text-Image Retrieval for Remote Sensing Images (2022.09) [pdf]
CLIP-RS: A Cross-modal Remote Sensing Image Retrieval Based on CLIP, a Northern Virginia Case Study (2022.05) [pdf]
Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval (2022.04) [pdf]
Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information (2022.04) [pdf]
Fine tuning CLIP with Remote Sensing (Satellite) images and captions (2021.10) [pdf]
6 followers · starred Dec 2023
Collection of Remote Sensing Vision-Language models and papers
To add your work to this repo, feel free to submit the request or contact me at zilun.zhang@zju.edu.cn
EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answering (2023.12) [pdf]
A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval (2023.10) [pdf]
A Fine-Grained Semantic Alignment Method Specific to Aggregate Multi-Scale Information for Cross-Modal Remote Sensing Image Retrieval (2023.10) [pdf]
Multilanguage Transformer for Improved Text to Remote Sensing Image Retrieval (2023.10) [pdf]
A Fusion Encoder with Multi-Task Guidance for Cross-Modal Text–Image Retrieval in Remote Sensing (2023.09) [pdf]
Parameter-Efficient Transfer Learning for Remote Sensing Image-Text Retrieval (2023.09) [pdf]
Hypersphere-based remote sensing cross-modal text–image retrieval via curriculum learning (2023.09) [pdf]
RS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation Model (2023.06) [pdf]
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing (2023.06) [pdf]
Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval (2023.06) [pdf]
Vision-Language Models in Remote Sensing: Current Progress and Future Trends (2023.05) [pdf]
MCRN: A Multi-source Cross-modal Retrieval Network for remote sensing (2022.12) [pdf]
RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data (2022.10) [pdf]
Learning to Evaluate Performance of Multi-modal Semantic Localization (2022.09) [pdf]
Knowledge-Aware Cross-Modal Text-Image Retrieval for Remote Sensing Images (2022.09) [pdf]
CLIP-RS: A Cross-modal Remote Sensing Image Retrieval Based on CLIP, a Northern Virginia Case Study (2022.05) [pdf]
Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval (2022.04) [pdf]
Remote Sensing Cross-Modal Text-Image Retrieval Based on Global and Local Information (2022.04) [pdf]
Fine tuning CLIP with Remote Sensing (Satellite) images and captions (2021.10) [pdf]
6 followers · starred Dec 2023