Chen-Yang-Liu/Awesome-RS-SpatioTemporal-VLMs

[IEEE GRSM 2025 🔥] Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey

226

41 commits

updated Oct 29, 2025

See the code

README

Awesome PR's Welcome


Chenyang Liu · Jiafan Zhang · Keyan Chen · Man Wang · Zhengxia Zou · Zhenwei Shi*✉

GRSM PDF Arxiv PDF


This repo is used for recording, and tracking recent Remote Sensing Spatio-Temporal Vision-Language Models (RS-STVLMs). If you find any work missing or have any suggestions (papers, implementations, and other resources), feel free to pull requests.

:star: Share us a :star:

Share us a :star: if you're interested in this repo. We will continue to track relevant progress and update this repository.

🙌 Add Your Paper in our Repo and Survey!

  • You are welcome to give us an issue or PR for your RS-STVLM work !!!!! We will record it for next version update of our survey

🥳 News

🔥🔥🔥 The rep is updating 🔥🔥🔥

✨ Highlight!!

✅ The first survey for Remote Sensing Spatio-Temporal Vision-Language Models.

✅ Some public datasets and code links are provided.

✅ We will continue to track related work in this repository.

📖 Introduction

Timeline of RS-STVLMs:

Alt Text

📖 Table of Contents

📚 Remote Sensing Spatio-Temporal Vision-language Tasks and Methods

Change Captioning

TimeModel NamePaper TitleVisual EncoderLanguage DecoderCode/Project
2021.10CNN-RNNCaptioning changes in bi-temporal remote sensing imagesVGG-16RNNN/A
2022.08CC-RNN/SVMChange captioning: A new paradigm for multitemporal remote sensing image analysisVGG-16RNN,SVMN/A
2022.11RSICCformerRemote sensing image change captioning with dual-branch transformers: A new method and a large scale datasetResNet-101Transformer DecoderStar
2023.07PSNetProgressive Scale-aware Network for Remote sensing Image Change CaptioningViT-B/32Transformer DecoderStar
2023.10PromptCCA Decoupling Paradigm with Prompt Learning for Remote Sensing Image Change CaptioningViT-B/32GPT-2Star
2023.11Chg2CapChanges to Captions: An Attentive Network for Remote Sensing Change CaptioningResNet-101Transformer DecoderStar
2023.11ICT-NetInteractive Change-Aware Transformer Network for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderStar
2024.03SITS-CCChange Caption for Satellite Images Time SeriesResNet-101Transformer DecoderStar
2024.05RSCaMaRSCaMa: Remote Sensing Image Change Captioning with State Space ModelViT-B/32Mamba, Transformer Decoder, GPT-2Star
2024.05SparseFocusA Lightweight Sparse Focus Transformer for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderStar
2024.05SENSingle-stream Extractor Network with Contrastive Pre-training for Remote Sensing Change CaptioningResNet with 6-channelTransformer DecoderStar
2024.05Diffusion-RSCCDiffusion model for learning cross-modal data distributionResNet-101DiffusionStar
2024.05CARDContext-aware Difference Distilling for Multi-change CaptioningResNet-101Transformer DecoderStar
2024.06ChangeRetCapTowards a multimodal framework for remote sensing image change retrieval and captioningResNet-101Transformer DecoderStar
2024.06Intelli-ChangeIntelli-Change Remote Sensing - A Novel Transformer ApproachResNet-101Transformer DecoderN/A
2024.06ChangeExpTowards Temporal Change Explanations from Bi-Temporal Satellite ImagesLLaVA-1.5LLaVA-1.5N/A
2024.07MAF-NetMulti-scale Attentive Fusion Network for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderN/A
2024.07SFENScale-wised feature enhancement network for change captioning of remote sensing imagesWideResNetTransformer DecoderN/A
2024.09MfrNetMfrNet: A New Multi-Scale Feature Refining Method for Remote Sensing Image Change CaptioningResNet-18Transformer DecoderN/A
2024.09SEIFNetInter-Temporal Interaction and Symmetric Difference Learning for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderStar
2024.10MV-CCMV-CC: Mask Enhanced Video Model for Remote Sensing Change CaptionInternVideo2Transformer DecoderStar
2024.10ChareptionChareption: Change-Aware Adaption Empowers Large Language Model for Effective Remote Sensing Image Change CaptioningCLIP ViT-L/14LLaMA-7BN/A
2024.11MADiffCCRemote Sensing Image Change Captioning Using Multi-Attentive Network with Diffusion ModelDiffusionTransformer DecoderN/A
2024.11CCExpertCCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational DatasetDiffusionTransformer DecoderStar
2024.12---Data Augmentation in Remote Sensing Image Change CaptioningViT-B/32Transformer DecoderN/A
2024.12Mask Approx NetMask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change CaptioningResNetTransformer DecoderLink
2025.01SAT-CapChange Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer ApproachResNet-101Transformer DecoderStar
2025.01MModalCCRobust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC FrameworkResNet-101Transformer DecoderStar
2025.01SGD-RSCCNScene Graph and Dependency Grammar Enhanced Remote Sensing Change Caption Network (SGD-RSCCN)ResNet-101Transformer DecoderN/A
2025.02TGIPGImage Editing based on Diffusion Model for Remote Sensing Image Change Captioning////N/A
2025.03Change3DChange3D: Revisiting Change Detection and Captioning from A Video Modeling PerspectiveX3D-L(video)Transformer DecoderStar
2025.03CD4CCD4C: Change Detection for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderN/A
2025.04RDD+ACRRegion-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningResNet-101Transformer DecoderN/A
2025.04FST-NetFrequency–Spatial–Temporal Domain Fusion Network for Remote Sensing Image Change CaptioningSegformerTransformer DecoderN/A
2025.05CTSD-NetA Cross-Spatial Differential Localization Network for Remote Sensing Change CaptioningSegFormerTransformer DecoderN/A
2025.06CTMCross-Temporal Remote Sensing Image Change Captioning: A Manifold Mapping and Bayesian Diffusion Approach for Land Use MonitoringCLIPTransformer DecoderN/A
2025.06IHM-SNetIHM-SNet: An Interactive Hierarchical Mamba-Based Screening Network for Remote Sensing Image Change CaptioningCLIP-ViTTransformer DecoderN/A
2025.07MTI-CCCross-layer Attention Enhanced Remote Sensing Image Change Captioning via Mamba-Transformer InteractionCLIP-ViTTransformer DecoderN/A
2025.08CI-NetRestricted supervised Cascade Information Network for remote sensing change captioning with serial sentencesAsymmetric Siamese NetworkCascade Linguistic ModuleN/A
2025.08SCCNetSCCNet: Siamese Networks for Selective Change Captioning in Bi-Temporal Remote Sensing ImagesViTTransformer DecoderN/A
2025.08--Text-Augmented Semantic Feature Extraction and Difference Information Learning for Remote Sensing Image Change CaptioningFastSAM+CLIPTransformer DecoderStar
2025.08C3aptionerC3aptioner: Improving Change Captioning by Leveraging Momentum Cross-view and Cross-modality Contrastive LearningResNet-101Transformer DecoderN/A
2025.08ICL-CCScalable Remote Sensing Image Change Captioning using In-Context LearningStar

| ........

Multitask Learning of Change Detection and Change Captioning

TimeModel NamePaper TitleVisual EncoderLanguage DecoderCode/Project
2024.01Pix4CapPixel-Level Change Detection Pseudo-Label Learning for Remote Sensing Change CaptioningViT-B/32Transformer DecoderN/A
2024.03Change-AgentChange-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and AnalysisViT-B/32Transformer DecoderStar
2024.07Semantic-CCSemantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic GuidanceSAMVicunaN/A
2024.09DetACC *Detection Assisted Change Captioning for Remote Sensing ImageResNet-101Transformer DecoderN/A
2024.09KCFIEnhancing Perception of Key Changes in Remote Sensing Image Change CaptioningViTQwenStar
2024.10MV-CC *MV-CC: Mask Enhanced Video Model for Remote Sensing Change CaptionInternVideo2Transformer DecoderStar
2024.10ChangeMindsChangeMinds: Multi-task Framework for Detecting and Describing Changes in Remote SensingSwin TransformerTransformer DecoderStar
2024.10CTMTNetA Multi-Task Network and Two Large Scale Datasets for Change Detection and Captioning in Remote Sensing ImagesResNet-101Transformer DecoderN/A
2024.12Mask Approx NetMask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change CaptioningResNetTransformer DecoderLink
2025.01MModalCC *Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC FrameworkResNet-101Transformer DecoderStar
2025.03CD4C *CD4C: Change Detection for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderN/A
2025.04FST-NetFrequency–Spatial–Temporal Domain Fusion Network for Remote Sensing Image Change CaptioningSegformerTransformer DecoderN/A
......

Change Question Answering

Text-driven Temporal Images Retrieval

Change Grounding

Text-driven Temporal Images Generation

👨‍🏫 Large Language Models Meets Temporal Images

LLM-driven Task-Specific Spatio-Temporal VLMs

Unified Spatio-Temporal Vision-Language Foundation Models

LLM-driven Remote Sensing Vision-Language Agents

🛰️ Dataset

Matching Temporal Images, Text, and Masks

DatasetTimeImage SizeImage ResolutionImage PairsCaptions*MasksTemporal Image Data SourceAnno.Link
DUBAI CCD2022.0850×5030m5002,500-Landsat-7 imageryManualLink
LEVIR CCD2022.08256×2560.5m5002,500-LEVIR-CDManualLink
LEVIR-CC2022.11256×2560.5m10,07750,385-LEVIR-CDManualLink
CCExpert2024.11--200K1.2M-LEVIR-CC, CLVER-Change, ImageEdit, Spot-the-dif, STVchrono, Vismin, ChangeSim, SYSU-CD, SECONDAuto.Link
SECTION2025.07256×2560.3-3m4,05912,200-SECONDManualLink
LEVIR-MCI2024.03256×2560.5m10,07750,385building, roadLEVIR-CCManualLink
LEVIR-CDC2024.11256×2560.5m10,07750,385buildingLEVIR-CCManualLink
WHU-CDC2024.11256×2560.075m7,43437,170buildingWHU-CDManualLink
SECOND-CC2025.01256×2560.3∼3m6,04130,2056 classesSECONDManualLink

Matching Temporal Images, Instruction and Response

DatasetTimeInstruction SamplesNumber of ImagesTemporal LengthTemporal Image Data SourceAnno.Link
CDVQA2022.09122,0002,9682SECONDManualLink
ChangeChat-87k2024.0987,19510,0772LEVIR-CC, LEVIR-MCIAuto.Link
QAG-360K2024.10360,0006,8102Hi-UCD, SECOND, LEVIR-CDAuto.Link
GeoLLaVA2024.10100,000100,0002fMoWAuto.Link
TEOChatlas2024.10554,071-1~8xBD, S2Looking, QFabric, fMoWAuto.Link
EarthDial2024.1211.11 Million-1~4fMoW, TreeSatAI-Time-Series, MUDS, xBD, QuakeSetManual & Auto.Link
UniRS2024.12318.8 K-1~T (T>2)LEVIR-CC, ERA-VideoAuto.Link
Falcon_SFT2025.0378 Million5.6 Million1~2CDD, EGY-BCD, HRSCD, LEVIR-CD, MSBC, MSOSCD, NJDS, S2Looking, SYSU-CD, WHU-CDAuto.Link
DVL-Suite2025.0569,92615,0636.9 (Average)U.S. National Agriculture Imagery Program (NAIP)Manual & Auto.N/A
....

💻 Others

Some CLIP Models in Remote Sensing

🖊️ Citation

If you find our survey and repository useful for your research, please consider citing our paper:

@ARTICLE{liu2024RSSTVLMsurvey,
  author={Liu, Chenyang and Zhang, Jiafan and Chen, Keyan and Wang, Man and Zou, Zhengxia and Shi, Zhenwei},
  journal={IEEE Geoscience and Remote Sensing Magazine}, 
  title={Remote Sensing Spatiotemporal Vision–Language Models: A comprehensive survey}, 
  year={2025},
  volume={},
  number={},
  pages={2-42},
  doi={10.1109/MGRS.2025.3598283}}

🐲 Contact

liuchenyang@buaa.edu.cn
change-detetion
foundation-models
large-language-models
remote-sensing
spatio-temporal-analysis
vision-language

Contributors

Chen-Yang-Liu

40 commits

ULTA-Web

1 commits

Chen-Yang-Liu/Awesome-RS-SpatioTemporal-VLMs

[IEEE GRSM 2025 🔥] Remote Sensing SpatioTemporal Vision-Language Models: A Comprehensive Survey

226

41 commits

updated Oct 29, 2025

See the code

README

Awesome PR's Welcome


Chenyang Liu · Jiafan Zhang · Keyan Chen · Man Wang · Zhengxia Zou · Zhenwei Shi*✉

GRSM PDF Arxiv PDF


This repo is used for recording, and tracking recent Remote Sensing Spatio-Temporal Vision-Language Models (RS-STVLMs). If you find any work missing or have any suggestions (papers, implementations, and other resources), feel free to pull requests.

:star: Share us a :star:

Share us a :star: if you're interested in this repo. We will continue to track relevant progress and update this repository.

🙌 Add Your Paper in our Repo and Survey!

  • You are welcome to give us an issue or PR for your RS-STVLM work !!!!! We will record it for next version update of our survey

🥳 News

🔥🔥🔥 The rep is updating 🔥🔥🔥

✨ Highlight!!

✅ The first survey for Remote Sensing Spatio-Temporal Vision-Language Models.

✅ Some public datasets and code links are provided.

✅ We will continue to track related work in this repository.

📖 Introduction

Timeline of RS-STVLMs:

Alt Text

📖 Table of Contents

📚 Remote Sensing Spatio-Temporal Vision-language Tasks and Methods

Change Captioning

TimeModel NamePaper TitleVisual EncoderLanguage DecoderCode/Project
2021.10CNN-RNNCaptioning changes in bi-temporal remote sensing imagesVGG-16RNNN/A
2022.08CC-RNN/SVMChange captioning: A new paradigm for multitemporal remote sensing image analysisVGG-16RNN,SVMN/A
2022.11RSICCformerRemote sensing image change captioning with dual-branch transformers: A new method and a large scale datasetResNet-101Transformer DecoderStar
2023.07PSNetProgressive Scale-aware Network for Remote sensing Image Change CaptioningViT-B/32Transformer DecoderStar
2023.10PromptCCA Decoupling Paradigm with Prompt Learning for Remote Sensing Image Change CaptioningViT-B/32GPT-2Star
2023.11Chg2CapChanges to Captions: An Attentive Network for Remote Sensing Change CaptioningResNet-101Transformer DecoderStar
2023.11ICT-NetInteractive Change-Aware Transformer Network for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderStar
2024.03SITS-CCChange Caption for Satellite Images Time SeriesResNet-101Transformer DecoderStar
2024.05RSCaMaRSCaMa: Remote Sensing Image Change Captioning with State Space ModelViT-B/32Mamba, Transformer Decoder, GPT-2Star
2024.05SparseFocusA Lightweight Sparse Focus Transformer for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderStar
2024.05SENSingle-stream Extractor Network with Contrastive Pre-training for Remote Sensing Change CaptioningResNet with 6-channelTransformer DecoderStar
2024.05Diffusion-RSCCDiffusion model for learning cross-modal data distributionResNet-101DiffusionStar
2024.05CARDContext-aware Difference Distilling for Multi-change CaptioningResNet-101Transformer DecoderStar
2024.06ChangeRetCapTowards a multimodal framework for remote sensing image change retrieval and captioningResNet-101Transformer DecoderStar
2024.06Intelli-ChangeIntelli-Change Remote Sensing - A Novel Transformer ApproachResNet-101Transformer DecoderN/A
2024.06ChangeExpTowards Temporal Change Explanations from Bi-Temporal Satellite ImagesLLaVA-1.5LLaVA-1.5N/A
2024.07MAF-NetMulti-scale Attentive Fusion Network for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderN/A
2024.07SFENScale-wised feature enhancement network for change captioning of remote sensing imagesWideResNetTransformer DecoderN/A
2024.09MfrNetMfrNet: A New Multi-Scale Feature Refining Method for Remote Sensing Image Change CaptioningResNet-18Transformer DecoderN/A
2024.09SEIFNetInter-Temporal Interaction and Symmetric Difference Learning for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderStar
2024.10MV-CCMV-CC: Mask Enhanced Video Model for Remote Sensing Change CaptionInternVideo2Transformer DecoderStar
2024.10ChareptionChareption: Change-Aware Adaption Empowers Large Language Model for Effective Remote Sensing Image Change CaptioningCLIP ViT-L/14LLaMA-7BN/A
2024.11MADiffCCRemote Sensing Image Change Captioning Using Multi-Attentive Network with Diffusion ModelDiffusionTransformer DecoderN/A
2024.11CCExpertCCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational DatasetDiffusionTransformer DecoderStar
2024.12---Data Augmentation in Remote Sensing Image Change CaptioningViT-B/32Transformer DecoderN/A
2024.12Mask Approx NetMask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change CaptioningResNetTransformer DecoderLink
2025.01SAT-CapChange Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer ApproachResNet-101Transformer DecoderStar
2025.01MModalCCRobust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC FrameworkResNet-101Transformer DecoderStar
2025.01SGD-RSCCNScene Graph and Dependency Grammar Enhanced Remote Sensing Change Caption Network (SGD-RSCCN)ResNet-101Transformer DecoderN/A
2025.02TGIPGImage Editing based on Diffusion Model for Remote Sensing Image Change Captioning////N/A
2025.03Change3DChange3D: Revisiting Change Detection and Captioning from A Video Modeling PerspectiveX3D-L(video)Transformer DecoderStar
2025.03CD4CCD4C: Change Detection for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderN/A
2025.04RDD+ACRRegion-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change CaptioningResNet-101Transformer DecoderN/A
2025.04FST-NetFrequency–Spatial–Temporal Domain Fusion Network for Remote Sensing Image Change CaptioningSegformerTransformer DecoderN/A
2025.05CTSD-NetA Cross-Spatial Differential Localization Network for Remote Sensing Change CaptioningSegFormerTransformer DecoderN/A
2025.06CTMCross-Temporal Remote Sensing Image Change Captioning: A Manifold Mapping and Bayesian Diffusion Approach for Land Use MonitoringCLIPTransformer DecoderN/A
2025.06IHM-SNetIHM-SNet: An Interactive Hierarchical Mamba-Based Screening Network for Remote Sensing Image Change CaptioningCLIP-ViTTransformer DecoderN/A
2025.07MTI-CCCross-layer Attention Enhanced Remote Sensing Image Change Captioning via Mamba-Transformer InteractionCLIP-ViTTransformer DecoderN/A
2025.08CI-NetRestricted supervised Cascade Information Network for remote sensing change captioning with serial sentencesAsymmetric Siamese NetworkCascade Linguistic ModuleN/A
2025.08SCCNetSCCNet: Siamese Networks for Selective Change Captioning in Bi-Temporal Remote Sensing ImagesViTTransformer DecoderN/A
2025.08--Text-Augmented Semantic Feature Extraction and Difference Information Learning for Remote Sensing Image Change CaptioningFastSAM+CLIPTransformer DecoderStar
2025.08C3aptionerC3aptioner: Improving Change Captioning by Leveraging Momentum Cross-view and Cross-modality Contrastive LearningResNet-101Transformer DecoderN/A
2025.08ICL-CCScalable Remote Sensing Image Change Captioning using In-Context LearningStar

| ........

Multitask Learning of Change Detection and Change Captioning

TimeModel NamePaper TitleVisual EncoderLanguage DecoderCode/Project
2024.01Pix4CapPixel-Level Change Detection Pseudo-Label Learning for Remote Sensing Change CaptioningViT-B/32Transformer DecoderN/A
2024.03Change-AgentChange-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and AnalysisViT-B/32Transformer DecoderStar
2024.07Semantic-CCSemantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic GuidanceSAMVicunaN/A
2024.09DetACC *Detection Assisted Change Captioning for Remote Sensing ImageResNet-101Transformer DecoderN/A
2024.09KCFIEnhancing Perception of Key Changes in Remote Sensing Image Change CaptioningViTQwenStar
2024.10MV-CC *MV-CC: Mask Enhanced Video Model for Remote Sensing Change CaptionInternVideo2Transformer DecoderStar
2024.10ChangeMindsChangeMinds: Multi-task Framework for Detecting and Describing Changes in Remote SensingSwin TransformerTransformer DecoderStar
2024.10CTMTNetA Multi-Task Network and Two Large Scale Datasets for Change Detection and Captioning in Remote Sensing ImagesResNet-101Transformer DecoderN/A
2024.12Mask Approx NetMask Approximation Net: A Novel Diffusion Model Approach for Remote Sensing Change CaptioningResNetTransformer DecoderLink
2025.01MModalCC *Robust Change Captioning in Remote Sensing: SECOND-CC Dataset and MModalCC FrameworkResNet-101Transformer DecoderStar
2025.03CD4C *CD4C: Change Detection for Remote Sensing Image Change CaptioningResNet-101Transformer DecoderN/A
2025.04FST-NetFrequency–Spatial–Temporal Domain Fusion Network for Remote Sensing Image Change CaptioningSegformerTransformer DecoderN/A
......

Change Question Answering

Text-driven Temporal Images Retrieval

Change Grounding

Text-driven Temporal Images Generation

👨‍🏫 Large Language Models Meets Temporal Images

LLM-driven Task-Specific Spatio-Temporal VLMs

Unified Spatio-Temporal Vision-Language Foundation Models

LLM-driven Remote Sensing Vision-Language Agents

🛰️ Dataset

Matching Temporal Images, Text, and Masks

DatasetTimeImage SizeImage ResolutionImage PairsCaptions*MasksTemporal Image Data SourceAnno.Link
DUBAI CCD2022.0850×5030m5002,500-Landsat-7 imageryManualLink
LEVIR CCD2022.08256×2560.5m5002,500-LEVIR-CDManualLink
LEVIR-CC2022.11256×2560.5m10,07750,385-LEVIR-CDManualLink
CCExpert2024.11--200K1.2M-LEVIR-CC, CLVER-Change, ImageEdit, Spot-the-dif, STVchrono, Vismin, ChangeSim, SYSU-CD, SECONDAuto.Link
SECTION2025.07256×2560.3-3m4,05912,200-SECONDManualLink
LEVIR-MCI2024.03256×2560.5m10,07750,385building, roadLEVIR-CCManualLink
LEVIR-CDC2024.11256×2560.5m10,07750,385buildingLEVIR-CCManualLink
WHU-CDC2024.11256×2560.075m7,43437,170buildingWHU-CDManualLink
SECOND-CC2025.01256×2560.3∼3m6,04130,2056 classesSECONDManualLink

Matching Temporal Images, Instruction and Response

DatasetTimeInstruction SamplesNumber of ImagesTemporal LengthTemporal Image Data SourceAnno.Link
CDVQA2022.09122,0002,9682SECONDManualLink
ChangeChat-87k2024.0987,19510,0772LEVIR-CC, LEVIR-MCIAuto.Link
QAG-360K2024.10360,0006,8102Hi-UCD, SECOND, LEVIR-CDAuto.Link
GeoLLaVA2024.10100,000100,0002fMoWAuto.Link
TEOChatlas2024.10554,071-1~8xBD, S2Looking, QFabric, fMoWAuto.Link
EarthDial2024.1211.11 Million-1~4fMoW, TreeSatAI-Time-Series, MUDS, xBD, QuakeSetManual & Auto.Link
UniRS2024.12318.8 K-1~T (T>2)LEVIR-CC, ERA-VideoAuto.Link
Falcon_SFT2025.0378 Million5.6 Million1~2CDD, EGY-BCD, HRSCD, LEVIR-CD, MSBC, MSOSCD, NJDS, S2Looking, SYSU-CD, WHU-CDAuto.Link
DVL-Suite2025.0569,92615,0636.9 (Average)U.S. National Agriculture Imagery Program (NAIP)Manual & Auto.N/A
....

💻 Others

Some CLIP Models in Remote Sensing

🖊️ Citation

If you find our survey and repository useful for your research, please consider citing our paper:

@ARTICLE{liu2024RSSTVLMsurvey,
  author={Liu, Chenyang and Zhang, Jiafan and Chen, Keyan and Wang, Man and Zou, Zhengxia and Shi, Zhenwei},
  journal={IEEE Geoscience and Remote Sensing Magazine}, 
  title={Remote Sensing Spatiotemporal Vision–Language Models: A comprehensive survey}, 
  year={2025},
  volume={},
  number={},
  pages={2-42},
  doi={10.1109/MGRS.2025.3598283}}

🐲 Contact

liuchenyang@buaa.edu.cn
change-detetion
foundation-models
large-language-models
remote-sensing
spatio-temporal-analysis
vision-language

Contributors

Chen-Yang-Liu

40 commits

ULTA-Web

1 commits