zytx121/Awesome-VLGFM

A Survey on Vision-Language Geo-Foundation Models (VLGFMs)

180

64 commits

updated May 24, 2025

See the code

README

Awesome PR's Welcome

Towards Vision-Language Geo-Foundation Models: A Survey

arXiv, 2024
Yue Zhou · Litong Feng · Xue Jiang · Junchi Yan · Xue Yang · Wayne Zhang

arXiv PDF


This repo is used for recording, tracking, and benchmarking several recent vision-language geo-foundation models (VLGFM) to supplement our survey. If you find any work missing or have any suggestions (papers, implementations, and other resources), feel free to pull requests. We will add the missing papers to this repo as soon as possible.

🙌 Add Your Paper in our Repo and Survey!!!!!

  • You are welcome to give us an issue or PR for your VLGFM work !!!!!

  • Note that: Due to the huge paper in Arxiv, we are sorry to cover all in our survey. You can directly present a PR into this repo and we will record it for next version update of our survey.

🥳 New

🔥🔥🔥 Last Updated on 2025.05.24 🔥🔥🔥

  • 2025.05.24: Update ImageRAG.
  • 2025.02.16: Update UniRS, REO-VLM, GeoPixel.
  • 2025.01.14: Update GeoPix.
  • 2024.12.25: Update VHM (accepted by AAAI'2025), which is a new version of H2RSVLM.
  • 2024.12.21: Update EarthDial.

✨ Highlight!!

  • The first survey for vision-language geo-foundation models, including contrastive/conversational/generative geo-foundation models.

  • It also contains several related works, including exploration and application of some downstream tasks.

  • We list detailed results for the most representative works and give a fairer and clearer comparison of different approaches.

📖 Introduction

This survey presents the first detailed survey on remote sensing vision language foundation models, including Contrastive/Conversational/Generative VLGFMs.

Alt Text

📗 Summary of Contents

📚 Methods: A Survey

Keywords

  • clip: Use CLIP
  • llm: Use LLM (Large Language Model)
  • sam: Use SAM (Segment Anything Model)
  • i-t: Annotate using image-text tuples
  • v-t: Annotate using video-text tuples
  • i-t-b: Annotate using image-text-box triplets
  • i-t-m: Annotate using image-text-mask triplets

image-caption-mask triplets

Contrastive VLGFMs

Conversational VLGFMs

YearVenueKeywordsPaper TitleCode/Project
2023arXivllmRsgpt: A remote sensing vision language model and benchmarkCode
2024CVPRllmGeoChat: Grounded Large Vision-Language Model for Remote SensingCode
2024arXivllmSkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language ModelCode
2024TGRSllmEarthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domainN/A
2024ECCVllmLHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language ModelCode
2024arXivllmPopeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing ImageryN/A
2024arXivllmLarge Language Models for Captioning and Retrieving Remote Sensing ImagesN/A
2024arXivllmH2RSVLM: Towards Helpful and Honest Remote Sensing Large Vision Language ModelN/A
2024RSllmRS-LLaVA: A Large Vision-Language Model for Joint Captioning and Question Answering in Remote Sensing ImageryCode
2024arXivllmSkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language UnderstandingCode
2024arXivllmEarthMarker: A Visual Prompt Learning Framework for Region-level and Point-level Remote Sensing Imagery ComprehensionCode
2024arXivllmTEOChat: A Large Vision-Language Assistant for Temporal Earth Observation DataCode
2024arXivllmAquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image ComprehensionN/A
2024arXivllmGeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual GroundingCode
2024arXivllmLHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language InterpretationCode
2024arXivllmGeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote SensingN/A
2024arXivllmRS-MoE: Mixture of Experts for Remote Sensing Image Captioning and Visual Question AnsweringN/A
2024TGRSllmRingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and grounded tasksN/A
2024arXivllmRSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of ExpertsCode
2024arXivllmEarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesN/A
2024arXivllmUniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language ModelsN/A
2024arXivllmREO-VLM: Transforming VLM to Meet Regression Challenges in Earth ObservationN/A
2025AAAIllmVHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisCode
2025arXivllmGeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote SensingN/A
2025arXivllmGeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingCode
2025GRSMllmImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAGCode

Generative VLGFMs

Datasets & Benchmark

YearVenueKeywordsNameCode/ProjectDownload
2016CITSi-tSydney-Captions & UCM-Captions[N/A]link,link2
2017TGRSi-tRSICDProjectLink
2020TGRSi-tRSVQA-LR & RSVQA-HRProjectlink1,link2
2021IGARSSi-tRSVQAxBENProjectlink
2021Accessi-tFloodNetProjectlink
2021TGRSi-tRSITMDCodelink
2021TGRSi-tRSIVQACodelink
2022TGRSi-tNWPU-CaptionsProjectlink
2022TGRSi-tCRSVQAProjectlink
2022TGRSi-tLEVIR-CCProjectlink
2022TGRSi-tCDVQAProjectlink
2022TGRSi-tUAV-CaptionsN/AN/A
2022MMi-t-bRSVGProjectlink
2022RSv-tCapERAProjectlink
2023TGRSi-t-bDIOR-RSVGProjectlink
2023arXivi-tRemoteCountCodeN/A
2023arXivi-tRS5MCodelink
2023arXivi-tRSICap & RSIEvalCodeN/A
2023arXivi-tLAION-EON/Alink
2023ICCVWi-tSATINProjectlink
2024ICLRi-tNAIP-OSMProjectN/A
2024AAAIi-tSkyScriptCodelink
2024AAAIi-t-mEarthVQAProjectN/A
2024TGRSi-t-mRRSISCodelink
2024CVPRi-tGeoChat-Instruct & GeoChat-BenchCodelink
2024CVPRi-t-mRRSIS-DCodelink
2024ECCVi-tGeoTextProjectlink
2024arXivi-tSkyEye-968kCodeN/A
2024arXivi-tMMRS-1MProjectN/A
2024arXivi-tLHRS-Align & LHRS-InstructCodeN/A
2024arXivi-t-mChatEarthNetprojectlink
2024arXivi-tVLEO-BenchCodelink
2024arXivi-tLuoJiaHOGN/AN/A
2024arXivi-t-mFineGripN/AN/A
2024arXivi-tRS-GPT4VN/AN/A
2024arXivi-tVRSBenchN/AN/A
2024arXivi-tRSTellerProjectlink
2024arXivi-tMME-RealWorldProjectlink
2024arXivi-tUrBenchProjectN/A
2024arXivi-tMMM-RSProjectN/A
2024arXivi-tDDFAVProjectN/A
2024arXivi-tCOREvalN/AN/A
2024arXivi-tGEOBench-VLMProjectN/A

🕹️ Application

Captioning

Visual Question Answering

Change Detection

Scene Classification

Referring Expression Segmentation (RES)

|2024|TGRS||RRSIS: Referring Remote Sensing Image Segmentation|Code| |2024|CVPR||Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation|Code|

Geospatial Localization

Retrieval

Segmentation

Object Detection

YearVenueKeywordsPaper TitleCode/Project
2023arXivclipStable Diffusion For Aerial Object DetectionN/A

Super-Resolution

YearVenueKeywordsPaper TitleCode/Project
2023arXivclipZooming Out on Zooming In: Advancing Super-Resolution for Remote SensingCode

📊 Exploration

👨‍🏫 Survey

🖊️ Citation

If you find our survey and repository useful for your research project, please consider citing our paper:

@article{zhou2024vlgfm,
  title={Towards Vision-Language Geo-Foundation Models: A Survey},
  author={Yue Zhou and Litong Feng and Yiping Ke and Xue Jiang and Junchi Yan and Xue Yang and Wayne Zhang},
  journal={arXiv preprint arXiv:2406.09385},
  year={2024}
}

🐲 Contact

yue.zhou@ntu.edu.sg
foundation-models
remote-sensing
survey
vision-language-model

Contributors

zytx121

64 commits

zytx121/Awesome-VLGFM

A Survey on Vision-Language Geo-Foundation Models (VLGFMs)

180

64 commits

updated May 24, 2025

See the code

README

Awesome PR's Welcome

Towards Vision-Language Geo-Foundation Models: A Survey

arXiv, 2024
Yue Zhou · Litong Feng · Xue Jiang · Junchi Yan · Xue Yang · Wayne Zhang

arXiv PDF


This repo is used for recording, tracking, and benchmarking several recent vision-language geo-foundation models (VLGFM) to supplement our survey. If you find any work missing or have any suggestions (papers, implementations, and other resources), feel free to pull requests. We will add the missing papers to this repo as soon as possible.

🙌 Add Your Paper in our Repo and Survey!!!!!

  • You are welcome to give us an issue or PR for your VLGFM work !!!!!

  • Note that: Due to the huge paper in Arxiv, we are sorry to cover all in our survey. You can directly present a PR into this repo and we will record it for next version update of our survey.

🥳 New

🔥🔥🔥 Last Updated on 2025.05.24 🔥🔥🔥

  • 2025.05.24: Update ImageRAG.
  • 2025.02.16: Update UniRS, REO-VLM, GeoPixel.
  • 2025.01.14: Update GeoPix.
  • 2024.12.25: Update VHM (accepted by AAAI'2025), which is a new version of H2RSVLM.
  • 2024.12.21: Update EarthDial.

✨ Highlight!!

  • The first survey for vision-language geo-foundation models, including contrastive/conversational/generative geo-foundation models.

  • It also contains several related works, including exploration and application of some downstream tasks.

  • We list detailed results for the most representative works and give a fairer and clearer comparison of different approaches.

📖 Introduction

This survey presents the first detailed survey on remote sensing vision language foundation models, including Contrastive/Conversational/Generative VLGFMs.

Alt Text

📗 Summary of Contents

📚 Methods: A Survey

Keywords

  • clip: Use CLIP
  • llm: Use LLM (Large Language Model)
  • sam: Use SAM (Segment Anything Model)
  • i-t: Annotate using image-text tuples
  • v-t: Annotate using video-text tuples
  • i-t-b: Annotate using image-text-box triplets
  • i-t-m: Annotate using image-text-mask triplets

image-caption-mask triplets

Contrastive VLGFMs

Conversational VLGFMs

YearVenueKeywordsPaper TitleCode/Project
2023arXivllmRsgpt: A remote sensing vision language model and benchmarkCode
2024CVPRllmGeoChat: Grounded Large Vision-Language Model for Remote SensingCode
2024arXivllmSkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language ModelCode
2024TGRSllmEarthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domainN/A
2024ECCVllmLHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language ModelCode
2024arXivllmPopeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing ImageryN/A
2024arXivllmLarge Language Models for Captioning and Retrieving Remote Sensing ImagesN/A
2024arXivllmH2RSVLM: Towards Helpful and Honest Remote Sensing Large Vision Language ModelN/A
2024RSllmRS-LLaVA: A Large Vision-Language Model for Joint Captioning and Question Answering in Remote Sensing ImageryCode
2024arXivllmSkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language UnderstandingCode
2024arXivllmEarthMarker: A Visual Prompt Learning Framework for Region-level and Point-level Remote Sensing Imagery ComprehensionCode
2024arXivllmTEOChat: A Large Vision-Language Assistant for Temporal Earth Observation DataCode
2024arXivllmAquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image ComprehensionN/A
2024arXivllmGeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual GroundingCode
2024arXivllmLHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language InterpretationCode
2024arXivllmGeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote SensingN/A
2024arXivllmRS-MoE: Mixture of Experts for Remote Sensing Image Captioning and Visual Question AnsweringN/A
2024TGRSllmRingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and grounded tasksN/A
2024arXivllmRSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of ExpertsCode
2024arXivllmEarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesN/A
2024arXivllmUniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language ModelsN/A
2024arXivllmREO-VLM: Transforming VLM to Meet Regression Challenges in Earth ObservationN/A
2025AAAIllmVHM: Versatile and Honest Vision Language Model for Remote Sensing Image AnalysisCode
2025arXivllmGeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote SensingN/A
2025arXivllmGeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingCode
2025GRSMllmImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAGCode

Generative VLGFMs

Datasets & Benchmark

YearVenueKeywordsNameCode/ProjectDownload
2016CITSi-tSydney-Captions & UCM-Captions[N/A]link,link2
2017TGRSi-tRSICDProjectLink
2020TGRSi-tRSVQA-LR & RSVQA-HRProjectlink1,link2
2021IGARSSi-tRSVQAxBENProjectlink
2021Accessi-tFloodNetProjectlink
2021TGRSi-tRSITMDCodelink
2021TGRSi-tRSIVQACodelink
2022TGRSi-tNWPU-CaptionsProjectlink
2022TGRSi-tCRSVQAProjectlink
2022TGRSi-tLEVIR-CCProjectlink
2022TGRSi-tCDVQAProjectlink
2022TGRSi-tUAV-CaptionsN/AN/A
2022MMi-t-bRSVGProjectlink
2022RSv-tCapERAProjectlink
2023TGRSi-t-bDIOR-RSVGProjectlink
2023arXivi-tRemoteCountCodeN/A
2023arXivi-tRS5MCodelink
2023arXivi-tRSICap & RSIEvalCodeN/A
2023arXivi-tLAION-EON/Alink
2023ICCVWi-tSATINProjectlink
2024ICLRi-tNAIP-OSMProjectN/A
2024AAAIi-tSkyScriptCodelink
2024AAAIi-t-mEarthVQAProjectN/A
2024TGRSi-t-mRRSISCodelink
2024CVPRi-tGeoChat-Instruct & GeoChat-BenchCodelink
2024CVPRi-t-mRRSIS-DCodelink
2024ECCVi-tGeoTextProjectlink
2024arXivi-tSkyEye-968kCodeN/A
2024arXivi-tMMRS-1MProjectN/A
2024arXivi-tLHRS-Align & LHRS-InstructCodeN/A
2024arXivi-t-mChatEarthNetprojectlink
2024arXivi-tVLEO-BenchCodelink
2024arXivi-tLuoJiaHOGN/AN/A
2024arXivi-t-mFineGripN/AN/A
2024arXivi-tRS-GPT4VN/AN/A
2024arXivi-tVRSBenchN/AN/A
2024arXivi-tRSTellerProjectlink
2024arXivi-tMME-RealWorldProjectlink
2024arXivi-tUrBenchProjectN/A
2024arXivi-tMMM-RSProjectN/A
2024arXivi-tDDFAVProjectN/A
2024arXivi-tCOREvalN/AN/A
2024arXivi-tGEOBench-VLMProjectN/A

🕹️ Application

Captioning

Visual Question Answering

Change Detection

Scene Classification

Referring Expression Segmentation (RES)

|2024|TGRS||RRSIS: Referring Remote Sensing Image Segmentation|Code| |2024|CVPR||Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation|Code|

Geospatial Localization

Retrieval

Segmentation

Object Detection

YearVenueKeywordsPaper TitleCode/Project
2023arXivclipStable Diffusion For Aerial Object DetectionN/A

Super-Resolution

YearVenueKeywordsPaper TitleCode/Project
2023arXivclipZooming Out on Zooming In: Advancing Super-Resolution for Remote SensingCode

📊 Exploration

👨‍🏫 Survey

🖊️ Citation

If you find our survey and repository useful for your research project, please consider citing our paper:

@article{zhou2024vlgfm,
  title={Towards Vision-Language Geo-Foundation Models: A Survey},
  author={Yue Zhou and Litong Feng and Yiping Ke and Xue Jiang and Junchi Yan and Xue Yang and Wayne Zhang},
  journal={arXiv preprint arXiv:2406.09385},
  year={2024}
}

🐲 Contact

yue.zhou@ntu.edu.sg
foundation-models
remote-sensing
survey
vision-language-model

Contributors

zytx121

64 commits