This is a repository for "A Survey on Remote Sensing Foundation Models: From Vision to Multimodality".
🌏 Please check out our survey paper: A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
You can see details of all papers and datasets here: homepage
This repository maintains the resource list associated with the survey. It is intended to help readers quickly find papers, datasets, models, benchmarks, and agentic remote sensing systems discussed in the survey.
Last curated update: 2026-06-24.
Corrections and additions are welcome. Please see CONTRIBUTING.md before opening a pull request.
The 2026 revision broadens the repository beyond early vision and vision-language foundation models. The additions below highlight recent resources on sensor-adaptive pretraining, multimodal reasoning, long-tail datasets, robustness benchmarks, open-vocabulary grounding, and agentic Earth observation.
| Name | Focus | Paper / Project |
|---|---|---|
| RAMEN | Resolution-adjustable multimodal encoder for Earth observation | CVPR 2026 |
| THOR | Versatile Earth observation foundation model for climate and society applications | CVPRW 2026 |
| TerraFlow | Multimodal and multitemporal Earth observation representation learning | arXiv 2026 |
| SpectralEarth-FM | Hyperspectral imagery in multimodal Earth observation pretraining | arXiv 2026 |
| FLORO | Multimodal geospatial foundation model across sensors and scales | arXiv 2026 |
| SMARTIES | Spectrum-aware multi-sensor auto-encoder | ICCV 2025 |
| SIGMAE | Spectral-index-guided foundation model for multispectral remote sensing | arXiv 2026 |
| RingMoE | Mixture-of-modality-experts remote sensing foundation model | arXiv 2025 |
| SkySense V2 | Unified foundation model for multimodal remote sensing | arXiv 2025 |
| Falcon | Remote sensing vision-language foundation model | arXiv 2025 |
| GeoGround | Unified large vision-language model for remote sensing visual grounding | arXiv 2024 |
| SkyMoE | Vision-language foundation model with mixture of experts | AAAI 2026 |
| Earth-OneVision | MLLM extended to more remote sensing modalities and tasks | arXiv 2026 |
| TerraScope | Pixel-grounded visual reasoning for Earth observation | CVPR 2026 |
| SkyNative | Native multimodal framework for visual evidence reasoning | arXiv 2026 |
| GeoVLM-R1 | Reinforcement fine-tuning for remote sensing reasoning | arXiv 2025 |
| RemoteReasoner | Unified geospatial reasoning workflow | arXiv 2025 |
| RemoteZero | Geospatial reasoning with zero human annotations | arXiv 2026 |
| GeoX | Geospatial reasoning through self-play and verifiable rewards | arXiv 2026 |
| Name | Category | Paper / Project |
|---|---|---|
| PANGAEA | Global benchmark for geospatial foundation models | arXiv 2024 |
| SpectralEarth-MM | Multimodal, multisensor pretraining data | arXiv 2026 |
| BigEarthNet.txt | Large-scale multi-sensor image-text dataset and benchmark | arXiv 2026 |
| GeoSeg-1M | Open-world geospatial segmentation data | arXiv 2026 |
| UHR-CoZ | Ultra-high-resolution visual focusing / evidence-grounded understanding | arXiv 2026 |
| FusionRS | RGB-infrared remote sensing vision-language dataset | arXiv 2026 |
| Sky-VT-300k | Multimodal in-context segmentation data | CVPR 2026 |
| OpenEarthAgent Dataset | Agentic Earth observation instruction data | arXiv 2026 |
| ChronoEarth-492K | Long-horizon spatiotemporal hyperspectral dataset and benchmark | arXiv 2026 |
| SkyCap | Bitemporal VHR optical-SAR quartets | arXiv 2025 |
| Sentinel2Cap | Human-annotated Sentinel-2 image captioning benchmark | arXiv 2026 |
| VLRS-Bench | Vision-language reasoning benchmark for remote sensing | arXiv 2026 |
| OmniEarth | Geospatial VLM benchmark | arXiv 2026 |
| UHR-Micro | Ultra-high-resolution evidence localization benchmark | arXiv 2026 |
| EarthShift | Robustness benchmark for real-world distribution shifts | arXiv 2026 |
| GeoMMBench | Expert-level multimodal intelligence in geoscience and remote sensing | arXiv 2026 |
| LithoBench | Remote-sensing lithology interpretation benchmark | arXiv 2026 |
| GroundSet | Cadastral-grounded vector-data spatial understanding dataset | arXiv 2026 |
| Name | Focus | Paper / Project |
|---|---|---|
| RemoteAgent | RL-based agentic MLLMs for Earth observation | arXiv 2026 |
| EO-Gym | Interactive environment for Earth observation agents | arXiv 2026 |
| Earth-Agent | Geospatial agentic system | OpenReview 2026 |
| OpenEarthAgent | Open agent framework for Earth observation | arXiv 2026 |
| ThinkGeo | Remote sensing tool orchestration | arXiv 2025 |
| GeoDisaster | Agentic geospatial reasoning for disaster response | arXiv 2026 |
| IC-EO | Interpretable code-based assistant for Earth observation | arXiv 2026 |
| Bidirectional Semantic Complementary Tool Retrieval | Tool retrieval for remote sensing agents | arXiv 2026 |
| Risk-Aware LLM Agents for Geospatial Data Retrieval | Adversarial evaluation for unsafe tool use | arXiv 2026 |
| Atmospheric Retrieval Hijacking | Prompt-injection risk in remote sensing VLM-RAG | arXiv 2026 |
| No One Knows the State of the Art in Geospatial Foundation Models | Audit of comparability and reproducibility problems | arXiv 2026 |
| Dataset Name | Categories | Detailed Info |
|---|---|---|
| Sydney-Captions | Image-Text Pair | Link |
| UCM-Captions | Image-Text Pair | Link |
| RSICD | Image-Text Pair | Link |
| RSVQA-HR | VQA | Link |
| RSVQA-LR | VQA | Link |
| TextRS | Image-Text Pair | Link |
| BigEarthNet-MM | Image-Text Pair | Link |
| FloodNet | VQA | Link |
| RSIVQA | VQA | Link |
| RSVQA×BEN | VQA | Link |
| CDVQA | VQA | Link |
| GeoVG | Visual Localization | Link |
| LEVIR-CC | Image-Text Pair | Link |
| NWPU-Captions | Image-Text Pair | Link |
| RSITMD | Image-Text Pair | Link |
| SSL4EO-S12 | Multimodal Pre-training | Link |
| VQA-TextRS | VQA | Link |
| DIOR_RSVG | Visual Localization | Link |
| LAION-5B | Image-Text Pair | Link |
| OPT-RSVG | Visual Localization | Link |
| RS5M | Image-Text Pair | Link |
| RSICap | Image-Text Pair | Link |
| RSIEval | Image-Text Pair | Link |
| SkyScript | Image-Text Pair | Link |
| ChatEarthNet | Image-Text Pair | Link |
| MMEarth | Multimodal Pre-training | Link |
| Dataset Name | Categories | Detailed Info |
|---|---|---|
| UAV123 | Object Tracking | Link |
| VisDrone2019-MOT | Object Tracking | Link |
| VisDrone2019-SOT | Object Tracking | Link |
| VisDrone2019-VID | Detection | Link |
| 地空背景下红外图像弱小飞机目标检测跟踪数据集 | Object Tracking | Link |
| VISO-MOT | Object Tracking | Link |
| VISO-SOT | Object Tracking | Link |
| 复杂背景下红外弱小运动目标检测数据集 | Object Tracking | Link |
| CapERA | Video Caption | Link |
| Model Name | Paper Name | Published in | Detailed Info |
|---|---|---|---|
| RSICap | RSGPT A Remote Sensing Vision Language Model and Benchmark | Arxiv 2023 | Link |
| EarthGPT | EarthGPT A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain | Arxiv 2024 | Link |
| Geochat | GeoChat Grounded Large Vision-Language Model for Remote Sensing | CVPR 2024 | Link |
| H2RSVLM | H2RSVLM Towards Helpful and Honest Remote Sensing Large Vision Language Model | Arxiv 2024 | Link |
| LHRS-Instruct | LHRS-Bot Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model | Arxiv 2024 | Link |
| Popeye | Popeye A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery | Arxiv 2024 | Link |
| RS-LLaVA | RS-LLaVA Large Vision Language Model for Joint Captioning and Question Answering in Remote Sensing Imagery | RS 2024 | Link |
| SkyEyeGPT | SkyEyeGPT Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model | Arxiv 2024 | Link |
| Model Name | Paper Name | Published in | Detailed Info |
|---|---|---|---|
| Tree-GPT | Tree-GPT Modular Large Language Model Expert System for Forest Remote Sensing Image Understanding and Interactive Analysis | Arxiv 2023 | Link |
| Change-Agent | Change-Agent Towards Interactive Comprehensive Remote Sensing Change Interpretation and Analysis | Arxiv 2024 | Link |
| - | Evaluating Tool-Augmented Agents in Remote Sensing Platforms | ICLR 2024 | Link |
| - | GeoLLM-Engine A Realistic Environment for Building Geospatial Copilots | CVPR 2024 | Link |
| ChatGPT | Remote Sensing ChatGPT Solving Remote Sensing Tasks with ChatGPT and Visual Models | IGARSS 2024 | Link |
This is a repository for "A Survey on Remote Sensing Foundation Models: From Vision to Multimodality".
🌏 Please check out our survey paper: A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
You can see details of all papers and datasets here: homepage
This repository maintains the resource list associated with the survey. It is intended to help readers quickly find papers, datasets, models, benchmarks, and agentic remote sensing systems discussed in the survey.
Last curated update: 2026-06-24.
Corrections and additions are welcome. Please see CONTRIBUTING.md before opening a pull request.
The 2026 revision broadens the repository beyond early vision and vision-language foundation models. The additions below highlight recent resources on sensor-adaptive pretraining, multimodal reasoning, long-tail datasets, robustness benchmarks, open-vocabulary grounding, and agentic Earth observation.
| Name | Focus | Paper / Project |
|---|---|---|
| RAMEN | Resolution-adjustable multimodal encoder for Earth observation | CVPR 2026 |
| THOR | Versatile Earth observation foundation model for climate and society applications | CVPRW 2026 |
| TerraFlow | Multimodal and multitemporal Earth observation representation learning | arXiv 2026 |
| SpectralEarth-FM | Hyperspectral imagery in multimodal Earth observation pretraining | arXiv 2026 |
| FLORO | Multimodal geospatial foundation model across sensors and scales | arXiv 2026 |
| SMARTIES | Spectrum-aware multi-sensor auto-encoder | ICCV 2025 |
| SIGMAE | Spectral-index-guided foundation model for multispectral remote sensing | arXiv 2026 |
| RingMoE | Mixture-of-modality-experts remote sensing foundation model | arXiv 2025 |
| SkySense V2 | Unified foundation model for multimodal remote sensing | arXiv 2025 |
| Falcon | Remote sensing vision-language foundation model | arXiv 2025 |
| GeoGround | Unified large vision-language model for remote sensing visual grounding | arXiv 2024 |
| SkyMoE | Vision-language foundation model with mixture of experts | AAAI 2026 |
| Earth-OneVision | MLLM extended to more remote sensing modalities and tasks | arXiv 2026 |
| TerraScope | Pixel-grounded visual reasoning for Earth observation | CVPR 2026 |
| SkyNative | Native multimodal framework for visual evidence reasoning | arXiv 2026 |
| GeoVLM-R1 | Reinforcement fine-tuning for remote sensing reasoning | arXiv 2025 |
| RemoteReasoner | Unified geospatial reasoning workflow | arXiv 2025 |
| RemoteZero | Geospatial reasoning with zero human annotations | arXiv 2026 |
| GeoX | Geospatial reasoning through self-play and verifiable rewards | arXiv 2026 |
| Name | Category | Paper / Project |
|---|---|---|
| PANGAEA | Global benchmark for geospatial foundation models | arXiv 2024 |
| SpectralEarth-MM | Multimodal, multisensor pretraining data | arXiv 2026 |
| BigEarthNet.txt | Large-scale multi-sensor image-text dataset and benchmark | arXiv 2026 |
| GeoSeg-1M | Open-world geospatial segmentation data | arXiv 2026 |
| UHR-CoZ | Ultra-high-resolution visual focusing / evidence-grounded understanding | arXiv 2026 |
| FusionRS | RGB-infrared remote sensing vision-language dataset | arXiv 2026 |
| Sky-VT-300k | Multimodal in-context segmentation data | CVPR 2026 |
| OpenEarthAgent Dataset | Agentic Earth observation instruction data | arXiv 2026 |
| ChronoEarth-492K | Long-horizon spatiotemporal hyperspectral dataset and benchmark | arXiv 2026 |
| SkyCap | Bitemporal VHR optical-SAR quartets | arXiv 2025 |
| Sentinel2Cap | Human-annotated Sentinel-2 image captioning benchmark | arXiv 2026 |
| VLRS-Bench | Vision-language reasoning benchmark for remote sensing | arXiv 2026 |
| OmniEarth | Geospatial VLM benchmark | arXiv 2026 |
| UHR-Micro | Ultra-high-resolution evidence localization benchmark | arXiv 2026 |
| EarthShift | Robustness benchmark for real-world distribution shifts | arXiv 2026 |
| GeoMMBench | Expert-level multimodal intelligence in geoscience and remote sensing | arXiv 2026 |
| LithoBench | Remote-sensing lithology interpretation benchmark | arXiv 2026 |
| GroundSet | Cadastral-grounded vector-data spatial understanding dataset | arXiv 2026 |
| Name | Focus | Paper / Project |
|---|---|---|
| RemoteAgent | RL-based agentic MLLMs for Earth observation | arXiv 2026 |
| EO-Gym | Interactive environment for Earth observation agents | arXiv 2026 |
| Earth-Agent | Geospatial agentic system | OpenReview 2026 |
| OpenEarthAgent | Open agent framework for Earth observation | arXiv 2026 |
| ThinkGeo | Remote sensing tool orchestration | arXiv 2025 |
| GeoDisaster | Agentic geospatial reasoning for disaster response | arXiv 2026 |
| IC-EO | Interpretable code-based assistant for Earth observation | arXiv 2026 |
| Bidirectional Semantic Complementary Tool Retrieval | Tool retrieval for remote sensing agents | arXiv 2026 |
| Risk-Aware LLM Agents for Geospatial Data Retrieval | Adversarial evaluation for unsafe tool use | arXiv 2026 |
| Atmospheric Retrieval Hijacking | Prompt-injection risk in remote sensing VLM-RAG | arXiv 2026 |
| No One Knows the State of the Art in Geospatial Foundation Models | Audit of comparability and reproducibility problems | arXiv 2026 |
| Dataset Name | Categories | Detailed Info |
|---|---|---|
| Sydney-Captions | Image-Text Pair | Link |
| UCM-Captions | Image-Text Pair | Link |
| RSICD | Image-Text Pair | Link |
| RSVQA-HR | VQA | Link |
| RSVQA-LR | VQA | Link |
| TextRS | Image-Text Pair | Link |
| BigEarthNet-MM | Image-Text Pair | Link |
| FloodNet | VQA | Link |
| RSIVQA | VQA | Link |
| RSVQA×BEN | VQA | Link |
| CDVQA | VQA | Link |
| GeoVG | Visual Localization | Link |
| LEVIR-CC | Image-Text Pair | Link |
| NWPU-Captions | Image-Text Pair | Link |
| RSITMD | Image-Text Pair | Link |
| SSL4EO-S12 | Multimodal Pre-training | Link |
| VQA-TextRS | VQA | Link |
| DIOR_RSVG | Visual Localization | Link |
| LAION-5B | Image-Text Pair | Link |
| OPT-RSVG | Visual Localization | Link |
| RS5M | Image-Text Pair | Link |
| RSICap | Image-Text Pair | Link |
| RSIEval | Image-Text Pair | Link |
| SkyScript | Image-Text Pair | Link |
| ChatEarthNet | Image-Text Pair | Link |
| MMEarth | Multimodal Pre-training | Link |
| Dataset Name | Categories | Detailed Info |
|---|---|---|
| UAV123 | Object Tracking | Link |
| VisDrone2019-MOT | Object Tracking | Link |
| VisDrone2019-SOT | Object Tracking | Link |
| VisDrone2019-VID | Detection | Link |
| 地空背景下红外图像弱小飞机目标检测跟踪数据集 | Object Tracking | Link |
| VISO-MOT | Object Tracking | Link |
| VISO-SOT | Object Tracking | Link |
| 复杂背景下红外弱小运动目标检测数据集 | Object Tracking | Link |
| CapERA | Video Caption | Link |
| Model Name | Paper Name | Published in | Detailed Info |
|---|---|---|---|
| RSICap | RSGPT A Remote Sensing Vision Language Model and Benchmark | Arxiv 2023 | Link |
| EarthGPT | EarthGPT A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain | Arxiv 2024 | Link |
| Geochat | GeoChat Grounded Large Vision-Language Model for Remote Sensing | CVPR 2024 | Link |
| H2RSVLM | H2RSVLM Towards Helpful and Honest Remote Sensing Large Vision Language Model | Arxiv 2024 | Link |
| LHRS-Instruct | LHRS-Bot Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model | Arxiv 2024 | Link |
| Popeye | Popeye A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery | Arxiv 2024 | Link |
| RS-LLaVA | RS-LLaVA Large Vision Language Model for Joint Captioning and Question Answering in Remote Sensing Imagery | RS 2024 | Link |
| SkyEyeGPT | SkyEyeGPT Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model | Arxiv 2024 | Link |
| Model Name | Paper Name | Published in | Detailed Info |
|---|---|---|---|
| Tree-GPT | Tree-GPT Modular Large Language Model Expert System for Forest Remote Sensing Image Understanding and Interactive Analysis | Arxiv 2023 | Link |
| Change-Agent | Change-Agent Towards Interactive Comprehensive Remote Sensing Change Interpretation and Analysis | Arxiv 2024 | Link |
| - | Evaluating Tool-Augmented Agents in Remote Sensing Platforms | ICLR 2024 | Link |
| - | GeoLLM-Engine A Realistic Environment for Building Geospatial Copilots | CVPR 2024 | Link |
| ChatGPT | Remote Sensing ChatGPT Solving Remote Sensing Tasks with ChatGPT and Visual Models | IGARSS 2024 | Link |