IRIP-BUAA/A-Survey-on-Remote-Sensing-Foundation-Models-From-Vision-to-Multimodality

A review for remote sensing vision language models

76

46 commits

updated Jun 24, 2026

See the code

README

A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

This is a repository for "A Survey on Remote Sensing Foundation Models: From Vision to Multimodality".

🌏 Please check out our survey paper: A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

Detail page

You can see details of all papers and datasets here: homepage

Repository Status

This repository maintains the resource list associated with the survey. It is intended to help readers quickly find papers, datasets, models, benchmarks, and agentic remote sensing systems discussed in the survey.

Last curated update: 2026-06-24.

Corrections and additions are welcome. Please see CONTRIBUTING.md before opening a pull request.

Table of Contents

Recent Survey Update

The 2026 revision broadens the repository beyond early vision and vision-language foundation models. The additions below highlight recent resources on sensor-adaptive pretraining, multimodal reasoning, long-tail datasets, robustness benchmarks, open-vocabulary grounding, and agentic Earth observation.

Recent Models and Methods

NameFocusPaper / Project
RAMENResolution-adjustable multimodal encoder for Earth observationCVPR 2026
THORVersatile Earth observation foundation model for climate and society applicationsCVPRW 2026
TerraFlowMultimodal and multitemporal Earth observation representation learningarXiv 2026
SpectralEarth-FMHyperspectral imagery in multimodal Earth observation pretrainingarXiv 2026
FLOROMultimodal geospatial foundation model across sensors and scalesarXiv 2026
SMARTIESSpectrum-aware multi-sensor auto-encoderICCV 2025
SIGMAESpectral-index-guided foundation model for multispectral remote sensingarXiv 2026
RingMoEMixture-of-modality-experts remote sensing foundation modelarXiv 2025
SkySense V2Unified foundation model for multimodal remote sensingarXiv 2025
FalconRemote sensing vision-language foundation modelarXiv 2025
GeoGroundUnified large vision-language model for remote sensing visual groundingarXiv 2024
SkyMoEVision-language foundation model with mixture of expertsAAAI 2026
Earth-OneVisionMLLM extended to more remote sensing modalities and tasksarXiv 2026
TerraScopePixel-grounded visual reasoning for Earth observationCVPR 2026
SkyNativeNative multimodal framework for visual evidence reasoningarXiv 2026
GeoVLM-R1Reinforcement fine-tuning for remote sensing reasoningarXiv 2025
RemoteReasonerUnified geospatial reasoning workflowarXiv 2025
RemoteZeroGeospatial reasoning with zero human annotationsarXiv 2026
GeoXGeospatial reasoning through self-play and verifiable rewardsarXiv 2026

Recent Datasets and Benchmarks

NameCategoryPaper / Project
PANGAEAGlobal benchmark for geospatial foundation modelsarXiv 2024
SpectralEarth-MMMultimodal, multisensor pretraining dataarXiv 2026
BigEarthNet.txtLarge-scale multi-sensor image-text dataset and benchmarkarXiv 2026
GeoSeg-1MOpen-world geospatial segmentation dataarXiv 2026
UHR-CoZUltra-high-resolution visual focusing / evidence-grounded understandingarXiv 2026
FusionRSRGB-infrared remote sensing vision-language datasetarXiv 2026
Sky-VT-300kMultimodal in-context segmentation dataCVPR 2026
OpenEarthAgent DatasetAgentic Earth observation instruction dataarXiv 2026
ChronoEarth-492KLong-horizon spatiotemporal hyperspectral dataset and benchmarkarXiv 2026
SkyCapBitemporal VHR optical-SAR quartetsarXiv 2025
Sentinel2CapHuman-annotated Sentinel-2 image captioning benchmarkarXiv 2026
VLRS-BenchVision-language reasoning benchmark for remote sensingarXiv 2026
OmniEarthGeospatial VLM benchmarkarXiv 2026
UHR-MicroUltra-high-resolution evidence localization benchmarkarXiv 2026
EarthShiftRobustness benchmark for real-world distribution shiftsarXiv 2026
GeoMMBenchExpert-level multimodal intelligence in geoscience and remote sensingarXiv 2026
LithoBenchRemote-sensing lithology interpretation benchmarkarXiv 2026
GroundSetCadastral-grounded vector-data spatial understanding datasetarXiv 2026

Recent Agentic and Trustworthy EO Resources

NameFocusPaper / Project
RemoteAgentRL-based agentic MLLMs for Earth observationarXiv 2026
EO-GymInteractive environment for Earth observation agentsarXiv 2026
Earth-AgentGeospatial agentic systemOpenReview 2026
OpenEarthAgentOpen agent framework for Earth observationarXiv 2026
ThinkGeoRemote sensing tool orchestrationarXiv 2025
GeoDisasterAgentic geospatial reasoning for disaster responsearXiv 2026
IC-EOInterpretable code-based assistant for Earth observationarXiv 2026
Bidirectional Semantic Complementary Tool RetrievalTool retrieval for remote sensing agentsarXiv 2026
Risk-Aware LLM Agents for Geospatial Data RetrievalAdversarial evaluation for unsafe tool usearXiv 2026
Atmospheric Retrieval HijackingPrompt-injection risk in remote sensing VLM-RAGarXiv 2026
No One Knows the State of the Art in Geospatial Foundation ModelsAudit of comparability and reproducibility problemsarXiv 2026

Dataset

Image

Dataset NameCategoriesDetailed Info
TASDetectionLink
OIRDSDetectionLink
SZTAKI-INRIA AirChangeChange DetectionLink
UCMerced_LandUseClassificationLink
ISPRS PotsdamSegmentationLink
ISPRS VaihingenSegmentationLink
WHU RS19ClassificationLink
Massachusetts BuildsSegmentationLink
Massachusetts RoadsSegmentationLink
SPARCSSegmentationLink
RSSCN7ClassificationLink
SATClassificationLink
VEDAIDetectionLink
DLR3kDetectionLink
HRSC2016DetectionLink
NWPU-RESISC45ClassificationLink
RS_C11ClassificationLink
SIRI-WHUClassificationLink
Aerial to MapImage GenerationLink
AIDClassificationLink
AIST Building Change Detection(ABCD)Change DetectionLink
CITY-OSMSegmentationLink
Dstl Satellite Imagery Feature DetectionSegmentationLink
OpenSARShipDetectionLink
RSD46-WHUClassificationLink
RSI-CBClassificationLink
SateHaze1kImage GenerationLink
TGRS-HRRSDDetectionLink
2018 Open AI Tanzania Building FootprintSegmentationLink
AeroscapesSegmentationLink
AIRSSegmentationLink
CDD Dataset (season-varying)Change DetectionLink
DeepGlobe Land Cover ClassificatSegmentationLink
DeepGlobe Road Detection ChallenSegmentationLink
DLRSDSegmentationLink
DOTA1.0DetectionLink
EuroSATClassificationLink
fMoWDetectionLink
ITCVDDetectionLink
LEVIRDetectionLink
Mapping ChallengeSegmentationLink
MASATIDetectionLink
Onera Satellite Change Detection (OSCD)Change DetectionLink
PatternNetClassificationLink
RIT-18SegmentationLink
Urban Drone Dataset(UDD)SegmentationLink
VisDrone2019-DETDetectionLink
WHDLDSegmentationLink
WHU Building Change Detection DatasetChange DetectionLink
The "Fine" part in WHU_GIDSegmentationLink
The "Large-scale" part in WHU_GIDSegmentationLink
The "secenClass" part in WHU_GIDClassificationLink
xViewDetectionLink
38-CloudSegmentationLink
Bijie Landslide DatasetSegmentationLink
Bridge DatasetDetectionLink
Change Detection Dataset(CDD)Change DetectionLink
DOTA1.5DetectionLink
DroneDeploySegmentationLink
DSIFN DatasetChange DetectionLink
GF2 Dataset for 3DFGCSegmentationLink
High Resolution Semantic Change (HRSCD)Change DetectionLink
iSAIDDetectionLink
Multi-temporal Scene WuHan (MtS-WH)Change DetectionLink
OPTIMAL-31ClassificationLink
ORSSDSegmentationLink
RoadTracerSegmentationLink
Semantic Drone DatasetSegmentationLink
SEN12MSSegmentationLink
WHU Building DatasetSegmentationLink
95-CloudSegmentationLink
BDCI2020SegmentationLink
BH-POOLSSegmentationLink
BH-WATERTANKSSegmentationLink
Change-Detection-dataset-for-High-resolution-Satellite-ImageChange DetectionLink
Continual Learning Benchmark for RemoteClassificationLink
DroneCrowdDetectionLink
EORSSDSegmentationLink
HRSIDDetectionLink
landcover_aiSegmentationLink
LEVIR-CDChange DetectionLink
MLRSNetClassificationLink
The "AiRound" part in Multi-View DatasetsClassificationLink
The "CV-BrCT" part in Multi-View DatasetsClassificationLink
RarePlanesDetectionLink
SEmantic Change detectiON Data(SECOND)Change DetectionLink
SenseEarth ChangeDetectionChange DetectionLink
Sentinel-2 Cloud Mask CatalogueSegmentationLink
Sentinel-2 Multitemporal Cities PairsChange DetectionLink
UAVidSegmentationLink
WHU Cloud DatasetSegmentationLink
WHU Multi-view DatasetImage GenerationLink
WHU Stereo DatasetImage GenerationLink
xBDChange DetectionLink
全国人工智能大赛AI遥感影像SegmentationLink
CASIA-aircraftDetectionLink
CASIA-ShipDetectionLink
DOTA2.0DetectionLink
FAIR1MDetectionLink
LEVIR-CD2Change DetectionLink
LoveDASegmentationLink
MillionAIDClassificationLink
NaSC-TG2ClassificationLink
S2LookingChange DetectionLink
SeCoSingle-modal Pre-trainingLink
Sun Yat-Sen University (SYSU)-CDChange DetectionLink
TG1HRSSCClassificationLink
VISO-DetectionDetectionLink
WHU TCL SatMVS dataset1.0Image GenerationLink
WHU TCL SatMVS dataset2.0Image GenerationLink
MiniFrance-DFC22SegmentationLink
SAMRSSegmentationLink
TOV-RS-balancedClassificationLink
CACoSingle-modal Pre-trainingLink
GEO-BenchEvaluationLink
SATINClassificationLink
SatlasPretrainClassificationLink
WHU-Mix (vector) building datasetSegmentationLink
ConstellationDetectionLink
Earth Parser DatasetSegmentationLink
EarthViewSingle-modal Pre-trainingLink
GeoPile-2Multimodal Pre-trainingLink
HSODBIT-V1DetectionLink
SARDet-100KDetectionLink
SARSimDetection-

Image+Text

Dataset NameCategoriesDetailed Info
Sydney-CaptionsImage-Text PairLink
UCM-CaptionsImage-Text PairLink
RSICDImage-Text PairLink
RSVQA-HRVQALink
RSVQA-LRVQALink
TextRSImage-Text PairLink
BigEarthNet-MMImage-Text PairLink
FloodNetVQALink
RSIVQAVQALink
RSVQA×BENVQALink
CDVQAVQALink
GeoVGVisual LocalizationLink
LEVIR-CCImage-Text PairLink
NWPU-CaptionsImage-Text PairLink
RSITMDImage-Text PairLink
SSL4EO-S12Multimodal Pre-trainingLink
VQA-TextRSVQALink
DIOR_RSVGVisual LocalizationLink
LAION-5BImage-Text PairLink
OPT-RSVGVisual LocalizationLink
RS5MImage-Text PairLink
RSICapImage-Text PairLink
RSIEvalImage-Text PairLink
SkyScriptImage-Text PairLink
ChatEarthNetImage-Text PairLink
MMEarthMultimodal Pre-trainingLink

Video

Dataset NameCategoriesDetailed Info
UAV123Object TrackingLink
VisDrone2019-MOTObject TrackingLink
VisDrone2019-SOTObject TrackingLink
VisDrone2019-VIDDetectionLink
地空背景下红外图像弱小飞机目标检测跟踪数据集Object TrackingLink
VISO-MOTObject TrackingLink
VISO-SOTObject TrackingLink
复杂背景下红外弱小运动目标检测数据集Object TrackingLink
CapERAVideo CaptionLink

Model

Pretrain

Model NamePaper NamePublished inDetailed Info
fMoWFunctional Map of the WorldCVPR 2018Link
-BIGEARTHNET A LARGE-SCALE BENCHMARK ARCHIVE FOR REMOTE SENSINGIMAGE UNDERSTANDINGIGARSS 2019Link
-Tile2Vec Unsupervised representation learning for spatially distributed dataAAAI 2019Link
-Remote Sensing Image Scene Classification with Self-Supervised Paradigm under Limited Labeled SamplesGRSL 2020Link
GASSLGeography-aware self-supervised learningICCV 2021Link
MillionAIDOn Creating Benchmark Dataset for Aerial Image Interpretation Reviews, Guidances, and Million-AIDJSTARS 2021Link
SecoSeasonal ContrastUnsupervised Pre-Training from Uncurated Remote Sensing DataCVPR 2021Link
CMC-RSSRSelf-Supervised Learning of Remote Sensing Scene Representations Using Contrastive Multiview CodingCVPR 2021Link
-Advancing plain vision transformer toward remote sensing foundation modelTGRS 2022Link
CSPTConsecutive Pre-Training A Knowledge Transfer Learning Strategy with Relevant Unlabeled Data for Remote Sensing DomainRemote Sensing 2022Link
Levir-KRGeographical Knowledge-Driven RepresentationLearning for Remote Sensing ImagesTGRS 2022Link
-Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing ImagesTGRS 2022Link
SatMAESatMAE Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryNeurIPS 2022Link
RSBYOLSelf-Supervised Learning for Invariant Representations from Multi-Spectral and SAR ImagesJSTARS 2022Link
MATTERSelf-Supervised Material and Texture Representation Learning for Remote Sensing TasksCVPR 2022Link
DINO-MMSelf-supervised vision transformers for joint sar-optical representation learningArxiv 2022Link
-Semantic segmentation of remote sensing images with self-supervised semantic-aware inpaintingGRSL 2022Link
-A billion-scale foundation model for remote sensing imagesArxiv 2023Link
-A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain FusionIGARSS 2023Link
-An Empirical Study of Remote Sensing PretrainingTGRS 2023Link
CACoChange-Aware Sampling and Contrastive Learning for Satellite ImagesCVPR 2023Link
CMIDCMID A Unified Self-Supervised Learning Framework for Remote Sensing Image UnderstandingTGRS 2023Link
CROMACROMA Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersNeurIPS 2023Link
MAECross-Scale MAE A Tale of Multiscale Exploitation in Remote SensingNeurIPS 2023Link
CtxMIMCtxMIM Context-Enhanced Masked Image Modeling for Remote Sensing Image UnderstandingArxiv 2023Link
DeCURDeCUR decoupling common & unique representations for multimodal self-supervisionArxiv 2023Link
DINO-MCExtending Global-local View Alignment for Self-supervised Learning with Remote Sensing ImageryArxiv 2023Link
EarthPTEarthPT a foundation model for Earth ObservationNeurIPS CCAI workshop 2023Link
FG-MAEFeature Guided Masked Autoencoder for Self-supervised Learning in Remote SensingArxiv 2023Link
-FoMo-Bench a multi-modal, multi-scale and multi-task Forest Monitoring Benchmark for remote sensing foundation modelsArxiv 2023Link
PrithviFoundation Models for Generalist Geospatial Artificial IntelligenceArxiv 2023Link
PrestoLightweight, Pre-trained Transformers for Remote Sensing TimeseriesArxiv 2023Link
IaI-SimCLRMulti Modal Multi Objective Contrastive Learning for Sentinel 1-2 ImageryCVPRW 2023Link
SAR-JEPAPredicting Gradient is Better Exploring Self-Supervised Learning for SAR ATR with a Joint-Embedding Predictive ArchitectureArxiv 2023Link
RingMo-liteRingMo-lite A Remote Sensing Multi-task Lightweight Network with CNN-Transformer Hybrid FrameworkArxiv 2023Link
RSPrompterRsprompter Learning to prompt for remote sensing instance segmentation based on visual foundation modelArxiv 2023Link
SatlasPretrainSatlasPretrain A Large-Scale Dataset for Remote Sensing Image UnderstandingICCV 2023Link
-Scale-MAE A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningICCV 2023Link
TOV-RSTOV The original vision model for optical remote sensing image understanding via self-supervised learningJSTARS 2023Link
TOV-RSTOV The Original Vision Model for Optical Remote Sensing Image Understanding via Self-Supervised LearningJSTARS 2023Link
GFMTowards Geospatial Foundation Models via Continual PretrainingICCV 2023Link
USatUSat A Unified Self-Supervised Encoder for Multi-Sensor Satellite ImageryArxiv 2023Link
GeoSenseGenerative ConvNet Foundation Model With Sparse Modeling and Low-Frequency Reconstruction for Remote Sensing Image InterpretationTGRS 2024Link
GeRSPGeneric Knowledge Boosted Pre-training ForRemote Sensing ImagesTGRS 2024Link
MTPMTP Advancing Remote Sensing FoundationModel via Multi-Task PretrainingArxiv 2024Link
OFA-NetOne for All Toward Unified Foundation Models for Earth VisionArxiv 2024Link
-Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryCVPR 2024Link
RingMoRingMo A Remote Sensing Foundation Model With Masked Image ModelingTGRS 2024Link
S2MAES2MAE A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataCVPR 2024Link
U-BARNSelf-Supervised Spatio-Temporal Representation Learning of Satellite Image Time SeriesJSTARS 2024Link
SkySenseSkySense A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryCVPR 2024Link
SpectralGPTSpectralGPT Spectral Foundation ModelTPAMI 2024Link
SwiMDiffSwiMDiff Scene-wide Matching Contrastive Learning with Diffusion Constraint for Remote Sensing ImageArxiv 2024Link
-Multi-source remote sensing pretraining based on contrastive self-supervised learningRemote Sensing 22Link
RSICapRSGPT: A Remote Sensing Vision Language Model and BenchmarkArxiv 2023-
HyperSIGMAHyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelIEEE TPAMI 2025Link
RS-CLIPRS-CLIP: Zero shot remote sensing scene classification via contrastive vision-language supervisionInternational Journal of Applied Earth Observation and Geoinformation 2023Link
StreetCLIPLearning generalized zero-shot learners for open-domain image geolocalizationArxiv 2024-
GeoCLAPLearning Tri-modal Embeddings for Zero-Shot Soundscape MappingArxiv 2024-
RSDiffRsdiff: Remote sensing image generation from text using diffusion modelArxiv 2024-
DiffusionSatDiffusionsat: A generative foundation model for satellite imageryArxiv 2024-
CRS-DiffCrs-diff: Controllable generative remote sensing foundation modelArxiv 2024-
MetaEarthMetaearth: A generative foundation model for global-scale remote sensing image generationArxiv 2024-
SkysensegptSkysensegpt: A fine-grained instruction tuning dataset and model for remote sensing vision-language understandingArxiv 2024-
RS-AgentRS-Agent: Automating Remote Sensing Tasks through Intelligent AgentsArxiv 2024-
SSLTransformerRSSelf-Supervised Vision Transformers for Land-Cover Segmentation and ClassificationArxiv 2023-
RingMo-SenseRingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution DisentanglingArxiv 2023-
A2-MAEA2-MAE: A spatial-temporal-spectral unified remote sensing pre-training method based on anchor-aware masked autoencoderArxiv 2023-
MSFEIGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing SymposiumIGARSS 2023-
GeCoGeographical Supervision Correction for Remote Sensing Representation LearningArxiv 2023-
SoftConMulti-label Guided Soft Contrastive Learning for Efficient Earth Observation PretrainingCVPR 2023-
Sen12MSSEN12MS--A curated dataset of georeferenced multi-spectral sentinel-1/2 imagery for deep learning and data fusionIEEE TGRS 2019-
BigEarthNet-S2Bigearthnet: A large-scale benchmark archive for remote sensing image understandingIEEE TGRS 2019-
MSARLarge-scale multi-class SAR image target detection dataset-1.0Journal of Radars 2022-
SAR-ShipSAR target recognition based on cross-domain and cross-task transfer learningIEEE Access 2019-
SAMPLEA SAR dataset for ATR development: the Synthetic and Measured Paired Labeled Experiment (SAMPLE)SPIE 2019-
DIOR-RSVGRsvg: Exploring data and models for visual grounding on remote sensing dataIEEE TGRS 2023-
Xviewxview: Objects in context in overhead imageryArxiv 2018-
DIOR-RAnchor-free oriented proposal generator for object detectionIEEE TGRS 2022-
SpaceNetv1Spacenet: A remote sensing dataset and challenge seriesArxiv 2018-
DynamicEarthNet-S2Dynamicearthnet: Daily multi-spectral satellite dataset for semantic change segmentationCVPR 2022-

VLM

MLLM

Agent

Other

Model NamePaper NamePublished inDetailed Info
-Transforming remote sensing images to textual descriptionsINT J APPL EARTH OBS 2022Link
-CSP Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual RepresentationsICML 2023Link
GeoCLIPGeoCLIP Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationNeurIPS 2023Link
-Good at captioning, bad at counting Benchmarking GPT-4V on Earth observation dataArxiv 2023Link
-On the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning ApplicationsArxiv 2023Link
SatCLIPSatCLIP Global, General-Purpose Location Embeddings with Satellite ImageryArxiv 2023Link
-The Potential of Visual ChatGPT for Remote SensingRS 2023Link
-遥感基础模型发展综述与未来设想遥感学报 2023Link
msGFMBridging Remote Sensors with Multisensor Geospatial Foundation ModelsCVPR 2024Link
-Charting New Territories Exploring the Geographic and Geospatial Capabilities of Multimodal LLMsArxiv 2024Link
-GeoLLM Extracting Geospatial Knowledge from Large Language ModelsICLR 2024Link
LeMeViTLeMeViT Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image InterpretationIJCAI 2024Link
MMEarthMMEarth Exploring Multi-Modal Pretext Tasks For Geospatial Representation LearningArxiv 2024Link
DOFANeural Plasticity-Inspired Foundation Model for Observing the Earth Crossing ModalitiesArxiv 2024Link
-On the Foundations of Earth and Climate Foundation ModelsArxiv 2024Link
SARATR-XSARATR-X A Foundation Model for Synthetic Aperture Radar Images Target RecognitionArxiv 2024Link
-Changes to Captions An Attentive Network for Remote Sensing Change CaptioningArxiv 2023Link
-Multi-source interactive stair attention for remote sensing image captioningRS 2023Link

IRIP-BUAA/A-Survey-on-Remote-Sensing-Foundation-Models-From-Vision-to-Multimodality

A review for remote sensing vision language models

76

46 commits

updated Jun 24, 2026

See the code

README

A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

This is a repository for "A Survey on Remote Sensing Foundation Models: From Vision to Multimodality".

🌏 Please check out our survey paper: A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

Detail page

You can see details of all papers and datasets here: homepage

Repository Status

This repository maintains the resource list associated with the survey. It is intended to help readers quickly find papers, datasets, models, benchmarks, and agentic remote sensing systems discussed in the survey.

Last curated update: 2026-06-24.

Corrections and additions are welcome. Please see CONTRIBUTING.md before opening a pull request.

Table of Contents

Recent Survey Update

The 2026 revision broadens the repository beyond early vision and vision-language foundation models. The additions below highlight recent resources on sensor-adaptive pretraining, multimodal reasoning, long-tail datasets, robustness benchmarks, open-vocabulary grounding, and agentic Earth observation.

Recent Models and Methods

NameFocusPaper / Project
RAMENResolution-adjustable multimodal encoder for Earth observationCVPR 2026
THORVersatile Earth observation foundation model for climate and society applicationsCVPRW 2026
TerraFlowMultimodal and multitemporal Earth observation representation learningarXiv 2026
SpectralEarth-FMHyperspectral imagery in multimodal Earth observation pretrainingarXiv 2026
FLOROMultimodal geospatial foundation model across sensors and scalesarXiv 2026
SMARTIESSpectrum-aware multi-sensor auto-encoderICCV 2025
SIGMAESpectral-index-guided foundation model for multispectral remote sensingarXiv 2026
RingMoEMixture-of-modality-experts remote sensing foundation modelarXiv 2025
SkySense V2Unified foundation model for multimodal remote sensingarXiv 2025
FalconRemote sensing vision-language foundation modelarXiv 2025
GeoGroundUnified large vision-language model for remote sensing visual groundingarXiv 2024
SkyMoEVision-language foundation model with mixture of expertsAAAI 2026
Earth-OneVisionMLLM extended to more remote sensing modalities and tasksarXiv 2026
TerraScopePixel-grounded visual reasoning for Earth observationCVPR 2026
SkyNativeNative multimodal framework for visual evidence reasoningarXiv 2026
GeoVLM-R1Reinforcement fine-tuning for remote sensing reasoningarXiv 2025
RemoteReasonerUnified geospatial reasoning workflowarXiv 2025
RemoteZeroGeospatial reasoning with zero human annotationsarXiv 2026
GeoXGeospatial reasoning through self-play and verifiable rewardsarXiv 2026

Recent Datasets and Benchmarks

NameCategoryPaper / Project
PANGAEAGlobal benchmark for geospatial foundation modelsarXiv 2024
SpectralEarth-MMMultimodal, multisensor pretraining dataarXiv 2026
BigEarthNet.txtLarge-scale multi-sensor image-text dataset and benchmarkarXiv 2026
GeoSeg-1MOpen-world geospatial segmentation dataarXiv 2026
UHR-CoZUltra-high-resolution visual focusing / evidence-grounded understandingarXiv 2026
FusionRSRGB-infrared remote sensing vision-language datasetarXiv 2026
Sky-VT-300kMultimodal in-context segmentation dataCVPR 2026
OpenEarthAgent DatasetAgentic Earth observation instruction dataarXiv 2026
ChronoEarth-492KLong-horizon spatiotemporal hyperspectral dataset and benchmarkarXiv 2026
SkyCapBitemporal VHR optical-SAR quartetsarXiv 2025
Sentinel2CapHuman-annotated Sentinel-2 image captioning benchmarkarXiv 2026
VLRS-BenchVision-language reasoning benchmark for remote sensingarXiv 2026
OmniEarthGeospatial VLM benchmarkarXiv 2026
UHR-MicroUltra-high-resolution evidence localization benchmarkarXiv 2026
EarthShiftRobustness benchmark for real-world distribution shiftsarXiv 2026
GeoMMBenchExpert-level multimodal intelligence in geoscience and remote sensingarXiv 2026
LithoBenchRemote-sensing lithology interpretation benchmarkarXiv 2026
GroundSetCadastral-grounded vector-data spatial understanding datasetarXiv 2026

Recent Agentic and Trustworthy EO Resources

NameFocusPaper / Project
RemoteAgentRL-based agentic MLLMs for Earth observationarXiv 2026
EO-GymInteractive environment for Earth observation agentsarXiv 2026
Earth-AgentGeospatial agentic systemOpenReview 2026
OpenEarthAgentOpen agent framework for Earth observationarXiv 2026
ThinkGeoRemote sensing tool orchestrationarXiv 2025
GeoDisasterAgentic geospatial reasoning for disaster responsearXiv 2026
IC-EOInterpretable code-based assistant for Earth observationarXiv 2026
Bidirectional Semantic Complementary Tool RetrievalTool retrieval for remote sensing agentsarXiv 2026
Risk-Aware LLM Agents for Geospatial Data RetrievalAdversarial evaluation for unsafe tool usearXiv 2026
Atmospheric Retrieval HijackingPrompt-injection risk in remote sensing VLM-RAGarXiv 2026
No One Knows the State of the Art in Geospatial Foundation ModelsAudit of comparability and reproducibility problemsarXiv 2026

Dataset

Image

Dataset NameCategoriesDetailed Info
TASDetectionLink
OIRDSDetectionLink
SZTAKI-INRIA AirChangeChange DetectionLink
UCMerced_LandUseClassificationLink
ISPRS PotsdamSegmentationLink
ISPRS VaihingenSegmentationLink
WHU RS19ClassificationLink
Massachusetts BuildsSegmentationLink
Massachusetts RoadsSegmentationLink
SPARCSSegmentationLink
RSSCN7ClassificationLink
SATClassificationLink
VEDAIDetectionLink
DLR3kDetectionLink
HRSC2016DetectionLink
NWPU-RESISC45ClassificationLink
RS_C11ClassificationLink
SIRI-WHUClassificationLink
Aerial to MapImage GenerationLink
AIDClassificationLink
AIST Building Change Detection(ABCD)Change DetectionLink
CITY-OSMSegmentationLink
Dstl Satellite Imagery Feature DetectionSegmentationLink
OpenSARShipDetectionLink
RSD46-WHUClassificationLink
RSI-CBClassificationLink
SateHaze1kImage GenerationLink
TGRS-HRRSDDetectionLink
2018 Open AI Tanzania Building FootprintSegmentationLink
AeroscapesSegmentationLink
AIRSSegmentationLink
CDD Dataset (season-varying)Change DetectionLink
DeepGlobe Land Cover ClassificatSegmentationLink
DeepGlobe Road Detection ChallenSegmentationLink
DLRSDSegmentationLink
DOTA1.0DetectionLink
EuroSATClassificationLink
fMoWDetectionLink
ITCVDDetectionLink
LEVIRDetectionLink
Mapping ChallengeSegmentationLink
MASATIDetectionLink
Onera Satellite Change Detection (OSCD)Change DetectionLink
PatternNetClassificationLink
RIT-18SegmentationLink
Urban Drone Dataset(UDD)SegmentationLink
VisDrone2019-DETDetectionLink
WHDLDSegmentationLink
WHU Building Change Detection DatasetChange DetectionLink
The "Fine" part in WHU_GIDSegmentationLink
The "Large-scale" part in WHU_GIDSegmentationLink
The "secenClass" part in WHU_GIDClassificationLink
xViewDetectionLink
38-CloudSegmentationLink
Bijie Landslide DatasetSegmentationLink
Bridge DatasetDetectionLink
Change Detection Dataset(CDD)Change DetectionLink
DOTA1.5DetectionLink
DroneDeploySegmentationLink
DSIFN DatasetChange DetectionLink
GF2 Dataset for 3DFGCSegmentationLink
High Resolution Semantic Change (HRSCD)Change DetectionLink
iSAIDDetectionLink
Multi-temporal Scene WuHan (MtS-WH)Change DetectionLink
OPTIMAL-31ClassificationLink
ORSSDSegmentationLink
RoadTracerSegmentationLink
Semantic Drone DatasetSegmentationLink
SEN12MSSegmentationLink
WHU Building DatasetSegmentationLink
95-CloudSegmentationLink
BDCI2020SegmentationLink
BH-POOLSSegmentationLink
BH-WATERTANKSSegmentationLink
Change-Detection-dataset-for-High-resolution-Satellite-ImageChange DetectionLink
Continual Learning Benchmark for RemoteClassificationLink
DroneCrowdDetectionLink
EORSSDSegmentationLink
HRSIDDetectionLink
landcover_aiSegmentationLink
LEVIR-CDChange DetectionLink
MLRSNetClassificationLink
The "AiRound" part in Multi-View DatasetsClassificationLink
The "CV-BrCT" part in Multi-View DatasetsClassificationLink
RarePlanesDetectionLink
SEmantic Change detectiON Data(SECOND)Change DetectionLink
SenseEarth ChangeDetectionChange DetectionLink
Sentinel-2 Cloud Mask CatalogueSegmentationLink
Sentinel-2 Multitemporal Cities PairsChange DetectionLink
UAVidSegmentationLink
WHU Cloud DatasetSegmentationLink
WHU Multi-view DatasetImage GenerationLink
WHU Stereo DatasetImage GenerationLink
xBDChange DetectionLink
全国人工智能大赛AI遥感影像SegmentationLink
CASIA-aircraftDetectionLink
CASIA-ShipDetectionLink
DOTA2.0DetectionLink
FAIR1MDetectionLink
LEVIR-CD2Change DetectionLink
LoveDASegmentationLink
MillionAIDClassificationLink
NaSC-TG2ClassificationLink
S2LookingChange DetectionLink
SeCoSingle-modal Pre-trainingLink
Sun Yat-Sen University (SYSU)-CDChange DetectionLink
TG1HRSSCClassificationLink
VISO-DetectionDetectionLink
WHU TCL SatMVS dataset1.0Image GenerationLink
WHU TCL SatMVS dataset2.0Image GenerationLink
MiniFrance-DFC22SegmentationLink
SAMRSSegmentationLink
TOV-RS-balancedClassificationLink
CACoSingle-modal Pre-trainingLink
GEO-BenchEvaluationLink
SATINClassificationLink
SatlasPretrainClassificationLink
WHU-Mix (vector) building datasetSegmentationLink
ConstellationDetectionLink
Earth Parser DatasetSegmentationLink
EarthViewSingle-modal Pre-trainingLink
GeoPile-2Multimodal Pre-trainingLink
HSODBIT-V1DetectionLink
SARDet-100KDetectionLink
SARSimDetection-

Image+Text

Dataset NameCategoriesDetailed Info
Sydney-CaptionsImage-Text PairLink
UCM-CaptionsImage-Text PairLink
RSICDImage-Text PairLink
RSVQA-HRVQALink
RSVQA-LRVQALink
TextRSImage-Text PairLink
BigEarthNet-MMImage-Text PairLink
FloodNetVQALink
RSIVQAVQALink
RSVQA×BENVQALink
CDVQAVQALink
GeoVGVisual LocalizationLink
LEVIR-CCImage-Text PairLink
NWPU-CaptionsImage-Text PairLink
RSITMDImage-Text PairLink
SSL4EO-S12Multimodal Pre-trainingLink
VQA-TextRSVQALink
DIOR_RSVGVisual LocalizationLink
LAION-5BImage-Text PairLink
OPT-RSVGVisual LocalizationLink
RS5MImage-Text PairLink
RSICapImage-Text PairLink
RSIEvalImage-Text PairLink
SkyScriptImage-Text PairLink
ChatEarthNetImage-Text PairLink
MMEarthMultimodal Pre-trainingLink

Video

Dataset NameCategoriesDetailed Info
UAV123Object TrackingLink
VisDrone2019-MOTObject TrackingLink
VisDrone2019-SOTObject TrackingLink
VisDrone2019-VIDDetectionLink
地空背景下红外图像弱小飞机目标检测跟踪数据集Object TrackingLink
VISO-MOTObject TrackingLink
VISO-SOTObject TrackingLink
复杂背景下红外弱小运动目标检测数据集Object TrackingLink
CapERAVideo CaptionLink

Model

Pretrain

Model NamePaper NamePublished inDetailed Info
fMoWFunctional Map of the WorldCVPR 2018Link
-BIGEARTHNET A LARGE-SCALE BENCHMARK ARCHIVE FOR REMOTE SENSINGIMAGE UNDERSTANDINGIGARSS 2019Link
-Tile2Vec Unsupervised representation learning for spatially distributed dataAAAI 2019Link
-Remote Sensing Image Scene Classification with Self-Supervised Paradigm under Limited Labeled SamplesGRSL 2020Link
GASSLGeography-aware self-supervised learningICCV 2021Link
MillionAIDOn Creating Benchmark Dataset for Aerial Image Interpretation Reviews, Guidances, and Million-AIDJSTARS 2021Link
SecoSeasonal ContrastUnsupervised Pre-Training from Uncurated Remote Sensing DataCVPR 2021Link
CMC-RSSRSelf-Supervised Learning of Remote Sensing Scene Representations Using Contrastive Multiview CodingCVPR 2021Link
-Advancing plain vision transformer toward remote sensing foundation modelTGRS 2022Link
CSPTConsecutive Pre-Training A Knowledge Transfer Learning Strategy with Relevant Unlabeled Data for Remote Sensing DomainRemote Sensing 2022Link
Levir-KRGeographical Knowledge-Driven RepresentationLearning for Remote Sensing ImagesTGRS 2022Link
-Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing ImagesTGRS 2022Link
SatMAESatMAE Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryNeurIPS 2022Link
RSBYOLSelf-Supervised Learning for Invariant Representations from Multi-Spectral and SAR ImagesJSTARS 2022Link
MATTERSelf-Supervised Material and Texture Representation Learning for Remote Sensing TasksCVPR 2022Link
DINO-MMSelf-supervised vision transformers for joint sar-optical representation learningArxiv 2022Link
-Semantic segmentation of remote sensing images with self-supervised semantic-aware inpaintingGRSL 2022Link
-A billion-scale foundation model for remote sensing imagesArxiv 2023Link
-A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain FusionIGARSS 2023Link
-An Empirical Study of Remote Sensing PretrainingTGRS 2023Link
CACoChange-Aware Sampling and Contrastive Learning for Satellite ImagesCVPR 2023Link
CMIDCMID A Unified Self-Supervised Learning Framework for Remote Sensing Image UnderstandingTGRS 2023Link
CROMACROMA Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersNeurIPS 2023Link
MAECross-Scale MAE A Tale of Multiscale Exploitation in Remote SensingNeurIPS 2023Link
CtxMIMCtxMIM Context-Enhanced Masked Image Modeling for Remote Sensing Image UnderstandingArxiv 2023Link
DeCURDeCUR decoupling common & unique representations for multimodal self-supervisionArxiv 2023Link
DINO-MCExtending Global-local View Alignment for Self-supervised Learning with Remote Sensing ImageryArxiv 2023Link
EarthPTEarthPT a foundation model for Earth ObservationNeurIPS CCAI workshop 2023Link
FG-MAEFeature Guided Masked Autoencoder for Self-supervised Learning in Remote SensingArxiv 2023Link
-FoMo-Bench a multi-modal, multi-scale and multi-task Forest Monitoring Benchmark for remote sensing foundation modelsArxiv 2023Link
PrithviFoundation Models for Generalist Geospatial Artificial IntelligenceArxiv 2023Link
PrestoLightweight, Pre-trained Transformers for Remote Sensing TimeseriesArxiv 2023Link
IaI-SimCLRMulti Modal Multi Objective Contrastive Learning for Sentinel 1-2 ImageryCVPRW 2023Link
SAR-JEPAPredicting Gradient is Better Exploring Self-Supervised Learning for SAR ATR with a Joint-Embedding Predictive ArchitectureArxiv 2023Link
RingMo-liteRingMo-lite A Remote Sensing Multi-task Lightweight Network with CNN-Transformer Hybrid FrameworkArxiv 2023Link
RSPrompterRsprompter Learning to prompt for remote sensing instance segmentation based on visual foundation modelArxiv 2023Link
SatlasPretrainSatlasPretrain A Large-Scale Dataset for Remote Sensing Image UnderstandingICCV 2023Link
-Scale-MAE A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningICCV 2023Link
TOV-RSTOV The original vision model for optical remote sensing image understanding via self-supervised learningJSTARS 2023Link
TOV-RSTOV The Original Vision Model for Optical Remote Sensing Image Understanding via Self-Supervised LearningJSTARS 2023Link
GFMTowards Geospatial Foundation Models via Continual PretrainingICCV 2023Link
USatUSat A Unified Self-Supervised Encoder for Multi-Sensor Satellite ImageryArxiv 2023Link
GeoSenseGenerative ConvNet Foundation Model With Sparse Modeling and Low-Frequency Reconstruction for Remote Sensing Image InterpretationTGRS 2024Link
GeRSPGeneric Knowledge Boosted Pre-training ForRemote Sensing ImagesTGRS 2024Link
MTPMTP Advancing Remote Sensing FoundationModel via Multi-Task PretrainingArxiv 2024Link
OFA-NetOne for All Toward Unified Foundation Models for Earth VisionArxiv 2024Link
-Rethinking Transformers Pre-training for Multi-Spectral Satellite ImageryCVPR 2024Link
RingMoRingMo A Remote Sensing Foundation Model With Masked Image ModelingTGRS 2024Link
S2MAES2MAE A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing DataCVPR 2024Link
U-BARNSelf-Supervised Spatio-Temporal Representation Learning of Satellite Image Time SeriesJSTARS 2024Link
SkySenseSkySense A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryCVPR 2024Link
SpectralGPTSpectralGPT Spectral Foundation ModelTPAMI 2024Link
SwiMDiffSwiMDiff Scene-wide Matching Contrastive Learning with Diffusion Constraint for Remote Sensing ImageArxiv 2024Link
-Multi-source remote sensing pretraining based on contrastive self-supervised learningRemote Sensing 22Link
RSICapRSGPT: A Remote Sensing Vision Language Model and BenchmarkArxiv 2023-
HyperSIGMAHyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelIEEE TPAMI 2025Link
RS-CLIPRS-CLIP: Zero shot remote sensing scene classification via contrastive vision-language supervisionInternational Journal of Applied Earth Observation and Geoinformation 2023Link
StreetCLIPLearning generalized zero-shot learners for open-domain image geolocalizationArxiv 2024-
GeoCLAPLearning Tri-modal Embeddings for Zero-Shot Soundscape MappingArxiv 2024-
RSDiffRsdiff: Remote sensing image generation from text using diffusion modelArxiv 2024-
DiffusionSatDiffusionsat: A generative foundation model for satellite imageryArxiv 2024-
CRS-DiffCrs-diff: Controllable generative remote sensing foundation modelArxiv 2024-
MetaEarthMetaearth: A generative foundation model for global-scale remote sensing image generationArxiv 2024-
SkysensegptSkysensegpt: A fine-grained instruction tuning dataset and model for remote sensing vision-language understandingArxiv 2024-
RS-AgentRS-Agent: Automating Remote Sensing Tasks through Intelligent AgentsArxiv 2024-
SSLTransformerRSSelf-Supervised Vision Transformers for Land-Cover Segmentation and ClassificationArxiv 2023-
RingMo-SenseRingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution DisentanglingArxiv 2023-
A2-MAEA2-MAE: A spatial-temporal-spectral unified remote sensing pre-training method based on anchor-aware masked autoencoderArxiv 2023-
MSFEIGARSS 2023 - 2023 IEEE International Geoscience and Remote Sensing SymposiumIGARSS 2023-
GeCoGeographical Supervision Correction for Remote Sensing Representation LearningArxiv 2023-
SoftConMulti-label Guided Soft Contrastive Learning for Efficient Earth Observation PretrainingCVPR 2023-
Sen12MSSEN12MS--A curated dataset of georeferenced multi-spectral sentinel-1/2 imagery for deep learning and data fusionIEEE TGRS 2019-
BigEarthNet-S2Bigearthnet: A large-scale benchmark archive for remote sensing image understandingIEEE TGRS 2019-
MSARLarge-scale multi-class SAR image target detection dataset-1.0Journal of Radars 2022-
SAR-ShipSAR target recognition based on cross-domain and cross-task transfer learningIEEE Access 2019-
SAMPLEA SAR dataset for ATR development: the Synthetic and Measured Paired Labeled Experiment (SAMPLE)SPIE 2019-
DIOR-RSVGRsvg: Exploring data and models for visual grounding on remote sensing dataIEEE TGRS 2023-
Xviewxview: Objects in context in overhead imageryArxiv 2018-
DIOR-RAnchor-free oriented proposal generator for object detectionIEEE TGRS 2022-
SpaceNetv1Spacenet: A remote sensing dataset and challenge seriesArxiv 2018-
DynamicEarthNet-S2Dynamicearthnet: Daily multi-spectral satellite dataset for semantic change segmentationCVPR 2022-

VLM

MLLM

Agent

Other

Model NamePaper NamePublished inDetailed Info
-Transforming remote sensing images to textual descriptionsINT J APPL EARTH OBS 2022Link
-CSP Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual RepresentationsICML 2023Link
GeoCLIPGeoCLIP Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationNeurIPS 2023Link
-Good at captioning, bad at counting Benchmarking GPT-4V on Earth observation dataArxiv 2023Link
-On the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning ApplicationsArxiv 2023Link
SatCLIPSatCLIP Global, General-Purpose Location Embeddings with Satellite ImageryArxiv 2023Link
-The Potential of Visual ChatGPT for Remote SensingRS 2023Link
-遥感基础模型发展综述与未来设想遥感学报 2023Link
msGFMBridging Remote Sensors with Multisensor Geospatial Foundation ModelsCVPR 2024Link
-Charting New Territories Exploring the Geographic and Geospatial Capabilities of Multimodal LLMsArxiv 2024Link
-GeoLLM Extracting Geospatial Knowledge from Large Language ModelsICLR 2024Link
LeMeViTLeMeViT Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image InterpretationIJCAI 2024Link
MMEarthMMEarth Exploring Multi-Modal Pretext Tasks For Geospatial Representation LearningArxiv 2024Link
DOFANeural Plasticity-Inspired Foundation Model for Observing the Earth Crossing ModalitiesArxiv 2024Link
-On the Foundations of Earth and Climate Foundation ModelsArxiv 2024Link
SARATR-XSARATR-X A Foundation Model for Synthetic Aperture Radar Images Target RecognitionArxiv 2024Link
-Changes to Captions An Attentive Network for Remote Sensing Change CaptioningArxiv 2023Link
-Multi-source interactive stair attention for remote sensing image captioningRS 2023Link