A survey and paper list of current Spatio-Temporal Foundation Models from the pipeline perspective with awesome resources (paper, code, survey, etc.).
See the codeUnraveling Spatio-Temporal Foundation Models via the Pipeline lens with awesome resources (paper, code, survey, etc.), which aims to comprehensively and systematically summarize the recent advances to the best of our knowledge.
🙋 We will continue to update this repo with the newest resources. If you find any missed resources (paper/code) or errors, please feel free to open an issue or make a pull request.
Authors: Yuchen Fang, Hao Miao, Yuxuan Liang, Liwei Deng, Yue Cui, Ximu Zeng, Yuyang Xia, Yan Zhao∗, Torben Bach Pedersen, Christian S. Jensen (IEEE Fellow), Xiaofang Zhou (IEEE Fellow), and Kai Zheng∗.
![]() |
|---|
| Figure 1. The pipeline of spatio-temporal foundation models. |
✨ If you found this survey and repository useful, please consider to star this repository and cite our survey paper:
@article{fang2025unraveling,
title={Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review},
author={Fang, Yuchen and Miao, Hao and Liang, Yuxuan and Deng, Liwei and Cui, Yue and Zeng, Ximu and Xia, Yuyang and Zhao, Yan and Pedersen, Torben Bach and Jensen, Christian S and Zhou, Xiaofang and and Zheng, Kai},
journal={arXiv preprint arXiv:2506.01364},
year={2025},
}
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics (Traj), in ICDE 2023. [paper] [official-code]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces (Traj), in arXiv 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery (Traj), in arXiv 2025. [paper]
ClimaX: A foundation model for weather and climate (Grid), in ICML 2023. [paper] [official-code]
Accurate medium-range global weather forecasting with 3D neural networks (Grid), in Nature 2023. [paper] [official-code]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model (Grid), in CIKM 2023. [paper]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction (Grid), in KDD 2024. [paper] [official-code]
UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning (Grid), in arXiv 2024. [paper] [official-code]
VideoLLM: Modeling Video Sequence with Large Language Models (Video), in arXiv 2023. [paper] [official-code]
VideoChat: Chat-Centric Video Understanding (Video), in CVPR 2024. [paper] [official-code]
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding (Video), in CVPR 2024. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting (Graph), in KDD 2022. [paper] [official-code]
Brant: Foundation Model for Intracranial Neural Signal (Graph), in NeurIPS 2023. [paper] [official-code]
Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective (Graph), in KDD 2025. [paper] [official-code]
TERI: An Effective Framework for Trajectory Recovery with Irregular Time Intervals (Traj), in VLDB 2023. [paper] [official-code]
Pre-Training General Trajectory Embeddings With Maximum Multi-View Entropy Coding (Traj), in TKDE 2023. [paper] [official-code]
KGTS: Contrastive Trajectory Similarity Learning over Prompt Knowledge Graph Embedding (Traj), in AAAI 2024. [paper]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces (Traj), in arXiv 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery (Traj), in arXiv 2025. [paper]
Learning Dynamic Context Graphs for Predicting Social Events (Event), in KDD 2019. [paper] [official-code]
Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models (Event), in WWW 2024. [paper] [official-code]
ClimaX: A foundation model for weather and climate (Grid), in ICML 2023. [paper] [official-code]
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning (Grid), in ICCV 2023. [paper] [official-code]
Accurate medium-range global weather forecasting with 3D neural networks (Grid), in Nature 2023. [paper] [official-code]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction (Grid), in KDD 2024. [paper] [official-code]
WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning (Grid), in ICLR 2025. [paper] [official-code]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding (Video), in EMNLP 2023. [paper] [official-code]
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models (Video), in ICCV workshop 2023. [paper]
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction (Video), in ICLR 2024. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting (Graph), in KDD 2022. [paper] [official-code]
PPi: Pretraining Brain Signal Model for Patient-independent Seizure Detection (Graph), in NeurIPS 2023. [paper] [official-code]
Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting (Graph), in IJCAI 2024. [paper] [official-code]
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction (Graph), in arXiv 2024. [paper] [official-code]
G2PTL: A Geography-Graph Pre-trained Model (Graph), in CIKM 2024. [paper]
PTrajM: Efficient and Semantic-rich Trajectory Learning with Pretrained Trajectory-Mamba (Traj), in arXiv 2024. [paper]
ControlTraj: Controllable Trajectory Generation with Topology-Constrained Diffusion Model (Traj), in KDD 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery (Traj), in arXiv 2025. [paper]
Robust Event Forecasting with Spatiotemporal Confounder Learning (Event), in KDD 2022. [paper]
ONSEP: A Novel Online Neural-Symbolic Framework for Event Prediction Based on Large Language Model (Event), in ACL 2024. [paper] [official-code]
Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models (Event), in WWW 2024. [paper] [official-code]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model (Grid), in CIKM 2023. [paper]
UrbanGPT: Spatio-Temporal Large Language Models (Grid), in KDD 2024. [paper] [official-code]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models (Video), in NeurIPS 2022. [paper] [official-code]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding (Video), in EMNLP 2023. [paper] [official-code]
GenAD: Generalized Predictive Model for Autonomous Driving (Video), in CVPR 2024. [paper] [official-code]
PowerPM: Foundation Model for Power Systems (Graph), in NeurIPS 2024. [paper]
Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective (Graph), in KDD 2025. [paper] [official-code]
Please refer to the Table II in the paper.
ClimaX: A foundation model for weather and climate, in ICML 2023. [paper] [official-code]
Automated Spatio-Temporal Graph Contrastive Learning, in WWW 2023. [paper] [official-code]
Spatial Structure-Aware Road Network Embedding via Graph Contrastive Learning, in EDBT 2023. [paper] [official-code]
Spatial-Temporal Graph Learning with Adversarial Contrastive Adaptation, in ICML 2023. [paper] [official-code]
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning, in ICCV 2023. [paper] [official-code]
Accurate medium-range global weather forecasting with 3D neural networks, in Nature 2023. [paper] [official-code]
FourCastNet: Accelerating Global High-Resolution Weather Forecasting Using Adaptive Fourier Neural Operators, in PASC 2023. [paper] [official-code]
FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead, in arXiv 2023. [paper] [official-code]
Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint, in ICASSP 2023. [paper]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model, in CIKM 2023. [paper]
DiffCrime: A Multimodal Conditional Diffusion Model for Crime Risk Map Inference, in KDD 2024. [paper] [official-code]
G2PTL: A Geography-Graph Pre-trained Model, in CIKM 2024. [paper]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction, in KDD 2024. [paper] [official-code]
Earthfarseer: Versatile Spatio-Temporal Dynamical Systems Modeling in One Model, in AAAI 2024. [paper] [official-code]
Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model, in NeurIPS 2024. [paper]
WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning, in ICLR 2025. [paper] [official-code]
PhyDA: Physics-Guided Diffusion Models for Data Assimilation in Atmospheric Systems, in arXiv 2025. [paper]
Efficient Trajectory Similarity Computation with Contrastive Learning, in CIKM 2022. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting, in KDD 2022. [paper] [official-code]
Pre-Training General Trajectory Embeddings With Maximum Multi-View Entropy Coding, in TKDE 2023. [paper] [official-code]
Contrastive Trajectory Similarity Learning with Dual-Feature Attention, in ICDE 2023. [paper] [official-code]
PPi: Pretraining Brain Signal Model for Patient-independent Seizure Detection, in NeurIPS 2023. [paper] [official-code]
TrajFM: A Vehicle Trajectory Foundation Model for Region and Task Transferability, in arXiv 2024. [paper] [official-code]
ControlTraj: Controllable Trajectory Generation with Topology-Constrained Diffusion Model, in KDD 2024. [paper] [official-code]
EEGPT: Unleashing the Potential of EEG Generalist Foundation Model by Autoregressive Pre-training, in arXiv 2024. [paper]
Frequency-aware Generative Models for Multivariate Time Series Imputation, in NeurIPS 2024. [paper] [official-code]
Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCI, in ICLR 2024. [paper] [official-code]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces, in arXiv 2024. [paper] [official-code]
More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning, in WWW 2024. [paper] [official-code]
MTSCI: A Conditional Diffusion Model for Multivariate Time Series Consistent Imputation, in CIKM 2024. [paper] [official-code]
Score-CDM: Score-Weighted Convolutional Diffusion Model for Multivariate Time Series Imputation, in IJCAI 2024. [paper]
Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting, in IJCAI 2024. [paper] [official-code]
Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series, in NeurIPS 2024. [paper] [official-code]
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts, in ICLR 2025. [paper] [official-code]
Robust Event Forecasting with Spatiotemporal Confounder Learning, in KDD 2022. [paper]
When Do Contrastive Learning Signals Help Spatio-Temporal Graph Forecasting?, in SIGSPATIAL 2022. [paper] [official-code]
Unified Route Representation Learning for Multi-Modal Transportation Recommendation with Spatiotemporal Pre-Training, in VLDBJ 2022. [paper]
Jointly Contrastive Representation Learning on Road Network and Trajectory, in CIKM 2022. [paper] [official-code]
Spatial-Temporal Hypergraph Self-Supervised Learning for Crime Prediction, in ICDE 2022. [paper] [official-code]
DiffSTG: Probabilistic Spatio-Temporal Graph Forecasting with Denoising Diffusion Models, in SIGSPATIAL 2023. [paper] [official-code]
PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation, in ICDE 2023. [paper] [official-code]
Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Photovoltaic Data Imputation, in SIGMOD 2023. [paper] [official-code]
Brant: Foundation Model for Intracranial Neural Signal, in NeurIPS 2023. [paper] [official-code]
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics, in ICDE 2023. [paper] [official-code]
Spatio-Temporal Meta Contrastive Learning, in CIKM 2023. [paper] [official-code]
Spatio-Temporal Self-Supervised Learning for Traffic Flow Prediction, in AAAI 2023. [paper] [official-code]
Spatio-temporal Diffusion Point Processes, in KDD 2023. [paper] [official-code]
A Spatio-Temporal Diffusion Model for Missing and Real-Time Financial Data Inference, in CIKM 2024. [paper]
DiffLight: A Partial Rewards Conditioned Diffusion Model for Traffic Signal Control with Missing Data, in NeurIPS 2024. [paper] [official-code]
Diffstock: Probabilistic Relational Stock Market Predictions Using Diffusion Models, in ICASSP 2024. [paper]
KGTS: Contrastive Trajectory Similarity Learning over Prompt Knowledge Graph Embedding, in AAAI 2024. [paper]
MSTEM: Masked Spatiotemporal Event Series Modeling for Urban Undisciplined Events Forecasting, in CIKM 2024. [paper]
Multi-Modality Spatio-Temporal Forecasting via Self-Supervised Learning, in IJCAI 2024. [paper] [official-code]
Towards Unifying Diffusion Models for Probabilistic Spatio-Temporal Graph Learning, in SIGSPATIAL 2024. [paper] [official-code]
UrbanGPT: Spatio-Temporal Large Language Models, in KDD 2024. [paper] [official-code]
UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning, in arXiv 2024. [paper] [official-code]
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction, in arXiv 2024. [paper] [official-code]
Social Physics Informed Diffusion Model for Crowd Simulation, in AAAI 2024. [paper] [official-code]
PowerPM: Foundation Model for Power Systems, in NeurIPS 2024. [paper]
FusionSF: Fuse Heterogeneous Modalities in a Vector Quantized Framework for Robust Solar Power Forecasting, in KDD 2024. [paper] [official-code]
SaSDim:Self-adaptive Noise Scaling Diffusion Model for Spatial Time Series Imputation, in IJCAI 2024. [paper]
ClimaX: A foundation model for weather and climate, in ICML 2023. [paper] [official-code]
PointGPT: Auto-regressively Generative Pre-training from Point Clouds, in NeurIPS 2023. [paper] [official-code]
FourCastNet: Accelerating Global High-Resolution Weather Forecasting Using Adaptive Fourier Neural Operators, in PASC 2023. [paper] [official-code]
FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead, in arXiv 2023. [paper] [official-code]
TrajFM: A Vehicle Trajectory Foundation Model for Region and Task Transferability, in arXiv 2024. [paper] [official-code]
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction, in arXiv 2024. [paper] [official-code]
Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCI, in ICLR 2024. [paper] [official-code]
EEGPT: Unleashing the Potential of EEG Generalist Foundation Model by Autoregressive Pre-training, in arXiv 2024. [paper]
iVideoGPT: Interactive VideoGPTs are Scalable World Models, in NeurIPS 2024. [paper] [official-code]
Everything is a Video: Unifying Modalities through Next-Frame Prediction, in arXiv 2024. [paper]
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction, in NeurIPS 2024. [paper] [official-code]
UVTM: Universal Vehicle Trajectory Modeling with ST Feature Domain Generation, in TKDE 2025. [paper] [official-code]
TrajMoE: Spatially-Aware Mixture of Experts for Unified Human Mobility Modeling, in arXiv 2025. [paper]
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts, in ICLR 2025. [paper] [official-code]
Masked Autoencoders for Point Cloud Self-supervised Learning, in ECCV 2022. [paper] [official-code]
Masked Autoencoders As Spatiotemporal Learners, in NeurIPS 2022. [paper] [official-code]
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, in NeurIPS 2022. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting, in KDD 2022. [paper] [official-code]
MARLIN: Masked Autoencoder for Facial Video Representation Learning, in CVPR 2023. [paper] [official-code]
GPT-ST: Generative Pre-Training of Spatio-Temporal Graph Neural Networks, in NeurIPS 2023. [paper] [official-code]
Brant: Foundation Model for Intracranial Neural Signal, in NeurIPS 2023. [paper] [official-code]
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning, in ICCV 2023. [paper] [official-code]
PowerPM: Foundation Model for Power Systems, in NeurIPS 2024. [paper]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces, in arXiv 2024. [paper] [official-code]
G2PTL: A Geography-Graph Pre-trained Model, in CIKM 2024. [paper]
Temporal-Frequency Masked Autoencoders for Time Series Anomaly Detection, in ICDE 2024. [paper] [official-code]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction, in KDD 2024. [paper] [official-code]
Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting, in IJCAI 2024. [paper] [official-code]
Revealing the Power of Masked Autoencoders in Traffic Forecasting, in CIKM 2024. [paper] [official-code]
WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning, in ICLR 2025. [paper] [official-code]
A Universal Pre-Training and Prompting Framework for General Urban Spatio-Temporal Prediction, in TKDE 2025. [paper]
Efficient Spatio-Temporal Contrastive Learning for Skeleton-Based 3-D Action Recognition, in TMM 2021. [paper]
Dual Contrastive Learning for Spatio-temporal Representation, in ACM MM 2022. [paper]
Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision, in CVPR 2022. [paper] [official-code]
Efficient Trajectory Similarity Computation with Contrastive Learning, in CIKM 2022. [paper] [official-code]
Mining Spatio-temporal Relations via Self-paced Graph Contrastive Learning, in KDD 2022. [paper] [official-code]
When Do Contrastive Learning Signals Help Spatio-Temporal Graph Forecasting?, in SIGSPATIAL 2022. [paper] [official-code]
Spatial-Temporal Hypergraph Self-Supervised Learning for Crime Prediction, in ICDE 2022. [paper] [official-code]
Pre-Training General Trajectory Embeddings With Maximum Multi-View Entropy Coding, in TKDE 2023. [paper] [official-code]
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics, in ICDE 2023. [paper] [official-code]
Automated Spatio-Temporal Graph Contrastive Learning, in WWW 2023. [paper] [official-code]
Contrastive Trajectory Similarity Learning with Dual-Feature Attention, in ICDE 2023. [paper] [official-code]
Spatio-Temporal Self-Supervised Learning for Traffic Flow Prediction, in AAAI 2023. [paper] [official-code]
STWave+: A Multi-Scale Efficient Spectral Graph Attention Network With Long-Term Trends for Disentangled Traffic Flow Forecasting, in TKDE 2023. [paper] [official-code]
Spatio-Temporal Meta Contrastive Learning, in CIKM 2023. [paper] [official-code]
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web, in WWW 2024. [paper] [official-code]
PTrajM: Efficient and Semantic-rich Trajectory Learning with Pretrained Trajectory-Mamba, in arXiv 2024. [paper]
More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning, in WWW 2024. [paper] [official-code]
FlashST: A Simple and Universal Prompt-Tuning Framework for Traffic Prediction, in ICML 2024. [paper] [official-code]
MM-Path: Multi-modal, Multi-granularity Path Representation Learning, in KDD 2025. [paper] [official-code]
Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint, in ICASSP 2023. [paper]
DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecasting, in NeurIPS 2023. [paper] [official-code]
DiffSTG: Probabilistic Spatio-Temporal Graph Forecasting with Denoising Diffusion Models, in SIGSPATIAL 2023. [paper] [official-code]
DiffTraj: Generating GPS Trajectory with Diffusion Probabilistic Model, in NeurIPS 2023. [paper] [official-code]
PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation, in ICDE 2023. [paper] [official-code]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model, in CIKM 2023. [paper]
Spatio-temporal Diffusion Point Processes, in KDD 2023. [paper] [official-code]
DiffCrime: A Multimodal Conditional Diffusion Model for Crime Risk Map Inference, in KDD 2024. [paper] [official-code]
Social Physics Informed Diffusion Model for Crowd Simulation, in AAAI 2024. [paper] [official-code]
Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model, in NeurIPS 2024. [paper]
Towards Unifying Diffusion Models for Probabilistic Spatio-Temporal Graph Learning, in SIGSPATIAL 2024. [paper] [official-code]
Frequency-aware Generative Models for Multivariate Time Series Imputation, in NeurIPS 2024. [paper] [official-code]
ControlTraj: Controllable Trajectory Generation with Topology-Constrained Diffusion Model, in KDD 2024. [paper] [official-code]
Score-CDM: Score-Weighted Convolutional Diffusion Model for Multivariate Time Series Imputation, in IJCAI 2024. [paper]
SaSDim:Self-adaptive Noise Scaling Diffusion Model for Spatial Time Series Imputation, in IJCAI 2024. [paper]
DiffLight: A Partial Rewards Conditioned Diffusion Model for Traffic Signal Control with Missing Data, in NeurIPS 2024. [paper] [official-code]
PhyDA: Physics-Guided Diffusion Models for Data Assimilation in Atmospheric Systems, in arXiv 2025. [paper]
Dual Contrastive Learning for Spatio-temporal Representation, in ACM MM 2022. [paper]
Masked Autoencoders As Spatiotemporal Learners, in NeurIPS 2022. [paper] [official-code]
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, in NeurIPS 2022. [paper] [official-code]
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction, in ICLR 2024. [paper] [official-code]
LLMs are Good Action Recognizers, in CVPR 2024. [paper]
iVideoGPT: Interactive VideoGPTs are Scalable World Models, in NeurIPS 2024. [paper] [official-code]
GenAD: Generalized Predictive Model for Autonomous Driving, in CVPR 2024. [paper] [official-code]
VideoChat: Chat-Centric Video Understanding, in CVPR 2024. [paper] [official-code]
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, in CVPR 2024. [paper] [official-code]
Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model, in NeurIPS 2024. [paper]
Leveraging Language Foundation Models for Human Mobility Forecasting, in SIGSPATIAL 2022. [paper] [official-code]
A large language model for electronic health records, in NPJ 2022. [paper]
GATGPT: A Pre-trained Large Language Model with Graph Attention Network for Spatiotemporal Imputation, in arXiv 2023. [paper]
Language Models are Causal Knowledge Extractors for Zero-shot Video Question Answering, in CVPR Workshop 2023. [paper]
Language Models Can Improve Event Prediction by Few-Shot Abductive Reasoning, in NeurIPS 2023. [paper] [official-code]
Harnessing LLMs for Temporal Data - A Study on Explainable Financial Time Series Forecasting, in ACL 2023. [paper]
Health system-scale language models are all-purpose prediction engines, in Nature 2023. [paper]
GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding, in EMNLP 2023. [paper] [official-code]
ChatGPT Informed Graph Neural Network for Stock Movement Prediction, in SIGKDD Workshop 2023. [paper] [official-code]
UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction, in NeurIPS 2024. [paper] [official-code]
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation, in NeurIPS 2024. [paper] [official-code]
Chain-of-History Reasoning for Temporal Knowledge Graph Forecasting, in ACL 2024. [paper]
VideoChat: Chat-Centric Video Understanding, in CVPR 2024. [paper] [official-code]
Extracting Spatiotemporal Data from Gradients with Large Language Models, in arXiv 2024. [paper]
How Can Large Language Models Understand Spatial-Temporal Data?, in arXiv 2024. [paper]
Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming, in arXiv 2024. [paper]
Large Language Models for Next Point-of-Interest Recommendation, in SIGIR 2024. [paper] [official-code]
Latent Logic Tree Extraction for Event Sequence Explanation from LLMs, in ICML 2024. [paper]
Orca: Ocean Significant Wave Height Estimation with Spatio-temporally Aware Large Language Models, in CIKM 2024. [paper]
Semantic Trajectory Data Mining with LLM-Informed POI Classification, in ITSC 2024. [paper]
Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models, in WWW 2024. [paper] [official-code]
MAS4POI: a Multi-Agents Collaboration System for Next POI Recommendation, in arXiv 2024. [paper] [official-code]
Spatial-Temporal Large Language Model for Traffic Prediction, in MDM 2024. [paper] [official-code]
TPLLM: A Traffic Prediction Framework Based on Pretrained Large Language Models, in arXiv 2024. [paper]
ONSEP: A Novel Online Neural-Symbolic Framework for Event Prediction Based on Large Language Model, in ACL 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery, in arXiv 2025. [paper]
LLMLight: Large Language Models as Traffic Signal Control Agents, in KDD 2025. [paper] [official-code]
TimeCMA: Towards LLM-Empowered Time Series Forecasting via Cross-Modality Alignment, in AAAI 2025. [paper] [official-code]
ST-LLM+: Graph Enhanced Spatio-Temporal Large Language Models for Traffic Prediction, in TKDE 2025. [paper]
Path-LLM: A Multi-Modal Path Representation Learning by Aligning and Fusing with Large Language Models, in WWW 2025. [paper] [official-code]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models, in NeurIPS 2022. [paper] [official-code]
VideoLLM: Modeling Video Sequence with Large Language Models, in arXiv 2023. [paper] [official-code]
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models, in ICCV workshop 2023. [paper]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, in EMNLP 2023. [paper] [official-code]
Leveraging Vision-Language Models for Granular Market Change Prediction, in AAAI Workshop 2023. [paper]
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, in CVPR 2024. [paper] [official-code]
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web, in WWW 2024. [paper] [official-code]
VideoLLM-online: Online Video Large Language Model for Streaming Video, in CVPR 2024. [paper] [official-code]
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View, in AAAI 2024. [paper] [official-code]
Seer: Language Instructed Video Prediction with Latent Diffusion Models, in ICLR 2024. [paper] [official-code]
Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial Topology, in KDD 2024. [paper]
UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction, in AAAI 2025. [paper] [official-code]
Leveraging Language Foundation Models for Human Mobility Forecasting, in SIGSPATIAL 2022. [paper] [official-code]
Harnessing LLMs for Temporal Data - A Study on Explainable Financial Time Series Forecasting, in ACL 2023. [paper]
PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting, in TKDE 2023. [paper] [official-code]
Large Language Models are Few-Shot Health Learners, in arXiv 2023. [paper]
Where Would I Go Next? Large Language Models as Human Mobility Predictors, in arXiv 2024. [paper] [official-code]
TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series, in ICLR 2024. [paper] [official-code]
Time-LLM: Time Series Forecasting by Reprogramming Large Language Models, in ICLR 2024. [paper] [official-code]
Language Knowledge-Assisted Representation Learning for Skeleton-Based Action Recognition, in arXiv 2023. [paper] [official-code]
Prompt to Transfer: Sim-to-Real Transfer for Traffic Signal Control with Prompt Learning, in AAAI 2024. [paper] [official-code]
RealTCD: Temporal Causal Discovery from Interventional Data with Large Language Model, in CIKM 2024. [paper]
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models, in ICCV workshop 2023. [paper]
Leveraging Vision-Language Models for Granular Market Change Prediction, in AAAI Workshop 2023. [paper]
ChatGPT Informed Graph Neural Network for Stock Movement Prediction, in SIGKDD Workshop 2023. [paper] [official-code]
Extracting Spatiotemporal Data from Gradients with Large Language Models, in arXiv 2024. [paper]
Orca: Ocean Significant Wave Height Estimation with Spatio-temporally Aware Large Language Models, in CIKM 2024. [paper]
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation, in NeurIPS 2024. [paper] [official-code]
UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction, in NeurIPS 2024. [paper] [official-code]
MAS4POI: a Multi-Agents Collaboration System for Next POI Recommendation, in arXiv 2024. [paper] [official-code]
Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial Topology, in KDD 2024. [paper]
Large Language Models for Next Point-of-Interest Recommendation, in SIGIR 2024. [paper] [official-code]
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web, in WWW 2024. [paper] [official-code]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models, in NeurIPS 2022. [paper] [official-code]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, in EMNLP 2023. [paper] [official-code]
VideoLLM: Modeling Video Sequence with Large Language Models, in arXiv 2023. [paper] [official-code]
VideoLLM-online: Online Video Large Language Model for Streaming Video, in CVPR 2024. [paper] [official-code]
LLMs are Good Action Recognizers, in CVPR 2024. [paper]
Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming, in arXiv 2024. [paper]
VideoPoet: A Large Language Model for Zero-Shot Video Generation, in ICML 2024. [paper] [official-code]
How Can Large Language Models Understand Spatial-Temporal Data?, in arXiv 2024. [paper]
UrbanGPT: Spatio-Temporal Large Language Models, in KDD 2024. [paper] [official-code]
TimeCMA: Towards LLM-Empowered Time Series Forecasting via Cross-Modality Alignment, in AAAI 2025. [paper] [official-code]
Path-LLM: A Multi-Modal Path Representation Learning by Aligning and Fusing with Large Language Models, in WWW 2025. [paper] [official-code]
SpaBERT: A Pretrained Language Model from Geographic Data for Geo-Entity Representation, in EMNLP 2022. [paper]
Leveraging Language Foundation Models for Human Mobility Forecasting, in SIGSPATIAL 2022. [paper] [official-code]
GATGPT: A Pre-trained Large Language Model with Graph Attention Network for Spatiotemporal Imputation, in arXiv 2023. [paper]
GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding, in EMNLP 2023. [paper] [official-code]
Spatial-Temporal Large Language Model for Traffic Prediction, in MDM 2024. [paper] [official-code]
TPLLM: A Traffic Prediction Framework Based on Pretrained Large Language Models, in arXiv 2024. [paper]
Large Language Models for Next Point-of-Interest Recommendation, in SIGIR 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery, in arXiv 2025. [paper]
Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge Distillation, in ICDE 2025. [paper] [official-code]
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook, in arXiv 2023. [paper] [link]
Urban Foundation Models: A Survey, in KDD 2024. [paper] [link]
A Survey on Diffusion Models for Time Series and Spatio-Temporal Data, in arXiv 2024. [paper] [link]
Foundation Models for Time Series Analysis: A Tutorial and Survey, in KDD 2024. [paper]
Foundation Models for Spatio-Temporal Data Science: A Tutorial and Survey, in KDD 2025. [paper]
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities, in arXiv 2025. [paper]
10 commits
A survey and paper list of current Spatio-Temporal Foundation Models from the pipeline perspective with awesome resources (paper, code, survey, etc.).
See the codeUnraveling Spatio-Temporal Foundation Models via the Pipeline lens with awesome resources (paper, code, survey, etc.), which aims to comprehensively and systematically summarize the recent advances to the best of our knowledge.
🙋 We will continue to update this repo with the newest resources. If you find any missed resources (paper/code) or errors, please feel free to open an issue or make a pull request.
Authors: Yuchen Fang, Hao Miao, Yuxuan Liang, Liwei Deng, Yue Cui, Ximu Zeng, Yuyang Xia, Yan Zhao∗, Torben Bach Pedersen, Christian S. Jensen (IEEE Fellow), Xiaofang Zhou (IEEE Fellow), and Kai Zheng∗.
![]() |
|---|
| Figure 1. The pipeline of spatio-temporal foundation models. |
✨ If you found this survey and repository useful, please consider to star this repository and cite our survey paper:
@article{fang2025unraveling,
title={Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review},
author={Fang, Yuchen and Miao, Hao and Liang, Yuxuan and Deng, Liwei and Cui, Yue and Zeng, Ximu and Xia, Yuyang and Zhao, Yan and Pedersen, Torben Bach and Jensen, Christian S and Zhou, Xiaofang and and Zheng, Kai},
journal={arXiv preprint arXiv:2506.01364},
year={2025},
}
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics (Traj), in ICDE 2023. [paper] [official-code]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces (Traj), in arXiv 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery (Traj), in arXiv 2025. [paper]
ClimaX: A foundation model for weather and climate (Grid), in ICML 2023. [paper] [official-code]
Accurate medium-range global weather forecasting with 3D neural networks (Grid), in Nature 2023. [paper] [official-code]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model (Grid), in CIKM 2023. [paper]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction (Grid), in KDD 2024. [paper] [official-code]
UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning (Grid), in arXiv 2024. [paper] [official-code]
VideoLLM: Modeling Video Sequence with Large Language Models (Video), in arXiv 2023. [paper] [official-code]
VideoChat: Chat-Centric Video Understanding (Video), in CVPR 2024. [paper] [official-code]
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding (Video), in CVPR 2024. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting (Graph), in KDD 2022. [paper] [official-code]
Brant: Foundation Model for Intracranial Neural Signal (Graph), in NeurIPS 2023. [paper] [official-code]
Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective (Graph), in KDD 2025. [paper] [official-code]
TERI: An Effective Framework for Trajectory Recovery with Irregular Time Intervals (Traj), in VLDB 2023. [paper] [official-code]
Pre-Training General Trajectory Embeddings With Maximum Multi-View Entropy Coding (Traj), in TKDE 2023. [paper] [official-code]
KGTS: Contrastive Trajectory Similarity Learning over Prompt Knowledge Graph Embedding (Traj), in AAAI 2024. [paper]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces (Traj), in arXiv 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery (Traj), in arXiv 2025. [paper]
Learning Dynamic Context Graphs for Predicting Social Events (Event), in KDD 2019. [paper] [official-code]
Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models (Event), in WWW 2024. [paper] [official-code]
ClimaX: A foundation model for weather and climate (Grid), in ICML 2023. [paper] [official-code]
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning (Grid), in ICCV 2023. [paper] [official-code]
Accurate medium-range global weather forecasting with 3D neural networks (Grid), in Nature 2023. [paper] [official-code]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction (Grid), in KDD 2024. [paper] [official-code]
WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning (Grid), in ICLR 2025. [paper] [official-code]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding (Video), in EMNLP 2023. [paper] [official-code]
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models (Video), in ICCV workshop 2023. [paper]
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction (Video), in ICLR 2024. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting (Graph), in KDD 2022. [paper] [official-code]
PPi: Pretraining Brain Signal Model for Patient-independent Seizure Detection (Graph), in NeurIPS 2023. [paper] [official-code]
Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting (Graph), in IJCAI 2024. [paper] [official-code]
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction (Graph), in arXiv 2024. [paper] [official-code]
G2PTL: A Geography-Graph Pre-trained Model (Graph), in CIKM 2024. [paper]
PTrajM: Efficient and Semantic-rich Trajectory Learning with Pretrained Trajectory-Mamba (Traj), in arXiv 2024. [paper]
ControlTraj: Controllable Trajectory Generation with Topology-Constrained Diffusion Model (Traj), in KDD 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery (Traj), in arXiv 2025. [paper]
Robust Event Forecasting with Spatiotemporal Confounder Learning (Event), in KDD 2022. [paper]
ONSEP: A Novel Online Neural-Symbolic Framework for Event Prediction Based on Large Language Model (Event), in ACL 2024. [paper] [official-code]
Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models (Event), in WWW 2024. [paper] [official-code]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model (Grid), in CIKM 2023. [paper]
UrbanGPT: Spatio-Temporal Large Language Models (Grid), in KDD 2024. [paper] [official-code]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models (Video), in NeurIPS 2022. [paper] [official-code]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding (Video), in EMNLP 2023. [paper] [official-code]
GenAD: Generalized Predictive Model for Autonomous Driving (Video), in CVPR 2024. [paper] [official-code]
PowerPM: Foundation Model for Power Systems (Graph), in NeurIPS 2024. [paper]
Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective (Graph), in KDD 2025. [paper] [official-code]
Please refer to the Table II in the paper.
ClimaX: A foundation model for weather and climate, in ICML 2023. [paper] [official-code]
Automated Spatio-Temporal Graph Contrastive Learning, in WWW 2023. [paper] [official-code]
Spatial Structure-Aware Road Network Embedding via Graph Contrastive Learning, in EDBT 2023. [paper] [official-code]
Spatial-Temporal Graph Learning with Adversarial Contrastive Adaptation, in ICML 2023. [paper] [official-code]
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning, in ICCV 2023. [paper] [official-code]
Accurate medium-range global weather forecasting with 3D neural networks, in Nature 2023. [paper] [official-code]
FourCastNet: Accelerating Global High-Resolution Weather Forecasting Using Adaptive Fourier Neural Operators, in PASC 2023. [paper] [official-code]
FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead, in arXiv 2023. [paper] [official-code]
Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint, in ICASSP 2023. [paper]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model, in CIKM 2023. [paper]
DiffCrime: A Multimodal Conditional Diffusion Model for Crime Risk Map Inference, in KDD 2024. [paper] [official-code]
G2PTL: A Geography-Graph Pre-trained Model, in CIKM 2024. [paper]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction, in KDD 2024. [paper] [official-code]
Earthfarseer: Versatile Spatio-Temporal Dynamical Systems Modeling in One Model, in AAAI 2024. [paper] [official-code]
Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model, in NeurIPS 2024. [paper]
WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning, in ICLR 2025. [paper] [official-code]
PhyDA: Physics-Guided Diffusion Models for Data Assimilation in Atmospheric Systems, in arXiv 2025. [paper]
Efficient Trajectory Similarity Computation with Contrastive Learning, in CIKM 2022. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting, in KDD 2022. [paper] [official-code]
Pre-Training General Trajectory Embeddings With Maximum Multi-View Entropy Coding, in TKDE 2023. [paper] [official-code]
Contrastive Trajectory Similarity Learning with Dual-Feature Attention, in ICDE 2023. [paper] [official-code]
PPi: Pretraining Brain Signal Model for Patient-independent Seizure Detection, in NeurIPS 2023. [paper] [official-code]
TrajFM: A Vehicle Trajectory Foundation Model for Region and Task Transferability, in arXiv 2024. [paper] [official-code]
ControlTraj: Controllable Trajectory Generation with Topology-Constrained Diffusion Model, in KDD 2024. [paper] [official-code]
EEGPT: Unleashing the Potential of EEG Generalist Foundation Model by Autoregressive Pre-training, in arXiv 2024. [paper]
Frequency-aware Generative Models for Multivariate Time Series Imputation, in NeurIPS 2024. [paper] [official-code]
Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCI, in ICLR 2024. [paper] [official-code]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces, in arXiv 2024. [paper] [official-code]
More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning, in WWW 2024. [paper] [official-code]
MTSCI: A Conditional Diffusion Model for Multivariate Time Series Consistent Imputation, in CIKM 2024. [paper] [official-code]
Score-CDM: Score-Weighted Convolutional Diffusion Model for Multivariate Time Series Imputation, in IJCAI 2024. [paper]
Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting, in IJCAI 2024. [paper] [official-code]
Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series, in NeurIPS 2024. [paper] [official-code]
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts, in ICLR 2025. [paper] [official-code]
Robust Event Forecasting with Spatiotemporal Confounder Learning, in KDD 2022. [paper]
When Do Contrastive Learning Signals Help Spatio-Temporal Graph Forecasting?, in SIGSPATIAL 2022. [paper] [official-code]
Unified Route Representation Learning for Multi-Modal Transportation Recommendation with Spatiotemporal Pre-Training, in VLDBJ 2022. [paper]
Jointly Contrastive Representation Learning on Road Network and Trajectory, in CIKM 2022. [paper] [official-code]
Spatial-Temporal Hypergraph Self-Supervised Learning for Crime Prediction, in ICDE 2022. [paper] [official-code]
DiffSTG: Probabilistic Spatio-Temporal Graph Forecasting with Denoising Diffusion Models, in SIGSPATIAL 2023. [paper] [official-code]
PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation, in ICDE 2023. [paper] [official-code]
Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Photovoltaic Data Imputation, in SIGMOD 2023. [paper] [official-code]
Brant: Foundation Model for Intracranial Neural Signal, in NeurIPS 2023. [paper] [official-code]
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics, in ICDE 2023. [paper] [official-code]
Spatio-Temporal Meta Contrastive Learning, in CIKM 2023. [paper] [official-code]
Spatio-Temporal Self-Supervised Learning for Traffic Flow Prediction, in AAAI 2023. [paper] [official-code]
Spatio-temporal Diffusion Point Processes, in KDD 2023. [paper] [official-code]
A Spatio-Temporal Diffusion Model for Missing and Real-Time Financial Data Inference, in CIKM 2024. [paper]
DiffLight: A Partial Rewards Conditioned Diffusion Model for Traffic Signal Control with Missing Data, in NeurIPS 2024. [paper] [official-code]
Diffstock: Probabilistic Relational Stock Market Predictions Using Diffusion Models, in ICASSP 2024. [paper]
KGTS: Contrastive Trajectory Similarity Learning over Prompt Knowledge Graph Embedding, in AAAI 2024. [paper]
MSTEM: Masked Spatiotemporal Event Series Modeling for Urban Undisciplined Events Forecasting, in CIKM 2024. [paper]
Multi-Modality Spatio-Temporal Forecasting via Self-Supervised Learning, in IJCAI 2024. [paper] [official-code]
Towards Unifying Diffusion Models for Probabilistic Spatio-Temporal Graph Learning, in SIGSPATIAL 2024. [paper] [official-code]
UrbanGPT: Spatio-Temporal Large Language Models, in KDD 2024. [paper] [official-code]
UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning, in arXiv 2024. [paper] [official-code]
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction, in arXiv 2024. [paper] [official-code]
Social Physics Informed Diffusion Model for Crowd Simulation, in AAAI 2024. [paper] [official-code]
PowerPM: Foundation Model for Power Systems, in NeurIPS 2024. [paper]
FusionSF: Fuse Heterogeneous Modalities in a Vector Quantized Framework for Robust Solar Power Forecasting, in KDD 2024. [paper] [official-code]
SaSDim:Self-adaptive Noise Scaling Diffusion Model for Spatial Time Series Imputation, in IJCAI 2024. [paper]
ClimaX: A foundation model for weather and climate, in ICML 2023. [paper] [official-code]
PointGPT: Auto-regressively Generative Pre-training from Point Clouds, in NeurIPS 2023. [paper] [official-code]
FourCastNet: Accelerating Global High-Resolution Weather Forecasting Using Adaptive Fourier Neural Operators, in PASC 2023. [paper] [official-code]
FengWu: Pushing the Skillful Global Medium-range Weather Forecast beyond 10 Days Lead, in arXiv 2023. [paper] [official-code]
TrajFM: A Vehicle Trajectory Foundation Model for Region and Task Transferability, in arXiv 2024. [paper] [official-code]
OpenCity: Open Spatio-Temporal Foundation Models for Traffic Prediction, in arXiv 2024. [paper] [official-code]
Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCI, in ICLR 2024. [paper] [official-code]
EEGPT: Unleashing the Potential of EEG Generalist Foundation Model by Autoregressive Pre-training, in arXiv 2024. [paper]
iVideoGPT: Interactive VideoGPTs are Scalable World Models, in NeurIPS 2024. [paper] [official-code]
Everything is a Video: Unifying Modalities through Next-Frame Prediction, in arXiv 2024. [paper]
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction, in NeurIPS 2024. [paper] [official-code]
UVTM: Universal Vehicle Trajectory Modeling with ST Feature Domain Generation, in TKDE 2025. [paper] [official-code]
TrajMoE: Spatially-Aware Mixture of Experts for Unified Human Mobility Modeling, in arXiv 2025. [paper]
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts, in ICLR 2025. [paper] [official-code]
Masked Autoencoders for Point Cloud Self-supervised Learning, in ECCV 2022. [paper] [official-code]
Masked Autoencoders As Spatiotemporal Learners, in NeurIPS 2022. [paper] [official-code]
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, in NeurIPS 2022. [paper] [official-code]
Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting, in KDD 2022. [paper] [official-code]
MARLIN: Masked Autoencoder for Facial Video Representation Learning, in CVPR 2023. [paper] [official-code]
GPT-ST: Generative Pre-Training of Spatio-Temporal Graph Neural Networks, in NeurIPS 2023. [paper] [official-code]
Brant: Foundation Model for Intracranial Neural Signal, in NeurIPS 2023. [paper] [official-code]
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning, in ICCV 2023. [paper] [official-code]
PowerPM: Foundation Model for Power Systems, in NeurIPS 2024. [paper]
UniTraj: Learning a Universal Trajectory Foundation Model from Billion-Scale Worldwide Traces, in arXiv 2024. [paper] [official-code]
G2PTL: A Geography-Graph Pre-trained Model, in CIKM 2024. [paper]
Temporal-Frequency Masked Autoencoders for Time Series Anomaly Detection, in ICDE 2024. [paper] [official-code]
UniST: A Prompt-Empowered Universal Model for Urban Spatio-Temporal Prediction, in KDD 2024. [paper] [official-code]
Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting, in IJCAI 2024. [paper] [official-code]
Revealing the Power of Masked Autoencoders in Traffic Forecasting, in CIKM 2024. [paper] [official-code]
WeatherGFM: Learning a Weather Generalist Foundation Model via In-context Learning, in ICLR 2025. [paper] [official-code]
A Universal Pre-Training and Prompting Framework for General Urban Spatio-Temporal Prediction, in TKDE 2025. [paper]
Efficient Spatio-Temporal Contrastive Learning for Skeleton-Based 3-D Action Recognition, in TMM 2021. [paper]
Dual Contrastive Learning for Spatio-temporal Representation, in ACM MM 2022. [paper]
Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision, in CVPR 2022. [paper] [official-code]
Efficient Trajectory Similarity Computation with Contrastive Learning, in CIKM 2022. [paper] [official-code]
Mining Spatio-temporal Relations via Self-paced Graph Contrastive Learning, in KDD 2022. [paper] [official-code]
When Do Contrastive Learning Signals Help Spatio-Temporal Graph Forecasting?, in SIGSPATIAL 2022. [paper] [official-code]
Spatial-Temporal Hypergraph Self-Supervised Learning for Crime Prediction, in ICDE 2022. [paper] [official-code]
Pre-Training General Trajectory Embeddings With Maximum Multi-View Entropy Coding, in TKDE 2023. [paper] [official-code]
Self-supervised Trajectory Representation Learning with Temporal Regularities and Travel Semantics, in ICDE 2023. [paper] [official-code]
Automated Spatio-Temporal Graph Contrastive Learning, in WWW 2023. [paper] [official-code]
Contrastive Trajectory Similarity Learning with Dual-Feature Attention, in ICDE 2023. [paper] [official-code]
Spatio-Temporal Self-Supervised Learning for Traffic Flow Prediction, in AAAI 2023. [paper] [official-code]
STWave+: A Multi-Scale Efficient Spectral Graph Attention Network With Long-Term Trends for Disentangled Traffic Flow Forecasting, in TKDE 2023. [paper] [official-code]
Spatio-Temporal Meta Contrastive Learning, in CIKM 2023. [paper] [official-code]
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web, in WWW 2024. [paper] [official-code]
PTrajM: Efficient and Semantic-rich Trajectory Learning with Pretrained Trajectory-Mamba, in arXiv 2024. [paper]
More Than Routing: Joint GPS and Route Modeling for Refine Trajectory Representation Learning, in WWW 2024. [paper] [official-code]
FlashST: A Simple and Universal Prompt-Tuning Framework for Traffic Prediction, in ICML 2024. [paper] [official-code]
MM-Path: Multi-modal, Multi-granularity Path Representation Learning, in KDD 2025. [paper] [official-code]
Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint, in ICASSP 2023. [paper]
DYffusion: A Dynamics-informed Diffusion Model for Spatiotemporal Forecasting, in NeurIPS 2023. [paper] [official-code]
DiffSTG: Probabilistic Spatio-Temporal Graph Forecasting with Denoising Diffusion Models, in SIGSPATIAL 2023. [paper] [official-code]
DiffTraj: Generating GPS Trajectory with Diffusion Probabilistic Model, in NeurIPS 2023. [paper] [official-code]
PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation, in ICDE 2023. [paper] [official-code]
DiffUFlow: Robust Fine-grained Urban Flow Inference with Denoising Diffusion Model, in CIKM 2023. [paper]
Spatio-temporal Diffusion Point Processes, in KDD 2023. [paper] [official-code]
DiffCrime: A Multimodal Conditional Diffusion Model for Crime Risk Map Inference, in KDD 2024. [paper] [official-code]
Social Physics Informed Diffusion Model for Crowd Simulation, in AAAI 2024. [paper] [official-code]
Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model, in NeurIPS 2024. [paper]
Towards Unifying Diffusion Models for Probabilistic Spatio-Temporal Graph Learning, in SIGSPATIAL 2024. [paper] [official-code]
Frequency-aware Generative Models for Multivariate Time Series Imputation, in NeurIPS 2024. [paper] [official-code]
ControlTraj: Controllable Trajectory Generation with Topology-Constrained Diffusion Model, in KDD 2024. [paper] [official-code]
Score-CDM: Score-Weighted Convolutional Diffusion Model for Multivariate Time Series Imputation, in IJCAI 2024. [paper]
SaSDim:Self-adaptive Noise Scaling Diffusion Model for Spatial Time Series Imputation, in IJCAI 2024. [paper]
DiffLight: A Partial Rewards Conditioned Diffusion Model for Traffic Signal Control with Missing Data, in NeurIPS 2024. [paper] [official-code]
PhyDA: Physics-Guided Diffusion Models for Data Assimilation in Atmospheric Systems, in arXiv 2025. [paper]
Dual Contrastive Learning for Spatio-temporal Representation, in ACM MM 2022. [paper]
Masked Autoencoders As Spatiotemporal Learners, in NeurIPS 2022. [paper] [official-code]
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training, in NeurIPS 2022. [paper] [official-code]
SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction, in ICLR 2024. [paper] [official-code]
LLMs are Good Action Recognizers, in CVPR 2024. [paper]
iVideoGPT: Interactive VideoGPTs are Scalable World Models, in NeurIPS 2024. [paper] [official-code]
GenAD: Generalized Predictive Model for Autonomous Driving, in CVPR 2024. [paper] [official-code]
VideoChat: Chat-Centric Video Understanding, in CVPR 2024. [paper] [official-code]
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, in CVPR 2024. [paper] [official-code]
Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion Model, in NeurIPS 2024. [paper]
Leveraging Language Foundation Models for Human Mobility Forecasting, in SIGSPATIAL 2022. [paper] [official-code]
A large language model for electronic health records, in NPJ 2022. [paper]
GATGPT: A Pre-trained Large Language Model with Graph Attention Network for Spatiotemporal Imputation, in arXiv 2023. [paper]
Language Models are Causal Knowledge Extractors for Zero-shot Video Question Answering, in CVPR Workshop 2023. [paper]
Language Models Can Improve Event Prediction by Few-Shot Abductive Reasoning, in NeurIPS 2023. [paper] [official-code]
Harnessing LLMs for Temporal Data - A Study on Explainable Financial Time Series Forecasting, in ACL 2023. [paper]
Health system-scale language models are all-purpose prediction engines, in Nature 2023. [paper]
GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding, in EMNLP 2023. [paper] [official-code]
ChatGPT Informed Graph Neural Network for Stock Movement Prediction, in SIGKDD Workshop 2023. [paper] [official-code]
UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction, in NeurIPS 2024. [paper] [official-code]
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation, in NeurIPS 2024. [paper] [official-code]
Chain-of-History Reasoning for Temporal Knowledge Graph Forecasting, in ACL 2024. [paper]
VideoChat: Chat-Centric Video Understanding, in CVPR 2024. [paper] [official-code]
Extracting Spatiotemporal Data from Gradients with Large Language Models, in arXiv 2024. [paper]
How Can Large Language Models Understand Spatial-Temporal Data?, in arXiv 2024. [paper]
Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming, in arXiv 2024. [paper]
Large Language Models for Next Point-of-Interest Recommendation, in SIGIR 2024. [paper] [official-code]
Latent Logic Tree Extraction for Event Sequence Explanation from LLMs, in ICML 2024. [paper]
Orca: Ocean Significant Wave Height Estimation with Spatio-temporally Aware Large Language Models, in CIKM 2024. [paper]
Semantic Trajectory Data Mining with LLM-Informed POI Classification, in ITSC 2024. [paper]
Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models, in WWW 2024. [paper] [official-code]
MAS4POI: a Multi-Agents Collaboration System for Next POI Recommendation, in arXiv 2024. [paper] [official-code]
Spatial-Temporal Large Language Model for Traffic Prediction, in MDM 2024. [paper] [official-code]
TPLLM: A Traffic Prediction Framework Based on Pretrained Large Language Models, in arXiv 2024. [paper]
ONSEP: A Novel Online Neural-Symbolic Framework for Event Prediction Based on Large Language Model, in ACL 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery, in arXiv 2025. [paper]
LLMLight: Large Language Models as Traffic Signal Control Agents, in KDD 2025. [paper] [official-code]
TimeCMA: Towards LLM-Empowered Time Series Forecasting via Cross-Modality Alignment, in AAAI 2025. [paper] [official-code]
ST-LLM+: Graph Enhanced Spatio-Temporal Large Language Models for Traffic Prediction, in TKDE 2025. [paper]
Path-LLM: A Multi-Modal Path Representation Learning by Aligning and Fusing with Large Language Models, in WWW 2025. [paper] [official-code]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models, in NeurIPS 2022. [paper] [official-code]
VideoLLM: Modeling Video Sequence with Large Language Models, in arXiv 2023. [paper] [official-code]
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models, in ICCV workshop 2023. [paper]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, in EMNLP 2023. [paper] [official-code]
Leveraging Vision-Language Models for Granular Market Change Prediction, in AAAI Workshop 2023. [paper]
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, in CVPR 2024. [paper] [official-code]
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web, in WWW 2024. [paper] [official-code]
VideoLLM-online: Online Video Large Language Model for Streaming Video, in CVPR 2024. [paper] [official-code]
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View, in AAAI 2024. [paper] [official-code]
Seer: Language Instructed Video Prediction with Latent Diffusion Models, in ICLR 2024. [paper] [official-code]
Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial Topology, in KDD 2024. [paper]
UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction, in AAAI 2025. [paper] [official-code]
Leveraging Language Foundation Models for Human Mobility Forecasting, in SIGSPATIAL 2022. [paper] [official-code]
Harnessing LLMs for Temporal Data - A Study on Explainable Financial Time Series Forecasting, in ACL 2023. [paper]
PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting, in TKDE 2023. [paper] [official-code]
Large Language Models are Few-Shot Health Learners, in arXiv 2023. [paper]
Where Would I Go Next? Large Language Models as Human Mobility Predictors, in arXiv 2024. [paper] [official-code]
TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series, in ICLR 2024. [paper] [official-code]
Time-LLM: Time Series Forecasting by Reprogramming Large Language Models, in ICLR 2024. [paper] [official-code]
Language Knowledge-Assisted Representation Learning for Skeleton-Based Action Recognition, in arXiv 2023. [paper] [official-code]
Prompt to Transfer: Sim-to-Real Transfer for Traffic Signal Control with Prompt Learning, in AAAI 2024. [paper] [official-code]
RealTCD: Temporal Causal Discovery from Interventional Data with Large Language Model, in CIKM 2024. [paper]
Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models, in ICCV workshop 2023. [paper]
Leveraging Vision-Language Models for Granular Market Change Prediction, in AAAI Workshop 2023. [paper]
ChatGPT Informed Graph Neural Network for Stock Movement Prediction, in SIGKDD Workshop 2023. [paper] [official-code]
Extracting Spatiotemporal Data from Gradients with Large Language Models, in arXiv 2024. [paper]
Orca: Ocean Significant Wave Height Estimation with Spatio-temporally Aware Large Language Models, in CIKM 2024. [paper]
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation, in NeurIPS 2024. [paper] [official-code]
UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction, in NeurIPS 2024. [paper] [official-code]
MAS4POI: a Multi-Agents Collaboration System for Next POI Recommendation, in arXiv 2024. [paper] [official-code]
Profiling Urban Streets: A Semi-Supervised Prediction Model Based on Street View Imagery and Spatial Topology, in KDD 2024. [paper]
Large Language Models for Next Point-of-Interest Recommendation, in SIGIR 2024. [paper] [official-code]
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web, in WWW 2024. [paper] [official-code]
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models, in NeurIPS 2022. [paper] [official-code]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, in EMNLP 2023. [paper] [official-code]
VideoLLM: Modeling Video Sequence with Large Language Models, in arXiv 2023. [paper] [official-code]
VideoLLM-online: Online Video Large Language Model for Streaming Video, in CVPR 2024. [paper] [official-code]
LLMs are Good Action Recognizers, in CVPR 2024. [paper]
Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming, in arXiv 2024. [paper]
VideoPoet: A Large Language Model for Zero-Shot Video Generation, in ICML 2024. [paper] [official-code]
How Can Large Language Models Understand Spatial-Temporal Data?, in arXiv 2024. [paper]
UrbanGPT: Spatio-Temporal Large Language Models, in KDD 2024. [paper] [official-code]
TimeCMA: Towards LLM-Empowered Time Series Forecasting via Cross-Modality Alignment, in AAAI 2025. [paper] [official-code]
Path-LLM: A Multi-Modal Path Representation Learning by Aligning and Fusing with Large Language Models, in WWW 2025. [paper] [official-code]
SpaBERT: A Pretrained Language Model from Geographic Data for Geo-Entity Representation, in EMNLP 2022. [paper]
Leveraging Language Foundation Models for Human Mobility Forecasting, in SIGSPATIAL 2022. [paper] [official-code]
GATGPT: A Pre-trained Large Language Model with Graph Attention Network for Spatiotemporal Imputation, in arXiv 2023. [paper]
GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding, in EMNLP 2023. [paper] [official-code]
Spatial-Temporal Large Language Model for Traffic Prediction, in MDM 2024. [paper] [official-code]
TPLLM: A Traffic Prediction Framework Based on Pretrained Large Language Models, in arXiv 2024. [paper]
Large Language Models for Next Point-of-Interest Recommendation, in SIGIR 2024. [paper] [official-code]
PTR: A Pre-trained Language Model for Trajectory Recovery, in arXiv 2025. [paper]
Efficient Multivariate Time Series Forecasting via Calibrated Language Models with Privileged Knowledge Distillation, in ICDE 2025. [paper] [official-code]
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook, in arXiv 2023. [paper] [link]
Urban Foundation Models: A Survey, in KDD 2024. [paper] [link]
A Survey on Diffusion Models for Time Series and Spatio-Temporal Data, in arXiv 2024. [paper] [link]
Foundation Models for Time Series Analysis: A Tutorial and Survey, in KDD 2024. [paper]
Foundation Models for Spatio-Temporal Data Science: A Tutorial and Survey, in KDD 2025. [paper]
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities, in arXiv 2025. [paper]
10 commits