A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
3,441
170 commits
updated Sep 23, 2026
📜 A Curated List of Amazing Works in World Modeling, spanning applications in Embodied AI, Autonomous Driving, Natural Language Processing and Agents.
Based on Awesome-World-Model-for-Autonomous-Driving and Awesome-World-Model-for-Robotics.
Photo Credit: Gemini-Nano-Banana🍌.
Major updates and announcements are shown below. Scroll for full timeline.
🚀 [2025-11] 1k+ Stars ⭐️ Under 30 Days — 🌍 Awesome World Models reached 1k github stars within 30 days of initial release, let's go!!!
🗺️ [2025-10] Enhanced Visual Navigation — Introduced badge system for papers! All entries now display
for quick access to resources.
🔥 [2025-10] Repository Launch — Awesome World Models is now live! We're building a comprehensive collection spanning Embodied AI, Autonomous Driving, NLP, and more. See CONTRIBUTING.md for how to contribute.
💡 [Ongoing] Community Contributions Welcome — Help us maintain the most up-to-date world models resource! Submit papers via PR or contact us at email.
⭐ [Ongoing] Support This Project — If you find this useful, please cite our work and give us a star. Share with your research community!
World Models have become a hot topic in both research and industry, attracting unprecedented attention from the AI community and beyond. However, due to the interdisciplinary nature of the field (and because the term "world model" simply sounds amazing), the concept has been used with varying definitions across different domains.
This repository aims to:
Whether you're a researcher looking for related work, a practitioner seeking implementation references, or simply curious about world models, we hope this curated list helps you navigate the landscape!
While world models' outreach has been expanded again and again, it is widely adopted that the original sources of world models come from these two papers:
Some other great blogposts on world models include:
Pixel Space:
3D Mesh Space:
Refer to https://github.com/LMD0311/Awesome-World-Model for full list.
[!NOTE] 📢 [Call for Maintenance] The repo creator is no expert of autonomous driving, so this is a more-than-concise list of works without classification. We anticipate community effort on turning this section cleaner and more well-sorted.
PWM, "From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction".
Dream4Drive, "Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks".
SparseWorld, "SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries".
DriveVLA-W0: "DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving".
"Enhancing Physical Consistency in Lightweight World Models".
IRL-VLA: "IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model".
LiDARCrafter: "LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences".
FASTopoWM: "FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models".
Orbis: "Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models".
"World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving".
NRSeg: "NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models"
World4Drive: "World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model".
Epona: "Epona: Autoregressive Diffusion World Model for Autonomous Driving".
"Towards foundational LiDAR world models with efficient latent flow matching".
SceneDiffuser++: "SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model".
COME: "COME: Adding Scene-Centric Forecasting Control to Occupancy World Model"
STAGE: "STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation".
ReSim: "ReSim: Reliable World Simulation for Autonomous Driving".
"Ego-centric Learning of Communicative World Models for Autonomous Driving".
V2XCrafter: "V2XCrafter: Learning to Generate Driving Scene Across Agents".
Dreamland: "Dreamland: Controllable World Creation with Simulator and Generative Models".
LongDWM: "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model".
GeoDrive: "GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control".
FutureSightDrive: "FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving".
Raw2Drive: "Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)".
VL-SAFE: "VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving".
PosePilot: "PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth".
"World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks".
DriVerse: "DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment".
"End-to-End Driving with Online Trajectory Evaluation via BEV World Model".
"Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles".
MiLA: "MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving".
SimWorld: "SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model".
UniFuture: "Seeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and Perception".
EOT-WM: "Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space".
InDRiVE: "InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model".
MaskGWM: "MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction".
Dream to Drive: "Dream to Drive: Model-Based Vehicle Control Using Analytic World Models".
"Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving".
HERMES: "HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation".
AdaWM: "AdaWM: Adaptive World Model based Planning for Autonomous Driving".
AD-L-JEPA: "AD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data".
DrivingWorld: "DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT".
DrivingGPT: "DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers".
"An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training".
GEM: "GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control".
GaussianWorld: "GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction".
Doe-1: "Doe-1: Closed-Loop Autonomous Driving with Large World Model".
InfiniCube: "InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models".
InfinityDrive: "InfinityDrive: Breaking Time Limits in Driving World Models".
ReconDreamer: "ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration".
Imagine-2-Drive: "Imagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous Vehicles".
DynamicCity: "DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes".
DriveDreamer4D: "World Models Are Effective Data Machines for 4D Driving Scene Representation".
DOME: "Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model".
SSR: "Does End-to-End Autonomous Driving Really Need Perception Tasks?".
"Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models".
LatentDriver: "Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving".
OccLLaMA: "An Occupancy-Language-Action Generative World Model for Autonomous Driving".
DriveGenVLM: "Real-world Video Generation for Vision Language Model based Autonomous Driving".
Drive-OccWorld: "Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving".
CarFormer: "Self-Driving with Learned Object-Centric Representations".
BEVWorld: "A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space".
TOKEN: "Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving".
UMAD: "Unsupervised Mask-Level Anomaly Detection for Autonomous Driving".
AdaptiveDriver: "Planning with Adaptive World Models for Autonomous Driving".
UnO: "Unsupervised Occupancy Fields for Perception and Forecasting".
LAW: "Enhancing End-to-End Autonomous Driving with Latent World Model".
Delphi: "Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation".
OccSora: "4D Occupancy Generation Models as World Simulators for Autonomous Driving".
MagicDrive3D: "Controllable 3D Generation for Any-View Rendering in Street Scenes".
Vista: "A Generalizable Driving World Model with High Fidelity and Versatile Controllability".
CarDreamer: "Open-Source Learning Platform for World Model based Autonomous Driving".
DriveSim: "Probing Multimodal LLMs as World Models for Driving".
DriveWorld: "4D Pre-trained Scene Understanding via World Models for Autonomous Driving".
LidarDM: "Generative LiDAR Simulation in a Generated World".
SubjectDrive: "Scaling Generative Data in Autonomous Driving via Subject Control".
DriveDreamer-2: "LLM-Enhanced World Models for Diverse Driving Video Generation".
Think2Drive: "Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving".
MARL-CCE: "Modelling Competitive Behaviors in Autonomous Driving Under Generative World Model".
GenAD: "Generalized Predictive Model for Autonomous Driving".
NeMo: "Neural Volumetric World Models for Autonomous Driving".
MARL-CCE: "Modelling-Competitive-Behaviors-in-Autonomous-Driving-Under-Generative-World-Model".
ViDAR: "Visual Point Cloud Forecasting enables Scalable Autonomous Driving".
Drive-WM: "Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving".
Cam4DOCC: "Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications".
Panacea: "Panoramic and Controllable Video Generation for Autonomous Driving".
OccWorld: "Learning a 3D Occupancy World Model for Autonomous Driving".
DrivingDiffusion: "Layout-Guided multi-view driving scene video generation with latent diffusion model".
SafeDreamer: "Safe Reinforcement Learning with World Models".
MagicDrive: "Street View Generation with Diverse 3D Geometry Control".
DriveDreamer: "Towards Real-world-driven World Models for Autonomous Driving".
SEM2: "Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model".
Locomotion:
Loco-Manipulation:
Unifying World Models and VLAs in one model:
Combining World Models and VLAs:
This subsection focuses on general policy learning methods in embodied intelligence via leveraging world models.
Real-world policy evaluation is expensive and noisy. The promise of world models is by accurately capturing environment dynamics, it can serve as a surrogate evaluation environment with high correlation to the policy performance in the real world. Before world models, the role for that was simulators:
For World Model Evaluation:
Related resource:
Natural Science:
Social Science:
Interactive Video Generation:
3D Scene Generation:
Genie Series:
V-JEPA Series:
Cosmos Series:
World-Lab Projects:
Other Awesome Models:
The represents a "bottom-up" approach to achieving intelligence, sensorimotor before abstraction. In the 2D pixel space, world models often build upon pre-existing image/video generation approaches.
To what extent does Vision Intelligence exist in Video Generation Models:
Useful Approaches in Video Generation:
From Video Generation Models to World Models:
Pixel Space World Models:
3D Mesh is also a useful representaiton of the physical world, including benefits such as spatial consistency.
The represents a "top-down" approach to achieving intelligence, abstraction before sensorimotor.
Aiming to Advance LLM/VLM skills:
Aiming to enhance computer-use agent performance:
Symbolic World Models:
LLM-in-the-loop World Generation:
A recent trend of work is bridging highly-compressed semantic tokens (e.g. language) with information-sparse cues in the observation space (e.g. vision). This results in World Models that combine high-level and low-level intelligence.
While learning in the observation space (pixel, 3D mesh, language, etc.) is a common approach, for many applications (planning, policy evaluation, etc.) learning in latent space is sufficient or is believed to lead to even better performace.
JEPA is a special kind of learning in latent space, where the loss is put on the latent space, and the encoder & predictor are co-trained. However, the usage of JEPA is not only in world models (e.g. V-JEPA2-AC), but also representation learning (e.g. I-JEPA, V-JEPA), we provide representative works from both perspectives below.
Truncated — view the full README on GitHub.
(top 30 of 38)
A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
3,441
170 commits
updated Sep 23, 2026
📜 A Curated List of Amazing Works in World Modeling, spanning applications in Embodied AI, Autonomous Driving, Natural Language Processing and Agents.
Based on Awesome-World-Model-for-Autonomous-Driving and Awesome-World-Model-for-Robotics.
Photo Credit: Gemini-Nano-Banana🍌.
Major updates and announcements are shown below. Scroll for full timeline.
🚀 [2025-11] 1k+ Stars ⭐️ Under 30 Days — 🌍 Awesome World Models reached 1k github stars within 30 days of initial release, let's go!!!
🗺️ [2025-10] Enhanced Visual Navigation — Introduced badge system for papers! All entries now display
for quick access to resources.
🔥 [2025-10] Repository Launch — Awesome World Models is now live! We're building a comprehensive collection spanning Embodied AI, Autonomous Driving, NLP, and more. See CONTRIBUTING.md for how to contribute.
💡 [Ongoing] Community Contributions Welcome — Help us maintain the most up-to-date world models resource! Submit papers via PR or contact us at email.
⭐ [Ongoing] Support This Project — If you find this useful, please cite our work and give us a star. Share with your research community!
World Models have become a hot topic in both research and industry, attracting unprecedented attention from the AI community and beyond. However, due to the interdisciplinary nature of the field (and because the term "world model" simply sounds amazing), the concept has been used with varying definitions across different domains.
This repository aims to:
Whether you're a researcher looking for related work, a practitioner seeking implementation references, or simply curious about world models, we hope this curated list helps you navigate the landscape!
While world models' outreach has been expanded again and again, it is widely adopted that the original sources of world models come from these two papers:
Some other great blogposts on world models include:
Pixel Space:
3D Mesh Space:
Refer to https://github.com/LMD0311/Awesome-World-Model for full list.
[!NOTE] 📢 [Call for Maintenance] The repo creator is no expert of autonomous driving, so this is a more-than-concise list of works without classification. We anticipate community effort on turning this section cleaner and more well-sorted.
PWM, "From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction".
Dream4Drive, "Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks".
SparseWorld, "SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries".
DriveVLA-W0: "DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving".
"Enhancing Physical Consistency in Lightweight World Models".
IRL-VLA: "IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model".
LiDARCrafter: "LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences".
FASTopoWM: "FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models".
Orbis: "Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models".
"World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving".
NRSeg: "NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models"
World4Drive: "World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model".
Epona: "Epona: Autoregressive Diffusion World Model for Autonomous Driving".
"Towards foundational LiDAR world models with efficient latent flow matching".
SceneDiffuser++: "SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model".
COME: "COME: Adding Scene-Centric Forecasting Control to Occupancy World Model"
STAGE: "STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation".
ReSim: "ReSim: Reliable World Simulation for Autonomous Driving".
"Ego-centric Learning of Communicative World Models for Autonomous Driving".
V2XCrafter: "V2XCrafter: Learning to Generate Driving Scene Across Agents".
Dreamland: "Dreamland: Controllable World Creation with Simulator and Generative Models".
LongDWM: "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model".
GeoDrive: "GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control".
FutureSightDrive: "FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving".
Raw2Drive: "Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)".
VL-SAFE: "VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving".
PosePilot: "PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth".
"World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks".
DriVerse: "DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment".
"End-to-End Driving with Online Trajectory Evaluation via BEV World Model".
"Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles".
MiLA: "MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving".
SimWorld: "SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model".
UniFuture: "Seeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and Perception".
EOT-WM: "Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space".
InDRiVE: "InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model".
MaskGWM: "MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction".
Dream to Drive: "Dream to Drive: Model-Based Vehicle Control Using Analytic World Models".
"Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving".
HERMES: "HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation".
AdaWM: "AdaWM: Adaptive World Model based Planning for Autonomous Driving".
AD-L-JEPA: "AD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data".
DrivingWorld: "DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT".
DrivingGPT: "DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers".
"An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training".
GEM: "GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control".
GaussianWorld: "GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction".
Doe-1: "Doe-1: Closed-Loop Autonomous Driving with Large World Model".
InfiniCube: "InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models".
InfinityDrive: "InfinityDrive: Breaking Time Limits in Driving World Models".
ReconDreamer: "ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration".
Imagine-2-Drive: "Imagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous Vehicles".
DynamicCity: "DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes".
DriveDreamer4D: "World Models Are Effective Data Machines for 4D Driving Scene Representation".
DOME: "Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model".
SSR: "Does End-to-End Autonomous Driving Really Need Perception Tasks?".
"Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models".
LatentDriver: "Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving".
OccLLaMA: "An Occupancy-Language-Action Generative World Model for Autonomous Driving".
DriveGenVLM: "Real-world Video Generation for Vision Language Model based Autonomous Driving".
Drive-OccWorld: "Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving".
CarFormer: "Self-Driving with Learned Object-Centric Representations".
BEVWorld: "A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space".
TOKEN: "Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving".
UMAD: "Unsupervised Mask-Level Anomaly Detection for Autonomous Driving".
AdaptiveDriver: "Planning with Adaptive World Models for Autonomous Driving".
UnO: "Unsupervised Occupancy Fields for Perception and Forecasting".
LAW: "Enhancing End-to-End Autonomous Driving with Latent World Model".
Delphi: "Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation".
OccSora: "4D Occupancy Generation Models as World Simulators for Autonomous Driving".
MagicDrive3D: "Controllable 3D Generation for Any-View Rendering in Street Scenes".
Vista: "A Generalizable Driving World Model with High Fidelity and Versatile Controllability".
CarDreamer: "Open-Source Learning Platform for World Model based Autonomous Driving".
DriveSim: "Probing Multimodal LLMs as World Models for Driving".
DriveWorld: "4D Pre-trained Scene Understanding via World Models for Autonomous Driving".
LidarDM: "Generative LiDAR Simulation in a Generated World".
SubjectDrive: "Scaling Generative Data in Autonomous Driving via Subject Control".
DriveDreamer-2: "LLM-Enhanced World Models for Diverse Driving Video Generation".
Think2Drive: "Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving".
MARL-CCE: "Modelling Competitive Behaviors in Autonomous Driving Under Generative World Model".
GenAD: "Generalized Predictive Model for Autonomous Driving".
NeMo: "Neural Volumetric World Models for Autonomous Driving".
MARL-CCE: "Modelling-Competitive-Behaviors-in-Autonomous-Driving-Under-Generative-World-Model".
ViDAR: "Visual Point Cloud Forecasting enables Scalable Autonomous Driving".
Drive-WM: "Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving".
Cam4DOCC: "Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications".
Panacea: "Panoramic and Controllable Video Generation for Autonomous Driving".
OccWorld: "Learning a 3D Occupancy World Model for Autonomous Driving".
DrivingDiffusion: "Layout-Guided multi-view driving scene video generation with latent diffusion model".
SafeDreamer: "Safe Reinforcement Learning with World Models".
MagicDrive: "Street View Generation with Diverse 3D Geometry Control".
DriveDreamer: "Towards Real-world-driven World Models for Autonomous Driving".
SEM2: "Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model".
Locomotion:
Loco-Manipulation:
Unifying World Models and VLAs in one model:
Combining World Models and VLAs:
This subsection focuses on general policy learning methods in embodied intelligence via leveraging world models.
Real-world policy evaluation is expensive and noisy. The promise of world models is by accurately capturing environment dynamics, it can serve as a surrogate evaluation environment with high correlation to the policy performance in the real world. Before world models, the role for that was simulators:
For World Model Evaluation:
Related resource:
Natural Science:
Social Science:
Interactive Video Generation:
3D Scene Generation:
Genie Series:
V-JEPA Series:
Cosmos Series:
World-Lab Projects:
Other Awesome Models:
The represents a "bottom-up" approach to achieving intelligence, sensorimotor before abstraction. In the 2D pixel space, world models often build upon pre-existing image/video generation approaches.
To what extent does Vision Intelligence exist in Video Generation Models:
Useful Approaches in Video Generation:
From Video Generation Models to World Models:
Pixel Space World Models:
3D Mesh is also a useful representaiton of the physical world, including benefits such as spatial consistency.
The represents a "top-down" approach to achieving intelligence, abstraction before sensorimotor.
Aiming to Advance LLM/VLM skills:
Aiming to enhance computer-use agent performance:
Symbolic World Models:
LLM-in-the-loop World Generation:
A recent trend of work is bridging highly-compressed semantic tokens (e.g. language) with information-sparse cues in the observation space (e.g. vision). This results in World Models that combine high-level and low-level intelligence.
While learning in the observation space (pixel, 3D mesh, language, etc.) is a common approach, for many applications (planning, policy evaluation, etc.) learning in latent space is sufficient or is believed to lead to even better performace.
JEPA is a special kind of learning in latent space, where the loss is put on the latent space, and the encoder & predictor are co-trained. However, the usage of JEPA is not only in world models (e.g. V-JEPA2-AC), but also representation learning (e.g. I-JEPA, V-JEPA), we provide representative works from both perspectives below.
Truncated — view the full README on GitHub.
(top 30 of 38)