AIWorldLab/Awesome-Vision-World-Model

From Seeing to Knowing the World: A Survey of Vision World Models

193

38 commits

updated Aug 18, 2026

See the code

README

Logo Awesome Vision World Models

Awesome Logo Visitors arXiv PRs Welcome

This repository accompanies our survey From Seeing to Knowing the World: A Survey of Vision World Models and maintains a structured collection of Vision World Model resources:

  • It organizes papers across 7 Vision World Model designs and 5 dataset/benchmark categories, with their corresponding arXiv IDs, GitHub repositories, and project pages.
  • It also collects broader resources for world modeling, including theoretical analyses, top-tier conference workshops, interesting repositories, downstream-task applications, and other useful perspectives.

For more details, kindly refer to our paper :rocket:

:newspaper: News

  • 2026-07-07: Add 27 papers, including 19 world model methods and 8 datasets/benchmarks. Update the papers in the teaser figure through 2026-07-01.
  • 2026-05-23: Add 16 papers, 3 github repositories and 5 workshops.
More updates
  • 2026-05-18: Fix metadata for 2 papers.
  • 2026-04-08: Add 45 papers.

Table of Contents

1. Designs

1.1. Sequential Generation

Visual Autoregressive Modeling

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
VideoWorld 2VideoWorld 2: Learning Transferable Knowledge from Real-world VideosarXiv GitHub Website
iMoWMiMoWM: Taming Interactive Multi-Modal World Model for Robotic ManipulationarXivWebsite
PWMFrom Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionarXiv GitHub
SAMPOSAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsarXiv
RynnVLA-001RynnVLA-001: Using Human Demonstrations to Improve Robot ManipulationarXiv GitHub
OccTENSOccTENS: 3D Occupancy World Model via Temporal Next-Scale PredictionarXiv
Genie 3Genie 3: A new frontier for world modelsWebsite
I²-worldI²-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene ForecastingarXiv GitHub
UniVLAUnified Vision-Language-Action ModelarXiv GitHub
WorldVLAWorldVLA: Towards Autoregressive Action World ModelarXiv GitHub
RoboScapeRoboScape: Physics-informed Embodied World ModelarXiv GitHub
Xray2XrayXray2Xray: World Model from Chest X-rays with Volumetric ContextarXiv
RLVR-WorldRLVR-World: Training World Models with Reinforcement LearningarXiv GitHub
MineWorldMineWorld: a Real-Time and Open-Source Interactive World Model on MinecraftarXiv GitHub
UVAUnified Video Action ModelarXiv GitHub
SurgWMSurgical Vision World ModelarXiv GitHub
DWSPre-Trained Video Generative Models as World SimulatorsarXiv
VideoWorldVideoWorld: Exploring Knowledge Learning from Unlabeled VideosarXiv GitHub
DrivingworldDrivingworld: Constructing world model for autonomous driving via video GPTarXiv GitHub
MotoMoto: Latent motion token as the bridging language for learning robot manipulation from videoarXiv GitHub
WHALEWHALE: Towards Generalizable and Scalable World Models for Embodied Decision-makingarXiv
GR-2GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationarXiv Website
LatentDriverLearning Multiple Probabilistic Decisions from Latent World Model in Autonomous DrivingarXiv GitHub
RenderworldRenderworld: World model with self-supervised 3D labelarXiv
VidITVideo In-context Learning: Autoregressive Transformers are Zero-Shot Video ImitatorsarXiv Website
iVideoGPTiVideoGPT: Interactive VideoGPTs are Scalable World ModelsarXiv GitHub
GenieGenie: Generative Interactive EnvironmentsarXiv Website
LWMWorld Model on Million-Length Video And Language With Blockwise RingAttentionarXiv GitHub
WHAMWorld and human action models towards gameplay ideationWebsite
WorldDreamerWorldDreamer: Towards General World Models for Video Generation via Predicting Masked TokensarXiv GitHub
OccWorldOccWorld: Learning a 3D Occupancy World Model for Autonomous DrivingarXiv GitHub
GAIA-1Gaia-1: A generative world model for autonomous drivingarXiv Website

MLLM-guided Multimodal Autoregressive Model

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
WLAWorld-Language-Action Model for Unified World Modeling, Language Reasoning, and Action SynthesisarXiv GitHub
WEMWorld-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied TasksarXiv GitHub Website
HERMES++HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and GenerationarXiv GitHub Website
AstraNav-WorldAstraNav-World: World Model for Foresight Control and ConsistencyarXiv Website GitHub
RynnVLA-002RynnVLA-002: A Unified Vision-Language-Action and World ModelarXiv GitHub
SWMSemantic World ModelsarXiv Website
UniWMUnified World Models: Memory-Augmented Planning and Foresight for Visual NavigationarXiv GitHub
F1F1: A Vision-Language-Action Model Bridging Understanding and Generation to ActionsarXiv GitHub
OccVLAOccVLA: Vision-Language-Action Model with Implicit 3D Occupancy SupervisionarXiv
VLWMPlanning with Reasoning using Vision Language World ModelarXiv
DreamVLADreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgearXiv GitHub
World4OmniWorld4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic ManipulationarXivWebsite
WALL-E 2.0WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentarXivGitHub
GR00T N1GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsarXiv GitHub
HERMESHERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationarXiv GitHub
Doe-1Doe-1: Closed-Loop Autonomous Driving with Large World ModelarXivGitHub
Owl-1Owl-1: Omni World Model for Consistent Long Video GenerationarXiv GitHub
EvaEva: An Embodied World Model for Future Video AnticipationarXiv Website
PIVOT-RPIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic ManipulationarXiv GitHub
OccLLaMAOccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous DrivingarXiv Website
WorldGPTWorldGPT: Empowering LLM as Multimodal World ModelarXiv GitHub
3D-VLA3D-VLA: A 3D Vision-Language-Action Generative World ModelarXivWebsite
ADriver-IADriver-I: A General World Model for Autonomous DrivingarXiv

1.2. Diffusion-based Generation

Latent Diffusion

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
RynnWorld-4DRynnWorld-4D: 4D Embodied World Models for Robotic ManipulationarXiv GitHub Website
Holo-WorldHolo-World: Unified Camera, Object and Weather Control for Video World ModelarXiv GitHub Website
PAIWorldPAIWorld: A 3D-Consistent World Foundation Model for Robotic ManipulationarXiv
Qwen-RobotWorldQwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video GenerationarXiv Website
KairosKairos: A Regret-Aware Native World-Action Model Stack for Physical AIarXiv GitHub
GEM-4DGEM-4D: Geometry-Enhanced Video World Models for Robot ManipulationarXiv Website
RoboFlow4DRoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic ManipulationarXiv GitHub Website
ReactiveGWMReactiveGWM: Steering NPC in Reactive Game World ModelsarXiv GitHub Website
Kinema4DKinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied SimulationarXiv GitHub Website
HyDRAOut of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World ModelsarXivGitHub
PhyGenesisToward Physically Consistent Driving Video World Models under Challenging TrajectoriesarXiv GitHub
ABot-PhysWorldABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics AlignmentarXiv GitHub
EVAEVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics RewardsarXiv GitHub
InSpatio-WorldFMInSpatio-WorldFM: An Open-Source Real-Time Generative Frame ModelarXivGitHub
PlayWorldPlayWorld: Learning Robot World Models from Autonomous PlayarXiv GitHub
WorldCacheWorldCache: Accelerating World Models for Free via Heterogeneous TokenarXiv GitHub
DreamWorldDreamWorld: Unified World Modeling in Video GenerationarXiv GitHub
Olaf-WorldOlaf-World: Orienting Latent Actions for Video World ModelingarXiv GitHub
EgoWMWalk Through Paintings : Ego-centric World Models from Internet PriorsarXiv GitHub
ReWorldReWorld: Multi-Dimensional Reward Modeling for Embodied World ModelsarXiv
NeoVerseNeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosarXivGitHub
GrndCtrlGrndCtrl: Grounding World Models via Self-Supervised Reward AlignmentarXiv Website
C^3World Models That Know When They Don’t Know: Controllable Video Generation with Calibrated UncertaintyarXiv GitHub Website
GAIA-3GAIA-3: Scaling World Models to Power Safety and EvaluationWebsite
WristWorldWristWorld: Generating Wrist-Views via 4D World Models for Robotic ManipulationarXiv GitHub
WorldSplatWorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous DrivingarXiv GitHub
WorldForgeWorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free GuidancearXiv GitHub
PEWMLearning Primitive Embodied World Models: Towards Scalable Robotic LearningarXiv GitHub
Video PolicyVideo Generators are Robot PoliciesarXiv GitHub
Ego-PMEgo-centric Predictive Model Conditioned on Hand TrajectoriesarXiv GitHub
GWMGWM: Towards Scalable Gaussian World Models for Robotic ManipulationarXiv GitHub
M3arsSynthMartian World Models: Controllable Video Synthesis with Physically Accurate 3D ReconstructionsarXiv GitHub
AirScapeAirScape: An Aerial Generative World Model with Motion ControllabilityarXiv GitHub
robot4dgenGeometry-aware 4D Video Generation for Robot ManipulationarXiv GitHub
GenesisGenesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyarXiv GitHub
MinDMinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk AnalysisarXiv GitHub
RealPlayPreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingarXiv GitHub
COMECOME: Adding Scene-Centric Forecasting Control to Occupancy World ModelarXiv GitHub
MeWMMedical World Model: Generative Simulation of Tumor Evolution for Treatment PlanningarXiv GitHub
StateSpaceDiffuserStateSpaceDiffuser: Bringing Long Context to Diffusion World ModelsarXiv GitHub
GeoDriveGeoDrive: 3D Geometry-Informed Driving World Model with Precise Action ControlarXiv GitHub
3DPEWMLearning 3d persistent embodied world modelsarXiv
RoboTransferRoboTransfer: Geometry-Consistent Video Diffusion for Robotic Visual Policy TransferarXiv GitHub
FlowDreamerFlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot ManipulationarXiv GitHub
LaDi-WMLaDi-WM: A Latent Diffusion-based World Model for Predictive ManipulationarXiv GitHub
LangToMoPixel motion as universal representation for robot controlarXiv GitHub
DreamGenDreamGen: Unlocking Generalization in Robot Learning through Video World ModelsarXiv GitHub
Learning to DriveLearning to Drive from a World ModelarXiv Website
TesseractTesseract: Learning 4d embodied world modelsarXiv GitHub
UWMUnified world models: Coupling video and action diffusion for pretraining on large robotic datasetsarXiv GitHub
ViMoViMo: A Generative Visual GUI World Model for App AgentsarXiv GitHub
DiST-4DDiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationarXiv GitHub
AetherAether: Geometric-aware unified world modelingarXiv GitHub
Cosmos transfer1Cosmos transfer1: Conditional world generation with adaptive multimodal controlarXiv GitHub
GAIA-2GAIA-2: A Controllable Multi-View Generative World Model for Autonomous DrivingarXiv Website
EDELINEEDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence ModelingarXiv GitHub
MaskGWMMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionarXiv GitHub
DreamdriveDreamdrive: Generative 4d scene modeling from street view imagesarXiv GitHub
CosmosCosmos World Foundation Model Platform for Physical AIarXiv GitHub
VPPVideo prediction policy: A generalist robot policy with predictive visual representationsarXiv GitHub
Imagine-2-driveImagine-2-drive: High-fidelity world modeling in carla for autonomous vehiclesarXiv GitHub
GenExGenerative World ExplorerarXiv GitHub
DOMEDOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World ModelarXiv GitHub
Panacea+Panacea+: Panoramic and Controllable Video Generation for Autonomous DrivingarXiv GitHub
BevworldBevworld: A multimodal world model for autonomous driving via unified bev latent spacearXiv GitHub
DelphiUnleashing generalization of end-to-end autonomous driving with controllable long video generationarXiv GitHub
OccsoraOccsora: 4d occupancy generation models as world simulators for autonomous drivingarXiv GitHub
GenADGeneralized Predictive Model for Autonomous DrivingarXiv GitHub
WorldGPTWorldgpt: a sora-inspired video ai agent as rich world models from text and image inputsarXiv
DriveDreamer-2DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationarXiv GitHub
WoVoGenWoVoGen: World Volume-Aware Diffusion for Controllable Multi-Camera Driving Scene GenerationarXiv GitHub
Drive-WMDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingarXiv GitHub
PanaceaPanacea: Panoramic and controllable video generation for autonomous drivingarXiv GitHub
DrivingdiffusionDrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion modelarXiv GitHub
DriveDreamerDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivingarXiv GitHub
UniPiLearning universal policies via text-guided video generationarXiv Website

Autoregressive Diffusion

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
MemLearnerMemLearner: Learning to Query Context memory for Video World ModelsarXiv Website
ActWorldActWorld: From Explorable to Interactive World Model via Action-Aware MemoryarXiv Website
DreamX-WorldDreamX-World 1.0: A General-Purpose Interactive World ModelarXiv GitHub Website
MirageLatent Spatial Memory for Video World ModelsarXiv GitHub Website
Echo-MemoryEcho-Memory: A Controlled Study of Memory in Action World ModelsarXiv GitHub Website
Cosmos 3Cosmos 3: Omnimodal World Models for Physical AIarXiv GitHub Website
Gamma-WorldGamma-World: Generative Multi-Agent World Modeling Beyond Two PlayersarXiv GitHub Website
SANA-WMSANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion TransformerarXiv GitHub Website
Matrix-Game 3.0Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon MemoryarXiv GitHub Website
WorldCamWorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric RepresentationarXiv GitHub
SWMGrounding World Simulation Models in a Real-World MetropolisarXiv GitHub
LIVELIVE: Long-horizon Interactive Video World ModelingarXivGitHub
LingBot-WorldAdvancing Open-source World ModelsarXiv GitHub
UniDrive-WMUniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous DrivingarXiv Website
Yume1.5Yume1.5: A Text-Controlled Interactive World Generation ModelarXiv GitHub Website
HY-World 1.5HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric ConsistencyGitHub
AstraAstra: General Interactive World Model with Autoregressive DenoisingarXiv GitHub Website
RELICRELIC: Interactive Video World Model with Long-Horizon MemoryarXiv Website
ANWMAerial World Model for Long-horizon Visual Generation and Navigation in 3D SpacearXiv
WorldPackWorldPack: Compressed Memory Improves Spatial Consistency in Video World ModelingarXiv
Hunyuan-GameCraft-2Hunyuan-GameCraft-2: Instruction-following Interactive Game World ModelarXiv Website
PANPAN: A World Model for General, Interactable, and Long-Horizon World SimulationarXiv Website
Emu3.5Emu3.5: Native Multimodal Models are World LearnersarXiv GitHub
OmniNWMOmniNWM: Omniscient Driving Navigation World ModelsarXiv GitHub
WoWWoW: Towards a World omniscient World model Through Embodied InteractionarXiv GitHub
LongScapePreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingarXiv GitHub
Dreamer V4Training Agents Inside of Scalable World ModelsarXiv Website
Genie EnvisionerGenie Envisioner: A Unified World Foundation Platform for Robotic ManipulationarXiv GitHub
LiDARCrafterLiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesarXiv GitHub
YanYan: Foundational Interactive Video GenerationarXiv Website
Matrix-Game 2.0Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World ModelarXiv GitHub
YumeYume: An Interactive World Generation ModelarXiv GitHub
Geometry ForcingGeometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingarXiv GitHub
spmemVideo World Models with Long-term Spatial MemoryarXiv Website
STAGESTAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene SimulationarXiv Website
CaMContext as Memory: Scene-Consistent Interactive Long Video Generation with Memory RetrievalarXiv Website
DeepVerseDeepVerse: 4D Autoregressive Video Generation as a World ModelarXiv GitHub
EponaEpona: Autoregressive Diffusion World Model for Autonomous DrivingarXiv GitHub
SceneDiffuser++SceneDiffuser++: City-Scale Traffic Simulation via a Generative World ModelarXiv
VMemVMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View MemoryarXiv GitHub
Hunyuan-GameCraftHunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History ConditionarXiv GitHub
Matrix-GameMatrix-Game: Interactive World Foundation ModelarXiv GitHub
PEVAWhole-Body Conditioned Egocentric Video PredictionarXivWebsite
NFDPlaying with Transformer at 30+ FPS via Next-Frame DiffusionarXivWebsite
VRAGLearning World Models for Interactive Video GenerationarXiv Website
DriVerseDriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion AlignmentarXiv GitHub
WORLDMEMWORLDMEM: Long-term Consistent World Simulation with MemoryarXiv GitHub
AdaworldAdaworld: Learning adaptable world models with latent actionsarXiv GitHub
TSWMToward Stable World Models: Measuring and Addressing World Instability in Generative EnvironmentsarXiv
GamefactoryGamefactory: Creating new games with generative interactive videosarXiv GitHub
PlayGenPlayable Game GenerationarXiv GitHub
InfinityDriveInfinityDrive: Breaking Time Limits in Driving World ModelsarXiv Website
GEMGem: A generalizable ego-vision multimodal world model for fine-grained ego-motion, object dynamics, and scene composition controlarXiv GitHub
UniMLVGUniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingarXiv GitHub
The MatrixThe Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlarXivWebsite
NWMNavigation World ModelsarXiv GitHub
Copilot4DCopilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionarXiv Website
DrivingsphereDrivingsphere: Building a high-fidelity 4d world for closed-loop simulationarXiv GitHub
Gamegen-xGamegen-x: Interactive open-world game video generationarXiv GitHub
Genie 2Genie 2: A large-scale foundation world modelWebsite
oasisOasis: A Universe in a TransformerGitHub Website
GameNGenDiffusion models are real-time game enginesarXivWebsite
IRAsimIRASim: A Fine-Grained World Model for Robot ManipulationarXiv GitHub
diamondDiffusion for World Modeling: Visual Details Matter in AtariarXiv GitHub
VistaVista: A generalizable driving world model with high fidelity and versatile controllabilityarXiv GitHub
UniSimUniSim: Learning Interactive Real-World SimulatorsarXivWebsite

1.3. Embedding Prediction

JEPA

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
AdaJEPAAdaJEPA: An Adaptive Latent World ModelarXiv GitHub Website
FR3DFuture Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionarXiv Website
Sub-JEPASub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World ModelsarXiv GitHub Website
HWMHierarchical Planning with Latent World ModelsarXiv GitHub
LeWorldModelLeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from PixelsarXiv GitHub
V-JEPA 2.1V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningarXiv GitHub
Temporal StraighteningTemporal Straightening for Latent PlanningarXiv GitHub
C-JEPACausal-JEPA: Learning World Models through Object-Level Latent InterventionsarXiv GitHub Website
VLA-JEPAVLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelarXiv GitHub
DDP-WMDDP-WM: Disentangled Dynamics Prediction for Efficient World ModelsarXiv GitHub
DINO-worldBack to the Features: DINO as a Foundation for Video World ModelsarXiv
V-JEPA 2V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningarXiv GitHub
SIVWMSparse Imagination for Efficient Visual World Model PlanningarXiv
seq-JEPAseq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World ModelsarXiv GitHub
FlareFlare: Robot learning with implicit world modelingarXiv GitHub
OSVI-WMOSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory GenerationarXiv GitHub
EchoWorldEchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidancearXiv GitHub
AD-L-JEPAAD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR DataarXiv GitHub
DINO-ForesightDINO-Foresight: Looking into the Future with DINOarXiv GitHub
DINO-WMDINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningarXiv GitHub
LAWEnhancing End-to-End Autonomous Driving with Latent World ModelarXiv GitHub
V-JEPARevisiting Feature Prediction for Learning Visual Representations from VideoarXiv GitHub
IWMLearning and Leveraging World Models in Visual Representation LearningarXiv

1.4. State Transition

Latent State-Space Modeling

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
NE-DreamerNext Embedding Prediction Makes World Models StrongerarXiv GitHub
BOOMBootstrap Off-policy with World ModelarXiv GitHub
GASv2Visuomotor Grasping with World Models for Surgical RobotsarXiv
LPSLatent Policy Steering with Embodiment-Agnostic Pretrained World ModelsarXiv
EMERALDAccurate and Efficient World Modeling with Masked Latent TransformersarXiv GitHub
FOUNDERFOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision MakingarXiv Website
NavMorphNavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsarXiv GitHub
Raw2DriveRaw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous DrivingarXiv
SSVWMLong-Context State-Space Video World ModelsarXiv Website
PIN-WMPIN-WM: Learning Physics-Informed World Models for Non-Prehensile ManipulationarXiv GitHub
ReDRAWAdapting World Models with Latent-State Dynamics ResidualsarXiv Website
WoTEEnd-to-End Driving with Online Trajectory Evaluation via BEV World ModelarXiv GitHub
DyWADyWA: Dynamics-Adaptive World Action Model for Generalizable Non-Prehensile ManipulationarXiv GitHub
DMWMDMWM: Dual-Mind World Model with Long-Term ImaginationarXiv GitHub
SimulusUncovering Untapped Potential in Sample-Efficient World Model AgentsarXiv GitHub
S5WMAccelerating Model-Based Reinforcement Learning with State-Space World ModelsarXiv
AdaWMAdaWM: Adaptive World Model based Planning for Autonomous DrivingarXiv
RoboHorizonRoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic ManipulationarXiv
LS-ImagineOpen-World Reinforcement Learning over Long Short-Term ImaginationarXiv GitHub
-Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World ModelsarXiv
GenRLGenRL: Multimodal-Foundation World Models for Generalization in Embodied AgentsarXiv GitHub
DriveWorldDriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingarXiv
PuppeteerHierarchical World Models as Visual Whole-Body Humanoid ControllersarXiv GitHub
R2IMastering Memory Tasks with World ModelsarXiv GitHub
REMImproving Token-Based World Models with Parallel Observation PredictionarXiv GitHub
Think2DriveThink2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous DrivingarXiv GitHub
MUVOMUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric RepresentationsarXiv GitHub
STORMSTORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningarXiv GitHub
HarmonyDreamHarmonyDream: Task Harmonization Inside World ModelsarXiv GitHub
TD-MPC2TD-MPC2: Scalable, Robust World Models for Continuous ControlarXiv GitHub
MoDem-V2MoDem-V2: Visuo-Motor World Models for Real-World Robot ManipulationarXiv GitHub
SWIMStructured World Models from Human VideosarXiv Website
DynalangLearning to Model the World with LanguagearXiv GitHub
SafeDreamerSafeDreamer: Safe Reinforcement Learning with World ModelsarXiv GitHub
CoWorldMaking Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningarXiv GitHub
TWMTransformer-Based World Models Are Happy With 100k InteractionsarXiv GitHub
MV-MWMMulti-View Masked World Models for Visual Robotic ManipulationarXiv GitHub
DreamV3Mastering Diverse Domains through World ModelsarXiv GitHub
IRISTransformers are Sample-Efficient World ModelsarXiv GitHub
MWMMasked World Models for Visual ControlarXiv GitHub
DayDreamerDayDreamer: World Models for Physical Robot LearningarXiv GitHub
CADDYPlayable Video GenerationarXiv GitHub
DreamV2Mastering Atari with Discrete World ModelsarXiv GitHub
DreamV1Dream to Control: Learning Behaviors by Latent ImaginationarXiv GitHub
PlaNetLearning Latent Dynamics for Planning from PixelsarXiv GitHub

Object-Centric Modeling

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
LPWMLatent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingarXiv GitHub
FIOC-WMLearning Interactive World Model for Object-Centric Reinforcement LearningarXiv
Dyn-oDyn-O: Building Structured World Models with Object-Centric RepresentationsarXiv GitHub
SlotPiSlotPi: Physics-informed Object-centric Reasoning ModelsarXiv
-Object-Centric World Model for Language-Guided ManipulationarXiv
DisWMDisentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningarXiv GitHub
Objects matterObjects matter: object-centric world models improve reinforcement learning in visually complex environmentsarXiv
DreamweaverDreamweaver: Learning Compositional World Models from PixelsarXiv GitHub
MEADEfficient Exploration and Discriminative World Model Learning with an Object-Centric AbstractionarXiv
slotSSMSlot State Space ModelsarXiv GitHub
RoboDreamerRoboDreamer: Learning Compositional World Models for Robot ImaginationarXiv GitHub
SSWMSlot Structured World ModelsarXiv GitHub
cosmosNeurosymbolic Grounding for Compositional World ModelsarXiv GitHub
FOCUSFOCUS: Object-Centric World Models for Robotics ManipulationarXiv GitHub
SlotFormerSlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric ModelsarXiv GitHub
HOWMToward Compositional Generalization in Object-Oriented World ModelingarXiv GitHub
G-SWMImproving Generative Imagination in Object-Centric World ModelsarXiv GitHub

1.5. Other Architectures

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
PhysWorldPhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration SynthesisarXiv GitHub
FASTopoWMFASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World ModelsarXiv GitHub
GWMGraph World ModelarXiv GitHub
World4DriveWorld4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World ModelarXiv GitHub
OrbisOrbis: Overcoming Challenges of Long-Horizon Prediction in Driving World ModelsarXiv GitHub
LiDARWMTowards foundational LiDAR world models with efficient latent flow matchingarXiv GitHub
ManiGaussian++ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World ModelarXiv GitHub
GAFGAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic ManipulationarXiv
HWMHumanoid World Models: Open World Foundation Models for Humanoid RoboticsarXiv
DriveXDriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous DrivingarXiv
OccProphetOccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner FrameworkarXiv GitHub
FleetWMMulti-Task Interactive Robot Fleet Learning with Visual World ModelsarXiv GitHub
GaussianWorldGaussianWorld: Gaussian World Model for Streaming 3D Occupancy PredictionarXiv GitHub
NeMoNeural volumetric world models for autonomous drivingWebsite

2. Datasets & Benchmarks

2.1. Foundational World Modeling

General World Prediction and Simulation

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
MemoBenchMemoBench: Benchmarking World Modeling in Dynamically Changing EnvironmentsarXiv GitHub Website
WRBenchCurrent World Models Lack a Persistent State CorearXiv GitHub Website
MBenchMBench: A Comprehensive Benchmark on Memory Capability for Video World ModelsarXiv GitHub Website
WBenchWBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model EvaluationarXiv GitHub Website
iWorld-BenchiWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation FrameworkarXiv GitHub Website
WR-ArenaWorld Reasoning ArenaarXiv GitHub
CoW-BenchThe Trinity of Consistency as a Defining Principle for General World ModelsarXiv GitHub
MINDMIND: Benchmarking Memory Consistency and Action Control in World ModelsarXiv GitHub Website
DynamicVerseDynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingarXiv GitHub Website
4DWorldBench4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsarXiv Website
Gen-ViReCan World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning BenchmarkarXiv GitHub
World-in-WorldWorld-in-World: World Models in a Closed-Loop WorldarXiv GitHub
OmniWorldOmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingarXiv GitHub
SekaiSekai: A Video Dataset towards World ExplorationarXiv GitHub
WorldPredictionWorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural PlanningarXiv Website
WorldScoreWorldScore: A Unified Evaluation Benchmark for World GenerationarXiv GitHub
MM-ORMM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical EnvironmentsarXiv GitHub
WorldModelBenchWorldModelBench: Judging Video Generation Models As World ModelsarXiv GitHub
OpenHumanVidOpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video GenerationarXiv GitHub
EgoVid-5MEgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video GenerationarXiv GitHub
Ego-Exo4DEgo-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesarXiv Website
InternVidInternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationarXiv GitHub
Ego4DEgo4D: Around the World in 3,000 Hours of Egocentric VideoarXiv GitHub
WebVid-2MFrozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalarXiv GitHub
EPIC-KITCHENS-100Rescaling Egocentric VisionarXiv GitHub
HowTo100MHowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsarXiv GitHub
COINCOIN: A Large-scale Dataset for Comprehensive Instructional Video AnalysisarXiv GitHub
SSV2The "something something" video database for learning and evaluating visual common sensearXiv Website
KineticsThe Kinetics Human Action Video DatasetarXiv GitHub
YouTube-8MYouTube-8M: A Large-Scale Video Classification BenchmarkarXiv GitHub
UCF101UCF101: A Dataset of 101 Human Actions Classes From Videos in The WildarXiv Website

Physics and Causality Benchmark

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
PhysEditWorldPhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World ModelsarXiv GitHub Website
Omni-WorldBenchOmni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World ModelsarXiv GitHub
WorldBenchWorldBench: Disambiguating Physics for Diagnostic Evaluation of World ModelsarXiv
VideoVerseVideoVerse: How Far is Your T2V Generator from a World Model?arXiv GitHub
PAI-BenchPhysical ai bench: A comprehensive benchmark for physical ai generation and understandingGitHub
PhysVidBenchCan Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation ModelsarXiv GitHub
PBenchPBench: A benchmark for evaluating generative modelsWebsite
IntPhys 2IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic EnvironmentsarXiv GitHub
T2VPhysBenchT2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video GenerationarXiv
PisaBenchPISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff DroparXiv GitHub
WISA-32KWISA: World Simulator Assistant for Physics-Aware Text-to-Video GenerationarXiv GitHub
VideoPhy-2VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationarXiv GitHub
VBench-2.0VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic FaithfulnessarXiv GitHub
PhyCoBenchA Physical Coherence Benchmark for Evaluating Video Generation Models via Optical Flow-guided Frame PredictionarXiv GitHub
Physics-IQDo generative video models understand physical principles?arXiv GitHub
PhyGenBenchTowards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationarXiv GitHub
PhyBenchPhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image ModelsarXiv GitHub
VideoPhyVideoPhy: Evaluating Physical Commonsense for Video GenerationarXiv GitHub
Physion++Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical PropertiesarXiv Website
InfLevelBenchmarking Progress to Infant-Level Physical Reasoning in AIGitHub
VoEA Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event CategoriesarXiv
PhysionPhysion: Evaluating Physical Prediction from Vision in Humans and MachinesarXiv GitHub
CoPhyCoPhy: Counterfactual Learning of Physical DynamicsarXiv GitHub
IntPhysIntPhys: A Framework and Benchmark for Visual Intuitive Physics ReasoningarXiv GitHub

2.2. Domain-specific World Modeling

Embodied AI and Robotics

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
WMBench (GigaWorld-1)GigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationarXiv GitHub Website
WorldArena 2.0WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and PlatformarXiv GitHub Website
GE-Sim 2.0Genie Envisioner World Simulator 2.0Website
RoboWM-BenchRoboWM-Bench: A Benchmark for Evaluating World Models in Robotic ManipulationarXiv GitHub Website
WorldArenaWorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsarXiv GitHub
DreamDojoDreamDojo: A Generalist Robot World Model from Large-Scale Human VideosarXiv GitHub
WoW-World-EvalWow, wo, val! A Comprehensive Embodied World Model Evaluation Turing TestarXiv
GigaWorld-0GigaWorld-0: World Models as Data Engine to Empower Embodied AIarXiv GitHub Website
Target-BenchTarget-Bench: Can World Models Achieve Mapless Path Planning with Semantic TargetsarXiv GitHub Website
WoWWoW: Towards a world omniscient world model through embodied interactionarXiv GitHub
Meta-World+Meta-World+: An improved, standardized, rl benchmarkarXiv GitHub
AgiBot-WorldAgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsarXiv GitHub
RoboCasaRoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist RobotsarXiv GitHub
DROIDDROID: A Large-Scale In-The-Wild Robot Manipulation DatasetarXiv GitHub
BEHAVIOR-1KBEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic SimulationarXiv GitHub
MimicGenMimicGen: A Data Generation System for Scalable Robot Learning using Human DemonstrationsarXiv GitHub
OXEOpen X-Embodiment: Robotic Learning Datasets and RT-X ModelsarXiv GitHub
RH20TRH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-ShotarXiv GitHub
VP²A Control-Centric Benchmark for Video PredictionarXiv GitHub
RT-1RT-1: Robotics Transformer for Real-World Control at ScalearXiv GitHub
BC-ZBC-Z: Zero-Shot Task Generalization with Robotic Imitation LearningarXiv Website
CALVINCALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation TasksarXiv GitHub
BridgeData V2BridgeData V2: A Dataset for Robot Learning at ScalearXiv GitHub
LIBEROLIBERO: Benchmarking Knowledge Transfer for Lifelong Robot LearningarXiv GitHub
Isaac GymIsaac Gym: High Performance GPU-Based Physics Simulation For Robot LearningarXiv GitHub
RoboNetRoboNet: Large-Scale Multi-Robot LearningarXiv GitHub
Meta-WorldMeta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement LearningarXiv GitHub
RLBenchRLBench: The Robot Learning Benchmark & Learning EnvironmentarXiv GitHub

Autonomous Driving

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
WorldLensWorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldarXiv GitHub Website
ACT-BenchACT-Bench: Towards Action Controllable World Models for Autonomous DrivingarXiv GitHub
DrivingDojoDrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World ModelarXiv GitHub
DriveArenaDriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivingarXiv GitHub
NAVSIMNAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and BenchmarkingarXiv GitHub
CarDreamerCarDreamer: Open-Source Learning Platform for World Model based Autonomous DrivingarXiv GitHub
OpenDV-2KGenAD: Generalized Predictive Model for Autonomous DrivingarXiv GitHub
ZODZenseact Open Dataset: A Large-Scale and Diverse Multimodal Dataset for Autonomous DrivingarXiv GitHub
Occ3DOcc3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous DrivingarXiv GitHub
OpenOccupancyOpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy PerceptionarXiv GitHub
Argoverse 2Argoverse 2: Next Generation Datasets for Self-Driving Perception and ForecastingarXiv GitHub
KITTI-360KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3DarXiv GitHub
NuPlanNuPlan: A closed-loop ML-based planning benchmark for autonomous vehiclesarXiv GitHub
ONCEOne Million Scenes for Autonomous Driving: ONCE DatasetarXiv GitHub
WOMDLarge Scale Interactive Motion Forecasting for Autonomous Driving: The Waymo Open Motion DatasetarXiv GitHub
Lyft Level 5One Thousand and One Hours: Self-driving Motion Prediction DatasetarXiv GitHub
A2D2A2D2: Audi Autonomous Driving DatasetarXiv GitHub
WaymoScalability in Perception for Autonomous Driving: Waymo Open DatasetarXiv GitHub
ArgoverseArgoverse: 3D Tracking and Forecasting With Rich MapsarXiv GitHub
INTERACTIONINTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic MapsarXiv GitHub
SemanticKITTISemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesarXiv GitHub
nuScenesnuScenes: A Multimodal Dataset for Autonomous DrivingarXiv GitHub
BDD100KBDD100K: A Diverse Driving Dataset for Heterogeneous Multitask LearningarXiv GitHub
ApolloScapeThe ApolloScape Dataset for Autonomous DrivingarXiv GitHub
CARLACARLA: An Open Urban Driving SimulatorarXiv GitHub
KITTIAre we ready for autonomous driving? The KITTI vision benchmark suiteWebsite

Interactive Environments and Gaming

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
SCOPESCOPE: Simulating Cross-game Operations in Playable Environments for FPS World ModelsarXiv GitHub Website
WorldMarkWorldMark: A Unified Benchmark Suite for Interactive Video World ModelsarXiv Website
WildWorldWildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPGarXiv GitHub
Matrix-Game-MCMatrix-Game: Interactive World Foundation ModelarXiv GitHub
LOOPNAVToward Memory-Aided World Models: Benchmarking via Spatial ConsistencyarXiv GitHub
JARVIS-VLAJARVIS-VLA: Post-Training Large-Scale Visual Language Models to Play Visual Games with Keyboards and MousearXiv GitHub
GF-MinecraftGameFactory: Creating New Games with Generative Interactive VideosarXiv GitHub
The MatrixThe Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlarXiv Website
GameGen-XGameGen-X: Interactive Open-world Game Video GenerationarXiv GitHub
MarsMars: Situated Inductive Reasoning in an Open-World EnvironmentarXiv GitHub
CrafterBenchmarking the Spectrum of Agent CapabilitiesarXiv GitHub
CS-DeathmatchCounter-Strike Deathmatch with Large-Scale Behavioural CloningarXiv GitHub
MineRLThe MineRL 2020 Competition on Sample Efficient Reinforcement Learning using Human PriorsarXiv GitHub
NLEThe NetHack Learning EnvironmentarXiv GitHub
ProcgenLeveraging Procedural Generation to Benchmark Reinforcement LearningarXiv GitHub
DMCDeepMind Control SuitearXiv GitHub
ALEThe Arcade Learning Environment: An Evaluation Platform for General AgentsarXiv GitHub

3. Others

Survey

:timer_clock: In chronological order, from the latest to the earliest.

PaperLink
World Model for Robot Learning: A Comprehensive SurveyarXiv GitHub
OpenWorldLib: A Unified Codebase and Definition of Advanced World ModelsarXiv GitHub
Video Generation Models as World Models: Efficient Paradigms, Architectures and AlgorithmsarXiv
Learning to Model the World: A Survey of World Models in Artificial IntelligenceTechrXiv GitHub
Towards Generalist Embodied AI: A Survey on World Models for VLA AgentsTechrXiv
Video Generation Models in Robotics -- Applications, Research Challenges, Future DirectionsarXiv
Digital Twin AI: Opportunities and Challenges from Large Language Models to World ModelsarXiv
Progressive Robustness-Aware World Models in Autonomous Driving: A Review and OutlookTechrXiv
Simulating the Visual World with Artificial Intelligence: A RoadmaparXiv GitHub
A Step Toward World Models: A Survey on Robotic ManipulationarXiv
The Safety Challenge of World Models for Embodied AI Agents: A ReviewarXiv
A Comprehensive Survey on World Models for Embodied AIarXiv GitHub
3D and 4D World Modeling: A SurveyarXiv GitHub
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied LearningarXiv
A Survey: Learning Embodied Intelligence from Physical Simulators and World ModelsarXiv GitHub
From 2D to 3D Cognition: A Brief Survey of General World ModelsarXiv
A Survey on World Models Grounded in Acoustic Physical InformationarXiv GitHub
World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a ChildarXiv
The Role of World Models in Shaping Autonomous Driving: A Comprehensive SurveyarXiv GitHub
A Survey of World Models for Autonomous DrivingarXiv GitHub
Understanding World or Predicting Future? A Comprehensive Survey of World ModelsarXiv GitHub
Is Sora a World Simulator? A Comprehensive Survey on General World Models and BeyondarXiv GitHub
World Models for Autonomous Driving: An Initial SurveyarXiv
A Path Towards Autonomous Machine IntelligenceWebsite

GitHub Repo

This part overlaps somewhat with the Survey section.

RepoLink
WMFactory 0.5: One environment · One procedure · Eleven interactive world modelsGitHub
A minimalist repository for training video world models based on diffusion-forcingGitHub
Awesome Video World Models with AR DiffusionGitHub
Awesome From Video Generation to World ModelGitHub
A Curated List of Amazing Works in World Modeling, spanning applications in Embodied AI, Autonomous Driving, Natural Laguage Processing and AgentsGitHub
A curated list of papers for World Models for General Video Generation, Embodied AI, and Autonomous DrivingGitHub
World Models for Autonomous Driving (and Robotic) papersGitHub
Awesome 3D and 4D World ModelsGitHub
Simulating the Real World: Survey & ResourcesGitHub
A curated list of world models for autonomous drivingGitHub
A curated list of resources related to World Models for Autonomous Driving (WMAD)GitHub
A Survey: Learning Embodied Intelligence from Physical Simulators and World ModelsGitHub
A Comprehensive Survey on World Models for Embodied AIGitHub
A Survey on World Models Grounded in Acoustic Physical InformationGitHub
A Comprehensive Survey on General World Models and BeyondGitHub
Understanding World or Predicting Future? A Comprehensive Survey of World ModelsGitHub
From Masks to Worlds: A Hitchhiker's Guide to World ModelsGitHub

Workshop

VenueWorkshopLink
ECCV 2026How to Build Effective World Models for Embodied AIWebsite
CVPR 2026GigaBrain Challenge 2026: World Models & VLA for Embodied IntelligenceWebsite
CVPR 2026Video World ModelsWebsite
CVPR 2026World Models Meet Active Sensing and Closed-Loop PlanningWebsite
ICLR 2026Workshop on World Models: Understanding, Modelling and ScalingWebsite
NeurIPS 2025Workshop on Bridging Language, Agent, and World Models for Reasoning and PlanningWebsite
NeurIPS 2025Embodied World Models for Decision MakingWebsite
ICCV 2025Reliable and Interactable World Models: Geometry, Physics, Interactivity and Real-World GeneralizationWebsite
ICML 2025Building Physically Plausible World ModelsWebsite
CVPR 2025Benchmarking World ModelsWebsite
ICLR 2025World Models: Understanding, Modelling, and ScalingWebsite

Theory

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
OrcaOrca: The World is in Your MindarXiv GitHub Website
LoopWMLooped World ModelsarXiv
OrbiSimOrbiSim: World Models as Differentiable Physics Engines for Embodied IntelligencearXiv Website
-Physically Native World Models: A Hamiltonian Perspective on Generative World ModelingarXiv
ZWMZero-shot World Models Are Developmentally Efficient LearnersarXiv GitHub
-Human Cognition in Machines: A Unified Perspective of World ModelsarXiv
EB-JEPAA Lightweight Library for Energy-Based Joint-Embedding Predictive ArchitecturesarXiv
-Research on World Models Is Not Merely Injecting World Knowledge into Specific TasksarXiv
-What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?arXiv
-Closing the Train-Test Gap in World Models for Gradient-Based PlanningarXiv GitHub
WWMWeb World ModelarXiv GitHub Website
DWMDexterous World ModelsarXiv [GitHub](https://github.com/s

Truncated — view the full README on GitHub.

Contributors

XiaoYu-1123

33 commits

AIROBOTAI

2 commits

zyc781125

2 commits

WayneJin0918

1 commits

AIWorldLab/Awesome-Vision-World-Model

From Seeing to Knowing the World: A Survey of Vision World Models

193

38 commits

updated Aug 18, 2026

See the code

README

Logo Awesome Vision World Models

Awesome Logo Visitors arXiv PRs Welcome

This repository accompanies our survey From Seeing to Knowing the World: A Survey of Vision World Models and maintains a structured collection of Vision World Model resources:

  • It organizes papers across 7 Vision World Model designs and 5 dataset/benchmark categories, with their corresponding arXiv IDs, GitHub repositories, and project pages.
  • It also collects broader resources for world modeling, including theoretical analyses, top-tier conference workshops, interesting repositories, downstream-task applications, and other useful perspectives.

For more details, kindly refer to our paper :rocket:

:newspaper: News

  • 2026-07-07: Add 27 papers, including 19 world model methods and 8 datasets/benchmarks. Update the papers in the teaser figure through 2026-07-01.
  • 2026-05-23: Add 16 papers, 3 github repositories and 5 workshops.
More updates
  • 2026-05-18: Fix metadata for 2 papers.
  • 2026-04-08: Add 45 papers.

Table of Contents

1. Designs

1.1. Sequential Generation

Visual Autoregressive Modeling

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
VideoWorld 2VideoWorld 2: Learning Transferable Knowledge from Real-world VideosarXiv GitHub Website
iMoWMiMoWM: Taming Interactive Multi-Modal World Model for Robotic ManipulationarXivWebsite
PWMFrom Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionarXiv GitHub
SAMPOSAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsarXiv
RynnVLA-001RynnVLA-001: Using Human Demonstrations to Improve Robot ManipulationarXiv GitHub
OccTENSOccTENS: 3D Occupancy World Model via Temporal Next-Scale PredictionarXiv
Genie 3Genie 3: A new frontier for world modelsWebsite
I²-worldI²-World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene ForecastingarXiv GitHub
UniVLAUnified Vision-Language-Action ModelarXiv GitHub
WorldVLAWorldVLA: Towards Autoregressive Action World ModelarXiv GitHub
RoboScapeRoboScape: Physics-informed Embodied World ModelarXiv GitHub
Xray2XrayXray2Xray: World Model from Chest X-rays with Volumetric ContextarXiv
RLVR-WorldRLVR-World: Training World Models with Reinforcement LearningarXiv GitHub
MineWorldMineWorld: a Real-Time and Open-Source Interactive World Model on MinecraftarXiv GitHub
UVAUnified Video Action ModelarXiv GitHub
SurgWMSurgical Vision World ModelarXiv GitHub
DWSPre-Trained Video Generative Models as World SimulatorsarXiv
VideoWorldVideoWorld: Exploring Knowledge Learning from Unlabeled VideosarXiv GitHub
DrivingworldDrivingworld: Constructing world model for autonomous driving via video GPTarXiv GitHub
MotoMoto: Latent motion token as the bridging language for learning robot manipulation from videoarXiv GitHub
WHALEWHALE: Towards Generalizable and Scalable World Models for Embodied Decision-makingarXiv
GR-2GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationarXiv Website
LatentDriverLearning Multiple Probabilistic Decisions from Latent World Model in Autonomous DrivingarXiv GitHub
RenderworldRenderworld: World model with self-supervised 3D labelarXiv
VidITVideo In-context Learning: Autoregressive Transformers are Zero-Shot Video ImitatorsarXiv Website
iVideoGPTiVideoGPT: Interactive VideoGPTs are Scalable World ModelsarXiv GitHub
GenieGenie: Generative Interactive EnvironmentsarXiv Website
LWMWorld Model on Million-Length Video And Language With Blockwise RingAttentionarXiv GitHub
WHAMWorld and human action models towards gameplay ideationWebsite
WorldDreamerWorldDreamer: Towards General World Models for Video Generation via Predicting Masked TokensarXiv GitHub
OccWorldOccWorld: Learning a 3D Occupancy World Model for Autonomous DrivingarXiv GitHub
GAIA-1Gaia-1: A generative world model for autonomous drivingarXiv Website

MLLM-guided Multimodal Autoregressive Model

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
WLAWorld-Language-Action Model for Unified World Modeling, Language Reasoning, and Action SynthesisarXiv GitHub
WEMWorld-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied TasksarXiv GitHub Website
HERMES++HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and GenerationarXiv GitHub Website
AstraNav-WorldAstraNav-World: World Model for Foresight Control and ConsistencyarXiv Website GitHub
RynnVLA-002RynnVLA-002: A Unified Vision-Language-Action and World ModelarXiv GitHub
SWMSemantic World ModelsarXiv Website
UniWMUnified World Models: Memory-Augmented Planning and Foresight for Visual NavigationarXiv GitHub
F1F1: A Vision-Language-Action Model Bridging Understanding and Generation to ActionsarXiv GitHub
OccVLAOccVLA: Vision-Language-Action Model with Implicit 3D Occupancy SupervisionarXiv
VLWMPlanning with Reasoning using Vision Language World ModelarXiv
DreamVLADreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgearXiv GitHub
World4OmniWorld4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic ManipulationarXivWebsite
WALL-E 2.0WALL-E 2.0: World Alignment by NeuroSymbolic Learning Improves World Model-based LLM AgentarXivGitHub
GR00T N1GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsarXiv GitHub
HERMESHERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationarXiv GitHub
Doe-1Doe-1: Closed-Loop Autonomous Driving with Large World ModelarXivGitHub
Owl-1Owl-1: Omni World Model for Consistent Long Video GenerationarXiv GitHub
EvaEva: An Embodied World Model for Future Video AnticipationarXiv Website
PIVOT-RPIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic ManipulationarXiv GitHub
OccLLaMAOccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous DrivingarXiv Website
WorldGPTWorldGPT: Empowering LLM as Multimodal World ModelarXiv GitHub
3D-VLA3D-VLA: A 3D Vision-Language-Action Generative World ModelarXivWebsite
ADriver-IADriver-I: A General World Model for Autonomous DrivingarXiv

1.2. Diffusion-based Generation

Latent Diffusion

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
RynnWorld-4DRynnWorld-4D: 4D Embodied World Models for Robotic ManipulationarXiv GitHub Website
Holo-WorldHolo-World: Unified Camera, Object and Weather Control for Video World ModelarXiv GitHub Website
PAIWorldPAIWorld: A 3D-Consistent World Foundation Model for Robotic ManipulationarXiv
Qwen-RobotWorldQwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video GenerationarXiv Website
KairosKairos: A Regret-Aware Native World-Action Model Stack for Physical AIarXiv GitHub
GEM-4DGEM-4D: Geometry-Enhanced Video World Models for Robot ManipulationarXiv Website
RoboFlow4DRoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic ManipulationarXiv GitHub Website
ReactiveGWMReactiveGWM: Steering NPC in Reactive Game World ModelsarXiv GitHub Website
Kinema4DKinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied SimulationarXiv GitHub Website
HyDRAOut of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World ModelsarXivGitHub
PhyGenesisToward Physically Consistent Driving Video World Models under Challenging TrajectoriesarXiv GitHub
ABot-PhysWorldABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics AlignmentarXiv GitHub
EVAEVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics RewardsarXiv GitHub
InSpatio-WorldFMInSpatio-WorldFM: An Open-Source Real-Time Generative Frame ModelarXivGitHub
PlayWorldPlayWorld: Learning Robot World Models from Autonomous PlayarXiv GitHub
WorldCacheWorldCache: Accelerating World Models for Free via Heterogeneous TokenarXiv GitHub
DreamWorldDreamWorld: Unified World Modeling in Video GenerationarXiv GitHub
Olaf-WorldOlaf-World: Orienting Latent Actions for Video World ModelingarXiv GitHub
EgoWMWalk Through Paintings : Ego-centric World Models from Internet PriorsarXiv GitHub
ReWorldReWorld: Multi-Dimensional Reward Modeling for Embodied World ModelsarXiv
NeoVerseNeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosarXivGitHub
GrndCtrlGrndCtrl: Grounding World Models via Self-Supervised Reward AlignmentarXiv Website
C^3World Models That Know When They Don’t Know: Controllable Video Generation with Calibrated UncertaintyarXiv GitHub Website
GAIA-3GAIA-3: Scaling World Models to Power Safety and EvaluationWebsite
WristWorldWristWorld: Generating Wrist-Views via 4D World Models for Robotic ManipulationarXiv GitHub
WorldSplatWorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous DrivingarXiv GitHub
WorldForgeWorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free GuidancearXiv GitHub
PEWMLearning Primitive Embodied World Models: Towards Scalable Robotic LearningarXiv GitHub
Video PolicyVideo Generators are Robot PoliciesarXiv GitHub
Ego-PMEgo-centric Predictive Model Conditioned on Hand TrajectoriesarXiv GitHub
GWMGWM: Towards Scalable Gaussian World Models for Robotic ManipulationarXiv GitHub
M3arsSynthMartian World Models: Controllable Video Synthesis with Physically Accurate 3D ReconstructionsarXiv GitHub
AirScapeAirScape: An Aerial Generative World Model with Motion ControllabilityarXiv GitHub
robot4dgenGeometry-aware 4D Video Generation for Robot ManipulationarXiv GitHub
GenesisGenesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyarXiv GitHub
MinDMinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk AnalysisarXiv GitHub
RealPlayPreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingarXiv GitHub
COMECOME: Adding Scene-Centric Forecasting Control to Occupancy World ModelarXiv GitHub
MeWMMedical World Model: Generative Simulation of Tumor Evolution for Treatment PlanningarXiv GitHub
StateSpaceDiffuserStateSpaceDiffuser: Bringing Long Context to Diffusion World ModelsarXiv GitHub
GeoDriveGeoDrive: 3D Geometry-Informed Driving World Model with Precise Action ControlarXiv GitHub
3DPEWMLearning 3d persistent embodied world modelsarXiv
RoboTransferRoboTransfer: Geometry-Consistent Video Diffusion for Robotic Visual Policy TransferarXiv GitHub
FlowDreamerFlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot ManipulationarXiv GitHub
LaDi-WMLaDi-WM: A Latent Diffusion-based World Model for Predictive ManipulationarXiv GitHub
LangToMoPixel motion as universal representation for robot controlarXiv GitHub
DreamGenDreamGen: Unlocking Generalization in Robot Learning through Video World ModelsarXiv GitHub
Learning to DriveLearning to Drive from a World ModelarXiv Website
TesseractTesseract: Learning 4d embodied world modelsarXiv GitHub
UWMUnified world models: Coupling video and action diffusion for pretraining on large robotic datasetsarXiv GitHub
ViMoViMo: A Generative Visual GUI World Model for App AgentsarXiv GitHub
DiST-4DDiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationarXiv GitHub
AetherAether: Geometric-aware unified world modelingarXiv GitHub
Cosmos transfer1Cosmos transfer1: Conditional world generation with adaptive multimodal controlarXiv GitHub
GAIA-2GAIA-2: A Controllable Multi-View Generative World Model for Autonomous DrivingarXiv Website
EDELINEEDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence ModelingarXiv GitHub
MaskGWMMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionarXiv GitHub
DreamdriveDreamdrive: Generative 4d scene modeling from street view imagesarXiv GitHub
CosmosCosmos World Foundation Model Platform for Physical AIarXiv GitHub
VPPVideo prediction policy: A generalist robot policy with predictive visual representationsarXiv GitHub
Imagine-2-driveImagine-2-drive: High-fidelity world modeling in carla for autonomous vehiclesarXiv GitHub
GenExGenerative World ExplorerarXiv GitHub
DOMEDOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World ModelarXiv GitHub
Panacea+Panacea+: Panoramic and Controllable Video Generation for Autonomous DrivingarXiv GitHub
BevworldBevworld: A multimodal world model for autonomous driving via unified bev latent spacearXiv GitHub
DelphiUnleashing generalization of end-to-end autonomous driving with controllable long video generationarXiv GitHub
OccsoraOccsora: 4d occupancy generation models as world simulators for autonomous drivingarXiv GitHub
GenADGeneralized Predictive Model for Autonomous DrivingarXiv GitHub
WorldGPTWorldgpt: a sora-inspired video ai agent as rich world models from text and image inputsarXiv
DriveDreamer-2DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationarXiv GitHub
WoVoGenWoVoGen: World Volume-Aware Diffusion for Controllable Multi-Camera Driving Scene GenerationarXiv GitHub
Drive-WMDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingarXiv GitHub
PanaceaPanacea: Panoramic and controllable video generation for autonomous drivingarXiv GitHub
DrivingdiffusionDrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion modelarXiv GitHub
DriveDreamerDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivingarXiv GitHub
UniPiLearning universal policies via text-guided video generationarXiv Website

Autoregressive Diffusion

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
MemLearnerMemLearner: Learning to Query Context memory for Video World ModelsarXiv Website
ActWorldActWorld: From Explorable to Interactive World Model via Action-Aware MemoryarXiv Website
DreamX-WorldDreamX-World 1.0: A General-Purpose Interactive World ModelarXiv GitHub Website
MirageLatent Spatial Memory for Video World ModelsarXiv GitHub Website
Echo-MemoryEcho-Memory: A Controlled Study of Memory in Action World ModelsarXiv GitHub Website
Cosmos 3Cosmos 3: Omnimodal World Models for Physical AIarXiv GitHub Website
Gamma-WorldGamma-World: Generative Multi-Agent World Modeling Beyond Two PlayersarXiv GitHub Website
SANA-WMSANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion TransformerarXiv GitHub Website
Matrix-Game 3.0Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon MemoryarXiv GitHub Website
WorldCamWorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric RepresentationarXiv GitHub
SWMGrounding World Simulation Models in a Real-World MetropolisarXiv GitHub
LIVELIVE: Long-horizon Interactive Video World ModelingarXivGitHub
LingBot-WorldAdvancing Open-source World ModelsarXiv GitHub
UniDrive-WMUniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous DrivingarXiv Website
Yume1.5Yume1.5: A Text-Controlled Interactive World Generation ModelarXiv GitHub Website
HY-World 1.5HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric ConsistencyGitHub
AstraAstra: General Interactive World Model with Autoregressive DenoisingarXiv GitHub Website
RELICRELIC: Interactive Video World Model with Long-Horizon MemoryarXiv Website
ANWMAerial World Model for Long-horizon Visual Generation and Navigation in 3D SpacearXiv
WorldPackWorldPack: Compressed Memory Improves Spatial Consistency in Video World ModelingarXiv
Hunyuan-GameCraft-2Hunyuan-GameCraft-2: Instruction-following Interactive Game World ModelarXiv Website
PANPAN: A World Model for General, Interactable, and Long-Horizon World SimulationarXiv Website
Emu3.5Emu3.5: Native Multimodal Models are World LearnersarXiv GitHub
OmniNWMOmniNWM: Omniscient Driving Navigation World ModelsarXiv GitHub
WoWWoW: Towards a World omniscient World model Through Embodied InteractionarXiv GitHub
LongScapePreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingarXiv GitHub
Dreamer V4Training Agents Inside of Scalable World ModelsarXiv Website
Genie EnvisionerGenie Envisioner: A Unified World Foundation Platform for Robotic ManipulationarXiv GitHub
LiDARCrafterLiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesarXiv GitHub
YanYan: Foundational Interactive Video GenerationarXiv Website
Matrix-Game 2.0Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World ModelarXiv GitHub
YumeYume: An Interactive World Generation ModelarXiv GitHub
Geometry ForcingGeometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingarXiv GitHub
spmemVideo World Models with Long-term Spatial MemoryarXiv Website
STAGESTAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene SimulationarXiv Website
CaMContext as Memory: Scene-Consistent Interactive Long Video Generation with Memory RetrievalarXiv Website
DeepVerseDeepVerse: 4D Autoregressive Video Generation as a World ModelarXiv GitHub
EponaEpona: Autoregressive Diffusion World Model for Autonomous DrivingarXiv GitHub
SceneDiffuser++SceneDiffuser++: City-Scale Traffic Simulation via a Generative World ModelarXiv
VMemVMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View MemoryarXiv GitHub
Hunyuan-GameCraftHunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History ConditionarXiv GitHub
Matrix-GameMatrix-Game: Interactive World Foundation ModelarXiv GitHub
PEVAWhole-Body Conditioned Egocentric Video PredictionarXivWebsite
NFDPlaying with Transformer at 30+ FPS via Next-Frame DiffusionarXivWebsite
VRAGLearning World Models for Interactive Video GenerationarXiv Website
DriVerseDriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion AlignmentarXiv GitHub
WORLDMEMWORLDMEM: Long-term Consistent World Simulation with MemoryarXiv GitHub
AdaworldAdaworld: Learning adaptable world models with latent actionsarXiv GitHub
TSWMToward Stable World Models: Measuring and Addressing World Instability in Generative EnvironmentsarXiv
GamefactoryGamefactory: Creating new games with generative interactive videosarXiv GitHub
PlayGenPlayable Game GenerationarXiv GitHub
InfinityDriveInfinityDrive: Breaking Time Limits in Driving World ModelsarXiv Website
GEMGem: A generalizable ego-vision multimodal world model for fine-grained ego-motion, object dynamics, and scene composition controlarXiv GitHub
UniMLVGUniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingarXiv GitHub
The MatrixThe Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlarXivWebsite
NWMNavigation World ModelsarXiv GitHub
Copilot4DCopilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionarXiv Website
DrivingsphereDrivingsphere: Building a high-fidelity 4d world for closed-loop simulationarXiv GitHub
Gamegen-xGamegen-x: Interactive open-world game video generationarXiv GitHub
Genie 2Genie 2: A large-scale foundation world modelWebsite
oasisOasis: A Universe in a TransformerGitHub Website
GameNGenDiffusion models are real-time game enginesarXivWebsite
IRAsimIRASim: A Fine-Grained World Model for Robot ManipulationarXiv GitHub
diamondDiffusion for World Modeling: Visual Details Matter in AtariarXiv GitHub
VistaVista: A generalizable driving world model with high fidelity and versatile controllabilityarXiv GitHub
UniSimUniSim: Learning Interactive Real-World SimulatorsarXivWebsite

1.3. Embedding Prediction

JEPA

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
AdaJEPAAdaJEPA: An Adaptive Latent World ModelarXiv GitHub Website
FR3DFuture Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionarXiv Website
Sub-JEPASub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World ModelsarXiv GitHub Website
HWMHierarchical Planning with Latent World ModelsarXiv GitHub
LeWorldModelLeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from PixelsarXiv GitHub
V-JEPA 2.1V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised LearningarXiv GitHub
Temporal StraighteningTemporal Straightening for Latent PlanningarXiv GitHub
C-JEPACausal-JEPA: Learning World Models through Object-Level Latent InterventionsarXiv GitHub Website
VLA-JEPAVLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelarXiv GitHub
DDP-WMDDP-WM: Disentangled Dynamics Prediction for Efficient World ModelsarXiv GitHub
DINO-worldBack to the Features: DINO as a Foundation for Video World ModelsarXiv
V-JEPA 2V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningarXiv GitHub
SIVWMSparse Imagination for Efficient Visual World Model PlanningarXiv
seq-JEPAseq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World ModelsarXiv GitHub
FlareFlare: Robot learning with implicit world modelingarXiv GitHub
OSVI-WMOSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory GenerationarXiv GitHub
EchoWorldEchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidancearXiv GitHub
AD-L-JEPAAD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR DataarXiv GitHub
DINO-ForesightDINO-Foresight: Looking into the Future with DINOarXiv GitHub
DINO-WMDINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningarXiv GitHub
LAWEnhancing End-to-End Autonomous Driving with Latent World ModelarXiv GitHub
V-JEPARevisiting Feature Prediction for Learning Visual Representations from VideoarXiv GitHub
IWMLearning and Leveraging World Models in Visual Representation LearningarXiv

1.4. State Transition

Latent State-Space Modeling

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
NE-DreamerNext Embedding Prediction Makes World Models StrongerarXiv GitHub
BOOMBootstrap Off-policy with World ModelarXiv GitHub
GASv2Visuomotor Grasping with World Models for Surgical RobotsarXiv
LPSLatent Policy Steering with Embodiment-Agnostic Pretrained World ModelsarXiv
EMERALDAccurate and Efficient World Modeling with Masked Latent TransformersarXiv GitHub
FOUNDERFOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision MakingarXiv Website
NavMorphNavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsarXiv GitHub
Raw2DriveRaw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous DrivingarXiv
SSVWMLong-Context State-Space Video World ModelsarXiv Website
PIN-WMPIN-WM: Learning Physics-Informed World Models for Non-Prehensile ManipulationarXiv GitHub
ReDRAWAdapting World Models with Latent-State Dynamics ResidualsarXiv Website
WoTEEnd-to-End Driving with Online Trajectory Evaluation via BEV World ModelarXiv GitHub
DyWADyWA: Dynamics-Adaptive World Action Model for Generalizable Non-Prehensile ManipulationarXiv GitHub
DMWMDMWM: Dual-Mind World Model with Long-Term ImaginationarXiv GitHub
SimulusUncovering Untapped Potential in Sample-Efficient World Model AgentsarXiv GitHub
S5WMAccelerating Model-Based Reinforcement Learning with State-Space World ModelsarXiv
AdaWMAdaWM: Adaptive World Model based Planning for Autonomous DrivingarXiv
RoboHorizonRoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic ManipulationarXiv
LS-ImagineOpen-World Reinforcement Learning over Long Short-Term ImaginationarXiv GitHub
-Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World ModelsarXiv
GenRLGenRL: Multimodal-Foundation World Models for Generalization in Embodied AgentsarXiv GitHub
DriveWorldDriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingarXiv
PuppeteerHierarchical World Models as Visual Whole-Body Humanoid ControllersarXiv GitHub
R2IMastering Memory Tasks with World ModelsarXiv GitHub
REMImproving Token-Based World Models with Parallel Observation PredictionarXiv GitHub
Think2DriveThink2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous DrivingarXiv GitHub
MUVOMUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric RepresentationsarXiv GitHub
STORMSTORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningarXiv GitHub
HarmonyDreamHarmonyDream: Task Harmonization Inside World ModelsarXiv GitHub
TD-MPC2TD-MPC2: Scalable, Robust World Models for Continuous ControlarXiv GitHub
MoDem-V2MoDem-V2: Visuo-Motor World Models for Real-World Robot ManipulationarXiv GitHub
SWIMStructured World Models from Human VideosarXiv Website
DynalangLearning to Model the World with LanguagearXiv GitHub
SafeDreamerSafeDreamer: Safe Reinforcement Learning with World ModelsarXiv GitHub
CoWorldMaking Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningarXiv GitHub
TWMTransformer-Based World Models Are Happy With 100k InteractionsarXiv GitHub
MV-MWMMulti-View Masked World Models for Visual Robotic ManipulationarXiv GitHub
DreamV3Mastering Diverse Domains through World ModelsarXiv GitHub
IRISTransformers are Sample-Efficient World ModelsarXiv GitHub
MWMMasked World Models for Visual ControlarXiv GitHub
DayDreamerDayDreamer: World Models for Physical Robot LearningarXiv GitHub
CADDYPlayable Video GenerationarXiv GitHub
DreamV2Mastering Atari with Discrete World ModelsarXiv GitHub
DreamV1Dream to Control: Learning Behaviors by Latent ImaginationarXiv GitHub
PlaNetLearning Latent Dynamics for Planning from PixelsarXiv GitHub

Object-Centric Modeling

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
LPWMLatent Particle World Models: Self-supervised Object-centric Stochastic Dynamics ModelingarXiv GitHub
FIOC-WMLearning Interactive World Model for Object-Centric Reinforcement LearningarXiv
Dyn-oDyn-O: Building Structured World Models with Object-Centric RepresentationsarXiv GitHub
SlotPiSlotPi: Physics-informed Object-centric Reasoning ModelsarXiv
-Object-Centric World Model for Language-Guided ManipulationarXiv
DisWMDisentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningarXiv GitHub
Objects matterObjects matter: object-centric world models improve reinforcement learning in visually complex environmentsarXiv
DreamweaverDreamweaver: Learning Compositional World Models from PixelsarXiv GitHub
MEADEfficient Exploration and Discriminative World Model Learning with an Object-Centric AbstractionarXiv
slotSSMSlot State Space ModelsarXiv GitHub
RoboDreamerRoboDreamer: Learning Compositional World Models for Robot ImaginationarXiv GitHub
SSWMSlot Structured World ModelsarXiv GitHub
cosmosNeurosymbolic Grounding for Compositional World ModelsarXiv GitHub
FOCUSFOCUS: Object-Centric World Models for Robotics ManipulationarXiv GitHub
SlotFormerSlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric ModelsarXiv GitHub
HOWMToward Compositional Generalization in Object-Oriented World ModelingarXiv GitHub
G-SWMImproving Generative Imagination in Object-Centric World ModelsarXiv GitHub

1.5. Other Architectures

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
PhysWorldPhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration SynthesisarXiv GitHub
FASTopoWMFASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World ModelsarXiv GitHub
GWMGraph World ModelarXiv GitHub
World4DriveWorld4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World ModelarXiv GitHub
OrbisOrbis: Overcoming Challenges of Long-Horizon Prediction in Driving World ModelsarXiv GitHub
LiDARWMTowards foundational LiDAR world models with efficient latent flow matchingarXiv GitHub
ManiGaussian++ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World ModelarXiv GitHub
GAFGAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic ManipulationarXiv
HWMHumanoid World Models: Open World Foundation Models for Humanoid RoboticsarXiv
DriveXDriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous DrivingarXiv
OccProphetOccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner FrameworkarXiv GitHub
FleetWMMulti-Task Interactive Robot Fleet Learning with Visual World ModelsarXiv GitHub
GaussianWorldGaussianWorld: Gaussian World Model for Streaming 3D Occupancy PredictionarXiv GitHub
NeMoNeural volumetric world models for autonomous drivingWebsite

2. Datasets & Benchmarks

2.1. Foundational World Modeling

General World Prediction and Simulation

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
MemoBenchMemoBench: Benchmarking World Modeling in Dynamically Changing EnvironmentsarXiv GitHub Website
WRBenchCurrent World Models Lack a Persistent State CorearXiv GitHub Website
MBenchMBench: A Comprehensive Benchmark on Memory Capability for Video World ModelsarXiv GitHub Website
WBenchWBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model EvaluationarXiv GitHub Website
iWorld-BenchiWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation FrameworkarXiv GitHub Website
WR-ArenaWorld Reasoning ArenaarXiv GitHub
CoW-BenchThe Trinity of Consistency as a Defining Principle for General World ModelsarXiv GitHub
MINDMIND: Benchmarking Memory Consistency and Action Control in World ModelsarXiv GitHub Website
DynamicVerseDynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingarXiv GitHub Website
4DWorldBench4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsarXiv Website
Gen-ViReCan World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning BenchmarkarXiv GitHub
World-in-WorldWorld-in-World: World Models in a Closed-Loop WorldarXiv GitHub
OmniWorldOmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingarXiv GitHub
SekaiSekai: A Video Dataset towards World ExplorationarXiv GitHub
WorldPredictionWorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural PlanningarXiv Website
WorldScoreWorldScore: A Unified Evaluation Benchmark for World GenerationarXiv GitHub
MM-ORMM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical EnvironmentsarXiv GitHub
WorldModelBenchWorldModelBench: Judging Video Generation Models As World ModelsarXiv GitHub
OpenHumanVidOpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video GenerationarXiv GitHub
EgoVid-5MEgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video GenerationarXiv GitHub
Ego-Exo4DEgo-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesarXiv Website
InternVidInternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationarXiv GitHub
Ego4DEgo4D: Around the World in 3,000 Hours of Egocentric VideoarXiv GitHub
WebVid-2MFrozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalarXiv GitHub
EPIC-KITCHENS-100Rescaling Egocentric VisionarXiv GitHub
HowTo100MHowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsarXiv GitHub
COINCOIN: A Large-scale Dataset for Comprehensive Instructional Video AnalysisarXiv GitHub
SSV2The "something something" video database for learning and evaluating visual common sensearXiv Website
KineticsThe Kinetics Human Action Video DatasetarXiv GitHub
YouTube-8MYouTube-8M: A Large-Scale Video Classification BenchmarkarXiv GitHub
UCF101UCF101: A Dataset of 101 Human Actions Classes From Videos in The WildarXiv Website

Physics and Causality Benchmark

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
PhysEditWorldPhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World ModelsarXiv GitHub Website
Omni-WorldBenchOmni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World ModelsarXiv GitHub
WorldBenchWorldBench: Disambiguating Physics for Diagnostic Evaluation of World ModelsarXiv
VideoVerseVideoVerse: How Far is Your T2V Generator from a World Model?arXiv GitHub
PAI-BenchPhysical ai bench: A comprehensive benchmark for physical ai generation and understandingGitHub
PhysVidBenchCan Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation ModelsarXiv GitHub
PBenchPBench: A benchmark for evaluating generative modelsWebsite
IntPhys 2IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic EnvironmentsarXiv GitHub
T2VPhysBenchT2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video GenerationarXiv
PisaBenchPISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff DroparXiv GitHub
WISA-32KWISA: World Simulator Assistant for Physics-Aware Text-to-Video GenerationarXiv GitHub
VideoPhy-2VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationarXiv GitHub
VBench-2.0VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic FaithfulnessarXiv GitHub
PhyCoBenchA Physical Coherence Benchmark for Evaluating Video Generation Models via Optical Flow-guided Frame PredictionarXiv GitHub
Physics-IQDo generative video models understand physical principles?arXiv GitHub
PhyGenBenchTowards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video GenerationarXiv GitHub
PhyBenchPhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image ModelsarXiv GitHub
VideoPhyVideoPhy: Evaluating Physical Commonsense for Video GenerationarXiv GitHub
Physion++Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical PropertiesarXiv Website
InfLevelBenchmarking Progress to Infant-Level Physical Reasoning in AIGitHub
VoEA Benchmark for Modeling Violation-of-Expectation in Physical Reasoning Across Event CategoriesarXiv
PhysionPhysion: Evaluating Physical Prediction from Vision in Humans and MachinesarXiv GitHub
CoPhyCoPhy: Counterfactual Learning of Physical DynamicsarXiv GitHub
IntPhysIntPhys: A Framework and Benchmark for Visual Intuitive Physics ReasoningarXiv GitHub

2.2. Domain-specific World Modeling

Embodied AI and Robotics

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
WMBench (GigaWorld-1)GigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationarXiv GitHub Website
WorldArena 2.0WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and PlatformarXiv GitHub Website
GE-Sim 2.0Genie Envisioner World Simulator 2.0Website
RoboWM-BenchRoboWM-Bench: A Benchmark for Evaluating World Models in Robotic ManipulationarXiv GitHub Website
WorldArenaWorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsarXiv GitHub
DreamDojoDreamDojo: A Generalist Robot World Model from Large-Scale Human VideosarXiv GitHub
WoW-World-EvalWow, wo, val! A Comprehensive Embodied World Model Evaluation Turing TestarXiv
GigaWorld-0GigaWorld-0: World Models as Data Engine to Empower Embodied AIarXiv GitHub Website
Target-BenchTarget-Bench: Can World Models Achieve Mapless Path Planning with Semantic TargetsarXiv GitHub Website
WoWWoW: Towards a world omniscient world model through embodied interactionarXiv GitHub
Meta-World+Meta-World+: An improved, standardized, rl benchmarkarXiv GitHub
AgiBot-WorldAgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied SystemsarXiv GitHub
RoboCasaRoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist RobotsarXiv GitHub
DROIDDROID: A Large-Scale In-The-Wild Robot Manipulation DatasetarXiv GitHub
BEHAVIOR-1KBEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic SimulationarXiv GitHub
MimicGenMimicGen: A Data Generation System for Scalable Robot Learning using Human DemonstrationsarXiv GitHub
OXEOpen X-Embodiment: Robotic Learning Datasets and RT-X ModelsarXiv GitHub
RH20TRH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-ShotarXiv GitHub
VP²A Control-Centric Benchmark for Video PredictionarXiv GitHub
RT-1RT-1: Robotics Transformer for Real-World Control at ScalearXiv GitHub
BC-ZBC-Z: Zero-Shot Task Generalization with Robotic Imitation LearningarXiv Website
CALVINCALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation TasksarXiv GitHub
BridgeData V2BridgeData V2: A Dataset for Robot Learning at ScalearXiv GitHub
LIBEROLIBERO: Benchmarking Knowledge Transfer for Lifelong Robot LearningarXiv GitHub
Isaac GymIsaac Gym: High Performance GPU-Based Physics Simulation For Robot LearningarXiv GitHub
RoboNetRoboNet: Large-Scale Multi-Robot LearningarXiv GitHub
Meta-WorldMeta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement LearningarXiv GitHub
RLBenchRLBench: The Robot Learning Benchmark & Learning EnvironmentarXiv GitHub

Autonomous Driving

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
WorldLensWorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldarXiv GitHub Website
ACT-BenchACT-Bench: Towards Action Controllable World Models for Autonomous DrivingarXiv GitHub
DrivingDojoDrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World ModelarXiv GitHub
DriveArenaDriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivingarXiv GitHub
NAVSIMNAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and BenchmarkingarXiv GitHub
CarDreamerCarDreamer: Open-Source Learning Platform for World Model based Autonomous DrivingarXiv GitHub
OpenDV-2KGenAD: Generalized Predictive Model for Autonomous DrivingarXiv GitHub
ZODZenseact Open Dataset: A Large-Scale and Diverse Multimodal Dataset for Autonomous DrivingarXiv GitHub
Occ3DOcc3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous DrivingarXiv GitHub
OpenOccupancyOpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy PerceptionarXiv GitHub
Argoverse 2Argoverse 2: Next Generation Datasets for Self-Driving Perception and ForecastingarXiv GitHub
KITTI-360KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3DarXiv GitHub
NuPlanNuPlan: A closed-loop ML-based planning benchmark for autonomous vehiclesarXiv GitHub
ONCEOne Million Scenes for Autonomous Driving: ONCE DatasetarXiv GitHub
WOMDLarge Scale Interactive Motion Forecasting for Autonomous Driving: The Waymo Open Motion DatasetarXiv GitHub
Lyft Level 5One Thousand and One Hours: Self-driving Motion Prediction DatasetarXiv GitHub
A2D2A2D2: Audi Autonomous Driving DatasetarXiv GitHub
WaymoScalability in Perception for Autonomous Driving: Waymo Open DatasetarXiv GitHub
ArgoverseArgoverse: 3D Tracking and Forecasting With Rich MapsarXiv GitHub
INTERACTIONINTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic MapsarXiv GitHub
SemanticKITTISemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesarXiv GitHub
nuScenesnuScenes: A Multimodal Dataset for Autonomous DrivingarXiv GitHub
BDD100KBDD100K: A Diverse Driving Dataset for Heterogeneous Multitask LearningarXiv GitHub
ApolloScapeThe ApolloScape Dataset for Autonomous DrivingarXiv GitHub
CARLACARLA: An Open Urban Driving SimulatorarXiv GitHub
KITTIAre we ready for autonomous driving? The KITTI vision benchmark suiteWebsite

Interactive Environments and Gaming

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
SCOPESCOPE: Simulating Cross-game Operations in Playable Environments for FPS World ModelsarXiv GitHub Website
WorldMarkWorldMark: A Unified Benchmark Suite for Interactive Video World ModelsarXiv Website
WildWorldWildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPGarXiv GitHub
Matrix-Game-MCMatrix-Game: Interactive World Foundation ModelarXiv GitHub
LOOPNAVToward Memory-Aided World Models: Benchmarking via Spatial ConsistencyarXiv GitHub
JARVIS-VLAJARVIS-VLA: Post-Training Large-Scale Visual Language Models to Play Visual Games with Keyboards and MousearXiv GitHub
GF-MinecraftGameFactory: Creating New Games with Generative Interactive VideosarXiv GitHub
The MatrixThe Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlarXiv Website
GameGen-XGameGen-X: Interactive Open-world Game Video GenerationarXiv GitHub
MarsMars: Situated Inductive Reasoning in an Open-World EnvironmentarXiv GitHub
CrafterBenchmarking the Spectrum of Agent CapabilitiesarXiv GitHub
CS-DeathmatchCounter-Strike Deathmatch with Large-Scale Behavioural CloningarXiv GitHub
MineRLThe MineRL 2020 Competition on Sample Efficient Reinforcement Learning using Human PriorsarXiv GitHub
NLEThe NetHack Learning EnvironmentarXiv GitHub
ProcgenLeveraging Procedural Generation to Benchmark Reinforcement LearningarXiv GitHub
DMCDeepMind Control SuitearXiv GitHub
ALEThe Arcade Learning Environment: An Evaluation Platform for General AgentsarXiv GitHub

3. Others

Survey

:timer_clock: In chronological order, from the latest to the earliest.

PaperLink
World Model for Robot Learning: A Comprehensive SurveyarXiv GitHub
OpenWorldLib: A Unified Codebase and Definition of Advanced World ModelsarXiv GitHub
Video Generation Models as World Models: Efficient Paradigms, Architectures and AlgorithmsarXiv
Learning to Model the World: A Survey of World Models in Artificial IntelligenceTechrXiv GitHub
Towards Generalist Embodied AI: A Survey on World Models for VLA AgentsTechrXiv
Video Generation Models in Robotics -- Applications, Research Challenges, Future DirectionsarXiv
Digital Twin AI: Opportunities and Challenges from Large Language Models to World ModelsarXiv
Progressive Robustness-Aware World Models in Autonomous Driving: A Review and OutlookTechrXiv
Simulating the Visual World with Artificial Intelligence: A RoadmaparXiv GitHub
A Step Toward World Models: A Survey on Robotic ManipulationarXiv
The Safety Challenge of World Models for Embodied AI Agents: A ReviewarXiv
A Comprehensive Survey on World Models for Embodied AIarXiv GitHub
3D and 4D World Modeling: A SurveyarXiv GitHub
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied LearningarXiv
A Survey: Learning Embodied Intelligence from Physical Simulators and World ModelsarXiv GitHub
From 2D to 3D Cognition: A Brief Survey of General World ModelsarXiv
A Survey on World Models Grounded in Acoustic Physical InformationarXiv GitHub
World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a ChildarXiv
The Role of World Models in Shaping Autonomous Driving: A Comprehensive SurveyarXiv GitHub
A Survey of World Models for Autonomous DrivingarXiv GitHub
Understanding World or Predicting Future? A Comprehensive Survey of World ModelsarXiv GitHub
Is Sora a World Simulator? A Comprehensive Survey on General World Models and BeyondarXiv GitHub
World Models for Autonomous Driving: An Initial SurveyarXiv
A Path Towards Autonomous Machine IntelligenceWebsite

GitHub Repo

This part overlaps somewhat with the Survey section.

RepoLink
WMFactory 0.5: One environment · One procedure · Eleven interactive world modelsGitHub
A minimalist repository for training video world models based on diffusion-forcingGitHub
Awesome Video World Models with AR DiffusionGitHub
Awesome From Video Generation to World ModelGitHub
A Curated List of Amazing Works in World Modeling, spanning applications in Embodied AI, Autonomous Driving, Natural Laguage Processing and AgentsGitHub
A curated list of papers for World Models for General Video Generation, Embodied AI, and Autonomous DrivingGitHub
World Models for Autonomous Driving (and Robotic) papersGitHub
Awesome 3D and 4D World ModelsGitHub
Simulating the Real World: Survey & ResourcesGitHub
A curated list of world models for autonomous drivingGitHub
A curated list of resources related to World Models for Autonomous Driving (WMAD)GitHub
A Survey: Learning Embodied Intelligence from Physical Simulators and World ModelsGitHub
A Comprehensive Survey on World Models for Embodied AIGitHub
A Survey on World Models Grounded in Acoustic Physical InformationGitHub
A Comprehensive Survey on General World Models and BeyondGitHub
Understanding World or Predicting Future? A Comprehensive Survey of World ModelsGitHub
From Masks to Worlds: A Hitchhiker's Guide to World ModelsGitHub

Workshop

VenueWorkshopLink
ECCV 2026How to Build Effective World Models for Embodied AIWebsite
CVPR 2026GigaBrain Challenge 2026: World Models & VLA for Embodied IntelligenceWebsite
CVPR 2026Video World ModelsWebsite
CVPR 2026World Models Meet Active Sensing and Closed-Loop PlanningWebsite
ICLR 2026Workshop on World Models: Understanding, Modelling and ScalingWebsite
NeurIPS 2025Workshop on Bridging Language, Agent, and World Models for Reasoning and PlanningWebsite
NeurIPS 2025Embodied World Models for Decision MakingWebsite
ICCV 2025Reliable and Interactable World Models: Geometry, Physics, Interactivity and Real-World GeneralizationWebsite
ICML 2025Building Physically Plausible World ModelsWebsite
CVPR 2025Benchmarking World ModelsWebsite
ICLR 2025World Models: Understanding, Modelling, and ScalingWebsite

Theory

:timer_clock: In chronological order, from the latest to the earliest.

DesignPaperLink
OrcaOrca: The World is in Your MindarXiv GitHub Website
LoopWMLooped World ModelsarXiv
OrbiSimOrbiSim: World Models as Differentiable Physics Engines for Embodied IntelligencearXiv Website
-Physically Native World Models: A Hamiltonian Perspective on Generative World ModelingarXiv
ZWMZero-shot World Models Are Developmentally Efficient LearnersarXiv GitHub
-Human Cognition in Machines: A Unified Perspective of World ModelsarXiv
EB-JEPAA Lightweight Library for Energy-Based Joint-Embedding Predictive ArchitecturesarXiv
-Research on World Models Is Not Merely Injecting World Knowledge into Specific TasksarXiv
-What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?arXiv
-Closing the Train-Test Gap in World Models for Gradient-Based PlanningarXiv GitHub
WWMWeb World ModelarXiv GitHub Website
DWMDexterous World ModelsarXiv [GitHub](https://github.com/s

Truncated — view the full README on GitHub.

Contributors

XiaoYu-1123

33 commits

AIROBOTAI

2 commits

zyc781125

2 commits

WayneJin0918

1 commits