HaobroLi/Awesome-ACWM

A curated, continuously updated reading list, paper blogs, and resources for Action Conditioned World Models

46

5 commits

updated Sep 3, 2026

See the code

README

Awesome Action-Conditioned World Models

Awesome GitHub stars GitHub issues Last commit

Action-conditioned video generation, game worlds, embodied simulation, and latent imagination.

A curated, community-driven reading list of visual world models that respond to actions,
roll out controllable futures, and support planning, policy learning, and evaluation.

Explore the list · Report an issue


About this list

This repository tracks action-conditioned visual world models (ACWMs) across three connected use cases:

🎮 Game World Models
Controllable game-like environments and long-horizon interactive rollouts.
🤖 Embodied Simulation
World models for robot policy planning, imitation, reinforcement learning, and evaluation.
🧠 Latent Control
Pixel, token, feature, and latent dynamics used for visual control.

The list includes models that generate future pixels/video as well as models that roll the world forward in a learned visual latent space. Last metadata check: 2026-08-14.

Contents

Scope

We use action-conditioned world model (ACWM) for a model that learns controlled visual dynamics:

future visual states · rewards · termination
= f(observation history, actions, context)

Here, future states may be pixels, video tokens, object/3D states, or learned visual features; actions may be observed controls or learned latent actions; and context may include language, goals, scene information, or embodiment details.

Included work should satisfy these criteria:

  1. Visual observations are modeled in pixel space, a renderable 3D/4D representation, discrete visual tokens, or a learned visual latent space.
  2. Actions or controls affect the transition model. Keyboard/mouse inputs, robot commands, trajectories, action chunks, camera controls, and learned latent actions all qualify.
  3. The learned transition is used for interactive rollout, simulation, planning, policy learning, evaluation, or control.

The sections are navigation aids rather than mutually exclusive theoretical categories. In particular, embodied papers are kept in one flat list and may carry several use tags.

Tag legend

Representation

  • Video — future RGB/RGB-D or multi-view observations.
  • Token — discrete visual tokens or autoregressive observation tokens.
  • Latent — recurrent or transformer latent states learned from visual observations.
  • Feature — dynamics in a pretrained visual feature space.
  • 3D/4D — occupancy, point, Gaussian, flow, or other renderable spatial states.
  • Unified — visual dynamics and action generation/decoding learned in one jointly trained model.

Downstream use

  • Game — game-playing or game-like interactive environments.
  • Interactive — controllable, multi-step visual rollouts.
  • Eval — policy evaluation, ranking, verification, or failure detection.
  • IL — imitation learning, behavior cloning, DAgger, or demonstration replay.
  • RL — reinforcement learning or reinforcement fine-tuning inside the model.
  • Plan — MPC, tree search, visual foresight, or goal-conditioned planning.
  • Policy — joint world/action learning or direct policy improvement.
  • Data — synthetic trajectories, observations, or demonstrations.

Surveys and perspectives

  • “From World Models to World Action Models: A Concise Tutorial for Robotics,” arXiv 2026.07. [Paper]
  • “Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends,” arXiv 2026.06. [Paper]
  • “World Model for Robot Learning: A Comprehensive Survey,” arXiv 2026.05. [Paper] [Project] [List]
  • “Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms,” arXiv 2026.03. [Paper]
  • “A Survey: Learning Embodied Intelligence from Physical Simulators and World Models,” arXiv 2025.07. [Paper] [List]

Game World Models

🎮 Game-like interactive environments. The learned model itself is a controllable game or game-like visual environment. Entries use the same table format as the embodied section and are ordered by first public release, newest first.

DateWorkRep.UsesLinks
2026.07ABot-World-0 — Infinite Interactive World Rollout on a Single Desktop GPUVideoGame InteractivePaper
2026.07AlayaWorld — Interactive Long-Horizon World ModelingVideoGame InteractivePaper
2026.07LingBot-World 2.0 (Infinity) — Infinite Worlds with Versatile InteractionsVideoGame InteractivePaper · Project · Code
2026.06ActWorld — From Explorable to Interactive World Model via Action-Aware MemoryVideoGame InteractivePaper
2026.06DreamX-World 1.0 — A General-Purpose Interactive World ModelVideoGame InteractivePaper
2026.05minWM — A Full-Stack Open-Source Framework for Real-Time Interactive Video World ModelsVideoGame InteractivePaper · Code
2026.05WorldCraft — From Camera Navigation to Object Manipulation in Interactive Video World ModelsVideoGame InteractivePaper · Project
2026.04Matrix-Game 3.0 — Real-Time and Streaming Interactive World Model with Long-Horizon MemoryVideoGame InteractivePaper · Project
2026.02LIVE — Long-horizon Interactive Video World ModelingVideoGame InteractivePaper · Project
2026.01LingBot-World — Advancing Open-source World ModelsVideoGame InteractivePaper · Project · Code
2025.12Yume-1.5 — A Text-Controlled Interactive World Generation ModelVideoGame InteractivePaper · Project · Code
2025.12WorldPlay — Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingVideoGame InteractivePaper
2025.12Astra — General Interactive World Model with Autoregressive DenoisingVideoGame InteractivePaper · Project · Code
2025.12RELIC — Interactive Video World Model with Long-Horizon MemoryVideoGame InteractivePaper · Project
2025.11Hunyuan-GameCraft-2 — Instruction-following Interactive Game World ModelVideoGame InteractivePaper · Project
2025.08Matrix-Game 2.0 — An Open-Source Real-Time and Streaming Interactive World ModelVideoGame InteractivePaper · Project
2025.08Genie 3 — A New Frontier for World ModelsVideoGame InteractiveBlog
2025.07Yume — An Interactive World Generation ModelVideoGame InteractivePaper · Project · Code
2025.06Matrix-Game — Interactive World Foundation ModelVideoGame InteractivePaper · Code
2025.06Hunyuan-GameCraft — High-dynamic Interactive Game Video Generation with Hybrid History ConditionVideoGame InteractivePaper · Project
2025.05VRAG — Learning World Models for Interactive Video GenerationVideoGame InteractivePaper
2025.05Vid2World — Crafting Video Diffusion Models to Interactive World ModelsVideoGame InteractivePaper · Project
2025.04MineWorld — A Real-Time and Open-Source Interactive World Model on MinecraftTokenGame InteractivePaper · Project · Code
2024.12Genie 2 — A Large-Scale Foundation World ModelVideoGame InteractiveBlog
2024.10OASIS — A Universe in a TransformerVideoGame InteractiveProject
2024.08GameNGen — Diffusion Models Are Real-Time Game EnginesVideoGame InteractivePaper · Project
2024.06Pandora — Towards General World Model with Natural Language Actions and Video StatesVideoGame InteractivePaper · Code
2024.05iVideoGPT — Interactive VideoGPTs are Scalable World ModelsTokenGame InteractivePaper · Code
2024.05DIAMOND — Diffusion for World Modeling: Visual Details Matter in AtariVideoGame InteractivePaper · Code
2024.02Genie — Generative Interactive EnvironmentsTokenGame InteractivePaper · Project

Embodied Simulation

🤖 Robot-facing world models. This is deliberately one flat list spanning video, latent, 3D/4D, and unified world models. The Uses column allows a paper to be simultaneously classified as evaluation, imitation learning, RL, planning, policy learning, and/or data generation.

DateWorkRep.UsesLinks
2026.08Hydra-0 — Action Flow for Generalist World Modeling and ControlVideoEval IL Plan Policy DataPaper · Project
2026.08WALL-SS — Action-conditioned world model for embodied simulationVideoEval IL RL Plan PolicyProject
2026.07CheckVLA — Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile ManipulationLatentEval PlanPaper
2026.06Recurrent Generative Replay — World Action Models Enable Continual Imitation Learning with Recurrent Generative ReplaysVideoIL DataPaper
2026.06WAM-RL — World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFTVideoRL PolicyPaper
2026.06PiL-World — A Chunk-Wise World Model for VLA Policy-in-the-Loop EvaluationVideoEvalPaper
2026.05GE-Sim 2.0 — A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic ManipulationVideoEval IL RL DataPaper · Project
2026.05Sword — Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-TrainingVideoRLPaper
2026.04X-WAM — Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising3D/4D UnifiedIL Plan PolicyPaper · Project
2026.04dWorldEval — Scalable Robotic Policy Evaluation via Discrete Diffusion World ModelTokenEvalPaper
2026.04Hi-WM — Human-in-the-World-Model for Scalable Robot Post-TrainingVideoIL RL PolicyPaper · Project
2026.04WM-DAgger — Enabling Efficient Data Aggregation for Imitation Learning with World ModelsVideoIL DataPaper · Code
2026.03DreamPlan — Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World ModelsVideoRL PlanPaper · Project
2026.03Kinema4D — Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation3D/4DEval DataPaper · Project · Code
2026.03Interactive World Simulator — Interactive World Simulator for Robot Policy Training and EvaluationVideoEval IL RLPaper · Project
2026.02GigaBrain-0.5M* — A VLA That Learns From World Model-Based Reinforcement LearningVideoRL PolicyPaper · Project
2026.02VLAW — Iterative Co-Improvement of Vision-Language-Action Policy and World ModelVideoIL Policy DataPaper · Project
2026.02RISE — Self-Improving Robot Policy with Compositional World ModelVideoIL RL PolicyPaper · Project
2026.02Say, Dream, and Act — Learning Video World Models for Instruction-Driven Robot ManipulationVideoPlan PolicyPaper
2026.02DreamDojo — A Generalist Robot World Model from Large-Scale Human VideosVideoIL Policy DataPaper · Project
2026.02World-VLA-Loop — Closed-Loop Learning of Video World Model and VLA PolicyVideoIL Policy DataPaper · Project
2026.02WoVR — World Models as Reliable Simulators for Post-Training VLA Policies with RLVideoRLPaper · Project
2026.02World-Gymnast — Training Robots with Reinforcement Learning in a World ModelVideoRLPaper · Project
2026.01lingbot-va — Causal World Modeling for Robot ControlVideoPlan PolicyPaper · Project · Code
2026.01PointWorld — Scaling 3D World Models for In-The-Wild Robotic Manipulation3D/4DPlan PolicyPaper · Project · Code · Model
2025.12Veo World Simulator — Evaluating Gemini Robotics Policies in a Veo World SimulatorVideoEvalPaper
2025.11Scalable Policy Evaluation — Scalable Policy Evaluation with Video World ModelsVideoEvalPaper
2025.10Ctrl-World — A Controllable Generative World Model for Robot ManipulationVideoPlan PolicyPaper · Project · Code
2025.10VLA-RFT — Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World SimulatorsVideoRLPaper · Project
2025.09World-Env — Leveraging World Model as a Virtual Environment for VLA Post-TrainingVideoRLPaper · Code
2025.09World4RL — Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic ManipulationVideoRLPaper
2025.08Genie Envisioner — A Unified World Foundation Platform for Robotic ManipulationVideoEval IL Plan PolicyPaper · Project
2025.06ParticleFormer — A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation3D/4DPlanPaper · Project
2025.06WorldVLA — Towards Autoregressive Action World ModelVideoPolicy PlanPaper · Code
2025.05LaDi-WM — A Latent Diffusion-based World Model for Predictive ManipulationLatentPlan PolicyPaper
2025.05FLARE — Robot Learning with Implicit World ModelingLatentPolicy DataPaper · Project · Code
2025.05WorldEval — World Model as Real-World Robot Policies EvaluatorVideoEvalPaper · Project
2025.05DreamGen — Unlocking Generalization in Robot Learning through Video World ModelsVideoIL DataPaper · Code
2025.04TesserAct — Learning 4D Embodied World Models3D/4DPlan PolicyPaper · Project
2025.04PIN-WM — Learning Physics-INformed World Models for Non-Prehensile Manipulation3D/4DPlanPaper
2025.04UWM — Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsVideo UnifiedIL Policy DataPaper · Project
2024.12Dream to Manipulate — Compositional World Models Empowering Robot Imitation Learning with ImaginationVideoIL Plan DataPaper · Project
2024.10EVA — An Embodied World Model for Future Video AnticipationVideoPlan PolicyPaper · Project
2024.06IRASim — A Fine-Grained World Model for Robot ManipulationVideoEval PlanPaper · Project · Code
2024.04RoboDreamer — Learning Compositional World Models for Robot ImaginationVideoPlan PolicyPaper · Project
2024.033D-VLA — A 3D Vision-Language-Action Generative World Model3D/4DPlan PolicyPaper
2024.03ManiGaussian — Dynamic Gaussian Splatting for Multi-task Robotic Manipulation3D/4DPlan PolicyPaper · Project
2023.10UniSim — Learning Interactive Real-World SimulatorsVideoIL RL Plan Policy DataPaper · Project

Latent Visual World Models and Control

🧠 Latent imagination and visual control. This section covers general visual control and model-based RL work whose action-conditioned dynamics live partly or entirely in learned latent space. Pixel-generative methods are included when they are foundational to this line.

DateWorkRep.UsesLinks
2025.06V-JEPA 2 / V-JEPA 2-AC — Self-Supervised Video Models Enable Understanding, Prediction and PlanningFeatureAction-conditioned planningPaper · Project · Code
2024.11DINO-WM — World Models on Pre-trained Visual Features enable Zero-shot PlanningFeatureZero-shot planningPaper · Project · Code
2024.10AVID — Adapting Video Diffusion Models to World ModelsVideoAction-conditioned diffusionPaper · Code
2024.06Delta-IRIS — Efficient World Models with Context-Aware TokenizationTokenVisual RLPaper · Code
2023.10TD-MPC2 — Scalable, Robust World Models for Continuous ControlLatentContinuous controlPaper · Code
2023.01DreamerV3 — Mastering Diverse Domains through World ModelsLatentGeneral RLPaper · Code
2022.09IRIS — Transformers are Sample-Efficient World ModelsTokenAtari controlPaper · Code
2022.06MWM — Masked World Models for Visual ControlLatentVisual RLPaper · Code
2022.06DayDreamer — World Models for Physical Robot LearningLatentReal-robot RLPaper · Code
2022.03TD-MPC — Temporal Difference Learning for Model Predictive ControlLatentMPC and controlPaper · Code
2020.10DreamerV2 — Mastering Atari with Discrete World ModelsLatentAtari controlPaper · Code
2019.12Dreamer — Dream to Control: Learning Behaviors by Latent ImaginationLatentContinuous controlPaper · Code
2019.11MuZero — Mastering Atari, Go, Chess and Shogi by Planning with a Learned ModelLatentTree-search planningPaper
2019.03SimPLe — Model-Based Reinforcement Learning for AtariVideoAtari controlPaper
2018.11PlaNet — Learning Latent Dynamics for Planning from PixelsLatentMPC and controlPaper · Code
2018.03World Models — World ModelsLatentControl by imaginationPaper · Project
2016.10Deep Visual Foresight — Deep Visual Foresight for Planning Robot MotionVideoVisual MPCPaper
2016.05Unsupervised Physical Interaction — Unsupervised Learning for Physical Interaction through Video PredictionVideoRobot controlPaper
2015.07Action-Conditional Video Prediction — Action-Conditional Video Prediction using Deep Networks in Atari GamesVideoAtari predictionPaper

Benchmarks

DateBenchmarkDomainLinks
2026.08WorldSimProbe — Action-following evaluation for action-conditioned world modelsAction-conditioned world modelsProject
2026.06WorldRoamBench — Long-horizon stability of interactive world modelsInteractivePaper
2026.06RoboTrustBench — Trustworthiness of video world models for robotic manipulationEmbodiedPaper · Project
2026.05MiraBench — Action-conditioned reliability in robotic world modelsEmbodiedPaper
2026.05WBench — Multi-turn interactive video world model evaluationInteractivePaper · Project
2026.05ACWM-Phys — Generalized physical interaction in action-conditioned video world modelsEmbodiedPaper · Project · Code
2026.05iWorld-Bench — Interactive world models with a unified action generation frameworkInteractivePaper
2026.04WorldMark — Unified benchmark suite for interactive video world modelsInteractivePaper
2026.04RoboWM-Bench — World models in robotic manipulationEmbodiedPaper · Project
2026.01Wow, wo, val! — Embodied world model evaluation Turing testEmbodiedPaper

Datasets

  • DROID: Large-scale, in-the-wild robot manipulation trajectories. [Project] [Code]
  • Open X-Embodiment: Cross-embodiment robot trajectories and RT-X models. [Project]
  • BridgeData V2: Broad manipulation data collected across many environments and tasks. [Project]
  • RoboNet: Multi-robot video data for learning visual dynamics. [Project]
  • BAIR Robot Pushing: Action-conditioned pushing videos used by early visual-foresight work. [Dataset]

License

This list is released under CC0 1.0.

HaobroLi/Awesome-ACWM

A curated, continuously updated reading list, paper blogs, and resources for Action Conditioned World Models

46

5 commits

updated Sep 3, 2026

See the code

README

Awesome Action-Conditioned World Models

Awesome GitHub stars GitHub issues Last commit

Action-conditioned video generation, game worlds, embodied simulation, and latent imagination.

A curated, community-driven reading list of visual world models that respond to actions,
roll out controllable futures, and support planning, policy learning, and evaluation.

Explore the list · Report an issue


About this list

This repository tracks action-conditioned visual world models (ACWMs) across three connected use cases:

🎮 Game World Models
Controllable game-like environments and long-horizon interactive rollouts.
🤖 Embodied Simulation
World models for robot policy planning, imitation, reinforcement learning, and evaluation.
🧠 Latent Control
Pixel, token, feature, and latent dynamics used for visual control.

The list includes models that generate future pixels/video as well as models that roll the world forward in a learned visual latent space. Last metadata check: 2026-08-14.

Contents

Scope

We use action-conditioned world model (ACWM) for a model that learns controlled visual dynamics:

future visual states · rewards · termination
= f(observation history, actions, context)

Here, future states may be pixels, video tokens, object/3D states, or learned visual features; actions may be observed controls or learned latent actions; and context may include language, goals, scene information, or embodiment details.

Included work should satisfy these criteria:

  1. Visual observations are modeled in pixel space, a renderable 3D/4D representation, discrete visual tokens, or a learned visual latent space.
  2. Actions or controls affect the transition model. Keyboard/mouse inputs, robot commands, trajectories, action chunks, camera controls, and learned latent actions all qualify.
  3. The learned transition is used for interactive rollout, simulation, planning, policy learning, evaluation, or control.

The sections are navigation aids rather than mutually exclusive theoretical categories. In particular, embodied papers are kept in one flat list and may carry several use tags.

Tag legend

Representation

  • Video — future RGB/RGB-D or multi-view observations.
  • Token — discrete visual tokens or autoregressive observation tokens.
  • Latent — recurrent or transformer latent states learned from visual observations.
  • Feature — dynamics in a pretrained visual feature space.
  • 3D/4D — occupancy, point, Gaussian, flow, or other renderable spatial states.
  • Unified — visual dynamics and action generation/decoding learned in one jointly trained model.

Downstream use

  • Game — game-playing or game-like interactive environments.
  • Interactive — controllable, multi-step visual rollouts.
  • Eval — policy evaluation, ranking, verification, or failure detection.
  • IL — imitation learning, behavior cloning, DAgger, or demonstration replay.
  • RL — reinforcement learning or reinforcement fine-tuning inside the model.
  • Plan — MPC, tree search, visual foresight, or goal-conditioned planning.
  • Policy — joint world/action learning or direct policy improvement.
  • Data — synthetic trajectories, observations, or demonstrations.

Surveys and perspectives

  • “From World Models to World Action Models: A Concise Tutorial for Robotics,” arXiv 2026.07. [Paper]
  • “Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends,” arXiv 2026.06. [Paper]
  • “World Model for Robot Learning: A Comprehensive Survey,” arXiv 2026.05. [Paper] [Project] [List]
  • “Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms,” arXiv 2026.03. [Paper]
  • “A Survey: Learning Embodied Intelligence from Physical Simulators and World Models,” arXiv 2025.07. [Paper] [List]

Game World Models

🎮 Game-like interactive environments. The learned model itself is a controllable game or game-like visual environment. Entries use the same table format as the embodied section and are ordered by first public release, newest first.

DateWorkRep.UsesLinks
2026.07ABot-World-0 — Infinite Interactive World Rollout on a Single Desktop GPUVideoGame InteractivePaper
2026.07AlayaWorld — Interactive Long-Horizon World ModelingVideoGame InteractivePaper
2026.07LingBot-World 2.0 (Infinity) — Infinite Worlds with Versatile InteractionsVideoGame InteractivePaper · Project · Code
2026.06ActWorld — From Explorable to Interactive World Model via Action-Aware MemoryVideoGame InteractivePaper
2026.06DreamX-World 1.0 — A General-Purpose Interactive World ModelVideoGame InteractivePaper
2026.05minWM — A Full-Stack Open-Source Framework for Real-Time Interactive Video World ModelsVideoGame InteractivePaper · Code
2026.05WorldCraft — From Camera Navigation to Object Manipulation in Interactive Video World ModelsVideoGame InteractivePaper · Project
2026.04Matrix-Game 3.0 — Real-Time and Streaming Interactive World Model with Long-Horizon MemoryVideoGame InteractivePaper · Project
2026.02LIVE — Long-horizon Interactive Video World ModelingVideoGame InteractivePaper · Project
2026.01LingBot-World — Advancing Open-source World ModelsVideoGame InteractivePaper · Project · Code
2025.12Yume-1.5 — A Text-Controlled Interactive World Generation ModelVideoGame InteractivePaper · Project · Code
2025.12WorldPlay — Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingVideoGame InteractivePaper
2025.12Astra — General Interactive World Model with Autoregressive DenoisingVideoGame InteractivePaper · Project · Code
2025.12RELIC — Interactive Video World Model with Long-Horizon MemoryVideoGame InteractivePaper · Project
2025.11Hunyuan-GameCraft-2 — Instruction-following Interactive Game World ModelVideoGame InteractivePaper · Project
2025.08Matrix-Game 2.0 — An Open-Source Real-Time and Streaming Interactive World ModelVideoGame InteractivePaper · Project
2025.08Genie 3 — A New Frontier for World ModelsVideoGame InteractiveBlog
2025.07Yume — An Interactive World Generation ModelVideoGame InteractivePaper · Project · Code
2025.06Matrix-Game — Interactive World Foundation ModelVideoGame InteractivePaper · Code
2025.06Hunyuan-GameCraft — High-dynamic Interactive Game Video Generation with Hybrid History ConditionVideoGame InteractivePaper · Project
2025.05VRAG — Learning World Models for Interactive Video GenerationVideoGame InteractivePaper
2025.05Vid2World — Crafting Video Diffusion Models to Interactive World ModelsVideoGame InteractivePaper · Project
2025.04MineWorld — A Real-Time and Open-Source Interactive World Model on MinecraftTokenGame InteractivePaper · Project · Code
2024.12Genie 2 — A Large-Scale Foundation World ModelVideoGame InteractiveBlog
2024.10OASIS — A Universe in a TransformerVideoGame InteractiveProject
2024.08GameNGen — Diffusion Models Are Real-Time Game EnginesVideoGame InteractivePaper · Project
2024.06Pandora — Towards General World Model with Natural Language Actions and Video StatesVideoGame InteractivePaper · Code
2024.05iVideoGPT — Interactive VideoGPTs are Scalable World ModelsTokenGame InteractivePaper · Code
2024.05DIAMOND — Diffusion for World Modeling: Visual Details Matter in AtariVideoGame InteractivePaper · Code
2024.02Genie — Generative Interactive EnvironmentsTokenGame InteractivePaper · Project

Embodied Simulation

🤖 Robot-facing world models. This is deliberately one flat list spanning video, latent, 3D/4D, and unified world models. The Uses column allows a paper to be simultaneously classified as evaluation, imitation learning, RL, planning, policy learning, and/or data generation.

DateWorkRep.UsesLinks
2026.08Hydra-0 — Action Flow for Generalist World Modeling and ControlVideoEval IL Plan Policy DataPaper · Project
2026.08WALL-SS — Action-conditioned world model for embodied simulationVideoEval IL RL Plan PolicyProject
2026.07CheckVLA — Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile ManipulationLatentEval PlanPaper
2026.06Recurrent Generative Replay — World Action Models Enable Continual Imitation Learning with Recurrent Generative ReplaysVideoIL DataPaper
2026.06WAM-RL — World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFTVideoRL PolicyPaper
2026.06PiL-World — A Chunk-Wise World Model for VLA Policy-in-the-Loop EvaluationVideoEvalPaper
2026.05GE-Sim 2.0 — A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic ManipulationVideoEval IL RL DataPaper · Project
2026.05Sword — Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-TrainingVideoRLPaper
2026.04X-WAM — Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising3D/4D UnifiedIL Plan PolicyPaper · Project
2026.04dWorldEval — Scalable Robotic Policy Evaluation via Discrete Diffusion World ModelTokenEvalPaper
2026.04Hi-WM — Human-in-the-World-Model for Scalable Robot Post-TrainingVideoIL RL PolicyPaper · Project
2026.04WM-DAgger — Enabling Efficient Data Aggregation for Imitation Learning with World ModelsVideoIL DataPaper · Code
2026.03DreamPlan — Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World ModelsVideoRL PlanPaper · Project
2026.03Kinema4D — Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation3D/4DEval DataPaper · Project · Code
2026.03Interactive World Simulator — Interactive World Simulator for Robot Policy Training and EvaluationVideoEval IL RLPaper · Project
2026.02GigaBrain-0.5M* — A VLA That Learns From World Model-Based Reinforcement LearningVideoRL PolicyPaper · Project
2026.02VLAW — Iterative Co-Improvement of Vision-Language-Action Policy and World ModelVideoIL Policy DataPaper · Project
2026.02RISE — Self-Improving Robot Policy with Compositional World ModelVideoIL RL PolicyPaper · Project
2026.02Say, Dream, and Act — Learning Video World Models for Instruction-Driven Robot ManipulationVideoPlan PolicyPaper
2026.02DreamDojo — A Generalist Robot World Model from Large-Scale Human VideosVideoIL Policy DataPaper · Project
2026.02World-VLA-Loop — Closed-Loop Learning of Video World Model and VLA PolicyVideoIL Policy DataPaper · Project
2026.02WoVR — World Models as Reliable Simulators for Post-Training VLA Policies with RLVideoRLPaper · Project
2026.02World-Gymnast — Training Robots with Reinforcement Learning in a World ModelVideoRLPaper · Project
2026.01lingbot-va — Causal World Modeling for Robot ControlVideoPlan PolicyPaper · Project · Code
2026.01PointWorld — Scaling 3D World Models for In-The-Wild Robotic Manipulation3D/4DPlan PolicyPaper · Project · Code · Model
2025.12Veo World Simulator — Evaluating Gemini Robotics Policies in a Veo World SimulatorVideoEvalPaper
2025.11Scalable Policy Evaluation — Scalable Policy Evaluation with Video World ModelsVideoEvalPaper
2025.10Ctrl-World — A Controllable Generative World Model for Robot ManipulationVideoPlan PolicyPaper · Project · Code
2025.10VLA-RFT — Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World SimulatorsVideoRLPaper · Project
2025.09World-Env — Leveraging World Model as a Virtual Environment for VLA Post-TrainingVideoRLPaper · Code
2025.09World4RL — Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic ManipulationVideoRLPaper
2025.08Genie Envisioner — A Unified World Foundation Platform for Robotic ManipulationVideoEval IL Plan PolicyPaper · Project
2025.06ParticleFormer — A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation3D/4DPlanPaper · Project
2025.06WorldVLA — Towards Autoregressive Action World ModelVideoPolicy PlanPaper · Code
2025.05LaDi-WM — A Latent Diffusion-based World Model for Predictive ManipulationLatentPlan PolicyPaper
2025.05FLARE — Robot Learning with Implicit World ModelingLatentPolicy DataPaper · Project · Code
2025.05WorldEval — World Model as Real-World Robot Policies EvaluatorVideoEvalPaper · Project
2025.05DreamGen — Unlocking Generalization in Robot Learning through Video World ModelsVideoIL DataPaper · Code
2025.04TesserAct — Learning 4D Embodied World Models3D/4DPlan PolicyPaper · Project
2025.04PIN-WM — Learning Physics-INformed World Models for Non-Prehensile Manipulation3D/4DPlanPaper
2025.04UWM — Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsVideo UnifiedIL Policy DataPaper · Project
2024.12Dream to Manipulate — Compositional World Models Empowering Robot Imitation Learning with ImaginationVideoIL Plan DataPaper · Project
2024.10EVA — An Embodied World Model for Future Video AnticipationVideoPlan PolicyPaper · Project
2024.06IRASim — A Fine-Grained World Model for Robot ManipulationVideoEval PlanPaper · Project · Code
2024.04RoboDreamer — Learning Compositional World Models for Robot ImaginationVideoPlan PolicyPaper · Project
2024.033D-VLA — A 3D Vision-Language-Action Generative World Model3D/4DPlan PolicyPaper
2024.03ManiGaussian — Dynamic Gaussian Splatting for Multi-task Robotic Manipulation3D/4DPlan PolicyPaper · Project
2023.10UniSim — Learning Interactive Real-World SimulatorsVideoIL RL Plan Policy DataPaper · Project

Latent Visual World Models and Control

🧠 Latent imagination and visual control. This section covers general visual control and model-based RL work whose action-conditioned dynamics live partly or entirely in learned latent space. Pixel-generative methods are included when they are foundational to this line.

DateWorkRep.UsesLinks
2025.06V-JEPA 2 / V-JEPA 2-AC — Self-Supervised Video Models Enable Understanding, Prediction and PlanningFeatureAction-conditioned planningPaper · Project · Code
2024.11DINO-WM — World Models on Pre-trained Visual Features enable Zero-shot PlanningFeatureZero-shot planningPaper · Project · Code
2024.10AVID — Adapting Video Diffusion Models to World ModelsVideoAction-conditioned diffusionPaper · Code
2024.06Delta-IRIS — Efficient World Models with Context-Aware TokenizationTokenVisual RLPaper · Code
2023.10TD-MPC2 — Scalable, Robust World Models for Continuous ControlLatentContinuous controlPaper · Code
2023.01DreamerV3 — Mastering Diverse Domains through World ModelsLatentGeneral RLPaper · Code
2022.09IRIS — Transformers are Sample-Efficient World ModelsTokenAtari controlPaper · Code
2022.06MWM — Masked World Models for Visual ControlLatentVisual RLPaper · Code
2022.06DayDreamer — World Models for Physical Robot LearningLatentReal-robot RLPaper · Code
2022.03TD-MPC — Temporal Difference Learning for Model Predictive ControlLatentMPC and controlPaper · Code
2020.10DreamerV2 — Mastering Atari with Discrete World ModelsLatentAtari controlPaper · Code
2019.12Dreamer — Dream to Control: Learning Behaviors by Latent ImaginationLatentContinuous controlPaper · Code
2019.11MuZero — Mastering Atari, Go, Chess and Shogi by Planning with a Learned ModelLatentTree-search planningPaper
2019.03SimPLe — Model-Based Reinforcement Learning for AtariVideoAtari controlPaper
2018.11PlaNet — Learning Latent Dynamics for Planning from PixelsLatentMPC and controlPaper · Code
2018.03World Models — World ModelsLatentControl by imaginationPaper · Project
2016.10Deep Visual Foresight — Deep Visual Foresight for Planning Robot MotionVideoVisual MPCPaper
2016.05Unsupervised Physical Interaction — Unsupervised Learning for Physical Interaction through Video PredictionVideoRobot controlPaper
2015.07Action-Conditional Video Prediction — Action-Conditional Video Prediction using Deep Networks in Atari GamesVideoAtari predictionPaper

Benchmarks

DateBenchmarkDomainLinks
2026.08WorldSimProbe — Action-following evaluation for action-conditioned world modelsAction-conditioned world modelsProject
2026.06WorldRoamBench — Long-horizon stability of interactive world modelsInteractivePaper
2026.06RoboTrustBench — Trustworthiness of video world models for robotic manipulationEmbodiedPaper · Project
2026.05MiraBench — Action-conditioned reliability in robotic world modelsEmbodiedPaper
2026.05WBench — Multi-turn interactive video world model evaluationInteractivePaper · Project
2026.05ACWM-Phys — Generalized physical interaction in action-conditioned video world modelsEmbodiedPaper · Project · Code
2026.05iWorld-Bench — Interactive world models with a unified action generation frameworkInteractivePaper
2026.04WorldMark — Unified benchmark suite for interactive video world modelsInteractivePaper
2026.04RoboWM-Bench — World models in robotic manipulationEmbodiedPaper · Project
2026.01Wow, wo, val! — Embodied world model evaluation Turing testEmbodiedPaper

Datasets

  • DROID: Large-scale, in-the-wild robot manipulation trajectories. [Project] [Code]
  • Open X-Embodiment: Cross-embodiment robot trajectories and RT-X models. [Project]
  • BridgeData V2: Broad manipulation data collected across many environments and tasks. [Project]
  • RoboNet: Multi-robot video data for learning visual dynamics. [Project]
  • BAIR Robot Pushing: Action-conditioned pushing videos used by early visual-foresight work. [Dataset]

License

This list is released under CC0 1.0.