A curated, continuously updated reading list, paper blogs, and resources for Action Conditioned World Models
46
5 commits
updated Sep 3, 2026
Action-conditioned video generation, game worlds, embodied simulation, and latent imagination.
A curated, community-driven reading list of visual world models that respond to actions,
roll out controllable futures, and support planning, policy learning, and evaluation.
This repository tracks action-conditioned visual world models (ACWMs) across three connected use cases:
| 🎮 Game World Models Controllable game-like environments and long-horizon interactive rollouts. | 🤖 Embodied Simulation World models for robot policy planning, imitation, reinforcement learning, and evaluation. | 🧠 Latent Control Pixel, token, feature, and latent dynamics used for visual control. |
The list includes models that generate future pixels/video as well as models that roll the world forward in a learned visual latent space. Last metadata check: 2026-08-14.
We use action-conditioned world model (ACWM) for a model that learns controlled visual dynamics:
future visual states · rewards · termination
= f(observation history, actions, context)
Here, future states may be pixels, video tokens, object/3D states, or learned visual features; actions may be observed controls or learned latent actions; and context may include language, goals, scene information, or embodiment details.
Included work should satisfy these criteria:
The sections are navigation aids rather than mutually exclusive theoretical categories. In particular, embodied papers are kept in one flat list and may carry several use tags.
Video — future RGB/RGB-D or multi-view observations.Token — discrete visual tokens or autoregressive observation tokens.Latent — recurrent or transformer latent states learned from visual observations.Feature — dynamics in a pretrained visual feature space.3D/4D — occupancy, point, Gaussian, flow, or other renderable spatial states.Unified — visual dynamics and action generation/decoding learned in one jointly trained model.Game — game-playing or game-like interactive environments.Interactive — controllable, multi-step visual rollouts.Eval — policy evaluation, ranking, verification, or failure detection.IL — imitation learning, behavior cloning, DAgger, or demonstration replay.RL — reinforcement learning or reinforcement fine-tuning inside the model.Plan — MPC, tree search, visual foresight, or goal-conditioned planning.Policy — joint world/action learning or direct policy improvement.Data — synthetic trajectories, observations, or demonstrations.arXiv 2026.07. [Paper]arXiv 2026.06. [Paper]arXiv 2026.05. [Paper] [Project] [List]arXiv 2026.03. [Paper]arXiv 2025.07. [Paper] [List]🎮 Game-like interactive environments. The learned model itself is a controllable game or game-like visual environment. Entries use the same table format as the embodied section and are ordered by first public release, newest first.
| Date | Work | Rep. | Uses | Links |
|---|---|---|---|---|
| 2026.07 | ABot-World-0 — Infinite Interactive World Rollout on a Single Desktop GPU | Video | Game Interactive | Paper |
| 2026.07 | AlayaWorld — Interactive Long-Horizon World Modeling | Video | Game Interactive | Paper |
| 2026.07 | LingBot-World 2.0 (Infinity) — Infinite Worlds with Versatile Interactions | Video | Game Interactive | Paper · Project · Code |
| 2026.06 | ActWorld — From Explorable to Interactive World Model via Action-Aware Memory | Video | Game Interactive | Paper |
| 2026.06 | DreamX-World 1.0 — A General-Purpose Interactive World Model | Video | Game Interactive | Paper |
| 2026.05 | minWM — A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models | Video | Game Interactive | Paper · Code |
| 2026.05 | WorldCraft — From Camera Navigation to Object Manipulation in Interactive Video World Models | Video | Game Interactive | Paper · Project |
| 2026.04 | Matrix-Game 3.0 — Real-Time and Streaming Interactive World Model with Long-Horizon Memory | Video | Game Interactive | Paper · Project |
| 2026.02 | LIVE — Long-horizon Interactive Video World Modeling | Video | Game Interactive | Paper · Project |
| 2026.01 | LingBot-World — Advancing Open-source World Models | Video | Game Interactive | Paper · Project · Code |
| 2025.12 | Yume-1.5 — A Text-Controlled Interactive World Generation Model | Video | Game Interactive | Paper · Project · Code |
| 2025.12 | WorldPlay — Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling | Video | Game Interactive | Paper |
| 2025.12 | Astra — General Interactive World Model with Autoregressive Denoising | Video | Game Interactive | Paper · Project · Code |
| 2025.12 | RELIC — Interactive Video World Model with Long-Horizon Memory | Video | Game Interactive | Paper · Project |
| 2025.11 | Hunyuan-GameCraft-2 — Instruction-following Interactive Game World Model | Video | Game Interactive | Paper · Project |
| 2025.08 | Matrix-Game 2.0 — An Open-Source Real-Time and Streaming Interactive World Model | Video | Game Interactive | Paper · Project |
| 2025.08 | Genie 3 — A New Frontier for World Models | Video | Game Interactive | Blog |
| 2025.07 | Yume — An Interactive World Generation Model | Video | Game Interactive | Paper · Project · Code |
| 2025.06 | Matrix-Game — Interactive World Foundation Model | Video | Game Interactive | Paper · Code |
| 2025.06 | Hunyuan-GameCraft — High-dynamic Interactive Game Video Generation with Hybrid History Condition | Video | Game Interactive | Paper · Project |
| 2025.05 | VRAG — Learning World Models for Interactive Video Generation | Video | Game Interactive | Paper |
| 2025.05 | Vid2World — Crafting Video Diffusion Models to Interactive World Models | Video | Game Interactive | Paper · Project |
| 2025.04 | MineWorld — A Real-Time and Open-Source Interactive World Model on Minecraft | Token | Game Interactive | Paper · Project · Code |
| 2024.12 | Genie 2 — A Large-Scale Foundation World Model | Video | Game Interactive | Blog |
| 2024.10 | OASIS — A Universe in a Transformer | Video | Game Interactive | Project |
| 2024.08 | GameNGen — Diffusion Models Are Real-Time Game Engines | Video | Game Interactive | Paper · Project |
| 2024.06 | Pandora — Towards General World Model with Natural Language Actions and Video States | Video | Game Interactive | Paper · Code |
| 2024.05 | iVideoGPT — Interactive VideoGPTs are Scalable World Models | Token | Game Interactive | Paper · Code |
| 2024.05 | DIAMOND — Diffusion for World Modeling: Visual Details Matter in Atari | Video | Game Interactive | Paper · Code |
| 2024.02 | Genie — Generative Interactive Environments | Token | Game Interactive | Paper · Project |
🤖 Robot-facing world models. This is deliberately one flat list spanning video, latent, 3D/4D, and unified world models. The Uses column allows a paper to be simultaneously classified as evaluation, imitation learning, RL, planning, policy learning, and/or data generation.
| Date | Work | Rep. | Uses | Links |
|---|---|---|---|---|
| 2026.08 | Hydra-0 — Action Flow for Generalist World Modeling and Control | Video | Eval IL Plan Policy Data | Paper · Project |
| 2026.08 | WALL-SS — Action-conditioned world model for embodied simulation | Video | Eval IL RL Plan Policy | Project |
| 2026.07 | CheckVLA — Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation | Latent | Eval Plan | Paper |
| 2026.06 | Recurrent Generative Replay — World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays | Video | IL Data | Paper |
| 2026.06 | WAM-RL — World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT | Video | RL Policy | Paper |
| 2026.06 | PiL-World — A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation | Video | Eval | Paper |
| 2026.05 | GE-Sim 2.0 — A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation | Video | Eval IL RL Data | Paper · Project |
| 2026.05 | Sword — Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training | Video | RL | Paper |
| 2026.04 | X-WAM — Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising | 3D/4D Unified | IL Plan Policy | Paper · Project |
| 2026.04 | dWorldEval — Scalable Robotic Policy Evaluation via Discrete Diffusion World Model | Token | Eval | Paper |
| 2026.04 | Hi-WM — Human-in-the-World-Model for Scalable Robot Post-Training | Video | IL RL Policy | Paper · Project |
| 2026.04 | WM-DAgger — Enabling Efficient Data Aggregation for Imitation Learning with World Models | Video | IL Data | Paper · Code |
| 2026.03 | DreamPlan — Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models | Video | RL Plan | Paper · Project |
| 2026.03 | Kinema4D — Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation | 3D/4D | Eval Data | Paper · Project · Code |
| 2026.03 | Interactive World Simulator — Interactive World Simulator for Robot Policy Training and Evaluation | Video | Eval IL RL | Paper · Project |
| 2026.02 | GigaBrain-0.5M* — A VLA That Learns From World Model-Based Reinforcement Learning | Video | RL Policy | Paper · Project |
| 2026.02 | VLAW — Iterative Co-Improvement of Vision-Language-Action Policy and World Model | Video | IL Policy Data | Paper · Project |
| 2026.02 | RISE — Self-Improving Robot Policy with Compositional World Model | Video | IL RL Policy | Paper · Project |
| 2026.02 | Say, Dream, and Act — Learning Video World Models for Instruction-Driven Robot Manipulation | Video | Plan Policy | Paper |
| 2026.02 | DreamDojo — A Generalist Robot World Model from Large-Scale Human Videos | Video | IL Policy Data | Paper · Project |
| 2026.02 | World-VLA-Loop — Closed-Loop Learning of Video World Model and VLA Policy | Video | IL Policy Data | Paper · Project |
| 2026.02 | WoVR — World Models as Reliable Simulators for Post-Training VLA Policies with RL | Video | RL | Paper · Project |
| 2026.02 | World-Gymnast — Training Robots with Reinforcement Learning in a World Model | Video | RL | Paper · Project |
| 2026.01 | lingbot-va — Causal World Modeling for Robot Control | Video | Plan Policy | Paper · Project · Code |
| 2026.01 | PointWorld — Scaling 3D World Models for In-The-Wild Robotic Manipulation | 3D/4D | Plan Policy | Paper · Project · Code · Model |
| 2025.12 | Veo World Simulator — Evaluating Gemini Robotics Policies in a Veo World Simulator | Video | Eval | Paper |
| 2025.11 | Scalable Policy Evaluation — Scalable Policy Evaluation with Video World Models | Video | Eval | Paper |
| 2025.10 | Ctrl-World — A Controllable Generative World Model for Robot Manipulation | Video | Plan Policy | Paper · Project · Code |
| 2025.10 | VLA-RFT — Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators | Video | RL | Paper · Project |
| 2025.09 | World-Env — Leveraging World Model as a Virtual Environment for VLA Post-Training | Video | RL | Paper · Code |
| 2025.09 | World4RL — Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation | Video | RL | Paper |
| 2025.08 | Genie Envisioner — A Unified World Foundation Platform for Robotic Manipulation | Video | Eval IL Plan Policy | Paper · Project |
| 2025.06 | ParticleFormer — A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation | 3D/4D | Plan | Paper · Project |
| 2025.06 | WorldVLA — Towards Autoregressive Action World Model | Video | Policy Plan | Paper · Code |
| 2025.05 | LaDi-WM — A Latent Diffusion-based World Model for Predictive Manipulation | Latent | Plan Policy | Paper |
| 2025.05 | FLARE — Robot Learning with Implicit World Modeling | Latent | Policy Data | Paper · Project · Code |
| 2025.05 | WorldEval — World Model as Real-World Robot Policies Evaluator | Video | Eval | Paper · Project |
| 2025.05 | DreamGen — Unlocking Generalization in Robot Learning through Video World Models | Video | IL Data | Paper · Code |
| 2025.04 | TesserAct — Learning 4D Embodied World Models | 3D/4D | Plan Policy | Paper · Project |
| 2025.04 | PIN-WM — Learning Physics-INformed World Models for Non-Prehensile Manipulation | 3D/4D | Plan | Paper |
| 2025.04 | UWM — Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets | Video Unified | IL Policy Data | Paper · Project |
| 2024.12 | Dream to Manipulate — Compositional World Models Empowering Robot Imitation Learning with Imagination | Video | IL Plan Data | Paper · Project |
| 2024.10 | EVA — An Embodied World Model for Future Video Anticipation | Video | Plan Policy | Paper · Project |
| 2024.06 | IRASim — A Fine-Grained World Model for Robot Manipulation | Video | Eval Plan | Paper · Project · Code |
| 2024.04 | RoboDreamer — Learning Compositional World Models for Robot Imagination | Video | Plan Policy | Paper · Project |
| 2024.03 | 3D-VLA — A 3D Vision-Language-Action Generative World Model | 3D/4D | Plan Policy | Paper |
| 2024.03 | ManiGaussian — Dynamic Gaussian Splatting for Multi-task Robotic Manipulation | 3D/4D | Plan Policy | Paper · Project |
| 2023.10 | UniSim — Learning Interactive Real-World Simulators | Video | IL RL Plan Policy Data | Paper · Project |
🧠 Latent imagination and visual control. This section covers general visual control and model-based RL work whose action-conditioned dynamics live partly or entirely in learned latent space. Pixel-generative methods are included when they are foundational to this line.
| Date | Work | Rep. | Uses | Links |
|---|---|---|---|---|
| 2025.06 | V-JEPA 2 / V-JEPA 2-AC — Self-Supervised Video Models Enable Understanding, Prediction and Planning | Feature | Action-conditioned planning | Paper · Project · Code |
| 2024.11 | DINO-WM — World Models on Pre-trained Visual Features enable Zero-shot Planning | Feature | Zero-shot planning | Paper · Project · Code |
| 2024.10 | AVID — Adapting Video Diffusion Models to World Models | Video | Action-conditioned diffusion | Paper · Code |
| 2024.06 | Delta-IRIS — Efficient World Models with Context-Aware Tokenization | Token | Visual RL | Paper · Code |
| 2023.10 | TD-MPC2 — Scalable, Robust World Models for Continuous Control | Latent | Continuous control | Paper · Code |
| 2023.01 | DreamerV3 — Mastering Diverse Domains through World Models | Latent | General RL | Paper · Code |
| 2022.09 | IRIS — Transformers are Sample-Efficient World Models | Token | Atari control | Paper · Code |
| 2022.06 | MWM — Masked World Models for Visual Control | Latent | Visual RL | Paper · Code |
| 2022.06 | DayDreamer — World Models for Physical Robot Learning | Latent | Real-robot RL | Paper · Code |
| 2022.03 | TD-MPC — Temporal Difference Learning for Model Predictive Control | Latent | MPC and control | Paper · Code |
| 2020.10 | DreamerV2 — Mastering Atari with Discrete World Models | Latent | Atari control | Paper · Code |
| 2019.12 | Dreamer — Dream to Control: Learning Behaviors by Latent Imagination | Latent | Continuous control | Paper · Code |
| 2019.11 | MuZero — Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model | Latent | Tree-search planning | Paper |
| 2019.03 | SimPLe — Model-Based Reinforcement Learning for Atari | Video | Atari control | Paper |
| 2018.11 | PlaNet — Learning Latent Dynamics for Planning from Pixels | Latent | MPC and control | Paper · Code |
| 2018.03 | World Models — World Models | Latent | Control by imagination | Paper · Project |
| 2016.10 | Deep Visual Foresight — Deep Visual Foresight for Planning Robot Motion | Video | Visual MPC | Paper |
| 2016.05 | Unsupervised Physical Interaction — Unsupervised Learning for Physical Interaction through Video Prediction | Video | Robot control | Paper |
| 2015.07 | Action-Conditional Video Prediction — Action-Conditional Video Prediction using Deep Networks in Atari Games | Video | Atari prediction | Paper |
| Date | Benchmark | Domain | Links |
|---|---|---|---|
| 2026.08 | WorldSimProbe — Action-following evaluation for action-conditioned world models | Action-conditioned world models | Project |
| 2026.06 | WorldRoamBench — Long-horizon stability of interactive world models | Interactive | Paper |
| 2026.06 | RoboTrustBench — Trustworthiness of video world models for robotic manipulation | Embodied | Paper · Project |
| 2026.05 | MiraBench — Action-conditioned reliability in robotic world models | Embodied | Paper |
| 2026.05 | WBench — Multi-turn interactive video world model evaluation | Interactive | Paper · Project |
| 2026.05 | ACWM-Phys — Generalized physical interaction in action-conditioned video world models | Embodied | Paper · Project · Code |
| 2026.05 | iWorld-Bench — Interactive world models with a unified action generation framework | Interactive | Paper |
| 2026.04 | WorldMark — Unified benchmark suite for interactive video world models | Interactive | Paper |
| 2026.04 | RoboWM-Bench — World models in robotic manipulation | Embodied | Paper · Project |
| 2026.01 | Wow, wo, val! — Embodied world model evaluation Turing test | Embodied | Paper |
This list is released under CC0 1.0.
A curated, continuously updated reading list, paper blogs, and resources for Action Conditioned World Models
46
5 commits
updated Sep 3, 2026
Action-conditioned video generation, game worlds, embodied simulation, and latent imagination.
A curated, community-driven reading list of visual world models that respond to actions,
roll out controllable futures, and support planning, policy learning, and evaluation.
This repository tracks action-conditioned visual world models (ACWMs) across three connected use cases:
| 🎮 Game World Models Controllable game-like environments and long-horizon interactive rollouts. | 🤖 Embodied Simulation World models for robot policy planning, imitation, reinforcement learning, and evaluation. | 🧠 Latent Control Pixel, token, feature, and latent dynamics used for visual control. |
The list includes models that generate future pixels/video as well as models that roll the world forward in a learned visual latent space. Last metadata check: 2026-08-14.
We use action-conditioned world model (ACWM) for a model that learns controlled visual dynamics:
future visual states · rewards · termination
= f(observation history, actions, context)
Here, future states may be pixels, video tokens, object/3D states, or learned visual features; actions may be observed controls or learned latent actions; and context may include language, goals, scene information, or embodiment details.
Included work should satisfy these criteria:
The sections are navigation aids rather than mutually exclusive theoretical categories. In particular, embodied papers are kept in one flat list and may carry several use tags.
Video — future RGB/RGB-D or multi-view observations.Token — discrete visual tokens or autoregressive observation tokens.Latent — recurrent or transformer latent states learned from visual observations.Feature — dynamics in a pretrained visual feature space.3D/4D — occupancy, point, Gaussian, flow, or other renderable spatial states.Unified — visual dynamics and action generation/decoding learned in one jointly trained model.Game — game-playing or game-like interactive environments.Interactive — controllable, multi-step visual rollouts.Eval — policy evaluation, ranking, verification, or failure detection.IL — imitation learning, behavior cloning, DAgger, or demonstration replay.RL — reinforcement learning or reinforcement fine-tuning inside the model.Plan — MPC, tree search, visual foresight, or goal-conditioned planning.Policy — joint world/action learning or direct policy improvement.Data — synthetic trajectories, observations, or demonstrations.arXiv 2026.07. [Paper]arXiv 2026.06. [Paper]arXiv 2026.05. [Paper] [Project] [List]arXiv 2026.03. [Paper]arXiv 2025.07. [Paper] [List]🎮 Game-like interactive environments. The learned model itself is a controllable game or game-like visual environment. Entries use the same table format as the embodied section and are ordered by first public release, newest first.
| Date | Work | Rep. | Uses | Links |
|---|---|---|---|---|
| 2026.07 | ABot-World-0 — Infinite Interactive World Rollout on a Single Desktop GPU | Video | Game Interactive | Paper |
| 2026.07 | AlayaWorld — Interactive Long-Horizon World Modeling | Video | Game Interactive | Paper |
| 2026.07 | LingBot-World 2.0 (Infinity) — Infinite Worlds with Versatile Interactions | Video | Game Interactive | Paper · Project · Code |
| 2026.06 | ActWorld — From Explorable to Interactive World Model via Action-Aware Memory | Video | Game Interactive | Paper |
| 2026.06 | DreamX-World 1.0 — A General-Purpose Interactive World Model | Video | Game Interactive | Paper |
| 2026.05 | minWM — A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models | Video | Game Interactive | Paper · Code |
| 2026.05 | WorldCraft — From Camera Navigation to Object Manipulation in Interactive Video World Models | Video | Game Interactive | Paper · Project |
| 2026.04 | Matrix-Game 3.0 — Real-Time and Streaming Interactive World Model with Long-Horizon Memory | Video | Game Interactive | Paper · Project |
| 2026.02 | LIVE — Long-horizon Interactive Video World Modeling | Video | Game Interactive | Paper · Project |
| 2026.01 | LingBot-World — Advancing Open-source World Models | Video | Game Interactive | Paper · Project · Code |
| 2025.12 | Yume-1.5 — A Text-Controlled Interactive World Generation Model | Video | Game Interactive | Paper · Project · Code |
| 2025.12 | WorldPlay — Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling | Video | Game Interactive | Paper |
| 2025.12 | Astra — General Interactive World Model with Autoregressive Denoising | Video | Game Interactive | Paper · Project · Code |
| 2025.12 | RELIC — Interactive Video World Model with Long-Horizon Memory | Video | Game Interactive | Paper · Project |
| 2025.11 | Hunyuan-GameCraft-2 — Instruction-following Interactive Game World Model | Video | Game Interactive | Paper · Project |
| 2025.08 | Matrix-Game 2.0 — An Open-Source Real-Time and Streaming Interactive World Model | Video | Game Interactive | Paper · Project |
| 2025.08 | Genie 3 — A New Frontier for World Models | Video | Game Interactive | Blog |
| 2025.07 | Yume — An Interactive World Generation Model | Video | Game Interactive | Paper · Project · Code |
| 2025.06 | Matrix-Game — Interactive World Foundation Model | Video | Game Interactive | Paper · Code |
| 2025.06 | Hunyuan-GameCraft — High-dynamic Interactive Game Video Generation with Hybrid History Condition | Video | Game Interactive | Paper · Project |
| 2025.05 | VRAG — Learning World Models for Interactive Video Generation | Video | Game Interactive | Paper |
| 2025.05 | Vid2World — Crafting Video Diffusion Models to Interactive World Models | Video | Game Interactive | Paper · Project |
| 2025.04 | MineWorld — A Real-Time and Open-Source Interactive World Model on Minecraft | Token | Game Interactive | Paper · Project · Code |
| 2024.12 | Genie 2 — A Large-Scale Foundation World Model | Video | Game Interactive | Blog |
| 2024.10 | OASIS — A Universe in a Transformer | Video | Game Interactive | Project |
| 2024.08 | GameNGen — Diffusion Models Are Real-Time Game Engines | Video | Game Interactive | Paper · Project |
| 2024.06 | Pandora — Towards General World Model with Natural Language Actions and Video States | Video | Game Interactive | Paper · Code |
| 2024.05 | iVideoGPT — Interactive VideoGPTs are Scalable World Models | Token | Game Interactive | Paper · Code |
| 2024.05 | DIAMOND — Diffusion for World Modeling: Visual Details Matter in Atari | Video | Game Interactive | Paper · Code |
| 2024.02 | Genie — Generative Interactive Environments | Token | Game Interactive | Paper · Project |
🤖 Robot-facing world models. This is deliberately one flat list spanning video, latent, 3D/4D, and unified world models. The Uses column allows a paper to be simultaneously classified as evaluation, imitation learning, RL, planning, policy learning, and/or data generation.
| Date | Work | Rep. | Uses | Links |
|---|---|---|---|---|
| 2026.08 | Hydra-0 — Action Flow for Generalist World Modeling and Control | Video | Eval IL Plan Policy Data | Paper · Project |
| 2026.08 | WALL-SS — Action-conditioned world model for embodied simulation | Video | Eval IL RL Plan Policy | Project |
| 2026.07 | CheckVLA — Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation | Latent | Eval Plan | Paper |
| 2026.06 | Recurrent Generative Replay — World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays | Video | IL Data | Paper |
| 2026.06 | WAM-RL — World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT | Video | RL Policy | Paper |
| 2026.06 | PiL-World — A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation | Video | Eval | Paper |
| 2026.05 | GE-Sim 2.0 — A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation | Video | Eval IL RL Data | Paper · Project |
| 2026.05 | Sword — Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training | Video | RL | Paper |
| 2026.04 | X-WAM — Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising | 3D/4D Unified | IL Plan Policy | Paper · Project |
| 2026.04 | dWorldEval — Scalable Robotic Policy Evaluation via Discrete Diffusion World Model | Token | Eval | Paper |
| 2026.04 | Hi-WM — Human-in-the-World-Model for Scalable Robot Post-Training | Video | IL RL Policy | Paper · Project |
| 2026.04 | WM-DAgger — Enabling Efficient Data Aggregation for Imitation Learning with World Models | Video | IL Data | Paper · Code |
| 2026.03 | DreamPlan — Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models | Video | RL Plan | Paper · Project |
| 2026.03 | Kinema4D — Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation | 3D/4D | Eval Data | Paper · Project · Code |
| 2026.03 | Interactive World Simulator — Interactive World Simulator for Robot Policy Training and Evaluation | Video | Eval IL RL | Paper · Project |
| 2026.02 | GigaBrain-0.5M* — A VLA That Learns From World Model-Based Reinforcement Learning | Video | RL Policy | Paper · Project |
| 2026.02 | VLAW — Iterative Co-Improvement of Vision-Language-Action Policy and World Model | Video | IL Policy Data | Paper · Project |
| 2026.02 | RISE — Self-Improving Robot Policy with Compositional World Model | Video | IL RL Policy | Paper · Project |
| 2026.02 | Say, Dream, and Act — Learning Video World Models for Instruction-Driven Robot Manipulation | Video | Plan Policy | Paper |
| 2026.02 | DreamDojo — A Generalist Robot World Model from Large-Scale Human Videos | Video | IL Policy Data | Paper · Project |
| 2026.02 | World-VLA-Loop — Closed-Loop Learning of Video World Model and VLA Policy | Video | IL Policy Data | Paper · Project |
| 2026.02 | WoVR — World Models as Reliable Simulators for Post-Training VLA Policies with RL | Video | RL | Paper · Project |
| 2026.02 | World-Gymnast — Training Robots with Reinforcement Learning in a World Model | Video | RL | Paper · Project |
| 2026.01 | lingbot-va — Causal World Modeling for Robot Control | Video | Plan Policy | Paper · Project · Code |
| 2026.01 | PointWorld — Scaling 3D World Models for In-The-Wild Robotic Manipulation | 3D/4D | Plan Policy | Paper · Project · Code · Model |
| 2025.12 | Veo World Simulator — Evaluating Gemini Robotics Policies in a Veo World Simulator | Video | Eval | Paper |
| 2025.11 | Scalable Policy Evaluation — Scalable Policy Evaluation with Video World Models | Video | Eval | Paper |
| 2025.10 | Ctrl-World — A Controllable Generative World Model for Robot Manipulation | Video | Plan Policy | Paper · Project · Code |
| 2025.10 | VLA-RFT — Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators | Video | RL | Paper · Project |
| 2025.09 | World-Env — Leveraging World Model as a Virtual Environment for VLA Post-Training | Video | RL | Paper · Code |
| 2025.09 | World4RL — Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation | Video | RL | Paper |
| 2025.08 | Genie Envisioner — A Unified World Foundation Platform for Robotic Manipulation | Video | Eval IL Plan Policy | Paper · Project |
| 2025.06 | ParticleFormer — A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation | 3D/4D | Plan | Paper · Project |
| 2025.06 | WorldVLA — Towards Autoregressive Action World Model | Video | Policy Plan | Paper · Code |
| 2025.05 | LaDi-WM — A Latent Diffusion-based World Model for Predictive Manipulation | Latent | Plan Policy | Paper |
| 2025.05 | FLARE — Robot Learning with Implicit World Modeling | Latent | Policy Data | Paper · Project · Code |
| 2025.05 | WorldEval — World Model as Real-World Robot Policies Evaluator | Video | Eval | Paper · Project |
| 2025.05 | DreamGen — Unlocking Generalization in Robot Learning through Video World Models | Video | IL Data | Paper · Code |
| 2025.04 | TesserAct — Learning 4D Embodied World Models | 3D/4D | Plan Policy | Paper · Project |
| 2025.04 | PIN-WM — Learning Physics-INformed World Models for Non-Prehensile Manipulation | 3D/4D | Plan | Paper |
| 2025.04 | UWM — Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets | Video Unified | IL Policy Data | Paper · Project |
| 2024.12 | Dream to Manipulate — Compositional World Models Empowering Robot Imitation Learning with Imagination | Video | IL Plan Data | Paper · Project |
| 2024.10 | EVA — An Embodied World Model for Future Video Anticipation | Video | Plan Policy | Paper · Project |
| 2024.06 | IRASim — A Fine-Grained World Model for Robot Manipulation | Video | Eval Plan | Paper · Project · Code |
| 2024.04 | RoboDreamer — Learning Compositional World Models for Robot Imagination | Video | Plan Policy | Paper · Project |
| 2024.03 | 3D-VLA — A 3D Vision-Language-Action Generative World Model | 3D/4D | Plan Policy | Paper |
| 2024.03 | ManiGaussian — Dynamic Gaussian Splatting for Multi-task Robotic Manipulation | 3D/4D | Plan Policy | Paper · Project |
| 2023.10 | UniSim — Learning Interactive Real-World Simulators | Video | IL RL Plan Policy Data | Paper · Project |
🧠 Latent imagination and visual control. This section covers general visual control and model-based RL work whose action-conditioned dynamics live partly or entirely in learned latent space. Pixel-generative methods are included when they are foundational to this line.
| Date | Work | Rep. | Uses | Links |
|---|---|---|---|---|
| 2025.06 | V-JEPA 2 / V-JEPA 2-AC — Self-Supervised Video Models Enable Understanding, Prediction and Planning | Feature | Action-conditioned planning | Paper · Project · Code |
| 2024.11 | DINO-WM — World Models on Pre-trained Visual Features enable Zero-shot Planning | Feature | Zero-shot planning | Paper · Project · Code |
| 2024.10 | AVID — Adapting Video Diffusion Models to World Models | Video | Action-conditioned diffusion | Paper · Code |
| 2024.06 | Delta-IRIS — Efficient World Models with Context-Aware Tokenization | Token | Visual RL | Paper · Code |
| 2023.10 | TD-MPC2 — Scalable, Robust World Models for Continuous Control | Latent | Continuous control | Paper · Code |
| 2023.01 | DreamerV3 — Mastering Diverse Domains through World Models | Latent | General RL | Paper · Code |
| 2022.09 | IRIS — Transformers are Sample-Efficient World Models | Token | Atari control | Paper · Code |
| 2022.06 | MWM — Masked World Models for Visual Control | Latent | Visual RL | Paper · Code |
| 2022.06 | DayDreamer — World Models for Physical Robot Learning | Latent | Real-robot RL | Paper · Code |
| 2022.03 | TD-MPC — Temporal Difference Learning for Model Predictive Control | Latent | MPC and control | Paper · Code |
| 2020.10 | DreamerV2 — Mastering Atari with Discrete World Models | Latent | Atari control | Paper · Code |
| 2019.12 | Dreamer — Dream to Control: Learning Behaviors by Latent Imagination | Latent | Continuous control | Paper · Code |
| 2019.11 | MuZero — Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model | Latent | Tree-search planning | Paper |
| 2019.03 | SimPLe — Model-Based Reinforcement Learning for Atari | Video | Atari control | Paper |
| 2018.11 | PlaNet — Learning Latent Dynamics for Planning from Pixels | Latent | MPC and control | Paper · Code |
| 2018.03 | World Models — World Models | Latent | Control by imagination | Paper · Project |
| 2016.10 | Deep Visual Foresight — Deep Visual Foresight for Planning Robot Motion | Video | Visual MPC | Paper |
| 2016.05 | Unsupervised Physical Interaction — Unsupervised Learning for Physical Interaction through Video Prediction | Video | Robot control | Paper |
| 2015.07 | Action-Conditional Video Prediction — Action-Conditional Video Prediction using Deep Networks in Atari Games | Video | Atari prediction | Paper |
| Date | Benchmark | Domain | Links |
|---|---|---|---|
| 2026.08 | WorldSimProbe — Action-following evaluation for action-conditioned world models | Action-conditioned world models | Project |
| 2026.06 | WorldRoamBench — Long-horizon stability of interactive world models | Interactive | Paper |
| 2026.06 | RoboTrustBench — Trustworthiness of video world models for robotic manipulation | Embodied | Paper · Project |
| 2026.05 | MiraBench — Action-conditioned reliability in robotic world models | Embodied | Paper |
| 2026.05 | WBench — Multi-turn interactive video world model evaluation | Interactive | Paper · Project |
| 2026.05 | ACWM-Phys — Generalized physical interaction in action-conditioned video world models | Embodied | Paper · Project · Code |
| 2026.05 | iWorld-Bench — Interactive world models with a unified action generation framework | Interactive | Paper |
| 2026.04 | WorldMark — Unified benchmark suite for interactive video world models | Interactive | Paper |
| 2026.04 | RoboWM-Bench — World models in robotic manipulation | Embodied | Paper · Project |
| 2026.01 | Wow, wo, val! — Embodied world model evaluation Turing test | Embodied | Paper |
This list is released under CC0 1.0.