ziqihuangg/Awesome-From-Video-Generation-to-World-Model

A list of works on video generation towards world model

517

stars

292

commits

Mar 21, 2026

updated

README

Awesome From Video Generation to World Model

overall_structure

The field of video generation is undergoing a paradigm shift - from generating realistic and appealing visuals to constructing world models that can simulate interactive and navigable environments. These models are not just visual tools; they serve as testbeds for training and evaluating intelligent agents, such as robots, autonomous vehicles, or virtual avatars. A central goal is to enable agents to perceive, act, and plan within generated video scenarios as if they were interacting with the real world. We compile key works that push video generation toward actionable world modeling, focusing physical plausibility, and the capacity for agents to navigate, manipulate, and learn from these synthetic environments.

arXiv Project Page Awesome visitors PRs Welcome

Overview

This repository currently contains the paper list for "Video Generation towards World Model".

What You'll Find Here

We hope to support the research and industrial communities by systematically collecting and organizing influential works that drive progress in video generation for world modeling.

News :fire:

Updates

This repository is updated periodically. If you have suggestions for additional resources, updates on methodologies, or fixes for expiring links, please feel free to do any of the following:

  • raise an Issue,
  • nominate awesome related works with Pull Requests,
  • For other queries: email both Ziqi ZIQI002 at e dot ntu dot edu dot sg and Jingtong yuejingtong137 at gmail dot com.

Table of Contents

1. Generation 1: Faithfulness - Accurate Simulation of the Real World

1.1 Video Foundation Model

DateVenueAcronymPaperProjectRepo@GitHub
2025-03-04ArxivHeliosHelios: Real Real-Time Long Video Generation ModelWebsiteCode
2024-12-30ArxivLTX-VideoLTX-Video: Realtime Video Latent DiffusionWebsiteCode
2024-12-12ArxivOwl-1Owl-1: Omni World Model for Consistent Long Video GenerationCode
2024-12-10ArxivSTIVSTIV: Scalable Text and Image Conditioned Video Generation
2024-09-24JT-CVWebsite
2024-09Hailuo AIWebsite
2024-06-06VideoTetrisVideoTetris: Towards Compositional Text-to-Video GenerationWebsiteCode
2024-02-22Snap VideoSnap Video: Scaled Spatiotemporal Transformers for Text-to-Video SynthesisWebsite
2024-01-23ArxivLumiereLumiere: A Space-Time Diffusion Model for Video GenerationWebsite
2024-01-17CVPR24VideoCrafter2Videocrafter2:Overcoming data limitations for high-quality video diffusion modelsWebsiteCode
2024-01-09ArxivMagicVideo-V2MagicVideo-V2: Multi-Stage High-Aesthetic Video GenerationWebsite
2024-01-05TMLR25LatteLatte: Latent Diffusion Transformer for Video GenerationWebsiteCode
2023-12-07ArxivHiGenHierarchical Spatio-temporal Decoupling for Text-to-Video GenerationWebsiteCode
2023-11-25ArxivSVDStable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsCode
2023-11-07ArxivI2VGen-XLI2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsCode
2023-10-31ICLR24SEINESEINE: Short-to-Long Video Diffusion Model for Generative Transition and PredictionWebsiteCode
2023-10-30ArxivVideoCrafter1Videocrafter1: Open diffusion models for high-quality video generationWebsiteCode
2023-10-18ECCV24DynamiCrafterDynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsWebsiteCode
2023-10-09ICLR24MAGVIT-v2Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
2023-09-27IJCV24Show-1Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video WebsiteCode
2023-09-26IJCV24LaVieLAVIE: High-Quality Video Generation with Cascaded Latent Diffusion ModelsWebsiteCode
2023-09-01ArxivVideoGenVideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video GenerationWebsite
2023-08-12ArxivModelScopeModelScope Text-to-Video Technical ReportWebsite
2023-07-10ICLR24AnimateDiffAnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningWebsiteCode
2023-06-29ArxivPikaWebsite
2023-06-07Gen-2Website
2023-02Gen-1Gen-1: The Next Step Forward for Generative AlWebsite
2022-12-10CVPR23MAGVITMAGVIT:Masked Generative Video TransformerWebsiteCode
2022-11-20ArxivMagicVideoMagicVideo: Efficient Video Generation With Latent Diffusion ModelsWebsite
2022-10-05ArxivImagen VideoImagen Video: High Definition Video Generation with Diffusion ModelsWebsite
2022-09-29ArxivMake-A-VideoMake-A-Video: Text-to-Video Generation without Text-Video Data
2022-05-29ICLR23CogVideoCogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersCode

1.2 Other Video Generation Model

1.2.1 GAN Based Video Generation

1.2.2 U-Net Based Video Generation

1.2.3 DiT Based Video Generation

1.2.4 Autoregressive Based Video Generation

1.3 Conditioned World Model

1.3.1 Conditined World Model in General Scene

Geometry Condition

3D Condition

Physics Condition

Trajectory Navigation

Camera Motion Navigation

Instruction Navigation

Action Navigation

1.3.2 Conditined World Model in Robotics

Action Navigation

Instruction Navigation

Goal Navigation

Hybrid Navigation

1.3.3 Conditined World Model in Autonomous Driving

Layout Condition

Instruction Navigation

Action Navigation

Hybrid Navigation

Other Navigation

1.3.4 Conditined World Model in Gaming

Controller Navigation

Action Navigation

2. Generation 2: Interactiveness - Controllability and Interactive Dynamics

2.1 High-quality World Foundation Model

DateVenueAcronymPaperProjectRepo@GitHub
2026-03-04ArxivHeliosHelios: Real Real-Time Long Video Generation ModelWebsiteCode
2026-02-02ArxivCausal ForcingCausal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationWebsiteCode
2026-01-24ArxivSkyReels-V3SkyReels-V3 Technique ReportWebsiteCode
2025-12-23ArxivSemanticGenSemanticGen: Video Generation in Semantic SpaceWebsite
2025-12-18ArxivKling-OmniKling-Omni Technical ReportWebsite
2025-12-16ArxivMemFlowMemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video NarrativesWebsite
2025-06-18Hailuo 02Website
2025-06-10ArxivSeedance 1.0Seedance 1.0: Exploring the Boundaries of Video Generation ModelsWebsite
2025-06-09ArxivSelf ForcingSelf Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionWebsiteCode
2025-05-19ArxivMAGI-1MAGI-1: Autoregressive Video Generation at ScaleWebsiteCode
2025-05Veo 3Veo 3: AI Video Generation with Realistic SoundWebsite
2025-04-17ArxivSkyReels-V2SkyReels-V2: Infinite-length Film Generative ModelWebsiteCode
2025-04-07Nova ReelWebsite
2025-03-31Gen-4Website
2025-03-26ArxivWan 2.1Wan: Open and Advanced Large-Scale Video Generative ModelsWebsiteCode
2025-03-13Step-Video-T2VWebsite
2025-03-12ArxivOpen-Sora2.0Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200kWebsiteCode
2025-02-11ArxivMagic 1-For-1Magic 1-For-1: Generating One Minute Video Clips within One MinuteWebsiteCode
2025-01-21MiracleVisionWebsite
2025-01-15ArxivRepVideoRepVideo: Rethinking Cross-Layer Representation for Video GenerationWebsiteCode
2025-01-14ArxivVchitect-2.0Vchitect-2.0: Parallel transformer for scaling up video diffusion modelsWebsiteCode
2025-01-07ArxivCosmosCosmos World Foundation Model Platform for Physical AIWebsiteCode
2024-12-29ArxivOpen-SoraOpen-sora: Democratizing efficient video production for allWebsiteCode
2024-12-10CVPR25CausVidFrom Slow Bidirectional to Fast Autoregressive Video Diffusion ModelsWebsiteCode
2024-12-03ArxivHunyuanVideoHunyuanVideo: A Systematic Framework For Large Video Generative ModelsWebsiteCode
2024-11-28ArxivOpen-Sora PlanOpen-Sora Plan: Open-Source Large Video Generation ModelCode
2024-10-22Mochi-1WebsiteCode
2024-08-12ICLR25CogvideoxCogvideox:Text-to-video diffusion models with an expert transformerCode
2024-07-08ArxivMiraMiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsWebsiteCode
2024-06-17Gen-3Website
2024-06-13LumaWebsite
2024-06-06KlingWebsite
2024-05-29ArxivEasyAnimateEasyanimate: A high-performance long video generation method based on transformer architectureWebsiteCode
2024-05-09JimengWebsite
2024-05-07ArxivViduVidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion ModelsWebsite
2024-02-15SoraVideo generation models as world simulatorsWebsite

World Model Regulation Methods

Stabilization:

Inference-time Physics Alignment:

Efficiency:

Planning Optimization:

Long Video Generation Methods

2.2 Video Generation as World Model in General Scenes

2.2.1 Geometry Condition Prior World Model

2.2.2 3D Condition Prior World Model

2.2.3 Physical Prior World Model

2.2.4 Audio Driven World Model

2.2.5 Trajectory Navigation World Model

2.2.6 Camera Motion Navigation World Model

2.2.7 Instruction Navigation World Model

2.2.8 Action Navigation World Model

2.3 Video Generation as World Model in Robotics

2.3.1 Action Navigation World Model

2.3.2 Instruction Navigation World Model

2.3.3 Goal Navigation World Model

2.3.4 Hybrid Navigation World Model

2.3.5 Real-time Interactive World Model

2.4. Video Generation as World Model in Autonomous Driving

2.4.1 Layout Prior World Model

2.4.2 Instruction Navigation World Model

2.4.3 Trajectory Navigation World Model

2.4.4 Action Navigation World Model

2.4.5 Hybrid Navigation World Model

2.4.6 Other Navigation World Model

2.5 Video Generation as World Model in Gaming

2.5.1 Controller Navigation World Model

2.5.2 Action Navigation World Model

2.5.3 Hybrid Navigation World Model

3. Generation 3: Planning - Modeling the Future Evolution of Complex Systems

For Robotics:

Note: Action and goal navigation for robotics.

4. Generation 4: Counterfactual and Outlier Modeling

4.1 Macroscopic Scale World Model

4.2 Mesoscopic Scale World Model

4.3 Microscopic Scale World Model

5. Evaluation and Datasets

5.1 Evaluation Metrics of Video Generation

5.2 Evaluation Metrics of World Model

5.3 Datasets

6. Study and Rethinking

6.1 Survey

6.2 Position & Perspective

7. Downstream Tasks for World Modeling

7.1 World Models as Data Generators

7.2 World Models as Reasoning Proxy

8. World Modele for Other Application

8.1 World Models for Medicine

Citation

If you find this paper useful, please consider citing:

@article{yue2025video,
  title={Simulating the World Model with Artificial Intelligence: A Roadmap},
  author={Jingtong Yue, Ziqi Huang, Zhaoxi Chen, Xintao Wang, Pengfei Wan, Ziwei Liu},
  journal={arXiv preprint arXiv:2511.08585},
  year={2025}
}

Contributors

Jingtong0527

286 commits

ziqihuangg

5 commits

SHYuanBest

1 commits

ziqihuangg/Awesome-From-Video-Generation-to-World-Model

A list of works on video generation towards world model

517

stars

292

commits

Mar 21, 2026

updated

README

Awesome From Video Generation to World Model

overall_structure

The field of video generation is undergoing a paradigm shift - from generating realistic and appealing visuals to constructing world models that can simulate interactive and navigable environments. These models are not just visual tools; they serve as testbeds for training and evaluating intelligent agents, such as robots, autonomous vehicles, or virtual avatars. A central goal is to enable agents to perceive, act, and plan within generated video scenarios as if they were interacting with the real world. We compile key works that push video generation toward actionable world modeling, focusing physical plausibility, and the capacity for agents to navigate, manipulate, and learn from these synthetic environments.

arXiv Project Page Awesome visitors PRs Welcome

Overview

This repository currently contains the paper list for "Video Generation towards World Model".

What You'll Find Here

We hope to support the research and industrial communities by systematically collecting and organizing influential works that drive progress in video generation for world modeling.

News :fire:

Updates

This repository is updated periodically. If you have suggestions for additional resources, updates on methodologies, or fixes for expiring links, please feel free to do any of the following:

  • raise an Issue,
  • nominate awesome related works with Pull Requests,
  • For other queries: email both Ziqi ZIQI002 at e dot ntu dot edu dot sg and Jingtong yuejingtong137 at gmail dot com.

Table of Contents

1. Generation 1: Faithfulness - Accurate Simulation of the Real World

1.1 Video Foundation Model

DateVenueAcronymPaperProjectRepo@GitHub
2025-03-04ArxivHeliosHelios: Real Real-Time Long Video Generation ModelWebsiteCode
2024-12-30ArxivLTX-VideoLTX-Video: Realtime Video Latent DiffusionWebsiteCode
2024-12-12ArxivOwl-1Owl-1: Omni World Model for Consistent Long Video GenerationCode
2024-12-10ArxivSTIVSTIV: Scalable Text and Image Conditioned Video Generation
2024-09-24JT-CVWebsite
2024-09Hailuo AIWebsite
2024-06-06VideoTetrisVideoTetris: Towards Compositional Text-to-Video GenerationWebsiteCode
2024-02-22Snap VideoSnap Video: Scaled Spatiotemporal Transformers for Text-to-Video SynthesisWebsite
2024-01-23ArxivLumiereLumiere: A Space-Time Diffusion Model for Video GenerationWebsite
2024-01-17CVPR24VideoCrafter2Videocrafter2:Overcoming data limitations for high-quality video diffusion modelsWebsiteCode
2024-01-09ArxivMagicVideo-V2MagicVideo-V2: Multi-Stage High-Aesthetic Video GenerationWebsite
2024-01-05TMLR25LatteLatte: Latent Diffusion Transformer for Video GenerationWebsiteCode
2023-12-07ArxivHiGenHierarchical Spatio-temporal Decoupling for Text-to-Video GenerationWebsiteCode
2023-11-25ArxivSVDStable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsCode
2023-11-07ArxivI2VGen-XLI2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsCode
2023-10-31ICLR24SEINESEINE: Short-to-Long Video Diffusion Model for Generative Transition and PredictionWebsiteCode
2023-10-30ArxivVideoCrafter1Videocrafter1: Open diffusion models for high-quality video generationWebsiteCode
2023-10-18ECCV24DynamiCrafterDynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsWebsiteCode
2023-10-09ICLR24MAGVIT-v2Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
2023-09-27IJCV24Show-1Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video WebsiteCode
2023-09-26IJCV24LaVieLAVIE: High-Quality Video Generation with Cascaded Latent Diffusion ModelsWebsiteCode
2023-09-01ArxivVideoGenVideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video GenerationWebsite
2023-08-12ArxivModelScopeModelScope Text-to-Video Technical ReportWebsite
2023-07-10ICLR24AnimateDiffAnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningWebsiteCode
2023-06-29ArxivPikaWebsite
2023-06-07Gen-2Website
2023-02Gen-1Gen-1: The Next Step Forward for Generative AlWebsite
2022-12-10CVPR23MAGVITMAGVIT:Masked Generative Video TransformerWebsiteCode
2022-11-20ArxivMagicVideoMagicVideo: Efficient Video Generation With Latent Diffusion ModelsWebsite
2022-10-05ArxivImagen VideoImagen Video: High Definition Video Generation with Diffusion ModelsWebsite
2022-09-29ArxivMake-A-VideoMake-A-Video: Text-to-Video Generation without Text-Video Data
2022-05-29ICLR23CogVideoCogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersCode

1.2 Other Video Generation Model

1.2.1 GAN Based Video Generation

1.2.2 U-Net Based Video Generation

1.2.3 DiT Based Video Generation

1.2.4 Autoregressive Based Video Generation

1.3 Conditioned World Model

1.3.1 Conditined World Model in General Scene

Geometry Condition

3D Condition

Physics Condition

Trajectory Navigation

Camera Motion Navigation

Instruction Navigation

Action Navigation

1.3.2 Conditined World Model in Robotics

Action Navigation

Instruction Navigation

Goal Navigation

Hybrid Navigation

1.3.3 Conditined World Model in Autonomous Driving

Layout Condition

Instruction Navigation

Action Navigation

Hybrid Navigation

Other Navigation

1.3.4 Conditined World Model in Gaming

Controller Navigation

Action Navigation

2. Generation 2: Interactiveness - Controllability and Interactive Dynamics

2.1 High-quality World Foundation Model

DateVenueAcronymPaperProjectRepo@GitHub
2026-03-04ArxivHeliosHelios: Real Real-Time Long Video Generation ModelWebsiteCode
2026-02-02ArxivCausal ForcingCausal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationWebsiteCode
2026-01-24ArxivSkyReels-V3SkyReels-V3 Technique ReportWebsiteCode
2025-12-23ArxivSemanticGenSemanticGen: Video Generation in Semantic SpaceWebsite
2025-12-18ArxivKling-OmniKling-Omni Technical ReportWebsite
2025-12-16ArxivMemFlowMemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video NarrativesWebsite
2025-06-18Hailuo 02Website
2025-06-10ArxivSeedance 1.0Seedance 1.0: Exploring the Boundaries of Video Generation ModelsWebsite
2025-06-09ArxivSelf ForcingSelf Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionWebsiteCode
2025-05-19ArxivMAGI-1MAGI-1: Autoregressive Video Generation at ScaleWebsiteCode
2025-05Veo 3Veo 3: AI Video Generation with Realistic SoundWebsite
2025-04-17ArxivSkyReels-V2SkyReels-V2: Infinite-length Film Generative ModelWebsiteCode
2025-04-07Nova ReelWebsite
2025-03-31Gen-4Website
2025-03-26ArxivWan 2.1Wan: Open and Advanced Large-Scale Video Generative ModelsWebsiteCode
2025-03-13Step-Video-T2VWebsite
2025-03-12ArxivOpen-Sora2.0Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200kWebsiteCode
2025-02-11ArxivMagic 1-For-1Magic 1-For-1: Generating One Minute Video Clips within One MinuteWebsiteCode
2025-01-21MiracleVisionWebsite
2025-01-15ArxivRepVideoRepVideo: Rethinking Cross-Layer Representation for Video GenerationWebsiteCode
2025-01-14ArxivVchitect-2.0Vchitect-2.0: Parallel transformer for scaling up video diffusion modelsWebsiteCode
2025-01-07ArxivCosmosCosmos World Foundation Model Platform for Physical AIWebsiteCode
2024-12-29ArxivOpen-SoraOpen-sora: Democratizing efficient video production for allWebsiteCode
2024-12-10CVPR25CausVidFrom Slow Bidirectional to Fast Autoregressive Video Diffusion ModelsWebsiteCode
2024-12-03ArxivHunyuanVideoHunyuanVideo: A Systematic Framework For Large Video Generative ModelsWebsiteCode
2024-11-28ArxivOpen-Sora PlanOpen-Sora Plan: Open-Source Large Video Generation ModelCode
2024-10-22Mochi-1WebsiteCode
2024-08-12ICLR25CogvideoxCogvideox:Text-to-video diffusion models with an expert transformerCode
2024-07-08ArxivMiraMiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsWebsiteCode
2024-06-17Gen-3Website
2024-06-13LumaWebsite
2024-06-06KlingWebsite
2024-05-29ArxivEasyAnimateEasyanimate: A high-performance long video generation method based on transformer architectureWebsiteCode
2024-05-09JimengWebsite
2024-05-07ArxivViduVidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion ModelsWebsite
2024-02-15SoraVideo generation models as world simulatorsWebsite

World Model Regulation Methods

Stabilization:

Inference-time Physics Alignment:

Efficiency:

Planning Optimization:

Long Video Generation Methods

2.2 Video Generation as World Model in General Scenes

2.2.1 Geometry Condition Prior World Model

2.2.2 3D Condition Prior World Model

2.2.3 Physical Prior World Model

2.2.4 Audio Driven World Model

2.2.5 Trajectory Navigation World Model

2.2.6 Camera Motion Navigation World Model

2.2.7 Instruction Navigation World Model

2.2.8 Action Navigation World Model

2.3 Video Generation as World Model in Robotics

2.3.1 Action Navigation World Model

2.3.2 Instruction Navigation World Model

2.3.3 Goal Navigation World Model

2.3.4 Hybrid Navigation World Model

2.3.5 Real-time Interactive World Model

2.4. Video Generation as World Model in Autonomous Driving

2.4.1 Layout Prior World Model

2.4.2 Instruction Navigation World Model

2.4.3 Trajectory Navigation World Model

2.4.4 Action Navigation World Model

2.4.5 Hybrid Navigation World Model

2.4.6 Other Navigation World Model

2.5 Video Generation as World Model in Gaming

2.5.1 Controller Navigation World Model

2.5.2 Action Navigation World Model

2.5.3 Hybrid Navigation World Model

3. Generation 3: Planning - Modeling the Future Evolution of Complex Systems

For Robotics:

Note: Action and goal navigation for robotics.

4. Generation 4: Counterfactual and Outlier Modeling

4.1 Macroscopic Scale World Model

4.2 Mesoscopic Scale World Model

4.3 Microscopic Scale World Model

5. Evaluation and Datasets

5.1 Evaluation Metrics of Video Generation

5.2 Evaluation Metrics of World Model

5.3 Datasets

6. Study and Rethinking

6.1 Survey

6.2 Position & Perspective

7. Downstream Tasks for World Modeling

7.1 World Models as Data Generators

7.2 World Models as Reasoning Proxy

8. World Modele for Other Application

8.1 World Models for Medicine

Citation

If you find this paper useful, please consider citing:

@article{yue2025video,
  title={Simulating the World Model with Artificial Intelligence: A Roadmap},
  author={Jingtong Yue, Ziqi Huang, Zhaoxi Chen, Xintao Wang, Pengfei Wan, Ziwei Liu},
  journal={arXiv preprint arXiv:2511.08585},
  year={2025}
}

Contributors

Jingtong0527

286 commits

ziqihuangg

5 commits

SHYuanBest

1 commits