[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the Real World: A Unified Survey of Multimodal Generative Models" (IEEE TPAMI, 2026) and Awesome-Text2X-Resources. Watch this repository for the latest updates! ๐ฅ
See the codeThis repository is divided into two main sections:
Our Survey Paper Collection - This section presents our survey, "Simulating the Real World: A Unified Survey of Multimodal Generative Models" (IEEE TPAMI, 2026), which systematically unify the study of 2D, video, 3D and 4D generation within a single framework.
Text2X Resources โ This section continues the original Awesome-Text2X-Resources, an open collection of state-of-the-art (SOTA) and novel Text-to-X (X can be everything) methods, including papers, codes, and datasets. The goal is to track the rapid progress in this field and provide researchers with up-to-date references.
โญ If you find this repository useful for your research or work, a star is highly appreciated!
๐ This repository is continuously updated. If you find relevant papers, blog posts, videos, or other resources that should be included, feel free to submit a pull request (PR) or open an issue. Community contributions are always welcome!
๐๐ข๐ฆ๐ฎ๐ฅ๐๐ญ๐ข๐ง๐ ๐ญ๐ก๐ ๐๐๐๐ฅ ๐๐จ๐ซ๐ฅ๐: ๐ ๐๐ง๐ข๐๐ข๐๐ ๐๐ฎ๐ซ๐ฏ๐๐ฒ ๐จ๐ ๐๐ฎ๐ฅ๐ญ๐ข๐ฆ๐จ๐๐๐ฅ ๐๐๐ง๐๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐จ๐๐๐ฅ๐ฌ
๐ฐ๐ฌ๐ฌ๐ฌ ๐ป๐ท๐จ๐ด๐ฐ, ๐๐๐๐
Abstract
Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the physical world, enabling more accurate simulations and meaningful interactions. However, current methods often treat different modalities, including 2D (images), videos, 3D, and 4D representations, as independent domains, overlooking their interdependencies. Additionally, these methods typically focus on isolated dimensions of reality without systematically integrating their connections. In this survey, we present a unified survey for multimodal generative models that investigate the progression of data dimensionality in real-world simulation. Specifically, this survey starts from 2D generation (appearance), then moves to video (appearance+dynamics) and 3D generation (appearance+geometry), and finally culminates in 4D generation that integrate all dimensions. To the best of our knowledge, this is the first attempt to systematically unify the study of 2D, video, 3D and 4D generation within a single framework. To guide future research, we provide a comprehensive review of datasets, evaluation metrics and future directions, and fostering insights for newcomers. This survey serves as a bridge to advance the study of multimodal generative models and real-world simulation within a unified framework.
โญ Citation
If you find this paper and repo helpful for your research, please cite it below:
@article{hu2026simulating,
title={Simulating the real world: A unified survey of multimodal generative models},
author={Hu, Yuqi and Wang, Longguang and Liu, Xian and Chen, Ling-Hao and Guo, Yuwei and Shi, Yukai and Liu, Ce and Rao, Anyi and Wang, Zeyu and Xiong, Hui},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
year={2026},
publisher={IEEE}
}
๐งญ Getting Started with Key Concepts
[!Note] If you are new to this field, you can find clear and concise definitions of essential technical terms and concepts, such as NeRF, 3DGS, SDS, and Diffusion Models in our Glossary.
[!TIP] Feel free to pull requests or contact us if you find any related papers that are not included here. The process to submit a pull request is as follows:
README.md using the following format:[Origin] **Paper Title** [[Paper](Paper Link)] [[GitHub](GitHub Link)] [[Project Page](Project Page Link)]
We present a unified framework connecting 2D, Video, 3D, and 4D generation through text-guided synthesis. This paradigm illustrates how higher-dimensional content is synthesized by extending foundational modalities along spatial and temporal axes. (1)2D->3D: Spatial lifting of 2D priors to achieve geometric consistency; (2)2D->Video: Temporal inflation of static features to capture motion dynamics; (3)Video->4D: Spatial reconstruction and stabilization of dynamic sequences; (4)3D->4D: Temporal animation and deformation of static geometry. This perspective underscores that higher-dimensional generation methodologies are derivatives of foundational lower-dimensional generative priors, adapted through specialized architectural extensions.
Here are some seminal papers and models.
Overview of text-to-video generation technologies categorized by three main approaches.
Text-to-video generation models adapt text-to-image frameworks to handle the additional dimension of dynamics in the real world. We classify these models into three categories based on different generative machine learning architectures.
Survey
(1) VAE- and GAN-based Approaches.
VAE-based Approaches.
GAN-based Approaches.
(2) Diffusion-based Approaches.
U-Net-based Architectures.
Transformer-based Architectures.
(3) Autoregressive-based Approaches.
Video Editing.
Novel View Synthesis.
Human Animation in Videos.
Recent text-to-3D, image-to-3D and video-to-3D generation methods.
Survey
Feedforward Approaches.
Optimization-based Approaches.
MVS-based Approaches.
Feedforward Approaches.
Optimization-based Approaches.
MVS-based Approaches.
Avatar Generation.
Scene Generation.
3D Editing.
Representative works of 4D generation methods. ''Rep'' stands for representations.
Feedforward Approaches.
Optimization-based Approaches.
4D Editing.
Human Animation.
Summary of the widely-used 2D, video, 3D and 4D generation datasets. [Link] directs to dataset websites.
Summary of common evaluation metrics.
Atlas: the world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D, World Labs, Sept 1st, 2026.
WorldFoundry([GitHub]): an open-source infrastructure for world models.
NVIDIA Cosmos ([GitHub] [Paper]): NVIDIA Cosmos is a world foundation model platform for accelerating the development of physical AI systems.
Genie3, Google Deepmind, August 5th, 2025.
๐ฏBack to Top - Our Survey Paper Collection
An open collection of state-of-the-art (SOTA), novel Text to X (X can be everything) methods (papers, codes and datasets), intended to keep pace with the anticipated surge of research.
2026.09.04 - rename previous T2V Subsection to Video Generation Subsection.2026.01.07 - update 2025 papers collection into docs.2025.12.03 - update several papers accepted by NeurIPS 2025, congrats to all ๐2025.05.08 - update new layout.2025.04.18 - update layout on section Related Resources.2025.03.10 - CVPR 2025 Accepted Papers๐2025.02.28 - update several papers status "CVPR 2025" to accepted papers, congrats to all ๐2025.01.23 - update several papers status "ICLR 2025" to accepted papers, congrats to all ๐2025.01.09 - update layout.2024.12.21 adjusted the layouts of several sections and Happy Winter Solstice โช๐ฅฃ.2024.09.26 - update several papers status "NeurIPS 2024" to accepted papers, congrats to all ๐2024.09.03 - add one new section 'text to model'.2024.06.30 - add one new section 'text to video'.2024.07.02 - update several papers status "ECCV 2024" to accepted papers, congrats to all ๐2024.06.21 - add one hot Topic about AIGC 4D Generation on the section of Suvery and Awesome Repos.2024.06.17 - an awesome repo for CVPR2024 Link ๐๐ป2024.04.05 adjusted the layout and added accepted lists and ArXiv lists to each section.2024.04.05 - an awesome repo for CVPR2024 on 3DGS and NeRF Link ๐๐ป2024.03.25 - add one new survey paper of 3D GS into the section of "Survey and Awesome Repos--Topic 1: 3D Gaussian Splatting".2024.03.12 - add a new section "Dynamic Gaussian Splatting", including Neural Deformable 3D Gaussians, 4D Gaussians, Dynamic 3D Gaussians.2024.03.11 - CVPR 2024 Accpeted Papers Link| Year | Title | Venue | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2026 | Geometry-aware 4D Video Generation for Robot Manipulation | ICLR 2026 | Link | Link | Link |
| 2026 | Turbo4DGen: Ultra-Fast Acceleration for 4D Generation | ICML 2026 | Link | Link | Link |
| 2026 | Code2Worlds: Empowering Coding LLMs for 4D World Generation | ICML 2026 | Link | Link | Link |
| 2026 | PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation | CVPR 2026 | Link | Link | Link |
| 2026 | AvatarPointillist: Autoregressive 4D Gaussian Avatarization | CVPR 2026 | Link | Link | Link |
| 2026 | Vista4D: Video Reshooting with 4D Point Clouds | CVPR 2026 | Link | Link | Link |
| 2026 | Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis | CVPR 2026 | Link | Link | Link |
| 2026 | NeuROK: Generative 4D Neural Object Kinematics | CVPR 2026 | Link | Coming Soon! | Link |
| 2026 | Choreographing a World of Dynamic Objects | CVPR 2026 | Link | Link | Link |
| 2026 | ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion | CVPR 2026 | Link | Link | Link |
| 2026 | MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing | ECCV 2026 | Link | Coming Soon! | Link |
| 2026 | VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward | ECCV 2026 | Link | -- | Link |
| 2026 | LivingWorld: Interactive 4D World Generation with Environmental Dynamics | ECCV 2026 | Link | Link | Link |
| 2026 | InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation | ECCV 2026 | Link | Datasets | Link |
| 2026 | MoGe4D: Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation | ECCV 2026 | Link | Link | Link |
| 2026 | Alignment Is All You Need For X-to-4D Generation | IEEE Transactions on Multimedia (TMM) 2026 | Link | -- | Link |
| 2026 | Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild | SIGGRAPH Asia 2026 | Link | Link | Link |
| 2026 | 4DAnyone: Create Anyone in 4D from a Casual Monocular Video | SIGGRAPH Asia 2026 | Link | Link | Link |
| 2026 | Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction | CVPR 2026 4DV Workshop | Link | -- | -- |
%accepted papers
@article{liu2025geometry,
title={Geometry-aware 4D Video Generation for Robot Manipulation},
author={Liu, Zeyi and Li, Shuang and Cousineau, Eric and Feng, Siyuan and Burchfiel, Benjamin and Song, Shuran},
journal={arXiv preprint arXiv:2507.01099},
year={2025}
}
@article{man2026turbo4dgen,
title={Turbo4DGen: Ultra-Fast Acceleration for 4D Generation},
author={Man, Yuanbin and Huang, Ying and Ren, Zhile and Yin, Miao},
journal={arXiv preprint arXiv:2603.29572},
year={2026}
}
@article{zhang2026code2worlds,
title={Code2worlds: Empowering coding llms for 4d world generation},
author={Zhang, Yi and Wang, Yunshuang and Zhang, Zeyu and Tang, Hao},
journal={arXiv preprint arXiv:2602.11757},
year={2026}
}
@article{zhan2026perpetualwonder,
title={PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation},
author={Zhan, Jiahao and Li, Zizhang and Yu, Hong-Xing and Wu, Jiajun},
journal={arXiv preprint arXiv:2602.04876},
year={2026}
}
@inproceedings{liu2026avatarpointillist,
title = {AvatarPointillist: Autoregressive 4D Gaussian Avatarization},
author = {Hongyu Liu and Xuan Wang and Yating Wang and Zijian Wu and Ziyu Wan and Yue Ma and Runtao Liu and Boyao Zhou and Yujun Shen and Qifeng Chen},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2026}
}
@misc{lin2026vista4dvideoreshooting4d,
title={Vista4D: Video Reshooting with 4D Point Clouds},
author={Kuan Heng Lin and Zhizheng Liu and Pablo Salamanca and Yash Kant and Ryan Burgert and Yuancheng Xu and Koichi Namekata and Yiwei Zhao and Bolei Zhou and Micah Goldblum and Paul Debevec and Ning Yu},
year={2026},
eprint={2604.21915},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.21915},
}
@inproceedings{chen2026motion,
title={Motion 3-to-4: 3d motion reconstruction for 4d synthesis},
author={Chen, Hongyuan and Chen, Xingyu and Xu, Zexiang and Chen, Anpei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={28947--28958},
year={2026}
}
@inproceedings{geng2026neurok,
title={NeuROK: Generative 4D Neural Object Kinematics},
author={Geng, Chen and He, Guangzhao and Gao, Yue and Zhang, Yunzhi and Wu, Shangzhe and Wu, Jiajun},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={39239--39251},
year={2026}
}
@article{lyu2026choreographing,
title={Choreographing a World of Dynamic Objects},
author={Lyu, Yanzhe and Geng, Chen and Dharmarajan, Karthik and Zhang, Yunzhi and Alzayer, Hadi and Wu, Shangzhe and Wu, Jiajun},
journal={arXiv preprint arXiv:2601.04194},
year={2026}
}
@article{sabathier2026actionmesh,
title={Actionmesh: Animated 3d mesh generation with temporal 3d diffusion},
author={Sabathier, Remy and Novotny, David and Mitra, Niloy J and Monnier, Tom},
journal={arXiv preprint arXiv:2601.16148},
year={2026}
}
@article{fiebelman2026mv,
title={MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing},
author={Fiebelman, Gal and Averbuch-Elor, Hadar and Benaim, Sagie},
journal={arXiv preprint arXiv:2607.05376},
year={2026}
}
@article{an2026vggrpo,
title={Vggrpo: Towards world-consistent video generation with 4d latent reward},
author={An, Zhaochong and Kupyn, Orest and Uscidda, Th{\'e}o and Colaco, Andrea and Ahuja, Karan and Belongie, Serge and Gonzalez-Franco, Mar and Gazulla, Marta Tintore},
journal={arXiv preprint arXiv:2603.26599},
year={2026}
}
@article{mun2026livingworld,
title={LivingWorld: Interactive 4D World Generation with Environmental Dynamics},
author={Mun, Hyeongju and Jin, In-Hwan and Kim, Sohyeong and Kong, Kyeongbo},
journal={arXiv preprint arXiv:2604.01641},
year={2026}
}
@article{peng2026interpet4d,
title={InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation},
author={Peng, Yichen and Song, Jyun-Ting and Liao, Chen-Chieh and Kitani, Kris and Koike, Hideki and Wu, Erwin},
journal={arXiv preprint arXiv:2607.10287},
year={2026}
}
@misc{zhang2026geometryawaresingleimage4dsynthesis,
title={Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation},
author={Yanran Zhang and Ziyi Wang and Wenzhao Zheng and Zheng Zhu and Jie Zhou and Jiwen Lu},
year={2026},
eprint={2512.05044},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.05044},
}
@article{miao2026alignment,
title={Alignment Is All You Need For X-to-4D Generation},
author={Miao, Qiaowei and Li, Kehan and Luo, Yawei and Yang, Yi},
journal={arXiv preprint arXiv:2607.02516},
year={2026}
}
@article{litman2026lift4d,
title={Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild},
author={Litman, Yehonathan and Ma, Xiaoxuan and Shah, Manan and Ugrinovic, Nicolas and Kitani, Kris and De la Torre, Fernando and Tulsiani, Shubham},
journal={arXiv preprint arXiv:2606.23688},
year={2026}
}
@article{jin2026fdanyone,
title={4DAnyone: Create Anyone in 4D from a Casual Monocular Video},
author={Jin, Yudong and Xie, Tao and Zhang, Qihang and Shen, Zehong and Xu, Zhen and Shen, Yujun and Bao, Hujun and Zhou, Xiaowei and Xu, Yinghao},
journal={arXiv preprint arXiv:2608.20335},
year={2026},
url={https://arxiv.org/abs/2608.20335}
}
@article{liu2026streaming4d,
title={Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction},
author={Liu, Xiaoyan and Liu, Jiaxin and Li, Kangrui and Zhou, Sifan},
journal={arXiv preprint arXiv:2609.00610},
year={2026}
}
Melonie de Almeida, Daniela Ivanova, Tong Shi, John H. Williamson, Paul Henderson (University of Glasgow)
Haonan Wang, Hanyu Zhou, Tao Gu, Luxin Yan
(Huazhong University of Science and Technology, National University of Singapore, Macquarie University)
Sunwoo Park, Taesung Kwon, Jong Chul Ye (KAIST AI)
Dvir Samuel, Yuval Atzmon, Gal Chechik, Yoni Kasten (NVIDIA Research, Bar-Ilan University)
Jiraphon Yenphraphai, Jianqi Chen, Jian Wang, Gordon Qian, Sergey Tulyakov, Rameen Abdal, Raymond A. Yeh, Peter Wonka, Chaoyang Wang
(Snap, Purdue University, KAUST)
Yiran Wang, Zeyu Zhang, Yuanming Li, Ziming Wang, Yang Zhao
(USYD, SpatialReal, ZJU, La Trobe)
Sai Kumar Dwivedi, Federica Bogo, Buฤra Tekin, Chenhongyi Yang, Nadine Bertsch, Tomas Hodan, Michael J. Black, Dimitrios Tzionas, Shreyas Hampali
(Meta, Max Planck Institute for Intelligent Systems, University of Amsterdam, Aristotle University of Thessaloniki)
JoungBin Lee, Jaewoo Jung, Jongmin Lee, Tongmin Kim, Hyunsung Kim, Takuya Narihira, Kazumi Fukuda, Jahyeok Koo, Jisang Han, Yuki Mitsufuji, Seungryong Kim
(KAIST AI, Sony AI, Sony Group Corporation)
Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
(DAMO Academy Alibaba Group, Hong Kong Embodied AI Lab, CUHK, Hupan Lab)
Hao Feng, Zhi Zuo, Jia-Hui Pan, Ka-Hei Hui, Zhengzhe Liu, Dian Zhang, Haoran Xie, Bin Sheng, Jingyu Hu
(Lingnan University, Chinese University of Hong Kong, Autodesk Research, Shanghai Jiao Tong University)
Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
(CASIA, UCAS, ShanghaiTech)
Yunpeng Bai, Haoxiang Li, Qixing Huang (UT Austin, Pixocial Technology)
Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang (Zhejiang University)
Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
(UCLA, Tsinghua University)
Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
(Peking University, Tencent Hunyuan)
Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
(The Hong Kong University of Science and Technology, ARC Lab Tencent IEG, The University of Hong Kong, The University of Texas at Austin)
| Year | Title | ArXiv Time | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2026 | Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians | 2 Jan 2026 | Link | -- | Link |
| 2026 | InSpatio-World | 20 Mar 2026 | Live Demo | Link | Link |
| 2026 | ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation | 8 May 2026 | Link | -- | -- |
| 2026 | Geometric 4D Stitching for Grounded 4D Generation | 11 May 2026 | Link | -- | -- |
| 2026 | Fast 4D Mesh Generation by Spatio-Temporal Attention Chains | 19 May 2026 | Link | -- | Link |
| 2026 | Helix4D: Complex 4D Mesh Generation | 25 May 2026 | Link | -- | Link |
| 2026 | SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction | 14 Jun 2026 | Link | -- | Link |
| 2026 | IMAGIN-4D: Image-Guided Controllable Interaction Generation | 22 Jun 2026 | Link | -- | Link |
| 2026 | MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation | 24 Jun 2026 | Link | Link | Link |
| 2026 | RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation | 7 Jul 2026 | Link | Link | Link |
| 2026 | SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation | 9 Jul 2026 | Link | -- | -- |
| 2026 | Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation | 15 Jul 2026 | Link | Link | Link |
| 2026 | PE-Field 4D: Video Generation Models as Canvas | 17 Jul 2026 | Link | -- | -- |
| 2026 | Beyond Pixels: From Video Priors to 4D Worlds | 11 Aug 2026 | Link | Link | Link |
| 2026 | Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models | 20 Aug 2026 | Link | -- | Link |
| 2026 | 4DStreamCtrl: Interactive Video Generation with Online 4D Control | 27 Aug 2026 | Link | -- | Link |
| 2026 | GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation | 21 Sep 2026 | Link | Link | Link |
%axiv papers
@article{de2026pixel,
title={Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians},
author={de Almeida, Melonie and Ivanova, Daniela and Shi, Tong and Williamson, John H and Henderson, Paul},
journal={arXiv preprint arXiv:2601.00678},
year={2026}
}
@misc{inspatio-world,
title={InSpatio-World},
author={InSpatio-World Contributors},
howpublished={\url{https://github.com/inspatio/inspatio-world}},
year={2025}
}
@article{wang2026st,
title={ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation},
author={Wang, Haonan and Zhou, Hanyu and Gu, Tao and Yan, Luxin},
journal={arXiv preprint arXiv:2605.07390},
year={2026}
}
@article{park2026geometric,
title={Geometric 4D Stitching for Grounded 4D Generation},
author={Park, Sunwoo and Kwon, Taesung and Ye, Jong Chul},
journal={arXiv preprint arXiv:2605.09984},
year={2026}
}
@article{samuel2026fast,
title={Fast 4D Mesh Generation by Spatio-Temporal Attention Chains},
author={Samuel, Dvir and Atzmon, Yuval and Chechik, Gal and Kasten, Yoni},
journal={arXiv preprint arXiv:2605.19786},
year={2026}
}
@article{yenphraphai2026helix4d,
title={Helix4D: Complex 4D Mesh Generation},
author={Yenphraphai, Jiraphon and Chen, Jianqi and Wang, Jian and Qian, Gordon and Tulyakov, Sergey and Abdal, Rameen and Yeh, Raymond A and Wonka, Peter and Wang, Chaoyang},
journal={arXiv preprint arXiv:2605.26109},
year={2026}
}
@article{wang2026spatialavatar,
title={SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction},
author={Wang, Yiran and Zhang, Zeyu and Li, Yuanming and Wang, Ziming and Zhao, Yang},
journal={arXiv preprint arXiv:2606.15659},
year={2026}
}
@misc{dwivedi2026imagin4dimageguidedcontrollableinteraction,
title={IMAGIN-4D: Image-Guided Controllable Interaction Generation},
author={Sai Kumar Dwivedi and Federica Bogo and Buฤra Tekin and Chenhongyi Yang and Nadine Bertsch and Tomas Hodan and Michael J. Black and Dimitrios Tzionas and Shreyas Hampali},
year={2026},
eprint={2606.23675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.23675},
}
@misc{lee2026mvtrack4genmultiviewpointtracking,
title={MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation},
author={JoungBin Lee and Jaewoo Jung and Jongmin Lee and Tongmin Kim and Hyunsung Kim and Takuya Narihira and Kazumi Fukuda and Jahyeok Koo and Jisang Han and Yuki Mitsufuji and Seungryong Kim},
year={2026},
eprint={2606.26087},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.26087},
}
@article{zhao2026rynnworld,
title={RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation},
author={Zhao, Haoyu and Zhao, Xingyue and Huang, Siteng and Li, Xin and Zhao, Deli and Li, Zhongyu},
journal={arXiv preprint arXiv:2607.06559},
year={2026}
}
@article{feng2026skelgen4d,
title={SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation},
author={Feng, Hao and Zuo, Zhi and Pan, Jia-Hui and Hui, Ka-Hei and Liu, Zhengzhe and Zhang, Dian and Xie, Haoran and Sheng, Bin and Hu, Jingyu},
journal={arXiv preprint arXiv:2607.08246},
year={2026}
}
@article{wang2026hallo4d,
title={Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation},
author={Wang, Hongbo and Huang, Huaibo and Cao, Jie and Liu, Jin and Tong, Haoyang and He, Ran},
journal={arXiv preprint arXiv:2607.12752},
year={2026}
}
@article{bai2026pe,
title={PE-Field 4D: Video Generation Models as Canvas},
author={Bai, Yunpeng and Li, Haoxiang and Huang, Qixing},
journal={arXiv preprint arXiv:2607.15667},
year={2026}
}
@misc{liu2026pixelsvideopriors4d,
title={Beyond Pixels: From Video Priors to 4D Worlds},
author={Zihao Liu and Xiaolong Shen and Zhenglin Zhou and Ruijie Quan and Yi Yang},
year={2026},
eprint={2608.10744},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.10744},
}
@misc{ban2026stream4d4dconsistencystreamingautoregressive,
title={Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models},
author={Yuanhao Ban and Jiaqi Feng and Hengguang Zhou and Xiaohuan Pei and Justin Cui and Cho-Jui Hsieh},
year={2026},
eprint={2608.19556},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.19556},
}
@article{li20264dstreamctrl,
title={4DStreamCtrl: Interactive Video Generation with Online 4D Control},
author={Li, Shiqian and Lin, Chenguo and Liu, Zhiguang and Tang, Yu and Ou, Jiarong and Chen, Rui and Zhu, Yixin},
journal={arXiv preprint arXiv:2608.25479},
year={2026}
}
@misc{lu2026gaelearninggeometrynativelatent,
title={GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation},
author={Jiahao Lu and Minghao Yin and Wenbo Hu and Hengyu Liu and Wang Zhao and Sai-Kit Yeung and Ying Shan and Yuan Liu},
year={2026},
eprint={2609.24981},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.24981},
}
For more details, please check the 2025 4D Papers, including 35 accepted papers, 6 arXiv papers and 2 arXiv surveys.
For more details, please check the 2024 4D Papers, including 27 accepted papers and 7 arXiv papers.
In 2023, tasks classified as text/Image to 4D and video to 4D generally involve producing four-dimensional data from text/Image or video input. For more details, please check the 2023 4D Papers, including 6 accepted papers and 3 arXiv papers.
Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
(CUHK-SZ, SLAI, NUS, CUHK, HKUST, HKUST-GZ, NVIDIA, UCLA, MSRA)
| Year | Title | ArXiv Time | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2026 | SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | 2 Sept 2026 | Link | Link | Link |
%axiv papers
@misc{huang2026solarwmopendatascalable,
title={SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models},
author={Junchao Huang and Guian Fang and Shengju Qian and Xianghao Kong and Zhuoran Zhao and Wei Huang and Yihua Du and Zixin Zhang and Justin Cui and Yuchao Gu and Yukang Chen and Xinting Hu and Tianyu He and Shaoshuai Shi and Zhuotao Tian and Xin Wang and Mike Zheng Shou and Li Jiang},
year={2026},
eprint={2609.02886},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.02886},
}
For more details, please check the 2025 T2V Papers, including 16 accepted papers and 11 arXiv papers.
For more details, please check the 2024 T2V Papers, including 25 accepted papers and 3 arXiv papers.
| Year | Title | Venue | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2025 | PanoDreamer: Optimization-Based Single Image to 360 3D Scene With Diffusion | SIGGRAPH Asia 2025 | Link | Link | Link |
%accepted papers
@inproceedings{paliwal2025panodreamer,
title={PanoDreamer: Optimization-Based Single Image to 360 3D Scene With Diffusion},
author={Paliwal, Avinash and Zhou, Xilong and Tsarov, Andrii and Kalantari, Nima},
booktitle={Proceedings of the SIGGRAPH Asia 2025 Conference Papers},
pages={1--10},
year={2025}
}
For more details, please check the 2025 3D Scene Papers, including 15 accepted papers, 10 arXiv papers and 2 arXiv surveys.
For more details, please check the 2023-2024 3D Scene Papers, including 25 accepted papers and 6 arXiv papers.
Awesome Repos
For more details, please check the 2025 Human Motion Papers, including 18 accepted papers and 4 arXiv papers.
For more details, please check the 2023-2024 Text to Human Motion Papers, including 37 accepted papers and 5 arXiv papers.
| Motion | Info | URL | Others |
|---|---|---|---|
| AIST | AIST Dance Motion Dataset | Link | -- |
| AIST++ | AIST++ Dance Motion Dataset | Link | dance video database with SMPL annotations |
| AMASS | optical marker-based motion capture datasets | Link | -- |
AMASS is a large database of human motion unifying different optical marker-based motion capture datasets by representing them within a common framework and parameterization. AMASS is readily useful for animation, visualization, and generating training data for deep learning.
Survey
For more details, please check the 2025 3D Human Papers, including 12 accepted papers and 1 arXiv papers.
For more details, please check the 2023-2024 3D Human Papers, including 22 accepted papers and 1 arXiv papers.
| Pretrained Models (human body) | Info | URL |
|---|---|---|
| SMPL | smpl model (smpl weights) | Link |
| SMPL-X | smpl model (smpl weights) | Link |
| human_body_prior | vposer model (smpl weights) | Link |
SMPL is an easy-to-use, realistic, model of the of the human body that is useful for animation and computer vision.
SMPL-X, that extends SMPL with fully articulated hands and facial expressions (55 joints, 10475 vertices)
๐ฏBack to Top - Text2X Resources
Here, other tasks refer to CAD, 3D modeling, music generation, and so on.
Text to CAD
Text to Music
Text to Model
Survey
Awesome Repos
Survey
Awesome Repos
Benchmark
Foundation Model
Survey
Awesome Repos
Survey
Neural Deformable 3D Gaussians
4D Gaussians
Dynamic 3D Gaussians
๐ฏBack to Top - Table of Contents
This repo is released under the MIT license.
โ๏ธ Any additions or suggestions, feel free to contact us.
[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the Real World: A Unified Survey of Multimodal Generative Models" (IEEE TPAMI, 2026) and Awesome-Text2X-Resources. Watch this repository for the latest updates! ๐ฅ
See the codeThis repository is divided into two main sections:
Our Survey Paper Collection - This section presents our survey, "Simulating the Real World: A Unified Survey of Multimodal Generative Models" (IEEE TPAMI, 2026), which systematically unify the study of 2D, video, 3D and 4D generation within a single framework.
Text2X Resources โ This section continues the original Awesome-Text2X-Resources, an open collection of state-of-the-art (SOTA) and novel Text-to-X (X can be everything) methods, including papers, codes, and datasets. The goal is to track the rapid progress in this field and provide researchers with up-to-date references.
โญ If you find this repository useful for your research or work, a star is highly appreciated!
๐ This repository is continuously updated. If you find relevant papers, blog posts, videos, or other resources that should be included, feel free to submit a pull request (PR) or open an issue. Community contributions are always welcome!
๐๐ข๐ฆ๐ฎ๐ฅ๐๐ญ๐ข๐ง๐ ๐ญ๐ก๐ ๐๐๐๐ฅ ๐๐จ๐ซ๐ฅ๐: ๐ ๐๐ง๐ข๐๐ข๐๐ ๐๐ฎ๐ซ๐ฏ๐๐ฒ ๐จ๐ ๐๐ฎ๐ฅ๐ญ๐ข๐ฆ๐จ๐๐๐ฅ ๐๐๐ง๐๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐จ๐๐๐ฅ๐ฌ
๐ฐ๐ฌ๐ฌ๐ฌ ๐ป๐ท๐จ๐ด๐ฐ, ๐๐๐๐
Abstract
Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the physical world, enabling more accurate simulations and meaningful interactions. However, current methods often treat different modalities, including 2D (images), videos, 3D, and 4D representations, as independent domains, overlooking their interdependencies. Additionally, these methods typically focus on isolated dimensions of reality without systematically integrating their connections. In this survey, we present a unified survey for multimodal generative models that investigate the progression of data dimensionality in real-world simulation. Specifically, this survey starts from 2D generation (appearance), then moves to video (appearance+dynamics) and 3D generation (appearance+geometry), and finally culminates in 4D generation that integrate all dimensions. To the best of our knowledge, this is the first attempt to systematically unify the study of 2D, video, 3D and 4D generation within a single framework. To guide future research, we provide a comprehensive review of datasets, evaluation metrics and future directions, and fostering insights for newcomers. This survey serves as a bridge to advance the study of multimodal generative models and real-world simulation within a unified framework.
โญ Citation
If you find this paper and repo helpful for your research, please cite it below:
@article{hu2026simulating,
title={Simulating the real world: A unified survey of multimodal generative models},
author={Hu, Yuqi and Wang, Longguang and Liu, Xian and Chen, Ling-Hao and Guo, Yuwei and Shi, Yukai and Liu, Ce and Rao, Anyi and Wang, Zeyu and Xiong, Hui},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
year={2026},
publisher={IEEE}
}
๐งญ Getting Started with Key Concepts
[!Note] If you are new to this field, you can find clear and concise definitions of essential technical terms and concepts, such as NeRF, 3DGS, SDS, and Diffusion Models in our Glossary.
[!TIP] Feel free to pull requests or contact us if you find any related papers that are not included here. The process to submit a pull request is as follows:
README.md using the following format:[Origin] **Paper Title** [[Paper](Paper Link)] [[GitHub](GitHub Link)] [[Project Page](Project Page Link)]
We present a unified framework connecting 2D, Video, 3D, and 4D generation through text-guided synthesis. This paradigm illustrates how higher-dimensional content is synthesized by extending foundational modalities along spatial and temporal axes. (1)2D->3D: Spatial lifting of 2D priors to achieve geometric consistency; (2)2D->Video: Temporal inflation of static features to capture motion dynamics; (3)Video->4D: Spatial reconstruction and stabilization of dynamic sequences; (4)3D->4D: Temporal animation and deformation of static geometry. This perspective underscores that higher-dimensional generation methodologies are derivatives of foundational lower-dimensional generative priors, adapted through specialized architectural extensions.
Here are some seminal papers and models.
Overview of text-to-video generation technologies categorized by three main approaches.
Text-to-video generation models adapt text-to-image frameworks to handle the additional dimension of dynamics in the real world. We classify these models into three categories based on different generative machine learning architectures.
Survey
(1) VAE- and GAN-based Approaches.
VAE-based Approaches.
GAN-based Approaches.
(2) Diffusion-based Approaches.
U-Net-based Architectures.
Transformer-based Architectures.
(3) Autoregressive-based Approaches.
Video Editing.
Novel View Synthesis.
Human Animation in Videos.
Recent text-to-3D, image-to-3D and video-to-3D generation methods.
Survey
Feedforward Approaches.
Optimization-based Approaches.
MVS-based Approaches.
Feedforward Approaches.
Optimization-based Approaches.
MVS-based Approaches.
Avatar Generation.
Scene Generation.
3D Editing.
Representative works of 4D generation methods. ''Rep'' stands for representations.
Feedforward Approaches.
Optimization-based Approaches.
4D Editing.
Human Animation.
Summary of the widely-used 2D, video, 3D and 4D generation datasets. [Link] directs to dataset websites.
Summary of common evaluation metrics.
Atlas: the world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D, World Labs, Sept 1st, 2026.
WorldFoundry([GitHub]): an open-source infrastructure for world models.
NVIDIA Cosmos ([GitHub] [Paper]): NVIDIA Cosmos is a world foundation model platform for accelerating the development of physical AI systems.
Genie3, Google Deepmind, August 5th, 2025.
๐ฏBack to Top - Our Survey Paper Collection
An open collection of state-of-the-art (SOTA), novel Text to X (X can be everything) methods (papers, codes and datasets), intended to keep pace with the anticipated surge of research.
2026.09.04 - rename previous T2V Subsection to Video Generation Subsection.2026.01.07 - update 2025 papers collection into docs.2025.12.03 - update several papers accepted by NeurIPS 2025, congrats to all ๐2025.05.08 - update new layout.2025.04.18 - update layout on section Related Resources.2025.03.10 - CVPR 2025 Accepted Papers๐2025.02.28 - update several papers status "CVPR 2025" to accepted papers, congrats to all ๐2025.01.23 - update several papers status "ICLR 2025" to accepted papers, congrats to all ๐2025.01.09 - update layout.2024.12.21 adjusted the layouts of several sections and Happy Winter Solstice โช๐ฅฃ.2024.09.26 - update several papers status "NeurIPS 2024" to accepted papers, congrats to all ๐2024.09.03 - add one new section 'text to model'.2024.06.30 - add one new section 'text to video'.2024.07.02 - update several papers status "ECCV 2024" to accepted papers, congrats to all ๐2024.06.21 - add one hot Topic about AIGC 4D Generation on the section of Suvery and Awesome Repos.2024.06.17 - an awesome repo for CVPR2024 Link ๐๐ป2024.04.05 adjusted the layout and added accepted lists and ArXiv lists to each section.2024.04.05 - an awesome repo for CVPR2024 on 3DGS and NeRF Link ๐๐ป2024.03.25 - add one new survey paper of 3D GS into the section of "Survey and Awesome Repos--Topic 1: 3D Gaussian Splatting".2024.03.12 - add a new section "Dynamic Gaussian Splatting", including Neural Deformable 3D Gaussians, 4D Gaussians, Dynamic 3D Gaussians.2024.03.11 - CVPR 2024 Accpeted Papers Link| Year | Title | Venue | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2026 | Geometry-aware 4D Video Generation for Robot Manipulation | ICLR 2026 | Link | Link | Link |
| 2026 | Turbo4DGen: Ultra-Fast Acceleration for 4D Generation | ICML 2026 | Link | Link | Link |
| 2026 | Code2Worlds: Empowering Coding LLMs for 4D World Generation | ICML 2026 | Link | Link | Link |
| 2026 | PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation | CVPR 2026 | Link | Link | Link |
| 2026 | AvatarPointillist: Autoregressive 4D Gaussian Avatarization | CVPR 2026 | Link | Link | Link |
| 2026 | Vista4D: Video Reshooting with 4D Point Clouds | CVPR 2026 | Link | Link | Link |
| 2026 | Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis | CVPR 2026 | Link | Link | Link |
| 2026 | NeuROK: Generative 4D Neural Object Kinematics | CVPR 2026 | Link | Coming Soon! | Link |
| 2026 | Choreographing a World of Dynamic Objects | CVPR 2026 | Link | Link | Link |
| 2026 | ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion | CVPR 2026 | Link | Link | Link |
| 2026 | MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing | ECCV 2026 | Link | Coming Soon! | Link |
| 2026 | VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward | ECCV 2026 | Link | -- | Link |
| 2026 | LivingWorld: Interactive 4D World Generation with Environmental Dynamics | ECCV 2026 | Link | Link | Link |
| 2026 | InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation | ECCV 2026 | Link | Datasets | Link |
| 2026 | MoGe4D: Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation | ECCV 2026 | Link | Link | Link |
| 2026 | Alignment Is All You Need For X-to-4D Generation | IEEE Transactions on Multimedia (TMM) 2026 | Link | -- | Link |
| 2026 | Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild | SIGGRAPH Asia 2026 | Link | Link | Link |
| 2026 | 4DAnyone: Create Anyone in 4D from a Casual Monocular Video | SIGGRAPH Asia 2026 | Link | Link | Link |
| 2026 | Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction | CVPR 2026 4DV Workshop | Link | -- | -- |
%accepted papers
@article{liu2025geometry,
title={Geometry-aware 4D Video Generation for Robot Manipulation},
author={Liu, Zeyi and Li, Shuang and Cousineau, Eric and Feng, Siyuan and Burchfiel, Benjamin and Song, Shuran},
journal={arXiv preprint arXiv:2507.01099},
year={2025}
}
@article{man2026turbo4dgen,
title={Turbo4DGen: Ultra-Fast Acceleration for 4D Generation},
author={Man, Yuanbin and Huang, Ying and Ren, Zhile and Yin, Miao},
journal={arXiv preprint arXiv:2603.29572},
year={2026}
}
@article{zhang2026code2worlds,
title={Code2worlds: Empowering coding llms for 4d world generation},
author={Zhang, Yi and Wang, Yunshuang and Zhang, Zeyu and Tang, Hao},
journal={arXiv preprint arXiv:2602.11757},
year={2026}
}
@article{zhan2026perpetualwonder,
title={PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation},
author={Zhan, Jiahao and Li, Zizhang and Yu, Hong-Xing and Wu, Jiajun},
journal={arXiv preprint arXiv:2602.04876},
year={2026}
}
@inproceedings{liu2026avatarpointillist,
title = {AvatarPointillist: Autoregressive 4D Gaussian Avatarization},
author = {Hongyu Liu and Xuan Wang and Yating Wang and Zijian Wu and Ziyu Wan and Yue Ma and Runtao Liu and Boyao Zhou and Yujun Shen and Qifeng Chen},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2026}
}
@misc{lin2026vista4dvideoreshooting4d,
title={Vista4D: Video Reshooting with 4D Point Clouds},
author={Kuan Heng Lin and Zhizheng Liu and Pablo Salamanca and Yash Kant and Ryan Burgert and Yuancheng Xu and Koichi Namekata and Yiwei Zhao and Bolei Zhou and Micah Goldblum and Paul Debevec and Ning Yu},
year={2026},
eprint={2604.21915},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.21915},
}
@inproceedings{chen2026motion,
title={Motion 3-to-4: 3d motion reconstruction for 4d synthesis},
author={Chen, Hongyuan and Chen, Xingyu and Xu, Zexiang and Chen, Anpei},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={28947--28958},
year={2026}
}
@inproceedings{geng2026neurok,
title={NeuROK: Generative 4D Neural Object Kinematics},
author={Geng, Chen and He, Guangzhao and Gao, Yue and Zhang, Yunzhi and Wu, Shangzhe and Wu, Jiajun},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={39239--39251},
year={2026}
}
@article{lyu2026choreographing,
title={Choreographing a World of Dynamic Objects},
author={Lyu, Yanzhe and Geng, Chen and Dharmarajan, Karthik and Zhang, Yunzhi and Alzayer, Hadi and Wu, Shangzhe and Wu, Jiajun},
journal={arXiv preprint arXiv:2601.04194},
year={2026}
}
@article{sabathier2026actionmesh,
title={Actionmesh: Animated 3d mesh generation with temporal 3d diffusion},
author={Sabathier, Remy and Novotny, David and Mitra, Niloy J and Monnier, Tom},
journal={arXiv preprint arXiv:2601.16148},
year={2026}
}
@article{fiebelman2026mv,
title={MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing},
author={Fiebelman, Gal and Averbuch-Elor, Hadar and Benaim, Sagie},
journal={arXiv preprint arXiv:2607.05376},
year={2026}
}
@article{an2026vggrpo,
title={Vggrpo: Towards world-consistent video generation with 4d latent reward},
author={An, Zhaochong and Kupyn, Orest and Uscidda, Th{\'e}o and Colaco, Andrea and Ahuja, Karan and Belongie, Serge and Gonzalez-Franco, Mar and Gazulla, Marta Tintore},
journal={arXiv preprint arXiv:2603.26599},
year={2026}
}
@article{mun2026livingworld,
title={LivingWorld: Interactive 4D World Generation with Environmental Dynamics},
author={Mun, Hyeongju and Jin, In-Hwan and Kim, Sohyeong and Kong, Kyeongbo},
journal={arXiv preprint arXiv:2604.01641},
year={2026}
}
@article{peng2026interpet4d,
title={InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation},
author={Peng, Yichen and Song, Jyun-Ting and Liao, Chen-Chieh and Kitani, Kris and Koike, Hideki and Wu, Erwin},
journal={arXiv preprint arXiv:2607.10287},
year={2026}
}
@misc{zhang2026geometryawaresingleimage4dsynthesis,
title={Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation},
author={Yanran Zhang and Ziyi Wang and Wenzhao Zheng and Zheng Zhu and Jie Zhou and Jiwen Lu},
year={2026},
eprint={2512.05044},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.05044},
}
@article{miao2026alignment,
title={Alignment Is All You Need For X-to-4D Generation},
author={Miao, Qiaowei and Li, Kehan and Luo, Yawei and Yang, Yi},
journal={arXiv preprint arXiv:2607.02516},
year={2026}
}
@article{litman2026lift4d,
title={Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild},
author={Litman, Yehonathan and Ma, Xiaoxuan and Shah, Manan and Ugrinovic, Nicolas and Kitani, Kris and De la Torre, Fernando and Tulsiani, Shubham},
journal={arXiv preprint arXiv:2606.23688},
year={2026}
}
@article{jin2026fdanyone,
title={4DAnyone: Create Anyone in 4D from a Casual Monocular Video},
author={Jin, Yudong and Xie, Tao and Zhang, Qihang and Shen, Zehong and Xu, Zhen and Shen, Yujun and Bao, Hujun and Zhou, Xiaowei and Xu, Yinghao},
journal={arXiv preprint arXiv:2608.20335},
year={2026},
url={https://arxiv.org/abs/2608.20335}
}
@article{liu2026streaming4d,
title={Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction},
author={Liu, Xiaoyan and Liu, Jiaxin and Li, Kangrui and Zhou, Sifan},
journal={arXiv preprint arXiv:2609.00610},
year={2026}
}
Melonie de Almeida, Daniela Ivanova, Tong Shi, John H. Williamson, Paul Henderson (University of Glasgow)
Haonan Wang, Hanyu Zhou, Tao Gu, Luxin Yan
(Huazhong University of Science and Technology, National University of Singapore, Macquarie University)
Sunwoo Park, Taesung Kwon, Jong Chul Ye (KAIST AI)
Dvir Samuel, Yuval Atzmon, Gal Chechik, Yoni Kasten (NVIDIA Research, Bar-Ilan University)
Jiraphon Yenphraphai, Jianqi Chen, Jian Wang, Gordon Qian, Sergey Tulyakov, Rameen Abdal, Raymond A. Yeh, Peter Wonka, Chaoyang Wang
(Snap, Purdue University, KAUST)
Yiran Wang, Zeyu Zhang, Yuanming Li, Ziming Wang, Yang Zhao
(USYD, SpatialReal, ZJU, La Trobe)
Sai Kumar Dwivedi, Federica Bogo, Buฤra Tekin, Chenhongyi Yang, Nadine Bertsch, Tomas Hodan, Michael J. Black, Dimitrios Tzionas, Shreyas Hampali
(Meta, Max Planck Institute for Intelligent Systems, University of Amsterdam, Aristotle University of Thessaloniki)
JoungBin Lee, Jaewoo Jung, Jongmin Lee, Tongmin Kim, Hyunsung Kim, Takuya Narihira, Kazumi Fukuda, Jahyeok Koo, Jisang Han, Yuki Mitsufuji, Seungryong Kim
(KAIST AI, Sony AI, Sony Group Corporation)
Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li
(DAMO Academy Alibaba Group, Hong Kong Embodied AI Lab, CUHK, Hupan Lab)
Hao Feng, Zhi Zuo, Jia-Hui Pan, Ka-Hei Hui, Zhengzhe Liu, Dian Zhang, Haoran Xie, Bin Sheng, Jingyu Hu
(Lingnan University, Chinese University of Hong Kong, Autodesk Research, Shanghai Jiao Tong University)
Hongbo Wang, Huaibo Huang, Jie Cao, Jin Liu, Haoyang Tong, Ran He
(CASIA, UCAS, ShanghaiTech)
Yunpeng Bai, Haoxiang Li, Qixing Huang (UT Austin, Pixocial Technology)
Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang (Zhejiang University)
Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
(UCLA, Tsinghua University)
Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
(Peking University, Tencent Hunyuan)
Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu
(The Hong Kong University of Science and Technology, ARC Lab Tencent IEG, The University of Hong Kong, The University of Texas at Austin)
| Year | Title | ArXiv Time | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2026 | Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians | 2 Jan 2026 | Link | -- | Link |
| 2026 | InSpatio-World | 20 Mar 2026 | Live Demo | Link | Link |
| 2026 | ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation | 8 May 2026 | Link | -- | -- |
| 2026 | Geometric 4D Stitching for Grounded 4D Generation | 11 May 2026 | Link | -- | -- |
| 2026 | Fast 4D Mesh Generation by Spatio-Temporal Attention Chains | 19 May 2026 | Link | -- | Link |
| 2026 | Helix4D: Complex 4D Mesh Generation | 25 May 2026 | Link | -- | Link |
| 2026 | SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction | 14 Jun 2026 | Link | -- | Link |
| 2026 | IMAGIN-4D: Image-Guided Controllable Interaction Generation | 22 Jun 2026 | Link | -- | Link |
| 2026 | MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation | 24 Jun 2026 | Link | Link | Link |
| 2026 | RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation | 7 Jul 2026 | Link | Link | Link |
| 2026 | SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation | 9 Jul 2026 | Link | -- | -- |
| 2026 | Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation | 15 Jul 2026 | Link | Link | Link |
| 2026 | PE-Field 4D: Video Generation Models as Canvas | 17 Jul 2026 | Link | -- | -- |
| 2026 | Beyond Pixels: From Video Priors to 4D Worlds | 11 Aug 2026 | Link | Link | Link |
| 2026 | Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models | 20 Aug 2026 | Link | -- | Link |
| 2026 | 4DStreamCtrl: Interactive Video Generation with Online 4D Control | 27 Aug 2026 | Link | -- | Link |
| 2026 | GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation | 21 Sep 2026 | Link | Link | Link |
%axiv papers
@article{de2026pixel,
title={Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians},
author={de Almeida, Melonie and Ivanova, Daniela and Shi, Tong and Williamson, John H and Henderson, Paul},
journal={arXiv preprint arXiv:2601.00678},
year={2026}
}
@misc{inspatio-world,
title={InSpatio-World},
author={InSpatio-World Contributors},
howpublished={\url{https://github.com/inspatio/inspatio-world}},
year={2025}
}
@article{wang2026st,
title={ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation},
author={Wang, Haonan and Zhou, Hanyu and Gu, Tao and Yan, Luxin},
journal={arXiv preprint arXiv:2605.07390},
year={2026}
}
@article{park2026geometric,
title={Geometric 4D Stitching for Grounded 4D Generation},
author={Park, Sunwoo and Kwon, Taesung and Ye, Jong Chul},
journal={arXiv preprint arXiv:2605.09984},
year={2026}
}
@article{samuel2026fast,
title={Fast 4D Mesh Generation by Spatio-Temporal Attention Chains},
author={Samuel, Dvir and Atzmon, Yuval and Chechik, Gal and Kasten, Yoni},
journal={arXiv preprint arXiv:2605.19786},
year={2026}
}
@article{yenphraphai2026helix4d,
title={Helix4D: Complex 4D Mesh Generation},
author={Yenphraphai, Jiraphon and Chen, Jianqi and Wang, Jian and Qian, Gordon and Tulyakov, Sergey and Abdal, Rameen and Yeh, Raymond A and Wonka, Peter and Wang, Chaoyang},
journal={arXiv preprint arXiv:2605.26109},
year={2026}
}
@article{wang2026spatialavatar,
title={SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction},
author={Wang, Yiran and Zhang, Zeyu and Li, Yuanming and Wang, Ziming and Zhao, Yang},
journal={arXiv preprint arXiv:2606.15659},
year={2026}
}
@misc{dwivedi2026imagin4dimageguidedcontrollableinteraction,
title={IMAGIN-4D: Image-Guided Controllable Interaction Generation},
author={Sai Kumar Dwivedi and Federica Bogo and Buฤra Tekin and Chenhongyi Yang and Nadine Bertsch and Tomas Hodan and Michael J. Black and Dimitrios Tzionas and Shreyas Hampali},
year={2026},
eprint={2606.23675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.23675},
}
@misc{lee2026mvtrack4genmultiviewpointtracking,
title={MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation},
author={JoungBin Lee and Jaewoo Jung and Jongmin Lee and Tongmin Kim and Hyunsung Kim and Takuya Narihira and Kazumi Fukuda and Jahyeok Koo and Jisang Han and Yuki Mitsufuji and Seungryong Kim},
year={2026},
eprint={2606.26087},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.26087},
}
@article{zhao2026rynnworld,
title={RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation},
author={Zhao, Haoyu and Zhao, Xingyue and Huang, Siteng and Li, Xin and Zhao, Deli and Li, Zhongyu},
journal={arXiv preprint arXiv:2607.06559},
year={2026}
}
@article{feng2026skelgen4d,
title={SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation},
author={Feng, Hao and Zuo, Zhi and Pan, Jia-Hui and Hui, Ka-Hei and Liu, Zhengzhe and Zhang, Dian and Xie, Haoran and Sheng, Bin and Hu, Jingyu},
journal={arXiv preprint arXiv:2607.08246},
year={2026}
}
@article{wang2026hallo4d,
title={Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation},
author={Wang, Hongbo and Huang, Huaibo and Cao, Jie and Liu, Jin and Tong, Haoyang and He, Ran},
journal={arXiv preprint arXiv:2607.12752},
year={2026}
}
@article{bai2026pe,
title={PE-Field 4D: Video Generation Models as Canvas},
author={Bai, Yunpeng and Li, Haoxiang and Huang, Qixing},
journal={arXiv preprint arXiv:2607.15667},
year={2026}
}
@misc{liu2026pixelsvideopriors4d,
title={Beyond Pixels: From Video Priors to 4D Worlds},
author={Zihao Liu and Xiaolong Shen and Zhenglin Zhou and Ruijie Quan and Yi Yang},
year={2026},
eprint={2608.10744},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.10744},
}
@misc{ban2026stream4d4dconsistencystreamingautoregressive,
title={Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models},
author={Yuanhao Ban and Jiaqi Feng and Hengguang Zhou and Xiaohuan Pei and Justin Cui and Cho-Jui Hsieh},
year={2026},
eprint={2608.19556},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.19556},
}
@article{li20264dstreamctrl,
title={4DStreamCtrl: Interactive Video Generation with Online 4D Control},
author={Li, Shiqian and Lin, Chenguo and Liu, Zhiguang and Tang, Yu and Ou, Jiarong and Chen, Rui and Zhu, Yixin},
journal={arXiv preprint arXiv:2608.25479},
year={2026}
}
@misc{lu2026gaelearninggeometrynativelatent,
title={GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation},
author={Jiahao Lu and Minghao Yin and Wenbo Hu and Hengyu Liu and Wang Zhao and Sai-Kit Yeung and Ying Shan and Yuan Liu},
year={2026},
eprint={2609.24981},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.24981},
}
For more details, please check the 2025 4D Papers, including 35 accepted papers, 6 arXiv papers and 2 arXiv surveys.
For more details, please check the 2024 4D Papers, including 27 accepted papers and 7 arXiv papers.
In 2023, tasks classified as text/Image to 4D and video to 4D generally involve producing four-dimensional data from text/Image or video input. For more details, please check the 2023 4D Papers, including 6 accepted papers and 3 arXiv papers.
Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
(CUHK-SZ, SLAI, NUS, CUHK, HKUST, HKUST-GZ, NVIDIA, UCLA, MSRA)
| Year | Title | ArXiv Time | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2026 | SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models | 2 Sept 2026 | Link | Link | Link |
%axiv papers
@misc{huang2026solarwmopendatascalable,
title={SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models},
author={Junchao Huang and Guian Fang and Shengju Qian and Xianghao Kong and Zhuoran Zhao and Wei Huang and Yihua Du and Zixin Zhang and Justin Cui and Yuchao Gu and Yukang Chen and Xinting Hu and Tianyu He and Shaoshuai Shi and Zhuotao Tian and Xin Wang and Mike Zheng Shou and Li Jiang},
year={2026},
eprint={2609.02886},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.02886},
}
For more details, please check the 2025 T2V Papers, including 16 accepted papers and 11 arXiv papers.
For more details, please check the 2024 T2V Papers, including 25 accepted papers and 3 arXiv papers.
| Year | Title | Venue | Paper | Code | Project Page |
|---|---|---|---|---|---|
| 2025 | PanoDreamer: Optimization-Based Single Image to 360 3D Scene With Diffusion | SIGGRAPH Asia 2025 | Link | Link | Link |
%accepted papers
@inproceedings{paliwal2025panodreamer,
title={PanoDreamer: Optimization-Based Single Image to 360 3D Scene With Diffusion},
author={Paliwal, Avinash and Zhou, Xilong and Tsarov, Andrii and Kalantari, Nima},
booktitle={Proceedings of the SIGGRAPH Asia 2025 Conference Papers},
pages={1--10},
year={2025}
}
For more details, please check the 2025 3D Scene Papers, including 15 accepted papers, 10 arXiv papers and 2 arXiv surveys.
For more details, please check the 2023-2024 3D Scene Papers, including 25 accepted papers and 6 arXiv papers.
Awesome Repos
For more details, please check the 2025 Human Motion Papers, including 18 accepted papers and 4 arXiv papers.
For more details, please check the 2023-2024 Text to Human Motion Papers, including 37 accepted papers and 5 arXiv papers.
| Motion | Info | URL | Others |
|---|---|---|---|
| AIST | AIST Dance Motion Dataset | Link | -- |
| AIST++ | AIST++ Dance Motion Dataset | Link | dance video database with SMPL annotations |
| AMASS | optical marker-based motion capture datasets | Link | -- |
AMASS is a large database of human motion unifying different optical marker-based motion capture datasets by representing them within a common framework and parameterization. AMASS is readily useful for animation, visualization, and generating training data for deep learning.
Survey
For more details, please check the 2025 3D Human Papers, including 12 accepted papers and 1 arXiv papers.
For more details, please check the 2023-2024 3D Human Papers, including 22 accepted papers and 1 arXiv papers.
| Pretrained Models (human body) | Info | URL |
|---|---|---|
| SMPL | smpl model (smpl weights) | Link |
| SMPL-X | smpl model (smpl weights) | Link |
| human_body_prior | vposer model (smpl weights) | Link |
SMPL is an easy-to-use, realistic, model of the of the human body that is useful for animation and computer vision.
SMPL-X, that extends SMPL with fully articulated hands and facial expressions (55 joints, 10475 vertices)
๐ฏBack to Top - Text2X Resources
Here, other tasks refer to CAD, 3D modeling, music generation, and so on.
Text to CAD
Text to Music
Text to Model
Survey
Awesome Repos
Survey
Awesome Repos
Benchmark
Foundation Model
Survey
Awesome Repos
Survey
Neural Deformable 3D Gaussians
4D Gaussians
Dynamic 3D Gaussians
๐ฏBack to Top - Table of Contents
This repo is released under the MIT license.
โ๏ธ Any additions or suggestions, feel free to contact us.