DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
12
123 commits
1 linked in READMEs
updated Jan 16, 2026
DynamicVerse is an integrated framework for dynamic scene understanding and 4D reconstruction. It combines advanced visual models such as Sa2VA, Qwen-VL, DAM, CameraBench, CoTracker, and UniDepth to achieve end-to-end processing from video to 4D scenes.
This repository hosts the processed datasets used in the DynamicVerse project. These data cover multiple mainstream dynamic scene datasets and have undergone keyframe extraction, multimodal analysis, dense segmentation, and 4D reconstruction processing.
| Attribute | Value |
|---|---|
| Total Scenes | 100k |
| Total Frames | 13.6 million |
| Storage Size | 3T |
| Scene Type | Mixed |
| Dynamic Type | Realistic |
| Real-world | Yes |
| Metric-scale | Yes |
The data storage structure of this repository is shown below. Data is classified according to the original dataset source and packaged into ZIP files.
For larger datasets (such as dynpose-100k), the data is distributed across multiple independent ZIP files (e.g., dynpose-0000.zip, dynpose-0001.zip, etc.). Each compressed package contains a specific number of independent Scenes.
βββ DynamicVerse
βββ DAVIS/ # Processed results for DAVIS dataset
β βββ DAVIS.zip
βββ dynamic_replica/ # Processed results for Dynamic Replica dataset
β βββ dynamic_replica-0000.zip
β βββ dynamic_replica-0001.zip
β βββ ...
β βββ dynamic_replica-0009.zip
βββ dynpose-100k/ # Processed results for DynPose-100k dataset
β βββ dynpose-0000.zip
β βββ dynpose-0001.zip
β βββ ...
β βββ dynpose-0089.zip
βββ MOSE/ # Processed results for MOSE dataset
β βββ MOSE-0000.zip
β βββ MOSE-0001.zip
βββ MVS-Synth/ # Processed results for MVS-Synth dataset
β βββ MVS-Synth.zip
βββ point_odyssey/ # Processed results for Point Odyssey dataset
β βββ point_odyssey-0000.zip
β βββ point_odyssey-0001.zip
βββ SAV/ # Processed results for SAV dataset
β βββ sav.zip
βββ spring/ # Processed results for Spring dataset
β βββ spring.zip
βββ uvo/ # Processed results for UVO dataset
β βββ uvo.zip
βββ VOST/ # Processed results for VOST dataset
β βββ VOST.zip
βββ youtube_vis/ # Processed results for YouTube-VIS dataset
β βββ youtube_vis.zip
βββ README.md
After decompressing the above ZIP files, each scene directory contains standardized data as follows:
<scene_id>/
βββ camera.npz # Camera parameters (poses, intrinsics)
βββ captions/ # Multimodal description files
β βββ camera_caption.json # Description of camera motion
β βββ object_caption.json # Description of objects
β βββ scene_caption.json # Overall scene description
βββ category/ # Category information
β βββ category.json
βββ depths/ # Depth map sequence (16-bit PNG, Metric Scale)
β βββ 00001.png
β βββ 00002.png
β βββ ... (n png files)
βββ mask/ # Segmentation mask sequence
β βββ 00001.png
β βββ 00002.png
β βββ ... (n png files)
βββ rgb/ # RGB image sequence
βββ 00001.jpg
βββ 00002.jpg
βββ ... (n jpg files)
You can use the Hugging Face CLI or directly download the required ZIP files. Since the files are independent, you can download parts of the data as needed.
The official repository provides the pipeline code for data generation. Here's a quick start guide to set up the environment and run the DynamicGen demo for processing a complete geometric scene pipeline.
git clone --recurse-submodules https://github.com/Dynamics-X/DynamicVerse.git
cd DynamicVerse
conda create -n dynamicverse python=3.10
conda activate dynamicverse
bash scripts/install.sh
bash scripts/download_weights.sh
This script will automatically download the following models:
Process a complete geometric scene pipeline:
cd dynamicgen
bash scripts/run_pipeline_demo.sh '' -all
This script executes the following steps:
For more detailed usage and configuration, including local Qwen2.5-VL deployment, refer to the GitHub repository.
If you find our dataset useful in your research, please citing the following paper:
@misc{wen2025dynamicverse,
title={DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling},
author={Kairun Wen and Yuzhi Huang and Runyu Chen and Hui Zheng and Yunlong Lin and Panwang Pan and Chenxin Li and Wenyan Cong and Jian Zhang and Junbin Lu and Chenguo Lin and Dilin Wang and Zhicheng Yan and Hongyu Xu and Justin Theiss and Yue Huang and Xinghao Ding and Rakesh Ranjan and Zhiwen Fan},
year={2025},
eprint={2512.03000},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.03000},
}
Apache-2.0
DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
12
123 commits
1 linked in READMEs
updated Jan 16, 2026
DynamicVerse is an integrated framework for dynamic scene understanding and 4D reconstruction. It combines advanced visual models such as Sa2VA, Qwen-VL, DAM, CameraBench, CoTracker, and UniDepth to achieve end-to-end processing from video to 4D scenes.
This repository hosts the processed datasets used in the DynamicVerse project. These data cover multiple mainstream dynamic scene datasets and have undergone keyframe extraction, multimodal analysis, dense segmentation, and 4D reconstruction processing.
| Attribute | Value |
|---|---|
| Total Scenes | 100k |
| Total Frames | 13.6 million |
| Storage Size | 3T |
| Scene Type | Mixed |
| Dynamic Type | Realistic |
| Real-world | Yes |
| Metric-scale | Yes |
The data storage structure of this repository is shown below. Data is classified according to the original dataset source and packaged into ZIP files.
For larger datasets (such as dynpose-100k), the data is distributed across multiple independent ZIP files (e.g., dynpose-0000.zip, dynpose-0001.zip, etc.). Each compressed package contains a specific number of independent Scenes.
βββ DynamicVerse
βββ DAVIS/ # Processed results for DAVIS dataset
β βββ DAVIS.zip
βββ dynamic_replica/ # Processed results for Dynamic Replica dataset
β βββ dynamic_replica-0000.zip
β βββ dynamic_replica-0001.zip
β βββ ...
β βββ dynamic_replica-0009.zip
βββ dynpose-100k/ # Processed results for DynPose-100k dataset
β βββ dynpose-0000.zip
β βββ dynpose-0001.zip
β βββ ...
β βββ dynpose-0089.zip
βββ MOSE/ # Processed results for MOSE dataset
β βββ MOSE-0000.zip
β βββ MOSE-0001.zip
βββ MVS-Synth/ # Processed results for MVS-Synth dataset
β βββ MVS-Synth.zip
βββ point_odyssey/ # Processed results for Point Odyssey dataset
β βββ point_odyssey-0000.zip
β βββ point_odyssey-0001.zip
βββ SAV/ # Processed results for SAV dataset
β βββ sav.zip
βββ spring/ # Processed results for Spring dataset
β βββ spring.zip
βββ uvo/ # Processed results for UVO dataset
β βββ uvo.zip
βββ VOST/ # Processed results for VOST dataset
β βββ VOST.zip
βββ youtube_vis/ # Processed results for YouTube-VIS dataset
β βββ youtube_vis.zip
βββ README.md
After decompressing the above ZIP files, each scene directory contains standardized data as follows:
<scene_id>/
βββ camera.npz # Camera parameters (poses, intrinsics)
βββ captions/ # Multimodal description files
β βββ camera_caption.json # Description of camera motion
β βββ object_caption.json # Description of objects
β βββ scene_caption.json # Overall scene description
βββ category/ # Category information
β βββ category.json
βββ depths/ # Depth map sequence (16-bit PNG, Metric Scale)
β βββ 00001.png
β βββ 00002.png
β βββ ... (n png files)
βββ mask/ # Segmentation mask sequence
β βββ 00001.png
β βββ 00002.png
β βββ ... (n png files)
βββ rgb/ # RGB image sequence
βββ 00001.jpg
βββ 00002.jpg
βββ ... (n jpg files)
You can use the Hugging Face CLI or directly download the required ZIP files. Since the files are independent, you can download parts of the data as needed.
The official repository provides the pipeline code for data generation. Here's a quick start guide to set up the environment and run the DynamicGen demo for processing a complete geometric scene pipeline.
git clone --recurse-submodules https://github.com/Dynamics-X/DynamicVerse.git
cd DynamicVerse
conda create -n dynamicverse python=3.10
conda activate dynamicverse
bash scripts/install.sh
bash scripts/download_weights.sh
This script will automatically download the following models:
Process a complete geometric scene pipeline:
cd dynamicgen
bash scripts/run_pipeline_demo.sh '' -all
This script executes the following steps:
For more detailed usage and configuration, including local Qwen2.5-VL deployment, refer to the GitHub repository.
If you find our dataset useful in your research, please citing the following paper:
@misc{wen2025dynamicverse,
title={DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling},
author={Kairun Wen and Yuzhi Huang and Runyu Chen and Hui Zheng and Yunlong Lin and Panwang Pan and Chenxin Li and Wenyan Cong and Jian Zhang and Junbin Lu and Chenguo Lin and Dilin Wang and Zhicheng Yan and Hongyu Xu and Justin Theiss and Yue Huang and Xinghao Ding and Rakesh Ranjan and Zhiwen Fan},
year={2025},
eprint={2512.03000},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.03000},
}
Apache-2.0