[NeurIPS 2025]"DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling"
See the codehttps://github.com/user-attachments/assets/61c818e1-9258-4bb1-9f9d-c656dfd5ce8a
DynamicVerse is an integrated framework for dynamic scene understanding and 4D reconstruction, combining advanced visual models such as Sa2VA, Qwen-VL, DAM, CameraBench, CoTracker, and UniDepth to achieve end-to-end processing from video to 4D scenes.
DynamicVerse/
βββ dynamicBA/ # 4D scene reconstruction module
β βββ unimatch/ # Optical flow and depth estimation
β βββ dataset_prepare/ # Data preprocessing tools
β βββ config/ # Configuration files
βββ data/ # Dataset directory
βββ scripts/ # Preprocessing scripts
βββ dynamicgen/ # Pipeline execution
β βββ scripts/ # DynamicGen pipeline
βββ Sa2VA/ # Vision-language multimodal model
βββ CoTracker/ # Point tracking model
βββ UniDepth/ # Monocular depth estimation
βββ ...
git clone --recurse-submodules https://github.com/Dynamics-X/DynamicVerse.git
cd DynamicVerse
conda create -n dynamicverse python=3.10
conda activate dynamicverse
bash scripts/install.sh
bash scripts/download_weights.sh
This script will automatically download the following models:
Process a complete geometric scene pipeline:
cd dynamicgen
bash scripts/run_pipeline_demo.sh '' -all
This script executes the following steps:
Qwen2.5-VL can be used in two ways:
For API service usage:
Set API Key: Set environment variable when running scripts
export DASHSCOPE_API_KEY=your_api_key
Or set it directly in dynamicgen/scripts/run_pipeline_demo.sh
Modify API Configuration: Edit dynamicgen/stage1_qwen.py
client = OpenAI(
api_key=api_key, # Use API key from environment variable
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1" # API service address
)
model="qvq-max-latest" # Or other Qwen models
For local deployment, modify dynamicgen/stage1_qwen.py to local service configuration:
client = OpenAI(
base_url="http://127.0.0.1:22002/v1", # Local service address
api_key="none" # Not needed for local service
)
# Specify model name
model="Qwen/Qwen2.5-VL-72B-Instruct"
Install Dependencies:
pip install accelerate
pip install qwen-vl-utils==0.0.14
uv pip install -U vllm # Requires vllm>=0.11.0
Start Local Service:
python -m vllm.entrypoints.openai.api_server \
--model <ckpt_path> \
--served-model-name Qwen/Qwen2.5-VL-72B-Instruct \
--tensor-parallel-size 4 \
--mm-encoder-tp-mode data \
--enable-expert-parallel \
--host 0.0.0.0 \
--port 22002 \
--dtype bfloat16 \
--gpu-memory-utilization 0.70 \
--quantization fp8 \
--distributed-executor-backend mp
For detailed deployment instructions, refer to Qwen-VL
Place videos or image sequences in the data/ directory
python motion_aware_key_frame_extract.py \
--input_root <input_path> \
--output_root <output_path> \
--flow_model 'unimatch'
python batch_process_qwen_pipeline.py \
<dataset_path> \
<output_path> \
--base_frame_dir <base_frame_dir> \
--key_frame_dir <key_frame_dir>
cd dynamicBA
python ./dynamicBA/run.py \
--config ./dynamicBA/config/config.yaml \
--experiment_name base \
--opt_intrinsics \
--workdir <workdir>
After processing, the following directory structure is generated:
data/
βββ key_frames/ # Keyframe extraction results
β βββ <dataset_name>/ # Dataset name
β βββ <scene_id>/ # Scene ID
β βββ frame_*.jpg
β βββ keyframe_info.json
βββ demo/ # Processed scene data
βββ <scene_id>/ # Scene ID directory
βββ videos/ # Original video files
β βββ <scene_id>.mp4
βββ rgb/ # Extracted RGB frames
β βββ 00001.jpg
β βββ 00002.jpg
β βββ ...
βββ analysis/ # Scene analysis results
β βββ dynamic_objects_<scene_id>.json # Dynamic object detection results
βββ qwen/ # Qwen model outputs
β βββ Annotations/ # Segmentation annotations
β βββ frame_00000.png
β βββ frame_00001.png
β βββ ...
βββ segmentation/ # Sa2VA segmentation results
β βββ frames/ # Frame-level segmentation results
β β βββ original/ # Original frames
β β βββ masks/ # Segmentation masks
β β βββ overlay/ # Overlay visualizations
β β βββ segmented/ # Segmented images
β βββ videos/ # Segmentation videos
β β βββ original.mp4 # Original video
β β βββ masks.mp4 # Mask video
β β βββ overlay.mp4 # Overlay video
β β βββ segmented.mp4 # Segmented video
β βββ instance_labels.json # Instance label information
β βββ result_summary.json # Segmentation result summary
βββ dynamicBA/ (Optional) # 4D reconstruction results
β βββ pose.npz # Intrinsic and Extrinsics
β βββ depth/ # Depth maps
β βββ flow/ # Optical flow data
βββ processing_log_<scene_id>.log # Processing log
We provide preprocessed datasets to reproduce Table 1 and 2 in our main paper.
You can also download our preprocessed data that we used for the quantitive results in our paper:
cd data
gdown https://drive.google.com/uc?id=1V1WIRvnJCJStL63rluwNZMPI2Gq4-yQy -O preprocessed.zip
unzip preprocessed.zip
We provide evaluation scripts for pose and depth metrics:
bash ./scripts/eval.sh
Our code is based on the following awesome repositories:
This project is built upon multiple open-source projects. Please refer to the license requirements of each submodule.
Issues and Pull Requests are welcome. Before submitting code, please ensure:
If you find our work useful in your research, please consider giving a star :star: and citing the following paper :pencil:.
@misc{wen2025dynamicverse,
title={DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling},
author={Kairun Wen and Yuzhi Huang and Runyu Chen and Hui Zheng and Yunlong Lin and Panwang Pan and Chenxin Li and Wenyan Cong and Jian Zhang and Junbin Lu and Chenguo Lin and Dilin Wang and Zhicheng Yan and Hongyu Xu and Justin Theiss and Yue Huang and Xinghao Ding and Rakesh Ranjan and Zhiwen Fan},
year={2025},
eprint={2512.03000},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.03000},
}
[NeurIPS 2025]"DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling"
See the codehttps://github.com/user-attachments/assets/61c818e1-9258-4bb1-9f9d-c656dfd5ce8a
DynamicVerse is an integrated framework for dynamic scene understanding and 4D reconstruction, combining advanced visual models such as Sa2VA, Qwen-VL, DAM, CameraBench, CoTracker, and UniDepth to achieve end-to-end processing from video to 4D scenes.
DynamicVerse/
βββ dynamicBA/ # 4D scene reconstruction module
β βββ unimatch/ # Optical flow and depth estimation
β βββ dataset_prepare/ # Data preprocessing tools
β βββ config/ # Configuration files
βββ data/ # Dataset directory
βββ scripts/ # Preprocessing scripts
βββ dynamicgen/ # Pipeline execution
β βββ scripts/ # DynamicGen pipeline
βββ Sa2VA/ # Vision-language multimodal model
βββ CoTracker/ # Point tracking model
βββ UniDepth/ # Monocular depth estimation
βββ ...
git clone --recurse-submodules https://github.com/Dynamics-X/DynamicVerse.git
cd DynamicVerse
conda create -n dynamicverse python=3.10
conda activate dynamicverse
bash scripts/install.sh
bash scripts/download_weights.sh
This script will automatically download the following models:
Process a complete geometric scene pipeline:
cd dynamicgen
bash scripts/run_pipeline_demo.sh '' -all
This script executes the following steps:
Qwen2.5-VL can be used in two ways:
For API service usage:
Set API Key: Set environment variable when running scripts
export DASHSCOPE_API_KEY=your_api_key
Or set it directly in dynamicgen/scripts/run_pipeline_demo.sh
Modify API Configuration: Edit dynamicgen/stage1_qwen.py
client = OpenAI(
api_key=api_key, # Use API key from environment variable
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1" # API service address
)
model="qvq-max-latest" # Or other Qwen models
For local deployment, modify dynamicgen/stage1_qwen.py to local service configuration:
client = OpenAI(
base_url="http://127.0.0.1:22002/v1", # Local service address
api_key="none" # Not needed for local service
)
# Specify model name
model="Qwen/Qwen2.5-VL-72B-Instruct"
Install Dependencies:
pip install accelerate
pip install qwen-vl-utils==0.0.14
uv pip install -U vllm # Requires vllm>=0.11.0
Start Local Service:
python -m vllm.entrypoints.openai.api_server \
--model <ckpt_path> \
--served-model-name Qwen/Qwen2.5-VL-72B-Instruct \
--tensor-parallel-size 4 \
--mm-encoder-tp-mode data \
--enable-expert-parallel \
--host 0.0.0.0 \
--port 22002 \
--dtype bfloat16 \
--gpu-memory-utilization 0.70 \
--quantization fp8 \
--distributed-executor-backend mp
For detailed deployment instructions, refer to Qwen-VL
Place videos or image sequences in the data/ directory
python motion_aware_key_frame_extract.py \
--input_root <input_path> \
--output_root <output_path> \
--flow_model 'unimatch'
python batch_process_qwen_pipeline.py \
<dataset_path> \
<output_path> \
--base_frame_dir <base_frame_dir> \
--key_frame_dir <key_frame_dir>
cd dynamicBA
python ./dynamicBA/run.py \
--config ./dynamicBA/config/config.yaml \
--experiment_name base \
--opt_intrinsics \
--workdir <workdir>
After processing, the following directory structure is generated:
data/
βββ key_frames/ # Keyframe extraction results
β βββ <dataset_name>/ # Dataset name
β βββ <scene_id>/ # Scene ID
β βββ frame_*.jpg
β βββ keyframe_info.json
βββ demo/ # Processed scene data
βββ <scene_id>/ # Scene ID directory
βββ videos/ # Original video files
β βββ <scene_id>.mp4
βββ rgb/ # Extracted RGB frames
β βββ 00001.jpg
β βββ 00002.jpg
β βββ ...
βββ analysis/ # Scene analysis results
β βββ dynamic_objects_<scene_id>.json # Dynamic object detection results
βββ qwen/ # Qwen model outputs
β βββ Annotations/ # Segmentation annotations
β βββ frame_00000.png
β βββ frame_00001.png
β βββ ...
βββ segmentation/ # Sa2VA segmentation results
β βββ frames/ # Frame-level segmentation results
β β βββ original/ # Original frames
β β βββ masks/ # Segmentation masks
β β βββ overlay/ # Overlay visualizations
β β βββ segmented/ # Segmented images
β βββ videos/ # Segmentation videos
β β βββ original.mp4 # Original video
β β βββ masks.mp4 # Mask video
β β βββ overlay.mp4 # Overlay video
β β βββ segmented.mp4 # Segmented video
β βββ instance_labels.json # Instance label information
β βββ result_summary.json # Segmentation result summary
βββ dynamicBA/ (Optional) # 4D reconstruction results
β βββ pose.npz # Intrinsic and Extrinsics
β βββ depth/ # Depth maps
β βββ flow/ # Optical flow data
βββ processing_log_<scene_id>.log # Processing log
We provide preprocessed datasets to reproduce Table 1 and 2 in our main paper.
You can also download our preprocessed data that we used for the quantitive results in our paper:
cd data
gdown https://drive.google.com/uc?id=1V1WIRvnJCJStL63rluwNZMPI2Gq4-yQy -O preprocessed.zip
unzip preprocessed.zip
We provide evaluation scripts for pose and depth metrics:
bash ./scripts/eval.sh
Our code is based on the following awesome repositories:
This project is built upon multiple open-source projects. Please refer to the license requirements of each submodule.
Issues and Pull Requests are welcome. Before submitting code, please ensure:
If you find our work useful in your research, please consider giving a star :star: and citing the following paper :pencil:.
@misc{wen2025dynamicverse,
title={DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling},
author={Kairun Wen and Yuzhi Huang and Runyu Chen and Hui Zheng and Yunlong Lin and Panwang Pan and Chenxin Li and Wenyan Cong and Jian Zhang and Junbin Lu and Chenguo Lin and Dilin Wang and Zhicheng Yan and Hongyu Xu and Justin Theiss and Yue Huang and Xinghao Ding and Rakesh Ranjan and Zhiwen Fan},
year={2025},
eprint={2512.03000},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.03000},
}