This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions.
Here is a data example:
| Generated Prompt: a woman and a child playing with a toy train. | ||
|
|
|
| Original Video | Segmented Subject | Augmented Subject |
|
|
|
| Depth Video | Mask Video | Motion Video |
We also release a testing dataset on HuggingFace at
https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Test
This dataset is intended to be used together with our code. Please refer to the GitHub repository below for more detailed instructions.
https://github.com/caiyuanhao1998/Open-OmniVCus
We also release three models based on Wan2.1-1.3B, Wan2.1-14B, and Wan2.2-14B in the following link:
https://huggingface.co/CaiYuanhao/OmniVCus
For more video customization results, please refer to our project page:
https://caiyuanhao1998.github.io/project/OmniVCus/
For more technical details, please refer to our NeurIPS 2025 paper:
https://arxiv.org/abs/2506.23361
If you find our code, data, and models useful, please consider citing our paper:
@inproceedings{omnivcus,
title={OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions},
author={Yuanhao Cai and He Zhang and Xi Chen and Jinbo Xing and Kai Zhang and Yiwei Hu and Yuqian Zhou and Zhifei Zhang and Soo Ye Kim and Tianyu Wang and Yulun Zhang and Xiaokang Yang and Zhe Lin and Alan Yuille},
booktitle={NeurIPS},
year={2025}
}
7 commits
1 commits
This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions.
Here is a data example:
| Generated Prompt: a woman and a child playing with a toy train. | ||
|
|
|
| Original Video | Segmented Subject | Augmented Subject |
|
|
|
| Depth Video | Mask Video | Motion Video |
We also release a testing dataset on HuggingFace at
https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Test
This dataset is intended to be used together with our code. Please refer to the GitHub repository below for more detailed instructions.
https://github.com/caiyuanhao1998/Open-OmniVCus
We also release three models based on Wan2.1-1.3B, Wan2.1-14B, and Wan2.2-14B in the following link:
https://huggingface.co/CaiYuanhao/OmniVCus
For more video customization results, please refer to our project page:
https://caiyuanhao1998.github.io/project/OmniVCus/
For more technical details, please refer to our NeurIPS 2025 paper:
https://arxiv.org/abs/2506.23361
If you find our code, data, and models useful, please consider citing our paper:
@inproceedings{omnivcus,
title={OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions},
author={Yuanhao Cai and He Zhang and Xi Chen and Jinbo Xing and Kai Zhang and Yiwei Hu and Yuqian Zhou and Zhifei Zhang and Soo Ye Kim and Tianyu Wang and Yulun Zhang and Xiaokang Yang and Zhe Lin and Alan Yuille},
booktitle={NeurIPS},
year={2025}
}
7 commits
1 commits