CaiYuanhao/OmniVCus-Train

Dataset

2

stars

8

commits

1

linked in READMEs

Dec 31, 2025

updated

README

[NeurIPS 2025] OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

Dataset Description

This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions.

Here is a data example:

Generated Prompt: a woman and a child playing with a toy train.
Original VideoSegmented SubjectAugmented Subject
Depth VideoMask VideoMotion Video

We also release a testing dataset on HuggingFace at

https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Test

This dataset is intended to be used together with our code. Please refer to the GitHub repository below for more detailed instructions.

https://github.com/caiyuanhao1998/Open-OmniVCus

We also release three models based on Wan2.1-1.3B, Wan2.1-14B, and Wan2.2-14B in the following link:

https://huggingface.co/CaiYuanhao/OmniVCus

For more video customization results, please refer to our project page:

https://caiyuanhao1998.github.io/project/OmniVCus/

For more technical details, please refer to our NeurIPS 2025 paper:

https://arxiv.org/abs/2506.23361

Citation

If you find our code, data, and models useful, please consider citing our paper:

@inproceedings{omnivcus,
  title={OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions},
  author={Yuanhao Cai and He Zhang and Xi Chen and Jinbo Xing and Kai Zhang and Yiwei Hu and Yuqian Zhou and Zhifei Zhang and Soo Ye Kim and Tianyu Wang and Yulun Zhang and Xiaokang Yang and Zhe Lin and Alan Yuille},
  booktitle={NeurIPS},
  year={2025}
}

Contributors

CaiYuanhao

7 commits

YU
yuanhaoc user

1 commits

CaiYuanhao/OmniVCus-Train

Dataset

2

stars

8

commits

1

linked in READMEs

Dec 31, 2025

updated

README

[NeurIPS 2025] OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

Dataset Description

This dataset supports multi-modal constrol video generation. It contains ~80K data samples processed from 140K videos by our VideoCus-Factory pipeline. Each data sample includes the original video, text prompts, subject reference image, depth video, mask video, and motion video conditions.

Here is a data example:

Generated Prompt: a woman and a child playing with a toy train.
Original VideoSegmented SubjectAugmented Subject
Depth VideoMask VideoMotion Video

We also release a testing dataset on HuggingFace at

https://huggingface.co/datasets/CaiYuanhao/OmniVCus-Test

This dataset is intended to be used together with our code. Please refer to the GitHub repository below for more detailed instructions.

https://github.com/caiyuanhao1998/Open-OmniVCus

We also release three models based on Wan2.1-1.3B, Wan2.1-14B, and Wan2.2-14B in the following link:

https://huggingface.co/CaiYuanhao/OmniVCus

For more video customization results, please refer to our project page:

https://caiyuanhao1998.github.io/project/OmniVCus/

For more technical details, please refer to our NeurIPS 2025 paper:

https://arxiv.org/abs/2506.23361

Citation

If you find our code, data, and models useful, please consider citing our paper:

@inproceedings{omnivcus,
  title={OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions},
  author={Yuanhao Cai and He Zhang and Xi Chen and Jinbo Xing and Kai Zhang and Yiwei Hu and Yuqian Zhou and Zhifei Zhang and Soo Ye Kim and Tianyu Wang and Yulun Zhang and Xiaokang Yang and Zhe Lin and Alan Yuille},
  booktitle={NeurIPS},
  year={2025}
}

Contributors

CaiYuanhao

7 commits

YU
yuanhaoc user

1 commits