Rethinking Video Generation Model for the Embodied World
5
26 commits
1 linked in READMEs
updated Jan 25, 2026
The benchmark is constructed from two complementary perspectives: task categories and robot embodiment types, covering a total of 650 image-text evaluation cases.
The task-oriented split contains 250 image-text pairs, with 50 samples per task, spanning five representative robotic task categories:
The embodiment-oriented split contains 400 image-text pairs, with 100 samples per embodiment, covering four mainstream robotic embodiment types:
This split evaluates whether generative models can correctly reflect embodiment-specific physical structures and action affordances.
Each evaluation sample is stored in JSON format and includes:
name: Unique sample identifierimage_path: Path to the reference imageprompt: Concise task descriptionrobotic manipulator / manipulated object: Key semantic entitiesview: Camera viewpoint (e.g., first-person)Images are provided in JPEG format.
This benchmark is intended for:
This dataset is released under the CC BY 4.0 License.
If you find this dataset useful, please cite our paper:
@article{deng2026rethinking,
title={Rethinking Video Generation Model for the Embodied World},
author={Deng, Yufan and Pan, Zilin and Zhang, Hongyu and Li, Xiaojie and Hu, Ruoqing and Ding, Yufei and Zou, Yiming and Zeng, Yan and Zhou, Daquan},
journal={arXiv preprint arXiv:2601.15282},
year={2026}
}
Rethinking Video Generation Model for the Embodied World
5
26 commits
1 linked in READMEs
updated Jan 25, 2026
The benchmark is constructed from two complementary perspectives: task categories and robot embodiment types, covering a total of 650 image-text evaluation cases.
The task-oriented split contains 250 image-text pairs, with 50 samples per task, spanning five representative robotic task categories:
The embodiment-oriented split contains 400 image-text pairs, with 100 samples per embodiment, covering four mainstream robotic embodiment types:
This split evaluates whether generative models can correctly reflect embodiment-specific physical structures and action affordances.
Each evaluation sample is stored in JSON format and includes:
name: Unique sample identifierimage_path: Path to the reference imageprompt: Concise task descriptionrobotic manipulator / manipulated object: Key semantic entitiesview: Camera viewpoint (e.g., first-person)Images are provided in JPEG format.
This benchmark is intended for:
This dataset is released under the CC BY 4.0 License.
If you find this dataset useful, please cite our paper:
@article{deng2026rethinking,
title={Rethinking Video Generation Model for the Embodied World},
author={Deng, Yufan and Pan, Zilin and Zhang, Hongyu and Li, Xiaojie and Hu, Ruoqing and Ding, Yufei and Zou, Yiming and Zeng, Yan and Zhou, Daquan},
journal={arXiv preprint arXiv:2601.15282},
year={2026}
}