WoW-1 Benchmark Samples is the official evaluation dataset released as part of the WoW (World-Omniscient World Model) project. This benchmark is designed to assess the physical consistency and causal reasoning capabilities of generative world models for robotics and embodied AI.
This dataset contains 612 natural language prompts representing real-world robot interaction tasks. These instructions are used to evaluate world models on their ability to understand and generate plausible, physically grounded responses in video or action space.
Each sample describes a short-term or long-horizon task involving:
This dataset is intended for:
{
"text": "Put the apples on the table into the basket."
}
Clean the table surfaceUse the right arm to grab the pearl and give it to the left armOpen the door of the red microwavePlace the tennis ball in the brown objectThis dataset is used for evaluating models such as:
WoW-1-DiT-2B, WoW-1-DiT-7BWoW-1-Wan-14BSOPHIA-guided generative modelsWoW: Towards a World omniscient World model Through Embodied Interaction
Xiaowei Chi et al., 2025 β arXiv:2509.22642
Please cite this paper if you use the dataset:
@article{chi2025wow,
title={WoW: Towards a World omniscient World model Through Embodied Interaction},
author={Chi, Xiaowei and Jia, Peidong and Fan, Chun-Kai and Ju, Xiaozhu and Mi, Weishi and Qin, Zhiyuan and Zhang, Kevin and Tian, Wanxin and Ge, Kuangzhi and Li, Hao and others},
journal={arXiv preprint arXiv:2509.22642},
year={2025}
}
This dataset is released under the MIT License.
π€ We encourage the community to explore, evaluate, and extend this benchmark. Contributions and feedback are welcome via GitHub or the project website.
8 commits
WoW-1 Benchmark Samples is the official evaluation dataset released as part of the WoW (World-Omniscient World Model) project. This benchmark is designed to assess the physical consistency and causal reasoning capabilities of generative world models for robotics and embodied AI.
This dataset contains 612 natural language prompts representing real-world robot interaction tasks. These instructions are used to evaluate world models on their ability to understand and generate plausible, physically grounded responses in video or action space.
Each sample describes a short-term or long-horizon task involving:
This dataset is intended for:
{
"text": "Put the apples on the table into the basket."
}
Clean the table surfaceUse the right arm to grab the pearl and give it to the left armOpen the door of the red microwavePlace the tennis ball in the brown objectThis dataset is used for evaluating models such as:
WoW-1-DiT-2B, WoW-1-DiT-7BWoW-1-Wan-14BSOPHIA-guided generative modelsWoW: Towards a World omniscient World model Through Embodied Interaction
Xiaowei Chi et al., 2025 β arXiv:2509.22642
Please cite this paper if you use the dataset:
@article{chi2025wow,
title={WoW: Towards a World omniscient World model Through Embodied Interaction},
author={Chi, Xiaowei and Jia, Peidong and Fan, Chun-Kai and Ju, Xiaozhu and Mi, Weishi and Qin, Zhiyuan and Zhang, Kevin and Tian, Wanxin and Ge, Kuangzhi and Li, Hao and others},
journal={arXiv preprint arXiv:2509.22642},
year={2025}
}
This dataset is released under the MIT License.
π€ We encourage the community to explore, evaluate, and extend this benchmark. Contributions and feedback are welcome via GitHub or the project website.
8 commits