🎬 PHYSION-EVAL: The First Human-Centered Benchmark for Physical Realism in AI-Generated Videos
19
12 commits
1 linked in READMEs
updated Jun 14, 2026
This dataset is developed by Physion Labs, a research team focused on advancing physical realism and reliability in multimodal generative AI.
We created this dataset to support physically grounded video generation, moving beyond visual realism toward true physical consistency. It enables research in:
By identifying where current models break physical rules, we aim to enable more reliable and trustworthy generative video systems.
Due to strong community demand and multiple email inquiries, we have decided to also open-source the caption associated with each video generation.
To reduce copyright risk in our initial video release, we removed the first five frames from each generated video, since the WISA dataset was crawled from online content. For users interested in the image prompt, we recommend using the first remaining frame of each uploaded video as a reasonable proxy for the initial image prompt used for the corresponding generation.
In this open-source version, we apply filtering procedures to reduce privacy and intellectual property risks:
This dataset is intended to support:
This dataset must not be used for:
If you use this dataset, please cite:
@article{zhang2026physion,
title={Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning},
author={Zhang, Qin and Jing, Peiyu and Yu, Hong-Xing and Ding, Fangqiang and Nie, Fan and Wang, Weimin and Du, Yilun and Zou, James and Wu, Jiajun and Shuai, Bing},
journal={arXiv preprint arXiv:2603.19607},
year={2026}
}
By using this dataset, you agree to:
12 commits
🎬 PHYSION-EVAL: The First Human-Centered Benchmark for Physical Realism in AI-Generated Videos
19
12 commits
1 linked in READMEs
updated Jun 14, 2026
This dataset is developed by Physion Labs, a research team focused on advancing physical realism and reliability in multimodal generative AI.
We created this dataset to support physically grounded video generation, moving beyond visual realism toward true physical consistency. It enables research in:
By identifying where current models break physical rules, we aim to enable more reliable and trustworthy generative video systems.
Due to strong community demand and multiple email inquiries, we have decided to also open-source the caption associated with each video generation.
To reduce copyright risk in our initial video release, we removed the first five frames from each generated video, since the WISA dataset was crawled from online content. For users interested in the image prompt, we recommend using the first remaining frame of each uploaded video as a reasonable proxy for the initial image prompt used for the corresponding generation.
In this open-source version, we apply filtering procedures to reduce privacy and intellectual property risks:
This dataset is intended to support:
This dataset must not be used for:
If you use this dataset, please cite:
@article{zhang2026physion,
title={Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning},
author={Zhang, Qin and Jing, Peiyu and Yu, Hong-Xing and Ding, Fangqiang and Nie, Fan and Wang, Weimin and Du, Yilun and Zou, James and Wu, Jiajun and Shuai, Bing},
journal={arXiv preprint arXiv:2603.19607},
year={2026}
}
By using this dataset, you agree to:
12 commits