PhysionLabs/Physion-Eval

Dataset

🎬 PHYSION-EVAL: The First Human-Centered Benchmark for Physical Realism in AI-Generated Videos

19

12 commits

1 linked in READMEs

updated Jun 14, 2026

See the code

README

🎬 PHYSION-EVAL: The First Human-Centered Benchmark for Physical Realism in AI-Generated Videos

✨ Overview

This dataset is developed by Physion Labs, a research team focused on advancing physical realism and reliability in multimodal generative AI.

We created this dataset to support physically grounded video generation, moving beyond visual realism toward true physical consistency. It enables research in:

  • Perceptual physical realism evaluation for AI-generated videos
  • Temporal consistency and causal reasoning
  • Human vs. model perception of physical plausibility

By identifying where current models break physical rules, we aim to enable more reliable and trustworthy generative video systems.

🆕 Caption Release Update

Due to strong community demand and multiple email inquiries, we have decided to also open-source the caption associated with each video generation.

To reduce copyright risk in our initial video release, we removed the first five frames from each generated video, since the WISA dataset was crawled from online content. For users interested in the image prompt, we recommend using the first remaining frame of each uploaded video as a reasonable proxy for the initial image prompt used for the corresponding generation.

🛡️ Content Filtering

In this open-source version, we apply filtering procedures to reduce privacy and intellectual property risks:

  • 🚫 Videos with identifiable human faces are excluded
  • 🚫 Videos with logos, brand marks, or trademarks are excluded

Permitted Use

This dataset is intended to support:

  • Academic and non-commercial research.
  • Benchmarking and evaluation of video generation models.
  • Studying physical realism and perceptual consistency in AI-generated videos.

This dataset must not be used for:

  • Surveillance or biometric identification.
  • Any application that violates privacy, publicity, or intellectual property rights.
  • High-risk or safety-critical decision-making systems.
  • Training, fine-tuning, or distillation of generative models without explicit permission

Citation

If you use this dataset, please cite:

@article{zhang2026physion,
  title={Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning},
  author={Zhang, Qin and Jing, Peiyu and Yu, Hong-Xing and Ding, Fangqiang and Nie, Fan and Wang, Weimin and Du, Yilun and Zou, James and Wu, Jiajun and Shuai, Bing},
  journal={arXiv preprint arXiv:2603.19607},
  year={2026}
}

📜 License

Custom Research Use Only License

By using this dataset, you agree to:

  • Use it for non-commercial research and evaluation purposes only
  • Not redistribute the dataset in any way that violates applicable laws or third-party rights
  • Comply with all relevant laws, regulations, and ethical guidelines

evaluation
multimodal
physical-reasoning
synthetic-data
video

Contributors

thephysioncorgi

12 commits

PhysionLabs/Physion-Eval

Dataset

🎬 PHYSION-EVAL: The First Human-Centered Benchmark for Physical Realism in AI-Generated Videos

19

12 commits

1 linked in READMEs

updated Jun 14, 2026

See the code

README

🎬 PHYSION-EVAL: The First Human-Centered Benchmark for Physical Realism in AI-Generated Videos

✨ Overview

This dataset is developed by Physion Labs, a research team focused on advancing physical realism and reliability in multimodal generative AI.

We created this dataset to support physically grounded video generation, moving beyond visual realism toward true physical consistency. It enables research in:

  • Perceptual physical realism evaluation for AI-generated videos
  • Temporal consistency and causal reasoning
  • Human vs. model perception of physical plausibility

By identifying where current models break physical rules, we aim to enable more reliable and trustworthy generative video systems.

🆕 Caption Release Update

Due to strong community demand and multiple email inquiries, we have decided to also open-source the caption associated with each video generation.

To reduce copyright risk in our initial video release, we removed the first five frames from each generated video, since the WISA dataset was crawled from online content. For users interested in the image prompt, we recommend using the first remaining frame of each uploaded video as a reasonable proxy for the initial image prompt used for the corresponding generation.

🛡️ Content Filtering

In this open-source version, we apply filtering procedures to reduce privacy and intellectual property risks:

  • 🚫 Videos with identifiable human faces are excluded
  • 🚫 Videos with logos, brand marks, or trademarks are excluded

Permitted Use

This dataset is intended to support:

  • Academic and non-commercial research.
  • Benchmarking and evaluation of video generation models.
  • Studying physical realism and perceptual consistency in AI-generated videos.

This dataset must not be used for:

  • Surveillance or biometric identification.
  • Any application that violates privacy, publicity, or intellectual property rights.
  • High-risk or safety-critical decision-making systems.
  • Training, fine-tuning, or distillation of generative models without explicit permission

Citation

If you use this dataset, please cite:

@article{zhang2026physion,
  title={Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning},
  author={Zhang, Qin and Jing, Peiyu and Yu, Hong-Xing and Ding, Fangqiang and Nie, Fan and Wang, Weimin and Du, Yilun and Zou, James and Wu, Jiajun and Shuai, Bing},
  journal={arXiv preprint arXiv:2603.19607},
  year={2026}
}

📜 License

Custom Research Use Only License

By using this dataset, you agree to:

  • Use it for non-commercial research and evaluation purposes only
  • Not redistribute the dataset in any way that violates applicable laws or third-party rights
  • Comply with all relevant laws, regulations, and ethical guidelines

evaluation
multimodal
physical-reasoning
synthetic-data
video

Contributors

thephysioncorgi

12 commits