The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our dataset is a benchmark designed to evaluate world models for Physical AI.
This dataset is ready for non-commercial use.
The use of this dataset is governed by CC BY-NC 4.0.
This benchmark dataset is intended to demonstrate and facilitate the understanding and evaluation of world models for Physical AI. It should primarily be used for educational and demonstration purposes.
This dataset focuses on the following areas: Autonomous Vehicle (AV) driving, Robotics, Industry (smart space), Physics, Human, Common Sense.
pbench/
βββ condition_image/ # Conditioning images for all domains
βββ vqa/ # Visual Question Answering pairs
βββ cosmos_predict2_bench_full_info.json # Complete dataset metadata
The dataset is stored in JSON files. The quantity, including the conditioning images, text prompts, and qa pairs, of the Pbench dataset is described in the table below.
| Domain | Quantity |
|---|---|
| AV | 118 |
| Common Sense | 239 |
| Human | 299 |
| Industry | 107 |
| Physics | 107 |
| Robotics | 174 |
| Total Storage Size | 226 MB |
If you use Physical AI Bench in your research, please cite:
@misc{zhou2025paibenchcomprehensivebenchmarkphysical,
title={PAI-Bench: A Comprehensive Benchmark For Physical AI},
author={Fengzhe Zhou and Jiannan Huang and Jialuo Li and Deva Ramanan and Humphrey Shi},
year={2025},
eprint={2512.01989},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.01989},
}
The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our dataset is a benchmark designed to evaluate world models for Physical AI.
This dataset is ready for non-commercial use.
The use of this dataset is governed by CC BY-NC 4.0.
This benchmark dataset is intended to demonstrate and facilitate the understanding and evaluation of world models for Physical AI. It should primarily be used for educational and demonstration purposes.
This dataset focuses on the following areas: Autonomous Vehicle (AV) driving, Robotics, Industry (smart space), Physics, Human, Common Sense.
pbench/
βββ condition_image/ # Conditioning images for all domains
βββ vqa/ # Visual Question Answering pairs
βββ cosmos_predict2_bench_full_info.json # Complete dataset metadata
The dataset is stored in JSON files. The quantity, including the conditioning images, text prompts, and qa pairs, of the Pbench dataset is described in the table below.
| Domain | Quantity |
|---|---|
| AV | 118 |
| Common Sense | 239 |
| Human | 299 |
| Industry | 107 |
| Physics | 107 |
| Robotics | 174 |
| Total Storage Size | 226 MB |
If you use Physical AI Bench in your research, please cite:
@misc{zhou2025paibenchcomprehensivebenchmarkphysical,
title={PAI-Bench: A Comprehensive Benchmark For Physical AI},
author={Fengzhe Zhou and Jiannan Huang and Jialuo Li and Deva Ramanan and Humphrey Shi},
year={2025},
eprint={2512.01989},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.01989},
}