shi-labs/physical-ai-bench-generation

Dataset

Physical AI Bench - Generation

5

4 commits

1 linked in READMEs

updated Dec 10, 2025

See the code

README

Physical AI Bench - Generation

Paper | Code

Dataset Description

The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our dataset is a benchmark designed to evaluate world models for Physical AI.

This dataset is ready for non-commercial use.

License/Terms of Use

The use of this dataset is governed by CC BY-NC 4.0.

Intended Usage

This benchmark dataset is intended to demonstrate and facilitate the understanding and evaluation of world models for Physical AI. It should primarily be used for educational and demonstration purposes.

Dataset Characterization

This dataset focuses on the following areas: Autonomous Vehicle (AV) driving, Robotics, Industry (smart space), Physics, Human, Common Sense.

Data Collection Method

  • AV: Automatic/Sensors
  • Industry: Automatic/Sensors
  • Physics: Automatic/Sensors
  • Robotics: Automatic/Sensors
  • Human: Automatic/Sensors
  • Common Sense: Human

Labeling Method

  • AV: Hybrid: Human, Automated
  • Industry: Hybrid: Human, Automated
  • Physics: Hybrid: Human, Automated
  • Robotics: Hybrid: Human, Automated
  • Human: Hybrid: Human, Automated
  • Common Sense: Hybrid: Human, Automated

Folder Structure

pbench/
β”œβ”€β”€ condition_image/                       # Conditioning images for all domains
β”œβ”€β”€ vqa/                                   # Visual Question Answering pairs
└── cosmos_predict2_bench_full_info.json   # Complete dataset metadata

Dataset Format

  • Modality: Image (jpg) and Text

Dataset Quantification

The dataset is stored in JSON files. The quantity, including the conditioning images, text prompts, and qa pairs, of the Pbench dataset is described in the table below.

DomainQuantity
AV118
Common Sense239
Human299
Industry107
Physics107
Robotics174
Total Storage Size226 MB

Citation

If you use Physical AI Bench in your research, please cite:

@misc{zhou2025paibenchcomprehensivebenchmarkphysical,
      title={PAI-Bench: A Comprehensive Benchmark For Physical AI}, 
      author={Fengzhe Zhou and Jiannan Huang and Jialuo Li and Deva Ramanan and Humphrey Shi},
      year={2025},
      eprint={2512.01989},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.01989}, 
}
benchmark
multimodal
physical-ai
world-models

shi-labs/physical-ai-bench-generation

Dataset

Physical AI Bench - Generation

5

4 commits

1 linked in READMEs

updated Dec 10, 2025

See the code

README

Physical AI Bench - Generation

Paper | Code

Dataset Description

The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our dataset is a benchmark designed to evaluate world models for Physical AI.

This dataset is ready for non-commercial use.

License/Terms of Use

The use of this dataset is governed by CC BY-NC 4.0.

Intended Usage

This benchmark dataset is intended to demonstrate and facilitate the understanding and evaluation of world models for Physical AI. It should primarily be used for educational and demonstration purposes.

Dataset Characterization

This dataset focuses on the following areas: Autonomous Vehicle (AV) driving, Robotics, Industry (smart space), Physics, Human, Common Sense.

Data Collection Method

  • AV: Automatic/Sensors
  • Industry: Automatic/Sensors
  • Physics: Automatic/Sensors
  • Robotics: Automatic/Sensors
  • Human: Automatic/Sensors
  • Common Sense: Human

Labeling Method

  • AV: Hybrid: Human, Automated
  • Industry: Hybrid: Human, Automated
  • Physics: Hybrid: Human, Automated
  • Robotics: Hybrid: Human, Automated
  • Human: Hybrid: Human, Automated
  • Common Sense: Hybrid: Human, Automated

Folder Structure

pbench/
β”œβ”€β”€ condition_image/                       # Conditioning images for all domains
β”œβ”€β”€ vqa/                                   # Visual Question Answering pairs
└── cosmos_predict2_bench_full_info.json   # Complete dataset metadata

Dataset Format

  • Modality: Image (jpg) and Text

Dataset Quantification

The dataset is stored in JSON files. The quantity, including the conditioning images, text prompts, and qa pairs, of the Pbench dataset is described in the table below.

DomainQuantity
AV118
Common Sense239
Human299
Industry107
Physics107
Robotics174
Total Storage Size226 MB

Citation

If you use Physical AI Bench in your research, please cite:

@misc{zhou2025paibenchcomprehensivebenchmarkphysical,
      title={PAI-Bench: A Comprehensive Benchmark For Physical AI}, 
      author={Fengzhe Zhou and Jiannan Huang and Jialuo Li and Deva Ramanan and Humphrey Shi},
      year={2025},
      eprint={2512.01989},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.01989}, 
}
benchmark
multimodal
physical-ai
world-models