shi-labs/physical-ai-bench-conditional-generation

Dataset

Physical AI Bench - Conditional Generation

0

7 commits

3 linked in READMEs

updated Dec 10, 2025

See the code

README

Physical AI Bench - Conditional Generation

Paper | Code

This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.

DatasetCategorySample Nums
Agibot WorldRobotics200
OpenDVAutonomous Driving200
Ego-Exo4DEgo-centric200

Dataset Summary

  • Dataset Size: 600 video samples
  • Video Format: MP4 files with various processing variants
  • Annotations: Text captions for each video
  • Processing Variants: Blur, Canny edge detection, Depth estimation, SAM2 segmentation

File Organization

physical-ai-bench-transfer/
β”œβ”€β”€ videos/          # Original video files
β”œβ”€β”€ blur/            # Blur-processed videos
β”œβ”€β”€ canny/           # Edge detection videos
β”œβ”€β”€ depth_vids/      # Depth estimation videos
β”œβ”€β”€ depth_npzs/      # Depth estimation numpy arrays
β”œβ”€β”€ sam2_vids/       # SAM2 segmentation videos
β”œβ”€β”€ sam2_pkls/       # SAM2 segmentation pickle files
└── captions/        # JSON files with video descriptions

Citation

If you use Physical AI Bench in your research, please cite:

@misc{zhou2025paibenchcomprehensivebenchmarkphysical,
      title={PAI-Bench: A Comprehensive Benchmark For Physical AI}, 
      author={Fengzhe Zhou and Jiannan Huang and Jialuo Li and Deva Ramanan and Humphrey Shi},
      year={2025},
      eprint={2512.01989},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.01989}, 
}

shi-labs/physical-ai-bench-conditional-generation

Dataset

Physical AI Bench - Conditional Generation

0

7 commits

3 linked in READMEs

updated Dec 10, 2025

See the code

README

Physical AI Bench - Conditional Generation

Paper | Code

This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.

DatasetCategorySample Nums
Agibot WorldRobotics200
OpenDVAutonomous Driving200
Ego-Exo4DEgo-centric200

Dataset Summary

  • Dataset Size: 600 video samples
  • Video Format: MP4 files with various processing variants
  • Annotations: Text captions for each video
  • Processing Variants: Blur, Canny edge detection, Depth estimation, SAM2 segmentation

File Organization

physical-ai-bench-transfer/
β”œβ”€β”€ videos/          # Original video files
β”œβ”€β”€ blur/            # Blur-processed videos
β”œβ”€β”€ canny/           # Edge detection videos
β”œβ”€β”€ depth_vids/      # Depth estimation videos
β”œβ”€β”€ depth_npzs/      # Depth estimation numpy arrays
β”œβ”€β”€ sam2_vids/       # SAM2 segmentation videos
β”œβ”€β”€ sam2_pkls/       # SAM2 segmentation pickle files
└── captions/        # JSON files with video descriptions

Citation

If you use Physical AI Bench in your research, please cite:

@misc{zhou2025paibenchcomprehensivebenchmarkphysical,
      title={PAI-Bench: A Comprehensive Benchmark For Physical AI}, 
      author={Fengzhe Zhou and Jiannan Huang and Jialuo Li and Deva Ramanan and Humphrey Shi},
      year={2025},
      eprint={2512.01989},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.01989}, 
}