LeoFan01/RoboBench

Dataset

7

stars

26

commits

1

linked in READMEs

Jul 30, 2026

updated

README

RoboBench Logo

RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

Paper GitHub Project Page Official Results License

πŸ“‹ Overview

RoboBench is a comprehensive evaluation benchmark designed to assess the capabilities of Multimodal Large Language Models (MLLMs) in embodied intelligence tasks. This benchmark provides a systematic framework for evaluating how well these models can understand and reason about robotic scenarios.

This repository contains the released RoboBench benchmark data. Official score tables and model-output JSON files are hosted separately at:

https://huggingface.co/datasets/lyl010221-pku/RoboBench-Results

The results repository includes CSV exports of the paper tables, coverage audits, and released model-output JSON files for Instruction Comprehension, Perception and Reasoning, Generalized Planning, Affordance Reasoning, and Error Analysis.

🎯 Key Features

  • 🧠 Comprehensive Evaluation: Covers multiple aspects of embodied intelligence
  • πŸ“Š Rich Dataset: Contains thousands of carefully curated examples
  • πŸ”¬ Scientific Rigor: Designed with research-grade evaluation metrics
  • 🌐 Multimodal: Supports text, images, and video data
  • πŸ€– Robotics Focus: Specifically tailored for robotic applications

πŸ“Š Dataset Statistics

CategoryCountDescription
Total Samples6092Comprehensive evaluation dataset
Image Samples1400High-quality visual data
Video Samples3142Temporal & Planning reasoning examples

πŸ—οΈ Dataset Structure

RoboBench/
β”œβ”€β”€ 1_instruction_comprehension/    # Instruction understanding tasks
β”œβ”€β”€ 2_perception_reasoning/         # Visual perception and reasoning
β”œβ”€β”€ 3_generalized_planning/         # Cross-domain planning tasks
β”œβ”€β”€ 4_affordance_reasoning/         # Object affordance understanding
β”œβ”€β”€ 5_error_analysis/               # Error analysis and debugging
└──system_prompt.json.              # Every task system prompts

πŸ”¬ Research Applications

This benchmark is designed for researchers working on:

  • Multimodal Large Language Models
  • Embodied AI Systems
  • Robotic Intelligence
  • Computer Vision
  • Natural Language Processing

πŸ“š Citation

If you use RoboBench in your research, please cite our paper:

@misc{luo2026robobenchcomprehensiveevaluationbenchmark,
      title={Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain}, 
      author={Yulin Luo and Chun-Kai Fan and Menghang Dong and Jiayu Shi and Xiangju Mi and Mengdi Zhao and Bo-Wen Zhang and Cheng Chi and Jiaming Liu and Gaole Dai and Rongyu Zhang and Ruichuan An and Kun Wu and Zhengping Che and Shaoxuan Xie and Guocai Yao and Zhongxia Zhao and Pengwei Wang and Guang Liu and Zhongyuan Wang and Tiejun Huang and Shanghang Zhang},
      year={2026},
      eprint={2510.17801},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2510.17801}, 
}

🀝 Contributing

We welcome contributions! Please see our Contributing Guidelines for more details.

πŸ“„ License

This dataset is released under the Creative Commons Attribution 4.0 International License.


Made with ❀️ by the RoboBench Team

Contributors

LeoFan01

25 commits

lyl010221-pku

1 commits

LeoFan01/RoboBench

Dataset

7

stars

26

commits

1

linked in READMEs

Jul 30, 2026

updated

README

RoboBench Logo

RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

Paper GitHub Project Page Official Results License

πŸ“‹ Overview

RoboBench is a comprehensive evaluation benchmark designed to assess the capabilities of Multimodal Large Language Models (MLLMs) in embodied intelligence tasks. This benchmark provides a systematic framework for evaluating how well these models can understand and reason about robotic scenarios.

This repository contains the released RoboBench benchmark data. Official score tables and model-output JSON files are hosted separately at:

https://huggingface.co/datasets/lyl010221-pku/RoboBench-Results

The results repository includes CSV exports of the paper tables, coverage audits, and released model-output JSON files for Instruction Comprehension, Perception and Reasoning, Generalized Planning, Affordance Reasoning, and Error Analysis.

🎯 Key Features

  • 🧠 Comprehensive Evaluation: Covers multiple aspects of embodied intelligence
  • πŸ“Š Rich Dataset: Contains thousands of carefully curated examples
  • πŸ”¬ Scientific Rigor: Designed with research-grade evaluation metrics
  • 🌐 Multimodal: Supports text, images, and video data
  • πŸ€– Robotics Focus: Specifically tailored for robotic applications

πŸ“Š Dataset Statistics

CategoryCountDescription
Total Samples6092Comprehensive evaluation dataset
Image Samples1400High-quality visual data
Video Samples3142Temporal & Planning reasoning examples

πŸ—οΈ Dataset Structure

RoboBench/
β”œβ”€β”€ 1_instruction_comprehension/    # Instruction understanding tasks
β”œβ”€β”€ 2_perception_reasoning/         # Visual perception and reasoning
β”œβ”€β”€ 3_generalized_planning/         # Cross-domain planning tasks
β”œβ”€β”€ 4_affordance_reasoning/         # Object affordance understanding
β”œβ”€β”€ 5_error_analysis/               # Error analysis and debugging
└──system_prompt.json.              # Every task system prompts

πŸ”¬ Research Applications

This benchmark is designed for researchers working on:

  • Multimodal Large Language Models
  • Embodied AI Systems
  • Robotic Intelligence
  • Computer Vision
  • Natural Language Processing

πŸ“š Citation

If you use RoboBench in your research, please cite our paper:

@misc{luo2026robobenchcomprehensiveevaluationbenchmark,
      title={Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain}, 
      author={Yulin Luo and Chun-Kai Fan and Menghang Dong and Jiayu Shi and Xiangju Mi and Mengdi Zhao and Bo-Wen Zhang and Cheng Chi and Jiaming Liu and Gaole Dai and Rongyu Zhang and Ruichuan An and Kun Wu and Zhengping Che and Shaoxuan Xie and Guocai Yao and Zhongxia Zhao and Pengwei Wang and Guang Liu and Zhongyuan Wang and Tiejun Huang and Shanghang Zhang},
      year={2026},
      eprint={2510.17801},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2510.17801}, 
}

🀝 Contributing

We welcome contributions! Please see our Contributing Guidelines for more details.

πŸ“„ License

This dataset is released under the Creative Commons Attribution 4.0 International License.


Made with ❀️ by the RoboBench Team

Contributors

LeoFan01

25 commits

lyl010221-pku

1 commits