FlagEval/ERQA

Dataset

Introduction

6

4 commits

2 linked in READMEs

updated Apr 22, 2025

See the code

README

Introduction

Disclaimer: This dataset is organized and adapted from embodiedreasoning/ERQA. The original data was provided in TFRecord format and has been converted here into a more accessible and easy-to-use format.

This evaluation benchmark covers a variety of topics related to spatial reasoning and world knowledge focused on real-world scenarios, particularly in the context of robotics. Please find more details and visualizations in the tech report.

Data Fields

Field NameTypeDescription
question_idstringUnique question ID
questionstringQuestion text
question_typestringType of question
answerstringAnswer
visual_indiceslist[int]List of visual indices
imageslist[Image]Image data
images_base64list[string]Image data in base64

Contributors

HelloGitHub

4 commits

FlagEval/ERQA

Dataset

Introduction

6

4 commits

2 linked in READMEs

updated Apr 22, 2025

See the code

README

Introduction

Disclaimer: This dataset is organized and adapted from embodiedreasoning/ERQA. The original data was provided in TFRecord format and has been converted here into a more accessible and easy-to-use format.

This evaluation benchmark covers a variety of topics related to spatial reasoning and world knowledge focused on real-world scenarios, particularly in the context of robotics. Please find more details and visualizations in the tech report.

Data Fields

Field NameTypeDescription
question_idstringUnique question ID
questionstringQuestion text
question_typestringType of question
answerstringAnswer
visual_indiceslist[int]List of visual indices
imageslist[Image]Image data
images_base64list[string]Image data in base64

Contributors

HelloGitHub

4 commits