This dataset contains only 7 videos, which were specifically used in this paper. These videos were selected from a larger set of 200 dashcam videos recorded in various cities across Peru, available as an extended dataset. The purpose of this dataset is to evaluate the performance of Vision-Language Models (VLMs) compared to human performance and to analyze their responses.

The dataset is organized into the following folders:
dataset/
│── videos/
│── human_responses/
│── vlm_responses/
│ │── one_response/
│ │── all_responses_cured/
│ │── all_responses_uncured/
│── IDs.csv # File containing video names and IDs
The dataset is intended for research on VLMs, specifically to evaluate how they respond to video sequences from Peru.
If you are interested in accessing the full dataset with 200 videos, please fill out the following form:
This dataset is shared under the CC-BY-NC 4.0 license. Users must provide attribution and are not allowed to use the dataset for commercial purposes.
If you use this dataset in your research, please cite it as follows:
@misc{cusipuma2025robusto1datasetcomparinghumans,
title={Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru},
author={Dunant Cusipuma and David Ortega and Victor Flores-Benites and Arturo Deza},
year={2025},
eprint={2503.07587},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.07587},
}
For questions or collaborations, please contact [info@artificio.org].
This dataset contains only 7 videos, which were specifically used in this paper. These videos were selected from a larger set of 200 dashcam videos recorded in various cities across Peru, available as an extended dataset. The purpose of this dataset is to evaluate the performance of Vision-Language Models (VLMs) compared to human performance and to analyze their responses.

The dataset is organized into the following folders:
dataset/
│── videos/
│── human_responses/
│── vlm_responses/
│ │── one_response/
│ │── all_responses_cured/
│ │── all_responses_uncured/
│── IDs.csv # File containing video names and IDs
The dataset is intended for research on VLMs, specifically to evaluate how they respond to video sequences from Peru.
If you are interested in accessing the full dataset with 200 videos, please fill out the following form:
This dataset is shared under the CC-BY-NC 4.0 license. Users must provide attribution and are not allowed to use the dataset for commercial purposes.
If you use this dataset in your research, please cite it as follows:
@misc{cusipuma2025robusto1datasetcomparinghumans,
title={Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru},
author={Dunant Cusipuma and David Ortega and Victor Flores-Benites and Arturo Deza},
year={2025},
eprint={2503.07587},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.07587},
}
For questions or collaborations, please contact [info@artificio.org].