ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
2
36 commits
1 linked in READMEs
updated Jun 4, 2026
The original videos used in this project were collected from publicly available online driving videos. Since online videos may be updated, removed, or become inaccessible over time, the released annotation files on Hugging Face are intended as reference materials and may not exactly reproduce the dataset used in our experiments.
Due to copyright and privacy constraints, we cannot release any original videos or extracted image frames. We recommend that users collect publicly available driving videos according to their own research needs and use the provided pipeline code and annotation utilities (see 5. Pipeline Code and Utilities) to construct new datasets.
The main contribution of this project is an open, scalable data annotation pipeline for large scale driving video data, together with reference annotations that can help the community build datasets for pretraining, fine tuning, and downstream autonomous driving tasks.
The code for the dataset construction pipeline is still under further maintenance and refinement. We appreciate your patience and understanding.
Figure 1: Overview of the ScenePilot-Bench benchmark and evaluation metrics.
ScenePilot-4K is a large-scale first-person driving dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of video and 27.7M front-view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene-level natural-language descriptions, risk assessment labels, key-participant annotations, ego trajectories, and camera parameters through a unified multi-stage annotation pipeline.
Building on this dataset, we establish ScenePilot-Bench, a standardized benchmark that evaluates vision-language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT-based semantic alignment. The benchmark includes fine-grained metrics and geographic generalization settings that expose model robustness under cross-region and cross-traffic domain shifts.
Baseline results on representative open-source and proprietary vision-language models show that current models remain competitive in high-level scene semantics but still exhibit substantial limitations in geometry-aware perception and planning-oriented reasoning.
The released files in this repository can be grouped into the following categories.
These two compressed files contain pretrained model weights obtained by training on a 200k-scale VQA training set constructed in this work.
Both models are trained using the same dataset and unified training pipeline, and are used in the main experiments and comparison studies.
VGGT.zip
Contains annotation data related to spatial perception and geometric reasoning, including:
This file is not the raw output of VGGT, but a post-processed version after trajectory cleaning.
Specifically, the annotation pipeline is as follows:
VGGT.py from pipeline_code.ziptraj_clean.py to remove physically implausible or noisy trajectoriesThe final annotations in this archive therefore correspond to cleaned and quality-controlled trajectory data, suitable for downstream tasks such as trajectory prediction and spatial reasoning.
YOLO.zip
Provides 2D object detection results for major traffic participants. All detections are generated by a unified detection model and are used as perception inputs for downstream VQA and risk assessment tasks.
scene_description.zip
Contains scene description results generated from the original data, including:
These descriptions are used for scene understanding and for constructing balanced dataset splits.
This file contains the original video-level dataset split, including:
All VQA datasets of different scales are constructed strictly based on this video-level split to avoid scene-level information leakage.
This archive contains all VQA data in JSON format. Files are organized according to training, validation, and test splits.
Examples include:
Deleted_2D_train_vqa_add_new.jsonDeleted_2D_train_vqa_new.jsonThe VQA data in this archive is generated using the original VQA generation pipeline and includes a total of 22 VQA categories (Q1–Q22):
After initial generation, parts of the dataset were refined and regenerated due to:
To support flexible usage, we provide:
classify.py (in pipeline_code.zip)Therefore, this archive contains a mixture of original and partially updated VQA data, and users are encouraged to use the provided tools to construct task-specific subsets.
This archive contains the 100k-scale VQA test datasets used in the experiments.
Deleted_2D_test_selected_vqa_100k_final.jsonAdditional test sets are provided for generalization studies:
europe, japan-and-korea, us, and other correspond to geographic generalization experiments.left correspond to left-hand traffic country experiments.Each test set contains 100k VQA samples.
This archive contains training datasets of different scales:
Additional subsets include:
china, used for geographic generalization experiments.right, used for right-hand traffic country experiments.This archive contains the updated VQA dataset with explicitly grounded target objects, focusing exclusively on spatial perception tasks.
It includes the following seven question categories:
These samples are designed to support more precise evaluation and training for object-grounded spatial perception in autonomous driving scenarios.
This archive contains a curated set of high-quality trajectory-related VQA samples obtained after trajectory filtering and cleaning.
It covers the following five trajectory-centric categories:
These samples are intended for motion planning and trajectory reasoning tasks, with improved annotation quality after trajectory validation and filtering.
This archive contains the full implementation of the dataset construction pipeline. The components cover data preprocessing, perception annotation, trajectory generation, VQA construction, and post-processing.
The main scripts are listed below:
clip.py
Extracts image frames from raw videos:
mask.py
Generates image masks based on 2D bounding boxes:
Old_vqa_Q1-19.py
Original VQA generation script:
Q1-6-10-11_new.py
Updated VQA generation logic for selected categories:
Region[0])Q20-21-22_new.py
Updated generation for additional spatial reasoning categories:
scene_description.py
Generates scene-level descriptions:
VGGT.py
Core perception annotation module:
traj_clean.py
Trajectory post-processing module:
classify.py
VQA classification and selection tool:
These scripts together define the complete and reproducible pipeline for building the ScenePilot-4K dataset, from raw video processing to structured multimodal annotations.
This file lists all videos used in the dataset along with their corresponding download links. It is provided to support dataset reproduction and access to the original video resources.
@misc{wang2026scenepilotbenchlargescaledatasetbenchmark,
title={ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving},
author={Yujin Wang and Yutong Zheng and Wenxian Fan and Tianyi Wang and Hongqing Chu and Li Zhang and Bingzhao Gao and Daxin Tian and Hong Chen},
year={2026},
eprint={2601.19582},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.19582},
}
ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
2
36 commits
1 linked in READMEs
updated Jun 4, 2026
The original videos used in this project were collected from publicly available online driving videos. Since online videos may be updated, removed, or become inaccessible over time, the released annotation files on Hugging Face are intended as reference materials and may not exactly reproduce the dataset used in our experiments.
Due to copyright and privacy constraints, we cannot release any original videos or extracted image frames. We recommend that users collect publicly available driving videos according to their own research needs and use the provided pipeline code and annotation utilities (see 5. Pipeline Code and Utilities) to construct new datasets.
The main contribution of this project is an open, scalable data annotation pipeline for large scale driving video data, together with reference annotations that can help the community build datasets for pretraining, fine tuning, and downstream autonomous driving tasks.
The code for the dataset construction pipeline is still under further maintenance and refinement. We appreciate your patience and understanding.
Figure 1: Overview of the ScenePilot-Bench benchmark and evaluation metrics.
ScenePilot-4K is a large-scale first-person driving dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of video and 27.7M front-view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene-level natural-language descriptions, risk assessment labels, key-participant annotations, ego trajectories, and camera parameters through a unified multi-stage annotation pipeline.
Building on this dataset, we establish ScenePilot-Bench, a standardized benchmark that evaluates vision-language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT-based semantic alignment. The benchmark includes fine-grained metrics and geographic generalization settings that expose model robustness under cross-region and cross-traffic domain shifts.
Baseline results on representative open-source and proprietary vision-language models show that current models remain competitive in high-level scene semantics but still exhibit substantial limitations in geometry-aware perception and planning-oriented reasoning.
The released files in this repository can be grouped into the following categories.
These two compressed files contain pretrained model weights obtained by training on a 200k-scale VQA training set constructed in this work.
Both models are trained using the same dataset and unified training pipeline, and are used in the main experiments and comparison studies.
VGGT.zip
Contains annotation data related to spatial perception and geometric reasoning, including:
This file is not the raw output of VGGT, but a post-processed version after trajectory cleaning.
Specifically, the annotation pipeline is as follows:
VGGT.py from pipeline_code.ziptraj_clean.py to remove physically implausible or noisy trajectoriesThe final annotations in this archive therefore correspond to cleaned and quality-controlled trajectory data, suitable for downstream tasks such as trajectory prediction and spatial reasoning.
YOLO.zip
Provides 2D object detection results for major traffic participants. All detections are generated by a unified detection model and are used as perception inputs for downstream VQA and risk assessment tasks.
scene_description.zip
Contains scene description results generated from the original data, including:
These descriptions are used for scene understanding and for constructing balanced dataset splits.
This file contains the original video-level dataset split, including:
All VQA datasets of different scales are constructed strictly based on this video-level split to avoid scene-level information leakage.
This archive contains all VQA data in JSON format. Files are organized according to training, validation, and test splits.
Examples include:
Deleted_2D_train_vqa_add_new.jsonDeleted_2D_train_vqa_new.jsonThe VQA data in this archive is generated using the original VQA generation pipeline and includes a total of 22 VQA categories (Q1–Q22):
After initial generation, parts of the dataset were refined and regenerated due to:
To support flexible usage, we provide:
classify.py (in pipeline_code.zip)Therefore, this archive contains a mixture of original and partially updated VQA data, and users are encouraged to use the provided tools to construct task-specific subsets.
This archive contains the 100k-scale VQA test datasets used in the experiments.
Deleted_2D_test_selected_vqa_100k_final.jsonAdditional test sets are provided for generalization studies:
europe, japan-and-korea, us, and other correspond to geographic generalization experiments.left correspond to left-hand traffic country experiments.Each test set contains 100k VQA samples.
This archive contains training datasets of different scales:
Additional subsets include:
china, used for geographic generalization experiments.right, used for right-hand traffic country experiments.This archive contains the updated VQA dataset with explicitly grounded target objects, focusing exclusively on spatial perception tasks.
It includes the following seven question categories:
These samples are designed to support more precise evaluation and training for object-grounded spatial perception in autonomous driving scenarios.
This archive contains a curated set of high-quality trajectory-related VQA samples obtained after trajectory filtering and cleaning.
It covers the following five trajectory-centric categories:
These samples are intended for motion planning and trajectory reasoning tasks, with improved annotation quality after trajectory validation and filtering.
This archive contains the full implementation of the dataset construction pipeline. The components cover data preprocessing, perception annotation, trajectory generation, VQA construction, and post-processing.
The main scripts are listed below:
clip.py
Extracts image frames from raw videos:
mask.py
Generates image masks based on 2D bounding boxes:
Old_vqa_Q1-19.py
Original VQA generation script:
Q1-6-10-11_new.py
Updated VQA generation logic for selected categories:
Region[0])Q20-21-22_new.py
Updated generation for additional spatial reasoning categories:
scene_description.py
Generates scene-level descriptions:
VGGT.py
Core perception annotation module:
traj_clean.py
Trajectory post-processing module:
classify.py
VQA classification and selection tool:
These scripts together define the complete and reproducible pipeline for building the ScenePilot-4K dataset, from raw video processing to structured multimodal annotations.
This file lists all videos used in the dataset along with their corresponding download links. It is provided to support dataset reproduction and access to the original video resources.
@misc{wang2026scenepilotbenchlargescaledatasetbenchmark,
title={ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving},
author={Yujin Wang and Yutong Zheng and Wenxian Fan and Tianyi Wang and Hongqing Chu and Li Zhang and Bingzhao Gao and Daxin Tian and Hong Chen},
year={2026},
eprint={2601.19582},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2601.19582},
}