ayeshaishaq/DriveLMMo1

Dataset

DriveLMM-o1 Dataset: Step-by-Step Reasoning for Autonomous Driving

6

5 commits

1 linked in READMEs

updated Mar 17, 2025

See the code

README

DriveLMM-o1 Dataset: Step-by-Step Reasoning for Autonomous Driving

The DriveLMM-o1 dataset is a benchmark designed to evaluate and train models on step-by-step reasoning in autonomous driving. It comprises over 18,000 visual question-answer pairs (VQAs) in the training set and more than 4,000 in the test set. Each example is enriched with manually curated reasoning annotations covering perception, prediction, and planning tasks.

Paper:

Key Features:

  • Multimodal Data: Incorporates both multiview images and LiDAR point clouds to capture a comprehensive real-world driving context.
  • Detailed Reasoning Annotations: Each Q&A pair is supported by multiple reasoning steps, enabling models to learn interpretable and logical decision-making.
  • Diverse Scenarios: Covers various urban and highway driving conditions for robust evaluation.
  • Driving-Specific Evaluation: Includes metrics for risk assessment, traffic rule adherence, and scene awareness to assess reasoning quality.

Data Preparation Instructions:

  • NuScenes Data Requirement:
    To use the DriveLMM-o1 dataset, you must download the images and LiDAR point clouds from the nuScenes trainval set (keyframes only)
  • Access and Licensing:
    Visit the nuScenes website for download instructions and licensing details.

Dataset Comparison:

The table below compares the DriveLMM-o1 dataset with other prominent autonomous driving benchmarks:

DatasetTrain FramesTrain QAsTest FramesTest QAsStep-by-Step ReasoningInput ModalitiesImage ViewsFinal AnnotationsSource
BDD-X [19]5,58823k6982,652✗Video1ManualBerkeley Deep Drive
NuScenes-QA [5]28k376k6,01983k✗Images, Points6AutomatedNuScenes
DriveLM [1]4,063377k79915k✗Images6Mostly-AutomatedNuScenes, CARLA
LingoQA [20]28k420k1001,000✗Video1Manual–
Reason2Drive [21]420k420k180k180k✗Video1AutomatedNuScenes, Waymo, OPEN
DrivingVQA [22]3,1423,142789789✓Image2ManualCode de la Route
DriveLMM-o1 (Ours)1,96218k5394,633✓Images, Points6ManualNuScenes
autonomous-driving
multimodal
reasoning

ayeshaishaq/DriveLMMo1

Dataset

DriveLMM-o1 Dataset: Step-by-Step Reasoning for Autonomous Driving

6

5 commits

1 linked in READMEs

updated Mar 17, 2025

See the code

README

DriveLMM-o1 Dataset: Step-by-Step Reasoning for Autonomous Driving

The DriveLMM-o1 dataset is a benchmark designed to evaluate and train models on step-by-step reasoning in autonomous driving. It comprises over 18,000 visual question-answer pairs (VQAs) in the training set and more than 4,000 in the test set. Each example is enriched with manually curated reasoning annotations covering perception, prediction, and planning tasks.

Paper:

Key Features:

  • Multimodal Data: Incorporates both multiview images and LiDAR point clouds to capture a comprehensive real-world driving context.
  • Detailed Reasoning Annotations: Each Q&A pair is supported by multiple reasoning steps, enabling models to learn interpretable and logical decision-making.
  • Diverse Scenarios: Covers various urban and highway driving conditions for robust evaluation.
  • Driving-Specific Evaluation: Includes metrics for risk assessment, traffic rule adherence, and scene awareness to assess reasoning quality.

Data Preparation Instructions:

  • NuScenes Data Requirement:
    To use the DriveLMM-o1 dataset, you must download the images and LiDAR point clouds from the nuScenes trainval set (keyframes only)
  • Access and Licensing:
    Visit the nuScenes website for download instructions and licensing details.

Dataset Comparison:

The table below compares the DriveLMM-o1 dataset with other prominent autonomous driving benchmarks:

DatasetTrain FramesTrain QAsTest FramesTest QAsStep-by-Step ReasoningInput ModalitiesImage ViewsFinal AnnotationsSource
BDD-X [19]5,58823k6982,652✗Video1ManualBerkeley Deep Drive
NuScenes-QA [5]28k376k6,01983k✗Images, Points6AutomatedNuScenes
DriveLM [1]4,063377k79915k✗Images6Mostly-AutomatedNuScenes, CARLA
LingoQA [20]28k420k1001,000✗Video1Manual–
Reason2Drive [21]420k420k180k180k✗Video1AutomatedNuScenes, Waymo, OPEN
DrivingVQA [22]3,1423,142789789✓Image2ManualCode de la Route
DriveLMM-o1 (Ours)1,96218k5394,633✓Images, Points6ManualNuScenes
autonomous-driving
multimodal
reasoning