SimLingo-Data is a large-scale autonomous driving CARLA 2.0 dataset containing sensor data, action labels, a wide range of simulator state information, and language labels for VQA, commentary and instruction following. The driving data is collected with the privileged rule-based expert PDM-Lite.
The dataset is organized hierarchically with the following main components:
data/: Raw sensor data (RGB, LiDAR, measurements, bounding boxes)commentary/: Natural language descriptions of driving decisionsdreamer/: Instruction following data with multiple instruction/action pairs per sampledrivelm/: VQA data, based on DriveLMThis dataset is chunked into groups of multiple routes for efficient download and processing.
# Clone the repository
git clone https://huggingface.co/datasets/RenzKa/simlingo
# Navigate to the directory
cd simlingo
# Pull the LFS files
git lfs pull
# Download individual files (replace with actual file URLs from Hugging Face)
wget https://huggingface.co/datasets/RenzKa/simlingo/resolve/main/[filename].tar.gz
# Create output directory
mkdir -p database/simlingo
# Extract all archives to the same directory
for file in *.tar.gz; do
echo "Extracting $file to database/simlingo/..."
tar -xzf "$file" -C database/simlingo/
done
Please refer to the license file for usage terms and conditions.
If you use this dataset in your research, please cite:
@inproceedings{renz2025simlingo,
title={SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment},
author={Renz, Katrin and Chen, Long and Arani, Elahe and Sinavski, Oleg},
booktitle={Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2025},
}
@inproceedings{sima2024drivelm,
title={DriveLM: Driving with Graph Visual Question Answering},
author={Chonghao Sima and Katrin Renz and Kashyap Chitta and Li Chen and Hanxue Zhang and Chengen Xie and Jens Beißwenger and Ping Luo and Andreas Geiger and Hongyang Li},
booktitle={European Conference on Computer Vision},
year={2024},
}
SimLingo-Data is a large-scale autonomous driving CARLA 2.0 dataset containing sensor data, action labels, a wide range of simulator state information, and language labels for VQA, commentary and instruction following. The driving data is collected with the privileged rule-based expert PDM-Lite.
The dataset is organized hierarchically with the following main components:
data/: Raw sensor data (RGB, LiDAR, measurements, bounding boxes)commentary/: Natural language descriptions of driving decisionsdreamer/: Instruction following data with multiple instruction/action pairs per sampledrivelm/: VQA data, based on DriveLMThis dataset is chunked into groups of multiple routes for efficient download and processing.
# Clone the repository
git clone https://huggingface.co/datasets/RenzKa/simlingo
# Navigate to the directory
cd simlingo
# Pull the LFS files
git lfs pull
# Download individual files (replace with actual file URLs from Hugging Face)
wget https://huggingface.co/datasets/RenzKa/simlingo/resolve/main/[filename].tar.gz
# Create output directory
mkdir -p database/simlingo
# Extract all archives to the same directory
for file in *.tar.gz; do
echo "Extracting $file to database/simlingo/..."
tar -xzf "$file" -C database/simlingo/
done
Please refer to the license file for usage terms and conditions.
If you use this dataset in your research, please cite:
@inproceedings{renz2025simlingo,
title={SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment},
author={Renz, Katrin and Chen, Long and Arani, Elahe and Sinavski, Oleg},
booktitle={Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2025},
}
@inproceedings{sima2024drivelm,
title={DriveLM: Driving with Graph Visual Question Answering},
author={Chonghao Sima and Katrin Renz and Kashyap Chitta and Li Chen and Hanxue Zhang and Chengen Xie and Jens Beißwenger and Ping Luo and Andreas Geiger and Hongyang Li},
booktitle={European Conference on Computer Vision},
year={2024},
}