> Unofficial redistribution of the KITTI 2D object detection benchmark's labelled data, at original resolution and PNG fidelity, split using the Chen et al. (2015) train/validation convention, under the original CC BY-NC-SA 3.0 license.
0
45 commits
2 linked in READMEs
updated Oct 1, 2026
Unofficial redistribution of the KITTI 2D object detection benchmark's labelled data, at original resolution and PNG fidelity, split using the Chen et al. (2015) train/validation convention, under the original CC BY-NC-SA 3.0 license.
This repository is not an official release of the KITTI dataset.
KITTI was created by Andreas Geiger, Philip Lenz, and Raquel Urtasun at the Karlsruhe Institute of Technology and the Toyota Technological Institute at Chicago, who retain all copyright and intellectual property rights. This repository does not claim ownership of any images, annotations, or metadata.
This repository is sourced directly from the official KITTI download (registration required at the official site) β the "left color images" and "training labels" archives β not from any third-party mirror or repackaging. Images are the original PNGs, byte-identical to the official release; no compression, resizing, or augmentation of any kind has been applied.
If you use this dataset, please respect the license terms below and cite the original paper, not this repository.
KITTI is a foundational autonomous-driving benchmark: street scenes captured from a moving vehicle in and around Karlsruhe, Germany, using a stereo camera rig, annotated for 2D object detection across 8 classes (Car, Cyclist, Misc, Pedestrian, Person_sitting, Tram, Truck, Van). It is one of the most widely cited datasets in autonomous-driving perception research.
Only the officially labelled portion is redistributable at all: KITTI's test split (used for the official leaderboard) has never had public ground-truth boxes, so this repository covers the labelled training release only.
No official train/validation split exists. KITTI's authors only ever published one labelled set (7,481 images) versus the unlabelled test set. This repository uses the split introduced by Chen, Kundu, Zhu, Berneshawi, Ma, Fidler, and Urtasun ("3DOP", NeurIPS 2015) β train 3,712 / valid 3,769 β which has become the de facto standard across the KITTI 3D and 2D detection literature (used by OpenPCDet, MMDetection3D, avod, SECOND, PointRCNN, and still the split reported in papers as recent as 2024β2025). Using this specific split, rather than an arbitrary one, means results on valid here are directly comparable to the wider published literature's "KITTI val" numbers.
.txt label file per image (type truncated occluded alpha left top right bottom h w l x y z ry, absolute-pixel box corners). This repository converts every box into YOLO's normalized class x_center y_center width height format, and into absolute-pixel COCO [x, y, w, h] boxes in a metadata.jsonl per split. Boxes were re-expressed, not resized or altered.DontCare regions dropped. KITTI's label files mark ambiguous or unlabelled areas with a DontCare entry β an ignore-region marker, not a real object class, meant to be excluded from scoring rather than predicted. These are not included as a class here (they were also never a detectable "thing").train/valid β there is no third split. Evaluate on valid.DontCare regions and (a small number of) zero-area degenerate boxes.<repo>/
βββ README.md
βββ kitti_banner.jpg
βββ data/
βββ data.yaml
βββ images/
β βββ train/ (*.png + metadata.jsonl)
β βββ valid/ (*.png + metadata.jsonl)
βββ labels/
βββ train/ (*.txt, mirrors images)
βββ valid/
where:
data/images/<split>/ contains the original, unmodified PNG images at native KITTI resolution (~1242x375, varies slightly by frame), plus a metadata.jsonl (file_name and objects.bbox as absolute-pixel COCO [x, y, w, h] with objects.categories) that drives the Hugging Face dataset viewer.data/labels/<split>/ contains one YOLO-format .txt annotation file per image (class x_center y_center width height, normalized), mirroring the image layout.data/data.yaml is the Ultralytics dataset configuration file (class names, split paths, relative to data/). It intentionally has no test: entry β see above.| id | class name | instances (all splits) | share |
|---|---|---|---|
| 0 | Car | 28,742 | 70.8% |
| 1 | Cyclist | 1,627 | 4.0% |
| 2 | Misc | 973 | 2.4% |
| 3 | Pedestrian | 4,487 | 11.1% |
| 4 | Person_sitting | 222 | 0.5% |
| 5 | Tram | 511 | 1.3% |
| 6 | Truck | 1,094 | 2.7% |
| 7 | Van | 2,914 | 7.2% |
Car dominates at 71% of all boxes; Person_sitting is very rare (0.5%). This mirrors the real-world distribution KITTI was recorded from (dense urban/suburban driving).
Are we ready for autonomous driving? The KITTI vision benchmark suite
Andreas Geiger, Philip Lenz, Raquel Urtasun
2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3354-3361.
DOI: 10.1109/CVPR.2012.6248074
ImageSets, which distributes this same split.All credit for collecting and annotating this dataset belongs entirely to the original KITTI authors: Andreas Geiger, Philip Lenz, and Raquel Urtasun, and the Karlsruhe Institute of Technology / Toyota Technological Institute at Chicago.
This repository only converts their annotations to a YOLO-compatible layout and applies the community-standard Chen et al. train/validation split on top of the official, unmodified images. It does not modify, reinterpret, or take credit for the underlying imagery or annotations.
If you use this dataset in your research, please cite the original publication below.
The official KITTI Vision Benchmark Suite is distributed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0) license, as stated on the official KITTI page.
Accordingly:
This repository is distributed under the same CC BY-NC-SA 3.0 license.
If you use this dataset, please cite:
@inproceedings{geiger2012kitti,
title={Are we ready for autonomous driving? The KITTI vision benchmark suite},
author={Geiger, Andreas and Lenz, Philip and Urtasun, Raquel},
booktitle={2012 IEEE Conference on Computer Vision and Pattern Recognition},
pages={3354--3361},
year={2012},
organization={IEEE},
doi={10.1109/CVPR.2012.6248074}
}
If you use the train/validation split specifically, please also credit:
@inproceedings{chen20153dop,
title={3D Object Proposals for Accurate Object Class Detection},
author={Chen, Xiaozhi and Kundu, Kaustav and Zhu, Yukun and Berneshawi, Andrew G and Ma, Huimin and Fidler, Sanja and Urtasun, Raquel},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2015},
doi={10.5555/2969239.2969287}
}
We sincerely thank Andreas Geiger, Philip Lenz, and Raquel Urtasun for creating and publicly releasing this foundational autonomous-driving benchmark, and Chen et al. for the train/validation split this repository adopts.
> Unofficial redistribution of the KITTI 2D object detection benchmark's labelled data, at original resolution and PNG fidelity, split using the Chen et al. (2015) train/validation convention, under the original CC BY-NC-SA 3.0 license.
0
45 commits
2 linked in READMEs
updated Oct 1, 2026
Unofficial redistribution of the KITTI 2D object detection benchmark's labelled data, at original resolution and PNG fidelity, split using the Chen et al. (2015) train/validation convention, under the original CC BY-NC-SA 3.0 license.
This repository is not an official release of the KITTI dataset.
KITTI was created by Andreas Geiger, Philip Lenz, and Raquel Urtasun at the Karlsruhe Institute of Technology and the Toyota Technological Institute at Chicago, who retain all copyright and intellectual property rights. This repository does not claim ownership of any images, annotations, or metadata.
This repository is sourced directly from the official KITTI download (registration required at the official site) β the "left color images" and "training labels" archives β not from any third-party mirror or repackaging. Images are the original PNGs, byte-identical to the official release; no compression, resizing, or augmentation of any kind has been applied.
If you use this dataset, please respect the license terms below and cite the original paper, not this repository.
KITTI is a foundational autonomous-driving benchmark: street scenes captured from a moving vehicle in and around Karlsruhe, Germany, using a stereo camera rig, annotated for 2D object detection across 8 classes (Car, Cyclist, Misc, Pedestrian, Person_sitting, Tram, Truck, Van). It is one of the most widely cited datasets in autonomous-driving perception research.
Only the officially labelled portion is redistributable at all: KITTI's test split (used for the official leaderboard) has never had public ground-truth boxes, so this repository covers the labelled training release only.
No official train/validation split exists. KITTI's authors only ever published one labelled set (7,481 images) versus the unlabelled test set. This repository uses the split introduced by Chen, Kundu, Zhu, Berneshawi, Ma, Fidler, and Urtasun ("3DOP", NeurIPS 2015) β train 3,712 / valid 3,769 β which has become the de facto standard across the KITTI 3D and 2D detection literature (used by OpenPCDet, MMDetection3D, avod, SECOND, PointRCNN, and still the split reported in papers as recent as 2024β2025). Using this specific split, rather than an arbitrary one, means results on valid here are directly comparable to the wider published literature's "KITTI val" numbers.
.txt label file per image (type truncated occluded alpha left top right bottom h w l x y z ry, absolute-pixel box corners). This repository converts every box into YOLO's normalized class x_center y_center width height format, and into absolute-pixel COCO [x, y, w, h] boxes in a metadata.jsonl per split. Boxes were re-expressed, not resized or altered.DontCare regions dropped. KITTI's label files mark ambiguous or unlabelled areas with a DontCare entry β an ignore-region marker, not a real object class, meant to be excluded from scoring rather than predicted. These are not included as a class here (they were also never a detectable "thing").train/valid β there is no third split. Evaluate on valid.DontCare regions and (a small number of) zero-area degenerate boxes.<repo>/
βββ README.md
βββ kitti_banner.jpg
βββ data/
βββ data.yaml
βββ images/
β βββ train/ (*.png + metadata.jsonl)
β βββ valid/ (*.png + metadata.jsonl)
βββ labels/
βββ train/ (*.txt, mirrors images)
βββ valid/
where:
data/images/<split>/ contains the original, unmodified PNG images at native KITTI resolution (~1242x375, varies slightly by frame), plus a metadata.jsonl (file_name and objects.bbox as absolute-pixel COCO [x, y, w, h] with objects.categories) that drives the Hugging Face dataset viewer.data/labels/<split>/ contains one YOLO-format .txt annotation file per image (class x_center y_center width height, normalized), mirroring the image layout.data/data.yaml is the Ultralytics dataset configuration file (class names, split paths, relative to data/). It intentionally has no test: entry β see above.| id | class name | instances (all splits) | share |
|---|---|---|---|
| 0 | Car | 28,742 | 70.8% |
| 1 | Cyclist | 1,627 | 4.0% |
| 2 | Misc | 973 | 2.4% |
| 3 | Pedestrian | 4,487 | 11.1% |
| 4 | Person_sitting | 222 | 0.5% |
| 5 | Tram | 511 | 1.3% |
| 6 | Truck | 1,094 | 2.7% |
| 7 | Van | 2,914 | 7.2% |
Car dominates at 71% of all boxes; Person_sitting is very rare (0.5%). This mirrors the real-world distribution KITTI was recorded from (dense urban/suburban driving).
Are we ready for autonomous driving? The KITTI vision benchmark suite
Andreas Geiger, Philip Lenz, Raquel Urtasun
2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3354-3361.
DOI: 10.1109/CVPR.2012.6248074
ImageSets, which distributes this same split.All credit for collecting and annotating this dataset belongs entirely to the original KITTI authors: Andreas Geiger, Philip Lenz, and Raquel Urtasun, and the Karlsruhe Institute of Technology / Toyota Technological Institute at Chicago.
This repository only converts their annotations to a YOLO-compatible layout and applies the community-standard Chen et al. train/validation split on top of the official, unmodified images. It does not modify, reinterpret, or take credit for the underlying imagery or annotations.
If you use this dataset in your research, please cite the original publication below.
The official KITTI Vision Benchmark Suite is distributed under the Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0) license, as stated on the official KITTI page.
Accordingly:
This repository is distributed under the same CC BY-NC-SA 3.0 license.
If you use this dataset, please cite:
@inproceedings{geiger2012kitti,
title={Are we ready for autonomous driving? The KITTI vision benchmark suite},
author={Geiger, Andreas and Lenz, Philip and Urtasun, Raquel},
booktitle={2012 IEEE Conference on Computer Vision and Pattern Recognition},
pages={3354--3361},
year={2012},
organization={IEEE},
doi={10.1109/CVPR.2012.6248074}
}
If you use the train/validation split specifically, please also credit:
@inproceedings{chen20153dop,
title={3D Object Proposals for Accurate Object Class Detection},
author={Chen, Xiaozhi and Kundu, Kaustav and Zhu, Yukun and Berneshawi, Andrew G and Ma, Huimin and Fidler, Sanja and Urtasun, Raquel},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2015},
doi={10.5555/2969239.2969287}
}
We sincerely thank Andreas Geiger, Philip Lenz, and Raquel Urtasun for creating and publicly releasing this foundational autonomous-driving benchmark, and Chen et al. for the train/validation split this repository adopts.