dronefreak/BDD100K-Semantic-Segmentation

Dataset

> Unofficial redistribution of BDD100K's semantic segmentation task (10K-image subset), under BDD100K's own data license (non-commercial/educational/research redistribution, with notice, explicitly permitted).

0

33 commits

2 linked in READMEs

updated Oct 1, 2026

See the code

README

BDD100K: Semantic Segmentation Dataset

BDD100K Semantic Segmentation Dataset Banner

Task Dataset Classes Splits License

Unofficial redistribution of BDD100K's semantic segmentation task (10K-image subset), under BDD100K's own data license (non-commercial/educational/research redistribution, with notice, explicitly permitted).

Dataset Description

Disclaimer

This repository is not an official release of BDD100K.

BDD100K was created by Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell at UC Berkeley (BAIR), who retain all copyright. This repository does not claim ownership of any images, annotations, or metadata.

This redistribution is sourced directly from BDD100K's official download (registration required at https://bdd-data.berkeley.edu/), not from any third-party mirror. Unlike this project's other BDD100K repositories, this one is not a re-sort of the 100K detection image set β€” semantic segmentation is a genuinely separate, smaller 10K-image subset with its own dedicated pixel-level annotations (pixel labeling is far more expensive than boxes or image-level tags, so only a sample of the full 100K images was ever annotated for this task).

Companion repositories cover BDD100K's other tasks: object detection, and the per-image attribute classification tasks β€” weather, scenario, and period/time-of-day.


Dataset Overview

BDD100K's semantic segmentation task provides dense, pixel-level class labels for a 10,000-image subset of the full dataset: 7,000 train + 1,000 validation images at 1280x720, each with a per-pixel class mask. The task uses a 19-class taxonomy explicitly designed to be compatible with Cityscapes (verified directly from the official bdd100k/label/label.py source), covering flat surfaces, construction, objects, nature, sky, humans, and vehicles. A reserved 255 ("ignore") value marks pixels that are unlabelled or belong to a class excluded from evaluation (e.g. ego-vehicle hood, dynamic/static clutter) β€” these should be excluded from loss and metric computation, not treated as a 20th real class.

No test split. BDD100K's official test images for this task (2,000 images) have no released masks (held out for the leaderboard) and are not included here.


Changes from the Official Release

  • No format conversion. Unlike this project's detection/classification BDD100K repositories, the official release already ships exactly what's needed β€” JPEG images, single-channel trainId-encoded PNG masks, and RGBA color-visualization PNGs β€” so this repository redistributes them unmodified, just reorganized and sharded for Hugging Face.
  • Three files per example, not one. Each example has an RGB photo, a grayscale class-id mask, and a color visualization mask, linked together via Hugging Face's multi-image *_file_name metadata convention (see Dataset Structure) rather than the single-image file_name convention this project's other cards use.
  • Sharded for Hub limits. Because each example contributes 3 files, shards here are capped at 3,000 examples (9,001 files) rather than ~8,000, to stay under Hugging Face's practical per-directory file-count ceiling (train: 3 shards, validation: 1 shard). This has no effect on the data itself.
  • An id2label.json was added at the repository root (not part of the official release) mapping all 19 trainId values β€” plus 255 β€” to their class names, for convenience. This mirrors the convention used by Hugging Face's own official segmentation sample datasets.
  • No image or mask pixel content was modified. No examples were added or removed relative to the official release.

Dataset Structure

<repo>/
β”œβ”€β”€ README.md
β”œβ”€β”€ bdd100k_segmentation_banner.jpg
β”œβ”€β”€ id2label.json
└── data/
    └── images/
        β”œβ”€β”€ train/
        β”‚   └── shard_000 … shard_002/   (3 files/example + metadata.jsonl)
        └── valid/
            └── shard_000/                (3 files/example + metadata.jsonl)

Each shard directory contains, for every example: <id>.jpg (the photo), <id>_train_id.png (the grayscale class mask), <id>_train_color.png (an RGBA color visualization of the same mask), and one metadata.jsonl with one line per example:

{"file_name": "<id>.jpg", "mask_file_name": "<id>_train_id.png", "color_mask_file_name": "<id>_train_color.png"}

Hugging Face's imagefolder loader auto-detects any metadata key ending in _file_name as an image reference, resolved relative to the metadata.jsonl it's listed in. This yields three image columns when loaded with datasets: image (the photo, from the bare file_name key), mask (the class-id PNG), and color_mask (the RGBA visualization) β€” verified directly against the datasets library's own loader source, not assumed.

Splits: train 7,000 images Β· validation 1,000 images (8,000 total β€” matches BDD100K's official semantic segmentation subset counts exactly).

Classes (19 + ignore)

Pixel share and per-image presence, computed directly from all 8,000 masks in this repository:

idclasstrain pixel sharetrain presenceval pixel shareval presence
0road21.49%6,877/7,00021.60%992/1,000
1sidewalk2.03%4,667/7,0002.05%732/1,000
2building13.26%6,189/7,00014.89%893/1,000
3wall0.48%1,077/7,0000.36%119/1,000
4fence1.03%2,141/7,0000.81%241/1,000
5pole0.92%6,648/7,0000.97%961/1,000
6traffic light0.18%3,293/7,0000.14%528/1,000
7traffic sign0.34%5,270/7,0000.23%776/1,000
8vegetation13.21%6,422/7,00015.42%958/1,000
9terrain1.03%2,563/7,0000.91%384/1,000
10sky17.30%6,635/7,00017.91%977/1,000
11person0.25%2,429/7,0000.27%379/1,000
12rider0.02%361/7,0000.01%39/1,000
13car8.11%6,811/7,0009.07%971/1,000
14truck0.97%2,137/7,0001.01%331/1,000
15bus0.56%1,051/7,0000.63%136/1,000
16train0.01%47/7,0000.01%7/1,000
17motorcycle0.02%266/7,0000.02%42/1,000
18bicycle0.05%447/7,0000.02%45/1,000
255unlabeled/ignore18.72%7,000/7,00013.67%1,000/1,000

road, sky, building, and vegetation together cover roughly two-thirds of labelled pixels; rider, train, and motorcycle are extremely rare both by pixel area and by image presence (train appears in only 47 of 7,000 training images). The 255 ignore value covers a substantial 14-19% of pixels β€” exclude it from loss/metric computation, it is not a 20th real class.


Dataset Sources

Original Paper

BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, Trevor Darrell

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pages 2633-2642. DOI: 10.1109/CVPR42600.2020.00271

Official Resources


Attribution

All credit for collecting and annotating this dataset belongs entirely to the original BDD100K authors: Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell, and UC Berkeley / BAIR.

This repository only reorganizes and shards the official files for Hugging Face compatibility; it does not modify, reinterpret, or take credit for the underlying imagery or annotations.

If you use this dataset in your research, please cite the original publication below.


License

BDD100K's code repository is BSD-3-Clause, but that license applies only to the code, not the data. The data and labels (downloaded from https://bdd-data.berkeley.edu/) carry their own license, reproduced here in full as required by its own terms:

Copyright Β©2018. The Regents of the University of California (Regents). All Rights Reserved.

THIS SOFTWARE AND/OR DATA WAS DEPOSITED IN THE BAIR OPEN RESEARCH COMMONS REPOSITORY ON 1/1/2021

Permission to use, copy, modify, and distribute this software and its documentation for educational, research, and not-for-profit purposes, without fee and without a signed licensing agreement; and permission to use, copy, modify and distribute this software for commercial purposes (such rights not subject to transfer) to BDD and BAIR Commons members and their affiliates, is hereby granted, provided that the above copyright notice, this paragraph and the following two paragraphs appear in all copies, modifications, and distributions. Contact The Office of Technology Licensing, UC Berkeley, 2150 Shattuck Avenue, Suite 510, Berkeley, CA 94720-1620, (510) 643-7201, otl@berkeley.edu, http://ipira.berkeley.edu/industry-info for commercial licensing opportunities.

IN NO EVENT SHALL REGENTS BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES, INCLUDING LOST PROFITS, ARISING OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF REGENTS HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

REGENTS SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE SOFTWARE AND ACCOMPANYING DOCUMENTATION, IF ANY, PROVIDED HEREUNDER IS PROVIDED "AS IS". REGENTS HAS NO OBLIGATION TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS.

Source: https://github.com/bdd100k/bdd100k/blob/master/doc/source/license.rst

Accordingly:

  • Redistribution for educational, research, and not-for-profit purposes is explicitly permitted, without fee, provided this notice is carried forward β€” which is what this repository does.
  • Commercial use and distribution is restricted to BDD and BAIR Commons members and their affiliates. Contact UC Berkeley's Office of Technology Licensing (contact details above) for commercial licensing.
  • This repository, and any further redistribution of it, must carry this same notice.

Citation

If you use this dataset, please cite:

@inproceedings{yu2020bdd100k,
  title={BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning},
  author={Yu, Fisher and Chen, Haofeng and Wang, Xin and Xian, Wenqi and Chen, Yingying and Liu, Fangchen and Madhavan, Vashisht and Darrell, Trevor},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={2633--2642},
  year={2020}
}

Acknowledgements

We sincerely thank Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell, and UC Berkeley / BAIR, for creating and publicly releasing this valuable driving-scene benchmark.

autonomous-driving
computer-vision
image-segmentation
scene-understanding
semantic-segmentation

dronefreak/BDD100K-Semantic-Segmentation

Dataset

> Unofficial redistribution of BDD100K's semantic segmentation task (10K-image subset), under BDD100K's own data license (non-commercial/educational/research redistribution, with notice, explicitly permitted).

0

33 commits

2 linked in READMEs

updated Oct 1, 2026

See the code

README

BDD100K: Semantic Segmentation Dataset

BDD100K Semantic Segmentation Dataset Banner

Task Dataset Classes Splits License

Unofficial redistribution of BDD100K's semantic segmentation task (10K-image subset), under BDD100K's own data license (non-commercial/educational/research redistribution, with notice, explicitly permitted).

Dataset Description

Disclaimer

This repository is not an official release of BDD100K.

BDD100K was created by Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell at UC Berkeley (BAIR), who retain all copyright. This repository does not claim ownership of any images, annotations, or metadata.

This redistribution is sourced directly from BDD100K's official download (registration required at https://bdd-data.berkeley.edu/), not from any third-party mirror. Unlike this project's other BDD100K repositories, this one is not a re-sort of the 100K detection image set β€” semantic segmentation is a genuinely separate, smaller 10K-image subset with its own dedicated pixel-level annotations (pixel labeling is far more expensive than boxes or image-level tags, so only a sample of the full 100K images was ever annotated for this task).

Companion repositories cover BDD100K's other tasks: object detection, and the per-image attribute classification tasks β€” weather, scenario, and period/time-of-day.


Dataset Overview

BDD100K's semantic segmentation task provides dense, pixel-level class labels for a 10,000-image subset of the full dataset: 7,000 train + 1,000 validation images at 1280x720, each with a per-pixel class mask. The task uses a 19-class taxonomy explicitly designed to be compatible with Cityscapes (verified directly from the official bdd100k/label/label.py source), covering flat surfaces, construction, objects, nature, sky, humans, and vehicles. A reserved 255 ("ignore") value marks pixels that are unlabelled or belong to a class excluded from evaluation (e.g. ego-vehicle hood, dynamic/static clutter) β€” these should be excluded from loss and metric computation, not treated as a 20th real class.

No test split. BDD100K's official test images for this task (2,000 images) have no released masks (held out for the leaderboard) and are not included here.


Changes from the Official Release

  • No format conversion. Unlike this project's detection/classification BDD100K repositories, the official release already ships exactly what's needed β€” JPEG images, single-channel trainId-encoded PNG masks, and RGBA color-visualization PNGs β€” so this repository redistributes them unmodified, just reorganized and sharded for Hugging Face.
  • Three files per example, not one. Each example has an RGB photo, a grayscale class-id mask, and a color visualization mask, linked together via Hugging Face's multi-image *_file_name metadata convention (see Dataset Structure) rather than the single-image file_name convention this project's other cards use.
  • Sharded for Hub limits. Because each example contributes 3 files, shards here are capped at 3,000 examples (9,001 files) rather than ~8,000, to stay under Hugging Face's practical per-directory file-count ceiling (train: 3 shards, validation: 1 shard). This has no effect on the data itself.
  • An id2label.json was added at the repository root (not part of the official release) mapping all 19 trainId values β€” plus 255 β€” to their class names, for convenience. This mirrors the convention used by Hugging Face's own official segmentation sample datasets.
  • No image or mask pixel content was modified. No examples were added or removed relative to the official release.

Dataset Structure

<repo>/
β”œβ”€β”€ README.md
β”œβ”€β”€ bdd100k_segmentation_banner.jpg
β”œβ”€β”€ id2label.json
└── data/
    └── images/
        β”œβ”€β”€ train/
        β”‚   └── shard_000 … shard_002/   (3 files/example + metadata.jsonl)
        └── valid/
            └── shard_000/                (3 files/example + metadata.jsonl)

Each shard directory contains, for every example: <id>.jpg (the photo), <id>_train_id.png (the grayscale class mask), <id>_train_color.png (an RGBA color visualization of the same mask), and one metadata.jsonl with one line per example:

{"file_name": "<id>.jpg", "mask_file_name": "<id>_train_id.png", "color_mask_file_name": "<id>_train_color.png"}

Hugging Face's imagefolder loader auto-detects any metadata key ending in _file_name as an image reference, resolved relative to the metadata.jsonl it's listed in. This yields three image columns when loaded with datasets: image (the photo, from the bare file_name key), mask (the class-id PNG), and color_mask (the RGBA visualization) β€” verified directly against the datasets library's own loader source, not assumed.

Splits: train 7,000 images Β· validation 1,000 images (8,000 total β€” matches BDD100K's official semantic segmentation subset counts exactly).

Classes (19 + ignore)

Pixel share and per-image presence, computed directly from all 8,000 masks in this repository:

idclasstrain pixel sharetrain presenceval pixel shareval presence
0road21.49%6,877/7,00021.60%992/1,000
1sidewalk2.03%4,667/7,0002.05%732/1,000
2building13.26%6,189/7,00014.89%893/1,000
3wall0.48%1,077/7,0000.36%119/1,000
4fence1.03%2,141/7,0000.81%241/1,000
5pole0.92%6,648/7,0000.97%961/1,000
6traffic light0.18%3,293/7,0000.14%528/1,000
7traffic sign0.34%5,270/7,0000.23%776/1,000
8vegetation13.21%6,422/7,00015.42%958/1,000
9terrain1.03%2,563/7,0000.91%384/1,000
10sky17.30%6,635/7,00017.91%977/1,000
11person0.25%2,429/7,0000.27%379/1,000
12rider0.02%361/7,0000.01%39/1,000
13car8.11%6,811/7,0009.07%971/1,000
14truck0.97%2,137/7,0001.01%331/1,000
15bus0.56%1,051/7,0000.63%136/1,000
16train0.01%47/7,0000.01%7/1,000
17motorcycle0.02%266/7,0000.02%42/1,000
18bicycle0.05%447/7,0000.02%45/1,000
255unlabeled/ignore18.72%7,000/7,00013.67%1,000/1,000

road, sky, building, and vegetation together cover roughly two-thirds of labelled pixels; rider, train, and motorcycle are extremely rare both by pixel area and by image presence (train appears in only 47 of 7,000 training images). The 255 ignore value covers a substantial 14-19% of pixels β€” exclude it from loss/metric computation, it is not a 20th real class.


Dataset Sources

Original Paper

BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, Trevor Darrell

Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pages 2633-2642. DOI: 10.1109/CVPR42600.2020.00271

Official Resources


Attribution

All credit for collecting and annotating this dataset belongs entirely to the original BDD100K authors: Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell, and UC Berkeley / BAIR.

This repository only reorganizes and shards the official files for Hugging Face compatibility; it does not modify, reinterpret, or take credit for the underlying imagery or annotations.

If you use this dataset in your research, please cite the original publication below.


License

BDD100K's code repository is BSD-3-Clause, but that license applies only to the code, not the data. The data and labels (downloaded from https://bdd-data.berkeley.edu/) carry their own license, reproduced here in full as required by its own terms:

Copyright Β©2018. The Regents of the University of California (Regents). All Rights Reserved.

THIS SOFTWARE AND/OR DATA WAS DEPOSITED IN THE BAIR OPEN RESEARCH COMMONS REPOSITORY ON 1/1/2021

Permission to use, copy, modify, and distribute this software and its documentation for educational, research, and not-for-profit purposes, without fee and without a signed licensing agreement; and permission to use, copy, modify and distribute this software for commercial purposes (such rights not subject to transfer) to BDD and BAIR Commons members and their affiliates, is hereby granted, provided that the above copyright notice, this paragraph and the following two paragraphs appear in all copies, modifications, and distributions. Contact The Office of Technology Licensing, UC Berkeley, 2150 Shattuck Avenue, Suite 510, Berkeley, CA 94720-1620, (510) 643-7201, otl@berkeley.edu, http://ipira.berkeley.edu/industry-info for commercial licensing opportunities.

IN NO EVENT SHALL REGENTS BE LIABLE TO ANY PARTY FOR DIRECT, INDIRECT, SPECIAL, INCIDENTAL, OR CONSEQUENTIAL DAMAGES, INCLUDING LOST PROFITS, ARISING OUT OF THE USE OF THIS SOFTWARE AND ITS DOCUMENTATION, EVEN IF REGENTS HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

REGENTS SPECIFICALLY DISCLAIMS ANY WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE. THE SOFTWARE AND ACCOMPANYING DOCUMENTATION, IF ANY, PROVIDED HEREUNDER IS PROVIDED "AS IS". REGENTS HAS NO OBLIGATION TO PROVIDE MAINTENANCE, SUPPORT, UPDATES, ENHANCEMENTS, OR MODIFICATIONS.

Source: https://github.com/bdd100k/bdd100k/blob/master/doc/source/license.rst

Accordingly:

  • Redistribution for educational, research, and not-for-profit purposes is explicitly permitted, without fee, provided this notice is carried forward β€” which is what this repository does.
  • Commercial use and distribution is restricted to BDD and BAIR Commons members and their affiliates. Contact UC Berkeley's Office of Technology Licensing (contact details above) for commercial licensing.
  • This repository, and any further redistribution of it, must carry this same notice.

Citation

If you use this dataset, please cite:

@inproceedings{yu2020bdd100k,
  title={BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning},
  author={Yu, Fisher and Chen, Haofeng and Wang, Xin and Xian, Wenqi and Chen, Yingying and Liu, Fangchen and Madhavan, Vashisht and Darrell, Trevor},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages={2633--2642},
  year={2020}
}

Acknowledgements

We sincerely thank Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell, and UC Berkeley / BAIR, for creating and publicly releasing this valuable driving-scene benchmark.

autonomous-driving
computer-vision
image-segmentation
scene-understanding
semantic-segmentation