EduardoPacheco/FoodSeg103

Dataset

17

stars

12

commits

3

linked in READMEs

Apr 29, 2024

updated

README

Dataset Card for FoodSeg103

Table of Contents

Dataset Description

Dataset Summary

FoodSeg103 is a large-scale benchmark for food image segmentation. It contains 103 food categories and 7118 images with ingredient level pixel-wise annotations. The dataset is a curated sample from Recipe1M and annotated and refined by human annotators. The dataset is split into 2 subsets: training set, validation set. The training set contains 4983 images and the validation set contains 2135 images.

Supported Tasks and Leaderboards

No leaderboard is available for this dataset at the moment.

Dataset Structure

Data categories

idingridient
0background
1candy
2egg tart
3french fries
4chocolate
5biscuit
6popcorn
7pudding
8ice cream
9cheese butter
10cake
11wine
12milkshake
13coffee
14juice
15milk
16tea
17almond
18red beans
19cashew
20dried cranberries
21soy
22walnut
23peanut
24egg
25apple
26date
27apricot
28avocado
29banana
30strawberry
31cherry
32blueberry
33raspberry
34mango
35olives
36peach
37lemon
38pear
39fig
40pineapple
41grape
42kiwi
43melon
44orange
45watermelon
46steak
47pork
48chicken duck
49sausage
50fried meat
51lamb
52sauce
53crab
54fish
55shellfish
56shrimp
57soup
58bread
59corn
60hamburg
61pizza
62hanamaki baozi
63wonton dumplings
64pasta
65noodles
66rice
67pie
68tofu
69eggplant
70potato
71garlic
72cauliflower
73tomato
74kelp
75seaweed
76spring onion
77rape
78ginger
79okra
80lettuce
81pumpkin
82cucumber
83white radish
84carrot
85asparagus
86bamboo shoots
87broccoli
88celery stick
89cilantro mint
90snow peas
91cabbage
92bean sprouts
93onion
94pepper
95green beans
96French beans
97king oyster mushroom
98shiitake
99enoki mushroom
100oyster mushroom
101white button mushroom
102salad
103other ingredients

Data Splits

This dataset only contains two splits. A training split and a validation split with 4983 and 2135 images respectively.

Dataset Creation

Curation Rationale

Select images from a large-scale recipe dataset and annotate them with pixel-wise segmentation masks.

Source Data

The dataset is a curated sample from Recipe1M.

Initial Data Collection and Normalization

After selecting the source of the data two more steps were added before image selection.

  1. Recipe1M contains 1.5k ingredient categoris, but only the top 124 categories were selected + a 'other' category (further became 103).
  2. Images should contain between 2 and 16 ingredients.
  3. Ingredients should be visible and easy to annotate.

Which then resulted in 7118 images.

Annotations

Annotation process

Third party annotators were hired to annotate the images respecting the following guidelines:

  1. Tag ingredients with appropriate categories.
  2. Draw pixel-wise masks for each ingredient.
  3. Ignore tiny regions (even if contains ingredients) with area covering less than 5% of the image.

Refinement process

The refinement process implemented the following steps:

  1. Correct mislabelled ingredients.
  2. Deleting unpopular categories that are assigned to less than 5 images (resulting in 103 categories in the final dataset).
  3. Merging visually similar ingredient categories (e.g. orange and citrus)

Who are the annotators?

A third party company that was not mentioned in the paper.

Additional Information

Dataset Curators

Authors of the paper A Large-Scale Benchmark for Food Image Segmentation.

Licensing Information

Apache 2.0 license.

Citation Information

@inproceedings{wu2021foodseg,
	title={A Large-Scale Benchmark for Food Image Segmentation},
	author={Wu, Xiongwei and Fu, Xin and Liu, Ying and Lim, Ee-Peng and Hoi, Steven CH and Sun, Qianru},
	booktitle={Proceedings of ACM international conference on Multimedia},
	year={2021}
}

Contributors

EduardoPacheco

11 commits

EP

EduardoPacheco/FoodSeg103

Dataset

17

stars

12

commits

3

linked in READMEs

Apr 29, 2024

updated

README

Dataset Card for FoodSeg103

Table of Contents

Dataset Description

Dataset Summary

FoodSeg103 is a large-scale benchmark for food image segmentation. It contains 103 food categories and 7118 images with ingredient level pixel-wise annotations. The dataset is a curated sample from Recipe1M and annotated and refined by human annotators. The dataset is split into 2 subsets: training set, validation set. The training set contains 4983 images and the validation set contains 2135 images.

Supported Tasks and Leaderboards

No leaderboard is available for this dataset at the moment.

Dataset Structure

Data categories

idingridient
0background
1candy
2egg tart
3french fries
4chocolate
5biscuit
6popcorn
7pudding
8ice cream
9cheese butter
10cake
11wine
12milkshake
13coffee
14juice
15milk
16tea
17almond
18red beans
19cashew
20dried cranberries
21soy
22walnut
23peanut
24egg
25apple
26date
27apricot
28avocado
29banana
30strawberry
31cherry
32blueberry
33raspberry
34mango
35olives
36peach
37lemon
38pear
39fig
40pineapple
41grape
42kiwi
43melon
44orange
45watermelon
46steak
47pork
48chicken duck
49sausage
50fried meat
51lamb
52sauce
53crab
54fish
55shellfish
56shrimp
57soup
58bread
59corn
60hamburg
61pizza
62hanamaki baozi
63wonton dumplings
64pasta
65noodles
66rice
67pie
68tofu
69eggplant
70potato
71garlic
72cauliflower
73tomato
74kelp
75seaweed
76spring onion
77rape
78ginger
79okra
80lettuce
81pumpkin
82cucumber
83white radish
84carrot
85asparagus
86bamboo shoots
87broccoli
88celery stick
89cilantro mint
90snow peas
91cabbage
92bean sprouts
93onion
94pepper
95green beans
96French beans
97king oyster mushroom
98shiitake
99enoki mushroom
100oyster mushroom
101white button mushroom
102salad
103other ingredients

Data Splits

This dataset only contains two splits. A training split and a validation split with 4983 and 2135 images respectively.

Dataset Creation

Curation Rationale

Select images from a large-scale recipe dataset and annotate them with pixel-wise segmentation masks.

Source Data

The dataset is a curated sample from Recipe1M.

Initial Data Collection and Normalization

After selecting the source of the data two more steps were added before image selection.

  1. Recipe1M contains 1.5k ingredient categoris, but only the top 124 categories were selected + a 'other' category (further became 103).
  2. Images should contain between 2 and 16 ingredients.
  3. Ingredients should be visible and easy to annotate.

Which then resulted in 7118 images.

Annotations

Annotation process

Third party annotators were hired to annotate the images respecting the following guidelines:

  1. Tag ingredients with appropriate categories.
  2. Draw pixel-wise masks for each ingredient.
  3. Ignore tiny regions (even if contains ingredients) with area covering less than 5% of the image.

Refinement process

The refinement process implemented the following steps:

  1. Correct mislabelled ingredients.
  2. Deleting unpopular categories that are assigned to less than 5 images (resulting in 103 categories in the final dataset).
  3. Merging visually similar ingredient categories (e.g. orange and citrus)

Who are the annotators?

A third party company that was not mentioned in the paper.

Additional Information

Dataset Curators

Authors of the paper A Large-Scale Benchmark for Food Image Segmentation.

Licensing Information

Apache 2.0 license.

Citation Information

@inproceedings{wu2021foodseg,
	title={A Large-Scale Benchmark for Food Image Segmentation},
	author={Wu, Xiongwei and Fu, Xin and Liu, Ying and Lim, Ee-Peng and Hoi, Steven CH and Sun, Qianru},
	booktitle={Proceedings of ACM international conference on Multimedia},
	year={2021}
}

Contributors

EduardoPacheco

11 commits

EP