HuggingFaceM4/charades

Dataset

Dataset Card for Charades

10

7 commits

2 linked in READMEs

updated Oct 20, 2022

See the code

README

Dataset Card for Charades

Table of Contents

Dataset Description

Dataset Summary

Charades is dataset composed of 9848 videos of daily indoors activities collected through Amazon Mechanical Turk. 267 different users were presented with a sentence, that includes objects and actions from a fixed vocabulary, and they recorded a video acting out the sentence (like in a game of Charades). The dataset contains 66,500 temporal annotations for 157 action classes, 41,104 labels for 46 object classes, and 27,847 textual descriptions of the videos

Supported Tasks and Leaderboards

  • multilabel-action-classification: The goal of this task is to classify actions happening in a video. This is a multilabel classification. The leaderboard is available here

Languages

The annotations in the dataset are in English.

Dataset Structure

Data Instances

{
  "video_id": "46GP8",
  "video": "/home/amanpreet_huggingface_co/.cache/huggingface/datasets/downloads/extracted/3f022da5305aaa189f09476dbf7d5e02f6fe12766b927c076707360d00deb44d/46GP8.mp4",
  "subject": "HR43",
  "scene": "Kitchen",
  "quality": 6,
  "relevance": 7,
  "verified": "Yes",
  "script": "A person cooking on a stove while watching something out a window.",
  "objects": ["food", "stove", "window"],
  "descriptions": [
    "A person cooks food on a stove before looking out of a window."
  ],
  "labels": [92, 147],
  "action_timings": [
    [11.899999618530273, 21.200000762939453],
    [0.0, 12.600000381469727]
  ],
  "length": 24.829999923706055
}

Data Fields

  • video_id: str Unique identifier for each video.
  • video: str Path to the video file
  • subject: str Unique identifier for each subject in the dataset
  • scene: str One of 15 indoor scenes in the dataset, such as Kitchen
  • quality: int The quality of the video judged by an annotator (7-point scale, 7=high quality), -100 if missing
  • relevance: int The relevance of the video to the script judged by an annotated (7-point scale, 7=very relevant), -100 if missing
  • verified: str 'Yes' if an annotator successfully verified that the video matches the script, else 'No'
  • script: str The human-generated script used to generate the video
  • descriptions: List[str] List of descriptions by annotators watching the video
  • labels: List[int] Multi-label actions found in the video. Indices from 0 to 156.
  • action_timings: List[Tuple[int, int]] Timing where each of the above actions happened.
  • length: float The length of the video in seconds
Click here to see the full list of Charades class labels mapping:
idClass
c000Holding some clothes
c001Putting clothes somewhere
c002Taking some clothes from somewhere
c003Throwing clothes somewhere
c004Tidying some clothes
c005Washing some clothes
c006Closing a door
c007Fixing a door
c008Opening a door
c009Putting something on a table
c010Sitting on a table
c011Sitting at a table
c012Tidying up a table
c013Washing a table
c014Working at a table
c015Holding a phone/camera
c016Playing with a phone/camera
c017Putting a phone/camera somewhere
c018Taking a phone/camera from somewhere
c019Talking on a phone/camera
c020Holding a bag
c021Opening a bag
c022Putting a bag somewhere
c023Taking a bag from somewhere
c024Throwing a bag somewhere
c025Closing a book
c026Holding a book
c027Opening a book
c028Putting a book somewhere
c029Smiling at a book
c030Taking a book from somewhere
c031Throwing a book somewhere
c032Watching/Reading/Looking at a book
c033Holding a towel/s
c034Putting a towel/s somewhere
c035Taking a towel/s from somewhere
c036Throwing a towel/s somewhere
c037Tidying up a towel/s
c038Washing something with a towel
c039Closing a box
c040Holding a box
c041Opening a box
c042Putting a box somewhere
c043Taking a box from somewhere
c044Taking something from a box
c045Throwing a box somewhere
c046Closing a laptop
c047Holding a laptop
c048Opening a laptop
c049Putting a laptop somewhere
c050Taking a laptop from somewhere
c051Watching a laptop or something on a laptop
c052Working/Playing on a laptop
c053Holding a shoe/shoes
c054Putting shoes somewhere
c055Putting on shoe/shoes
c056Taking shoes from somewhere
c057Taking off some shoes
c058Throwing shoes somewhere
c059Sitting in a chair
c060Standing on a chair
c061Holding some food
c062Putting some food somewhere
c063Taking food from somewhere
c064Throwing food somewhere
c065Eating a sandwich
c066Making a sandwich
c067Holding a sandwich
c068Putting a sandwich somewhere
c069Taking a sandwich from somewhere
c070Holding a blanket
c071Putting a blanket somewhere
c072Snuggling with a blanket
c073Taking a blanket from somewhere
c074Throwing a blanket somewhere
c075Tidying up a blanket/s
c076Holding a pillow
c077Putting a pillow somewhere
c078Snuggling with a pillow
c079Taking a pillow from somewhere
c080Throwing a pillow somewhere
c081Putting something on a shelf
c082Tidying a shelf or something on a shelf
c083Reaching for and grabbing a picture
c084Holding a picture
c085Laughing at a picture
c086Putting a picture somewhere
c087Taking a picture of something
c088Watching/looking at a picture
c089Closing a window
c090Opening a window
c091Washing a window
c092Watching/Looking outside of a window
c093Holding a mirror
c094Smiling in a mirror
c095Washing a mirror
c096Watching something/someone/themselves in a mirror
c097Walking through a doorway
c098Holding a broom
c099Putting a broom somewhere
c100Taking a broom from somewhere
c101Throwing a broom somewhere
c102Tidying up with a broom
c103Fixing a light
c104Turning on a light
c105Turning off a light
c106Drinking from a cup/glass/bottle
c107Holding a cup/glass/bottle of something
c108Pouring something into a cup/glass/bottle
c109Putting a cup/glass/bottle somewhere
c110Taking a cup/glass/bottle from somewhere
c111Washing a cup/glass/bottle
c112Closing a closet/cabinet
c113Opening a closet/cabinet
c114Tidying up a closet/cabinet
c115Someone is holding a paper/notebook
c116Putting their paper/notebook somewhere
c117Taking paper/notebook from somewhere
c118Holding a dish
c119Putting a dish/es somewhere
c120Taking a dish/es from somewhere
c121Wash a dish/dishes
c122Lying on a sofa/couch
c123Sitting on sofa/couch
c124Lying on the floor
c125Sitting on the floor
c126Throwing something on the floor
c127Tidying something on the floor
c128Holding some medicine
c129Taking/consuming some medicine
c130Putting groceries somewhere
c131Laughing at television
c132Watching television
c133Someone is awakening in bed
c134Lying on a bed
c135Sitting in a bed
c136Fixing a vacuum
c137Holding a vacuum
c138Taking a vacuum from somewhere
c139Washing their hands
c140Fixing a doorknob
c141Grasping onto a doorknob
c142Closing a refrigerator
c143Opening a refrigerator
c144Fixing their hair
c145Working on paper/notebook
c146Someone is awakening somewhere
c147Someone is cooking something
c148Someone is dressing
c149Someone is laughing
c150Someone is running somewhere
c151Someone is going from standing to sitting
c152Someone is smiling
c153Someone is sneezing
c154Someone is standing up from somewhere
c155Someone is undressing
c156Someone is eating something

Data Splits

trainvalidationtest
# of examples128116750000100000

Dataset Creation

Curation Rationale

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples of our daily dynamic scenes. While most of such scenes are not particularly exciting, they typically do not appear on YouTube, in movies or TV broadcasts. So how do we collect sufficiently many diverse but boring samples representing our lives? We propose a novel Hollywood in Homes approach to collect such data. Instead of shooting videos in the lab, we ensure diversity by distributing and crowdsourcing the whole process of video creation from script writing to video recording and annotation.

Source Data

Initial Data Collection and Normalization

Similar to filming, we have a three-step process for generating a video. The first step is generating the script of the indoor video. The key here is to allow workers to generate diverse scripts yet ensure that we have enough data for each category. The second step in the process is to use the script and ask workers to record a video of that sentence being acted out. In the final step, we ask the workers to verify if the recorded video corresponds to script, followed by an annotation procedure.

Who are the source language producers?

Amazon Mechnical Turk annotators

Annotations

Annotation process

Similar to filming, we have a three-step process for generating a video. The first step is generating the script of the indoor video. The key here is to allow workers to generate diverse scripts yet ensure that we have enough data for each category. The second step in the process is to use the script and ask workers to record a video of that sentence being acted out. In the final step, we ask the workers to verify if the recorded video corresponds to script, followed by an annotation procedure.

Who are the annotators?

Amazon Mechnical Turk annotators

Personal and Sensitive Information

Nothing specifically mentioned in the paper.

Considerations for Using the Data

Social Impact of Dataset

[More Information Needed]

Discussion of Biases

[More Information Needed]

Other Known Limitations

[More Information Needed]

Additional Information

Dataset Curators

AMT annotators

Licensing Information

License for Non-Commercial Use

If this software is redistributed, this license must be included. The term software includes any source files, documentation, executables, models, and data.

This software and data is available for general use by academic or non-profit, or government-sponsored researchers. It may also be used for evaluation purposes elsewhere. This license does not grant the right to use this software or any derivation of it in a for-profit enterprise. For commercial use, please contact The Allen Institute for Artificial Intelligence.

This license does not grant the right to modify and publicly release the data in any form.

This license does not grant the right to distribute the data to a third party in any form.

The subjects in this data should be treated with respect and dignity. This license only grants the right to publish short segments or still images in an academic publication where necessary to present examples, experimental results, or observations.

This software comes with no warranty or guarantee of any kind. By using this software, the user accepts full liability.

The Allen Institute for Artificial Intelligence (C) 2016.

Citation Information

@article{sigurdsson2016hollywood,
    author = {Gunnar A. Sigurdsson and G{\"u}l Varol and Xiaolong Wang and Ivan Laptev and Ali Farhadi and Abhinav Gupta},
    title = {Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding},
    journal = {ArXiv e-prints},
    eprint = {1604.01753}, 
    year = {2016},
    url = {http://arxiv.org/abs/1604.01753},
}

Contributions

Thanks to @apsdehal for adding this dataset.

HuggingFaceM4/charades

Dataset

Dataset Card for Charades

10

7 commits

2 linked in READMEs

updated Oct 20, 2022

See the code

README

Dataset Card for Charades

Table of Contents

Dataset Description

Dataset Summary

Charades is dataset composed of 9848 videos of daily indoors activities collected through Amazon Mechanical Turk. 267 different users were presented with a sentence, that includes objects and actions from a fixed vocabulary, and they recorded a video acting out the sentence (like in a game of Charades). The dataset contains 66,500 temporal annotations for 157 action classes, 41,104 labels for 46 object classes, and 27,847 textual descriptions of the videos

Supported Tasks and Leaderboards

  • multilabel-action-classification: The goal of this task is to classify actions happening in a video. This is a multilabel classification. The leaderboard is available here

Languages

The annotations in the dataset are in English.

Dataset Structure

Data Instances

{
  "video_id": "46GP8",
  "video": "/home/amanpreet_huggingface_co/.cache/huggingface/datasets/downloads/extracted/3f022da5305aaa189f09476dbf7d5e02f6fe12766b927c076707360d00deb44d/46GP8.mp4",
  "subject": "HR43",
  "scene": "Kitchen",
  "quality": 6,
  "relevance": 7,
  "verified": "Yes",
  "script": "A person cooking on a stove while watching something out a window.",
  "objects": ["food", "stove", "window"],
  "descriptions": [
    "A person cooks food on a stove before looking out of a window."
  ],
  "labels": [92, 147],
  "action_timings": [
    [11.899999618530273, 21.200000762939453],
    [0.0, 12.600000381469727]
  ],
  "length": 24.829999923706055
}

Data Fields

  • video_id: str Unique identifier for each video.
  • video: str Path to the video file
  • subject: str Unique identifier for each subject in the dataset
  • scene: str One of 15 indoor scenes in the dataset, such as Kitchen
  • quality: int The quality of the video judged by an annotator (7-point scale, 7=high quality), -100 if missing
  • relevance: int The relevance of the video to the script judged by an annotated (7-point scale, 7=very relevant), -100 if missing
  • verified: str 'Yes' if an annotator successfully verified that the video matches the script, else 'No'
  • script: str The human-generated script used to generate the video
  • descriptions: List[str] List of descriptions by annotators watching the video
  • labels: List[int] Multi-label actions found in the video. Indices from 0 to 156.
  • action_timings: List[Tuple[int, int]] Timing where each of the above actions happened.
  • length: float The length of the video in seconds
Click here to see the full list of Charades class labels mapping:
idClass
c000Holding some clothes
c001Putting clothes somewhere
c002Taking some clothes from somewhere
c003Throwing clothes somewhere
c004Tidying some clothes
c005Washing some clothes
c006Closing a door
c007Fixing a door
c008Opening a door
c009Putting something on a table
c010Sitting on a table
c011Sitting at a table
c012Tidying up a table
c013Washing a table
c014Working at a table
c015Holding a phone/camera
c016Playing with a phone/camera
c017Putting a phone/camera somewhere
c018Taking a phone/camera from somewhere
c019Talking on a phone/camera
c020Holding a bag
c021Opening a bag
c022Putting a bag somewhere
c023Taking a bag from somewhere
c024Throwing a bag somewhere
c025Closing a book
c026Holding a book
c027Opening a book
c028Putting a book somewhere
c029Smiling at a book
c030Taking a book from somewhere
c031Throwing a book somewhere
c032Watching/Reading/Looking at a book
c033Holding a towel/s
c034Putting a towel/s somewhere
c035Taking a towel/s from somewhere
c036Throwing a towel/s somewhere
c037Tidying up a towel/s
c038Washing something with a towel
c039Closing a box
c040Holding a box
c041Opening a box
c042Putting a box somewhere
c043Taking a box from somewhere
c044Taking something from a box
c045Throwing a box somewhere
c046Closing a laptop
c047Holding a laptop
c048Opening a laptop
c049Putting a laptop somewhere
c050Taking a laptop from somewhere
c051Watching a laptop or something on a laptop
c052Working/Playing on a laptop
c053Holding a shoe/shoes
c054Putting shoes somewhere
c055Putting on shoe/shoes
c056Taking shoes from somewhere
c057Taking off some shoes
c058Throwing shoes somewhere
c059Sitting in a chair
c060Standing on a chair
c061Holding some food
c062Putting some food somewhere
c063Taking food from somewhere
c064Throwing food somewhere
c065Eating a sandwich
c066Making a sandwich
c067Holding a sandwich
c068Putting a sandwich somewhere
c069Taking a sandwich from somewhere
c070Holding a blanket
c071Putting a blanket somewhere
c072Snuggling with a blanket
c073Taking a blanket from somewhere
c074Throwing a blanket somewhere
c075Tidying up a blanket/s
c076Holding a pillow
c077Putting a pillow somewhere
c078Snuggling with a pillow
c079Taking a pillow from somewhere
c080Throwing a pillow somewhere
c081Putting something on a shelf
c082Tidying a shelf or something on a shelf
c083Reaching for and grabbing a picture
c084Holding a picture
c085Laughing at a picture
c086Putting a picture somewhere
c087Taking a picture of something
c088Watching/looking at a picture
c089Closing a window
c090Opening a window
c091Washing a window
c092Watching/Looking outside of a window
c093Holding a mirror
c094Smiling in a mirror
c095Washing a mirror
c096Watching something/someone/themselves in a mirror
c097Walking through a doorway
c098Holding a broom
c099Putting a broom somewhere
c100Taking a broom from somewhere
c101Throwing a broom somewhere
c102Tidying up with a broom
c103Fixing a light
c104Turning on a light
c105Turning off a light
c106Drinking from a cup/glass/bottle
c107Holding a cup/glass/bottle of something
c108Pouring something into a cup/glass/bottle
c109Putting a cup/glass/bottle somewhere
c110Taking a cup/glass/bottle from somewhere
c111Washing a cup/glass/bottle
c112Closing a closet/cabinet
c113Opening a closet/cabinet
c114Tidying up a closet/cabinet
c115Someone is holding a paper/notebook
c116Putting their paper/notebook somewhere
c117Taking paper/notebook from somewhere
c118Holding a dish
c119Putting a dish/es somewhere
c120Taking a dish/es from somewhere
c121Wash a dish/dishes
c122Lying on a sofa/couch
c123Sitting on sofa/couch
c124Lying on the floor
c125Sitting on the floor
c126Throwing something on the floor
c127Tidying something on the floor
c128Holding some medicine
c129Taking/consuming some medicine
c130Putting groceries somewhere
c131Laughing at television
c132Watching television
c133Someone is awakening in bed
c134Lying on a bed
c135Sitting in a bed
c136Fixing a vacuum
c137Holding a vacuum
c138Taking a vacuum from somewhere
c139Washing their hands
c140Fixing a doorknob
c141Grasping onto a doorknob
c142Closing a refrigerator
c143Opening a refrigerator
c144Fixing their hair
c145Working on paper/notebook
c146Someone is awakening somewhere
c147Someone is cooking something
c148Someone is dressing
c149Someone is laughing
c150Someone is running somewhere
c151Someone is going from standing to sitting
c152Someone is smiling
c153Someone is sneezing
c154Someone is standing up from somewhere
c155Someone is undressing
c156Someone is eating something

Data Splits

trainvalidationtest
# of examples128116750000100000

Dataset Creation

Curation Rationale

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples of our daily dynamic scenes. While most of such scenes are not particularly exciting, they typically do not appear on YouTube, in movies or TV broadcasts. So how do we collect sufficiently many diverse but boring samples representing our lives? We propose a novel Hollywood in Homes approach to collect such data. Instead of shooting videos in the lab, we ensure diversity by distributing and crowdsourcing the whole process of video creation from script writing to video recording and annotation.

Source Data

Initial Data Collection and Normalization

Similar to filming, we have a three-step process for generating a video. The first step is generating the script of the indoor video. The key here is to allow workers to generate diverse scripts yet ensure that we have enough data for each category. The second step in the process is to use the script and ask workers to record a video of that sentence being acted out. In the final step, we ask the workers to verify if the recorded video corresponds to script, followed by an annotation procedure.

Who are the source language producers?

Amazon Mechnical Turk annotators

Annotations

Annotation process

Similar to filming, we have a three-step process for generating a video. The first step is generating the script of the indoor video. The key here is to allow workers to generate diverse scripts yet ensure that we have enough data for each category. The second step in the process is to use the script and ask workers to record a video of that sentence being acted out. In the final step, we ask the workers to verify if the recorded video corresponds to script, followed by an annotation procedure.

Who are the annotators?

Amazon Mechnical Turk annotators

Personal and Sensitive Information

Nothing specifically mentioned in the paper.

Considerations for Using the Data

Social Impact of Dataset

[More Information Needed]

Discussion of Biases

[More Information Needed]

Other Known Limitations

[More Information Needed]

Additional Information

Dataset Curators

AMT annotators

Licensing Information

License for Non-Commercial Use

If this software is redistributed, this license must be included. The term software includes any source files, documentation, executables, models, and data.

This software and data is available for general use by academic or non-profit, or government-sponsored researchers. It may also be used for evaluation purposes elsewhere. This license does not grant the right to use this software or any derivation of it in a for-profit enterprise. For commercial use, please contact The Allen Institute for Artificial Intelligence.

This license does not grant the right to modify and publicly release the data in any form.

This license does not grant the right to distribute the data to a third party in any form.

The subjects in this data should be treated with respect and dignity. This license only grants the right to publish short segments or still images in an academic publication where necessary to present examples, experimental results, or observations.

This software comes with no warranty or guarantee of any kind. By using this software, the user accepts full liability.

The Allen Institute for Artificial Intelligence (C) 2016.

Citation Information

@article{sigurdsson2016hollywood,
    author = {Gunnar A. Sigurdsson and G{\"u}l Varol and Xiaolong Wang and Ivan Laptev and Ali Farhadi and Abhinav Gupta},
    title = {Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding},
    journal = {ArXiv e-prints},
    eprint = {1604.01753}, 
    year = {2016},
    url = {http://arxiv.org/abs/1604.01753},
}

Contributions

Thanks to @apsdehal for adding this dataset.