Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
12
28 commits
3 linked in READMEs
updated Aug 25, 2024

Please download the multi-shot videos from OneDrive or HuggingFace.
We are excited to release a new video-text benchmark for multi-shot video understanding. This release contains a 134k version of our dataset. It includes detailed long summaries (human annotated + GPTV generated) for 134k videos and shot captions (human annotated) for 188k video shots.
Our 134k multi-shot videos come with detailed textual descriptions, consisting of 43k human annotation and 90k GPTV generation and covering over 548k video shots. The different files under data/annotations/:
Annotations are in JSON format, with each video as a JSON object:
Example:
[
{
"video": "video_name.mp4",
"image_id": "video_name.mp4",
"id": 0,
"whole_caption": "summary",
"whole_ASR": "ASR output",
"nvid": "video_name.mp4",
"video_names": ["shot_name1.mp4", "shot_name2.mp4"],
"audio_captions": ["narration1", "narration2"],
"captions": ["caption1", "caption2"],
"ASR": ["ASR shot1", "ASR shot2"]
},
...
]
We provide cached multi-shot videos at OneDrive and HuggingFace. It takes around 160GB of disk space and needs to extract video shots on your own.
Or, you can download on your own:
/134k_meta.csv, or you can download the update videos (in addition to 20k version) in ./data/annotations/114k_meta.csv../data/scripts/download_videos.py to download videos. Ensure you have necessary permissions../data/scripts/process_videos.py to prepare video clips and single-shot videos. As a prerequisite, please run data/scripts/get_existing_data.py to have all the downloaded raw videos for processing.We uphold the rights of individuals and copyright holders. If you are featured in any of our video annotations or hold copyright to a video and wish to have its annotation removed from our dataset, please reach out to us. Send an email to hanmingfei@bytedance.com with the subject line beginning with Shot2Story-optout, or raise an issue with the same title format. We commit to reviewing your request promptly and taking suitable action.
Our text annotations are licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License. They are available strictly for non-commercial research.
Users must refer to HD-VILA-100M for video access. By downloading our annotations, you agree to these terms. Respect for video copyright holders is paramount. Ensure your use of the videos aligns with the original source's terms.
If you find our work useful for your research, please consider citing the paper
@misc{han2023shot2story20k,
title={Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos},
author={Mingfei Han and Linjie Yang and Xiaojun Chang and Heng Wang},
year={2023},
eprint={2312.10300},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
We extend our thanks to the teams behind HD-VILA-100M and Whisper. Our work builds upon their valuable contributions. Please acknowledge these resources in your work.
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
12
28 commits
3 linked in READMEs
updated Aug 25, 2024

Please download the multi-shot videos from OneDrive or HuggingFace.
We are excited to release a new video-text benchmark for multi-shot video understanding. This release contains a 134k version of our dataset. It includes detailed long summaries (human annotated + GPTV generated) for 134k videos and shot captions (human annotated) for 188k video shots.
Our 134k multi-shot videos come with detailed textual descriptions, consisting of 43k human annotation and 90k GPTV generation and covering over 548k video shots. The different files under data/annotations/:
Annotations are in JSON format, with each video as a JSON object:
Example:
[
{
"video": "video_name.mp4",
"image_id": "video_name.mp4",
"id": 0,
"whole_caption": "summary",
"whole_ASR": "ASR output",
"nvid": "video_name.mp4",
"video_names": ["shot_name1.mp4", "shot_name2.mp4"],
"audio_captions": ["narration1", "narration2"],
"captions": ["caption1", "caption2"],
"ASR": ["ASR shot1", "ASR shot2"]
},
...
]
We provide cached multi-shot videos at OneDrive and HuggingFace. It takes around 160GB of disk space and needs to extract video shots on your own.
Or, you can download on your own:
/134k_meta.csv, or you can download the update videos (in addition to 20k version) in ./data/annotations/114k_meta.csv../data/scripts/download_videos.py to download videos. Ensure you have necessary permissions../data/scripts/process_videos.py to prepare video clips and single-shot videos. As a prerequisite, please run data/scripts/get_existing_data.py to have all the downloaded raw videos for processing.We uphold the rights of individuals and copyright holders. If you are featured in any of our video annotations or hold copyright to a video and wish to have its annotation removed from our dataset, please reach out to us. Send an email to hanmingfei@bytedance.com with the subject line beginning with Shot2Story-optout, or raise an issue with the same title format. We commit to reviewing your request promptly and taking suitable action.
Our text annotations are licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License. They are available strictly for non-commercial research.
Users must refer to HD-VILA-100M for video access. By downloading our annotations, you agree to these terms. Respect for video copyright holders is paramount. Ensure your use of the videos aligns with the original source's terms.
If you find our work useful for your research, please consider citing the paper
@misc{han2023shot2story20k,
title={Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos},
author={Mingfei Han and Linjie Yang and Xiaojun Chang and Heng Wang},
year={2023},
eprint={2312.10300},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
We extend our thanks to the teams behind HD-VILA-100M and Whisper. Our work builds upon their valuable contributions. Please acknowledge these resources in your work.