TencentARC/ShortVid-Bench

Dataset

2

stars

3

commits

2

linked in READMEs

Aug 5, 2025

updated

short video
video-audio understanding
video comprehension benchmark
video qa benchmark

README

ShortVid-Bench

arXiv Demo Code Static Badge Blog

Introduction

Existing benchmarks often fall short in capturing the nuanced complexities of user-generated content. To rigorously evaluate model’s ability to understand real-world short videos, we construct a specialized benchmark named ShortVid-Bench. Specifically, we develop an automated pipeline to generate multi-dimensional questions for each video, targeting capabilities that signify a deep, holistic comprehension through integrating both visual and audio cues. These dimensions include:

  • Temporal Reasoning and Localization
  • Affective Intent Classification
  • Creator Intent Taxonomy
  • Narrative Comprehension
  • Humor & Meme Deconstruction
  • Creative Innovation Analysis

For objective assessment, we employ a multiple-choice question (MCQ) format following previous work. Each question is carefully curated by human annotators who provide the ground-truth answer and design challenging, plausible distractors. Collectively, these dimensions with a total of 1,000 multiple-choice questions push the evaluation beyond mere descriptive captioning, demanding a genuine comprehension of the video’s context, intent, and narrative.

Model Performance

Modelfps#framesthinkShortVid-Bench
Qwen2.5-VL-7B-Instruct1.0150×69.3
Qwen2.5-Omni-7B1.0150×69.7
Keye-VL-8B1.015056.3
ARC-Hunyuan-Video-7B1.015073.0
Please note that the results in the table above are different from those in ARC-Hunyuan-Video-7B. This is because, after releasing the technical report, we expanded the benchmark dataset to 1,000 samples, whereas the results in the paper were based on 400 samples.

License

  • ShortVid-Bench is released under the Apache-2.0 license for academic purpose only.
  • All videos of the ShortVid-Bench are obtained from the Internet which are not property of our institutions. Our institution are not responsible for the content nor the meaning of these videos. The copyright remains with the original owners of the video.
  • If any video in our dataset infringes upon your rights, please contact us for removal.

Citation

If you find the work helpful, please consider citing:

@article{ge2025arc,
  title={ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts},
  author={Ge, Yuying and Ge, Yixiao and Li, Chen and Wang, Teng and Pu, Junfu and Li, Yizhuo and Qiu, Lu and Ma, Jin and Duan, Lisheng and Zuo, Xinyu and others},
  journal={arXiv preprint arXiv:2507.20939},
  year={2025}
}

Contributors

tttoaster

2 commits

RO
root

1 commits

TencentARC/ShortVid-Bench

Dataset

2

stars

3

commits

2

linked in READMEs

Aug 5, 2025

updated

short video
video-audio understanding
video comprehension benchmark
video qa benchmark

README

ShortVid-Bench

arXiv Demo Code Static Badge Blog

Introduction

Existing benchmarks often fall short in capturing the nuanced complexities of user-generated content. To rigorously evaluate model’s ability to understand real-world short videos, we construct a specialized benchmark named ShortVid-Bench. Specifically, we develop an automated pipeline to generate multi-dimensional questions for each video, targeting capabilities that signify a deep, holistic comprehension through integrating both visual and audio cues. These dimensions include:

  • Temporal Reasoning and Localization
  • Affective Intent Classification
  • Creator Intent Taxonomy
  • Narrative Comprehension
  • Humor & Meme Deconstruction
  • Creative Innovation Analysis

For objective assessment, we employ a multiple-choice question (MCQ) format following previous work. Each question is carefully curated by human annotators who provide the ground-truth answer and design challenging, plausible distractors. Collectively, these dimensions with a total of 1,000 multiple-choice questions push the evaluation beyond mere descriptive captioning, demanding a genuine comprehension of the video’s context, intent, and narrative.

Model Performance

Modelfps#framesthinkShortVid-Bench
Qwen2.5-VL-7B-Instruct1.0150×69.3
Qwen2.5-Omni-7B1.0150×69.7
Keye-VL-8B1.015056.3
ARC-Hunyuan-Video-7B1.015073.0
Please note that the results in the table above are different from those in ARC-Hunyuan-Video-7B. This is because, after releasing the technical report, we expanded the benchmark dataset to 1,000 samples, whereas the results in the paper were based on 400 samples.

License

  • ShortVid-Bench is released under the Apache-2.0 license for academic purpose only.
  • All videos of the ShortVid-Bench are obtained from the Internet which are not property of our institutions. Our institution are not responsible for the content nor the meaning of these videos. The copyright remains with the original owners of the video.
  • If any video in our dataset infringes upon your rights, please contact us for removal.

Citation

If you find the work helpful, please consider citing:

@article{ge2025arc,
  title={ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts},
  author={Ge, Yuying and Ge, Yixiao and Li, Chen and Wang, Teng and Pu, Junfu and Li, Yizhuo and Qiu, Lu and Ma, Jin and Duan, Lisheng and Zuo, Xinyu and others},
  journal={arXiv preprint arXiv:2507.20939},
  year={2025}
}

Contributors

tttoaster

2 commits

RO
root

1 commits