Existing benchmarks often fall short in capturing the nuanced complexities of user-generated content. To rigorously evaluate model’s ability to understand real-world short videos, we construct a specialized benchmark named ShortVid-Bench. Specifically, we develop an automated pipeline to generate multi-dimensional questions for each video, targeting capabilities that signify a deep, holistic comprehension through integrating both visual and audio cues. These dimensions include:
For objective assessment, we employ a multiple-choice question (MCQ) format following previous work. Each question is carefully curated by human annotators who provide the ground-truth answer and design challenging, plausible distractors. Collectively, these dimensions with a total of 1,000 multiple-choice questions push the evaluation beyond mere descriptive captioning, demanding a genuine comprehension of the video’s context, intent, and narrative.
| Model | fps | #frames | think | ShortVid-Bench |
|---|---|---|---|---|
| Qwen2.5-VL-7B-Instruct | 1.0 | 150 | × | 69.3 |
| Qwen2.5-Omni-7B | 1.0 | 150 | × | 69.7 |
| Keye-VL-8B | 1.0 | 150 | ✓ | 56.3 |
| ARC-Hunyuan-Video-7B | 1.0 | 150 | ✓ | 73.0 |
If you find the work helpful, please consider citing:
@article{ge2025arc,
title={ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts},
author={Ge, Yuying and Ge, Yixiao and Li, Chen and Wang, Teng and Pu, Junfu and Li, Yizhuo and Qiu, Lu and Ma, Jin and Duan, Lisheng and Zuo, Xinyu and others},
journal={arXiv preprint arXiv:2507.20939},
year={2025}
}
Existing benchmarks often fall short in capturing the nuanced complexities of user-generated content. To rigorously evaluate model’s ability to understand real-world short videos, we construct a specialized benchmark named ShortVid-Bench. Specifically, we develop an automated pipeline to generate multi-dimensional questions for each video, targeting capabilities that signify a deep, holistic comprehension through integrating both visual and audio cues. These dimensions include:
For objective assessment, we employ a multiple-choice question (MCQ) format following previous work. Each question is carefully curated by human annotators who provide the ground-truth answer and design challenging, plausible distractors. Collectively, these dimensions with a total of 1,000 multiple-choice questions push the evaluation beyond mere descriptive captioning, demanding a genuine comprehension of the video’s context, intent, and narrative.
| Model | fps | #frames | think | ShortVid-Bench |
|---|---|---|---|---|
| Qwen2.5-VL-7B-Instruct | 1.0 | 150 | × | 69.3 |
| Qwen2.5-Omni-7B | 1.0 | 150 | × | 69.7 |
| Keye-VL-8B | 1.0 | 150 | ✓ | 56.3 |
| ARC-Hunyuan-Video-7B | 1.0 | 150 | ✓ | 73.0 |
If you find the work helpful, please consider citing:
@article{ge2025arc,
title={ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts},
author={Ge, Yuying and Ge, Yixiao and Li, Chen and Wang, Teng and Pu, Junfu and Li, Yizhuo and Qiu, Lu and Ma, Jin and Duan, Lisheng and Zuo, Xinyu and others},
journal={arXiv preprint arXiv:2507.20939},
year={2025}
}