VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
12
78 commits
updated Sep 17, 2025
VideoMind is a video-centric omni-modal dataset, which enables the deep cognition of video content and enhances feature representations of multi-modal data. The VideoMind dataset contains 105K video samples (5K for test only), each of which is accompanied by audio, as well as systematic and detailed textual descriptions. Specifically, every video sample, together with its audio data, is described across three hierarchical layers (factual, abstract, and intent), progressing from the superficial to the profound. In total, more than 22 million words are included, with an average of approximately 225 words per sample. Compared with existing video-centric datasets, the distinguishing feature of VideoMind lies in providing intent expressions that are intuitively unattainable and must be speculated through the integration of context across the entire video. The Chain-of-Thought (COT) text generation manner is introduced, wherein the mLLM is prompted to derive deep-cognitive expressions under step-by-step guidance. Upon the detailed descriptions, various annotations, including subject, place, time, event, action, and intent, are marked, serving a series of downstream recognition tasks. More crucially, we establish a gold-standard benchmark comprising 5,000 meticulously manual-validated samples for the evaluation of deep-cognitive video understanding.
Examples of video clips and the corresponding factual layer, abstract layer and intent layer in the proposed VideoMind dataset.
Video statistics in VideoMind.
The word cloud of intent, audio style, subject, and place in the VideoMind dataset.
Based on the proposed VideoMind, we design a baseline model, Deep Multi-modal Embedder (DeME), which fully leverages hierarchically expressed texts.
Framework of the DeME to extract general embeddings for omni-modal data.
If you find this work useful for your research, please consider citing VideoMind. Your acknowledgement would greatly help us in continuing to contribute resources to the research community. 😊
@misc{yang2025videomindomnimodalvideodataset,
title={VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding},
author={Baoyao Yang and Wanyun Li and Dixin Chen and Junxiang Chen and Wenbin Yao and Haifeng Lin},
year={2025},
eprint={2507.18552},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.18552},
}
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
12
78 commits
updated Sep 17, 2025
VideoMind is a video-centric omni-modal dataset, which enables the deep cognition of video content and enhances feature representations of multi-modal data. The VideoMind dataset contains 105K video samples (5K for test only), each of which is accompanied by audio, as well as systematic and detailed textual descriptions. Specifically, every video sample, together with its audio data, is described across three hierarchical layers (factual, abstract, and intent), progressing from the superficial to the profound. In total, more than 22 million words are included, with an average of approximately 225 words per sample. Compared with existing video-centric datasets, the distinguishing feature of VideoMind lies in providing intent expressions that are intuitively unattainable and must be speculated through the integration of context across the entire video. The Chain-of-Thought (COT) text generation manner is introduced, wherein the mLLM is prompted to derive deep-cognitive expressions under step-by-step guidance. Upon the detailed descriptions, various annotations, including subject, place, time, event, action, and intent, are marked, serving a series of downstream recognition tasks. More crucially, we establish a gold-standard benchmark comprising 5,000 meticulously manual-validated samples for the evaluation of deep-cognitive video understanding.
Examples of video clips and the corresponding factual layer, abstract layer and intent layer in the proposed VideoMind dataset.
Video statistics in VideoMind.
The word cloud of intent, audio style, subject, and place in the VideoMind dataset.
Based on the proposed VideoMind, we design a baseline model, Deep Multi-modal Embedder (DeME), which fully leverages hierarchically expressed texts.
Framework of the DeME to extract general embeddings for omni-modal data.
If you find this work useful for your research, please consider citing VideoMind. Your acknowledgement would greatly help us in continuing to contribute resources to the research community. 😊
@misc{yang2025videomindomnimodalvideodataset,
title={VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding},
author={Baoyao Yang and Wanyun Li and Dixin Chen and Junxiang Chen and Wenbin Yao and Haifeng Lin},
year={2025},
eprint={2507.18552},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.18552},
}