cdx-cindy/VideoMind

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding

12

78 commits

updated Sep 17, 2025

See the code

README

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding [Paper]

:fire: News

  • We release the V1 version of the video annotations for VideoMind and a gold-standard benchmark which consists of 5000 meticulously manual-validated samples.(OpenDataLab | HuggingFace).

:book: Introduction

What is VideoMind?

VideoMind is a video-centric omni-modal dataset, which enables the deep cognition of video content and enhances feature representations of multi-modal data. The VideoMind dataset contains 105K video samples (5K for test only), each of which is accompanied by audio, as well as systematic and detailed textual descriptions. Specifically, every video sample, together with its audio data, is described across three hierarchical layers (factual, abstract, and intent), progressing from the superficial to the profound. In total, more than 22 million words are included, with an average of approximately 225 words per sample. Compared with existing video-centric datasets, the distinguishing feature of VideoMind lies in providing intent expressions that are intuitively unattainable and must be speculated through the integration of context across the entire video. The Chain-of-Thought (COT) text generation manner is introduced, wherein the mLLM is prompted to derive deep-cognitive expressions under step-by-step guidance. Upon the detailed descriptions, various annotations, including subject, place, time, event, action, and intent, are marked, serving a series of downstream recognition tasks. More crucially, we establish a gold-standard benchmark comprising 5,000 meticulously manual-validated samples for the evaluation of deep-cognitive video understanding.

examples for VideoMind Examples of video clips and the corresponding factual layer, abstract layer and intent layer in the proposed VideoMind dataset.

Data statistics

Sentence Length Video statistics in VideoMind.

Word Cloud The word cloud of intent, audio style, subject, and place in the VideoMind dataset.

Model

Based on the proposed VideoMind, we design a baseline model, Deep Multi-modal Embedder (DeME), which fully leverages hierarchically expressed texts.

multi-model

Framework of the DeME to extract general embeddings for omni-modal data.

Download

  1. You can download our video annotations from [OpenDataLab | HuggingFace ] or from Google drive
  2. You can download the videos of our benchmark from this link.

Citation

If you find this work useful for your research, please consider citing VideoMind. Your acknowledgement would greatly help us in continuing to contribute resources to the research community. 😊

@misc{yang2025videomindomnimodalvideodataset,
      title={VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding}, 
      author={Baoyao Yang and Wanyun Li and Dixin Chen and Junxiang Chen and Wenbin Yao and Haifeng Lin},
      year={2025},
      eprint={2507.18552},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2507.18552}, 
}

cdx-cindy/VideoMind

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding

12

78 commits

updated Sep 17, 2025

See the code

README

VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding [Paper]

:fire: News

  • We release the V1 version of the video annotations for VideoMind and a gold-standard benchmark which consists of 5000 meticulously manual-validated samples.(OpenDataLab | HuggingFace).

:book: Introduction

What is VideoMind?

VideoMind is a video-centric omni-modal dataset, which enables the deep cognition of video content and enhances feature representations of multi-modal data. The VideoMind dataset contains 105K video samples (5K for test only), each of which is accompanied by audio, as well as systematic and detailed textual descriptions. Specifically, every video sample, together with its audio data, is described across three hierarchical layers (factual, abstract, and intent), progressing from the superficial to the profound. In total, more than 22 million words are included, with an average of approximately 225 words per sample. Compared with existing video-centric datasets, the distinguishing feature of VideoMind lies in providing intent expressions that are intuitively unattainable and must be speculated through the integration of context across the entire video. The Chain-of-Thought (COT) text generation manner is introduced, wherein the mLLM is prompted to derive deep-cognitive expressions under step-by-step guidance. Upon the detailed descriptions, various annotations, including subject, place, time, event, action, and intent, are marked, serving a series of downstream recognition tasks. More crucially, we establish a gold-standard benchmark comprising 5,000 meticulously manual-validated samples for the evaluation of deep-cognitive video understanding.

examples for VideoMind Examples of video clips and the corresponding factual layer, abstract layer and intent layer in the proposed VideoMind dataset.

Data statistics

Sentence Length Video statistics in VideoMind.

Word Cloud The word cloud of intent, audio style, subject, and place in the VideoMind dataset.

Model

Based on the proposed VideoMind, we design a baseline model, Deep Multi-modal Embedder (DeME), which fully leverages hierarchically expressed texts.

multi-model

Framework of the DeME to extract general embeddings for omni-modal data.

Download

  1. You can download our video annotations from [OpenDataLab | HuggingFace ] or from Google drive
  2. You can download the videos of our benchmark from this link.

Citation

If you find this work useful for your research, please consider citing VideoMind. Your acknowledgement would greatly help us in continuing to contribute resources to the research community. 😊

@misc{yang2025videomindomnimodalvideodataset,
      title={VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding}, 
      author={Baoyao Yang and Wanyun Li and Dixin Chen and Junxiang Chen and Wenbin Yao and Haifeng Lin},
      year={2025},
      eprint={2507.18552},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2507.18552}, 
}