wangyueqian/MMDuetIT

Dataset

MMDuetIT

1

5 commits

5 linked in READMEs

updated Dec 19, 2024

See the code

README

MMDuetIT

Dataset Description

This repo contains the dataset MMDuetIT, which is used for training MMDuet, and benchmarks for evaluating MMDuet. The data distribution of MMDuetIT is as follows:

  • Dense Captioning
    • Shot2Story: 36949 examples from human_anno subset
    • COIN: 4574 examples from the train set with 2-4 minutes videos
  • Temporal Video Grounding
  • Multi-Answer Grounded Video Question Answering (MAGQA)
    • The proposed dataset for Multi-Answer Grounded Video Question Answering (MAGQA), Shot2Story-MAGQA-39k, is also included in this repository. Its training set is shot2story/annotations/magqa_train-0.25_0.5-earlier.json, and its test set is shot2story/annotations/magqa_test.json. The questions and answers are converted from Shot2Story human-annotated captions using GPT-4o.

Please refer to our paper for more details, and our github for the usage.

Citation

If you find this work useful in your research, please consider citing:

@misc{wang2024mmduet,
      title={VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format}, 
      author={Yueqian Wang and Xiaojun Meng and Yuxuan Wang and Jianxin Liang and Jiansheng Wei and Huishuai Zhang and Dongyan Zhao},
      year={2024},
      eprint={2411.17991},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2411.17991}, 
}

wangyueqian/MMDuetIT

Dataset

MMDuetIT

1

5 commits

5 linked in READMEs

updated Dec 19, 2024

See the code

README

MMDuetIT

Dataset Description

This repo contains the dataset MMDuetIT, which is used for training MMDuet, and benchmarks for evaluating MMDuet. The data distribution of MMDuetIT is as follows:

  • Dense Captioning
    • Shot2Story: 36949 examples from human_anno subset
    • COIN: 4574 examples from the train set with 2-4 minutes videos
  • Temporal Video Grounding
  • Multi-Answer Grounded Video Question Answering (MAGQA)
    • The proposed dataset for Multi-Answer Grounded Video Question Answering (MAGQA), Shot2Story-MAGQA-39k, is also included in this repository. Its training set is shot2story/annotations/magqa_train-0.25_0.5-earlier.json, and its test set is shot2story/annotations/magqa_test.json. The questions and answers are converted from Shot2Story human-annotated captions using GPT-4o.

Please refer to our paper for more details, and our github for the usage.

Citation

If you find this work useful in your research, please consider citing:

@misc{wang2024mmduet,
      title={VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format}, 
      author={Yueqian Wang and Xiaojun Meng and Yuxuan Wang and Jianxin Liang and Jiansheng Wei and Huishuai Zhang and Dongyan Zhao},
      year={2024},
      eprint={2411.17991},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2411.17991}, 
}