MBZUAI/VCGBench-Diverse

Dataset

๐Ÿ‘๏ธ VCGBench-Diverse Benchmarks

3

15 commits

1 linked in READMEs

updated Feb 9, 2026

See the code

README

๐Ÿ‘๏ธ VCGBench-Diverse Benchmarks


๐Ÿ“ Description

Recognizing the limited diversity in existing video conversation benchmarks, we introduce VCGBench-Diverse to comprehensively evaluate the generalization ability of video LMMs. While VCG-Bench provides an extensive evaluation protocol, it is limited to videos from the ActivityNet200 dataset. Our benchmark comprises a total of 877 videos, 18 broad video categories and 4,354 QA pairs, ensuring a robust evaluation framework.

Contributions

Dataset Contents

  1. vcgbench_diverse_qa.json - Contains VCGBench-Diverse question-answer pairs.
  2. videos.tar.gz - Contains the videos corresponding to vcgbench_diverse_qa.json.
  3. human_annotated_video_descriptions - Contains original human-annotated dense descriptions of the videos.
  4. gpt_evaluation_scripts - Contains the GPT-3.5-Turbo evaluation scripts to evaluate a model's predictions.
  5. sample_predictions - Contains the VideoGPT+ predictions on the VCGBench-Diverse. Compatible with gpt_evaluation_scripts.

In order to evaluate your model on VCGBench-Diverse, use question-answer pairs in vcgbench_diverse_qa.json to generate your model's predictions in format same as sample_predictions and then use gpt_evaluation_scripts for the evalution.

๐Ÿ’ป Download

To get started, follow these steps:

git lfs install
git clone https://huggingface.co/MBZUAI/VCGBench-Diverse

๐Ÿ“š Additional Resources

  • Paper: ArXiv.
  • GitHub Repository: For training and updates: GitHub.
  • HuggingFace Collection: For downloading the pretrained checkpoints, VCGBench-Diverse Benchmarks and Training data, visit HuggingFace Collection - VideoGPT+.

๐Ÿ“œ Citations and Acknowledgments

  @article{Maaz2024VideoGPT+,
      title={VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding},
      author={Maaz, Muhammad and Rasheed, Hanoona and Khan, Salman and Khan, Fahad Shahbaz},
      journal={arxiv},
      year={2024},
      url={https://arxiv.org/abs/2406.09418}
  }

MBZUAI/VCGBench-Diverse

Dataset

๐Ÿ‘๏ธ VCGBench-Diverse Benchmarks

3

15 commits

1 linked in READMEs

updated Feb 9, 2026

See the code

README

๐Ÿ‘๏ธ VCGBench-Diverse Benchmarks


๐Ÿ“ Description

Recognizing the limited diversity in existing video conversation benchmarks, we introduce VCGBench-Diverse to comprehensively evaluate the generalization ability of video LMMs. While VCG-Bench provides an extensive evaluation protocol, it is limited to videos from the ActivityNet200 dataset. Our benchmark comprises a total of 877 videos, 18 broad video categories and 4,354 QA pairs, ensuring a robust evaluation framework.

Contributions

Dataset Contents

  1. vcgbench_diverse_qa.json - Contains VCGBench-Diverse question-answer pairs.
  2. videos.tar.gz - Contains the videos corresponding to vcgbench_diverse_qa.json.
  3. human_annotated_video_descriptions - Contains original human-annotated dense descriptions of the videos.
  4. gpt_evaluation_scripts - Contains the GPT-3.5-Turbo evaluation scripts to evaluate a model's predictions.
  5. sample_predictions - Contains the VideoGPT+ predictions on the VCGBench-Diverse. Compatible with gpt_evaluation_scripts.

In order to evaluate your model on VCGBench-Diverse, use question-answer pairs in vcgbench_diverse_qa.json to generate your model's predictions in format same as sample_predictions and then use gpt_evaluation_scripts for the evalution.

๐Ÿ’ป Download

To get started, follow these steps:

git lfs install
git clone https://huggingface.co/MBZUAI/VCGBench-Diverse

๐Ÿ“š Additional Resources

  • Paper: ArXiv.
  • GitHub Repository: For training and updates: GitHub.
  • HuggingFace Collection: For downloading the pretrained checkpoints, VCGBench-Diverse Benchmarks and Training data, visit HuggingFace Collection - VideoGPT+.

๐Ÿ“œ Citations and Acknowledgments

  @article{Maaz2024VideoGPT+,
      title={VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding},
      author={Maaz, Muhammad and Rasheed, Hanoona and Khan, Salman and Khan, Fahad Shahbaz},
      journal={arxiv},
      year={2024},
      url={https://arxiv.org/abs/2406.09418}
  }