Recognizing the limited diversity in existing video conversation benchmarks, we introduce VCGBench-Diverse to comprehensively evaluate the generalization ability of video LMMs. While VCG-Bench provides an extensive evaluation protocol, it is limited to videos from the ActivityNet200 dataset. Our benchmark comprises a total of 877 videos, 18 broad video categories and 4,354 QA pairs, ensuring a robust evaluation framework.
vcgbench_diverse_qa.json - Contains VCGBench-Diverse question-answer pairs.videos.tar.gz - Contains the videos corresponding to vcgbench_diverse_qa.json.human_annotated_video_descriptions - Contains original human-annotated dense descriptions of the videos.gpt_evaluation_scripts - Contains the GPT-3.5-Turbo evaluation scripts to evaluate a model's predictions.sample_predictions - Contains the VideoGPT+ predictions on the VCGBench-Diverse. Compatible with gpt_evaluation_scripts.In order to evaluate your model on VCGBench-Diverse, use question-answer pairs in vcgbench_diverse_qa.json to generate your model's predictions in format same as
sample_predictions and then use gpt_evaluation_scripts for the evalution.
To get started, follow these steps:
git lfs install
git clone https://huggingface.co/MBZUAI/VCGBench-Diverse
@article{Maaz2024VideoGPT+,
title={VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding},
author={Maaz, Muhammad and Rasheed, Hanoona and Khan, Salman and Khan, Fahad Shahbaz},
journal={arxiv},
year={2024},
url={https://arxiv.org/abs/2406.09418}
}
Recognizing the limited diversity in existing video conversation benchmarks, we introduce VCGBench-Diverse to comprehensively evaluate the generalization ability of video LMMs. While VCG-Bench provides an extensive evaluation protocol, it is limited to videos from the ActivityNet200 dataset. Our benchmark comprises a total of 877 videos, 18 broad video categories and 4,354 QA pairs, ensuring a robust evaluation framework.
vcgbench_diverse_qa.json - Contains VCGBench-Diverse question-answer pairs.videos.tar.gz - Contains the videos corresponding to vcgbench_diverse_qa.json.human_annotated_video_descriptions - Contains original human-annotated dense descriptions of the videos.gpt_evaluation_scripts - Contains the GPT-3.5-Turbo evaluation scripts to evaluate a model's predictions.sample_predictions - Contains the VideoGPT+ predictions on the VCGBench-Diverse. Compatible with gpt_evaluation_scripts.In order to evaluate your model on VCGBench-Diverse, use question-answer pairs in vcgbench_diverse_qa.json to generate your model's predictions in format same as
sample_predictions and then use gpt_evaluation_scripts for the evalution.
To get started, follow these steps:
git lfs install
git clone https://huggingface.co/MBZUAI/VCGBench-Diverse
@article{Maaz2024VideoGPT+,
title={VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding},
author={Maaz, Muhammad and Rasheed, Hanoona and Khan, Salman and Khan, Fahad Shahbaz},
journal={arxiv},
year={2024},
url={https://arxiv.org/abs/2406.09418}
}