VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
4
11 commits
1 linked in READMEs
updated Mar 28, 2025
Rohit Bharadwaj*, Hanan Gani*, Muzammal Naseer, Fahad Khan, Salman Khan
*denotes equal contribution
VANE-Bench is a meticulously curated benchmark dataset designed to evaluate the performance of large multimodal models (LMMs) on video anomaly detection and understanding tasks. The dataset includes a diverse set of video clips categorized into AI-Generated and Real-World anomalies, having per-frame information and associated question-answer pairs to facilitate robust evaluation of model capabilities.
You can load the dataset in HuggingFace using the following code snippet:
from datasets import load_dataset
dataset = load_dataset("rohit901/VANE-Bench")
The above HF dataset has the following fields:
You can directly download the zip file from this repository.
The zip file has the below file structure:
VQA_Data/
|ββ Real World/
| |ββ UCFCrime
| | |ββ Arrest002
| | |ββ Arrest002_qa.txt
| | |ββ ... # remaining video-qa pairs
| |ββ UCSD-Ped1
| | |ββ Test_004
| | |ββ Test_004_qa.txt
| | |ββ ... # remaining video-qa pairs
... # remaining real-world anomaly dataset folders
|ββ AI-Generated/
| |ββ SORA
| | |ββ video_1_subset_2
| | |ββ video_1_subset_2_qa.txt
| | |ββ ... # remaining video-qa pairs
| |ββ opensora
| | |ββ 1
| | |ββ 1_qa.txt
| | |ββ ... # remaining video-qa pairs
... # remaining AI-generated anomaly dataset folders
Overall performance of Video-LMMs averaged across all the benchmark datasets.
Human vs Video-LMMs' performance on only SORA data.
The dataset is licensed under the Creative Commons Attribution Non Commercial Share Alike 4.0 License.
For any questions or issues, please reach out to the dataset maintainers: rohit.bharadwaj@mbzuai.ac.ae or hanan.ghani@mbzuai.ac.ae
@misc{bharadwaj2024vanebench,
title={VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs},
author={Rohit Bharadwaj and Hanan Gani and Muzammal Naseer and Fahad Shahbaz Khan and Salman Khan},
year={2024},
eprint={2406.10326},
archivePrefix={arXiv},
primaryClass={id='cs.CV' full_name='Computer Vision and Pattern Recognition' is_active=True alt_name=None in_archive='cs' is_general=False description='Covers image processing, computer vision, pattern recognition, and scene understanding. Roughly includes material in ACM Subject Classes I.2.10, I.4, and I.5.'}
}
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
4
11 commits
1 linked in READMEs
updated Mar 28, 2025
Rohit Bharadwaj*, Hanan Gani*, Muzammal Naseer, Fahad Khan, Salman Khan
*denotes equal contribution
VANE-Bench is a meticulously curated benchmark dataset designed to evaluate the performance of large multimodal models (LMMs) on video anomaly detection and understanding tasks. The dataset includes a diverse set of video clips categorized into AI-Generated and Real-World anomalies, having per-frame information and associated question-answer pairs to facilitate robust evaluation of model capabilities.
You can load the dataset in HuggingFace using the following code snippet:
from datasets import load_dataset
dataset = load_dataset("rohit901/VANE-Bench")
The above HF dataset has the following fields:
You can directly download the zip file from this repository.
The zip file has the below file structure:
VQA_Data/
|ββ Real World/
| |ββ UCFCrime
| | |ββ Arrest002
| | |ββ Arrest002_qa.txt
| | |ββ ... # remaining video-qa pairs
| |ββ UCSD-Ped1
| | |ββ Test_004
| | |ββ Test_004_qa.txt
| | |ββ ... # remaining video-qa pairs
... # remaining real-world anomaly dataset folders
|ββ AI-Generated/
| |ββ SORA
| | |ββ video_1_subset_2
| | |ββ video_1_subset_2_qa.txt
| | |ββ ... # remaining video-qa pairs
| |ββ opensora
| | |ββ 1
| | |ββ 1_qa.txt
| | |ββ ... # remaining video-qa pairs
... # remaining AI-generated anomaly dataset folders
Overall performance of Video-LMMs averaged across all the benchmark datasets.
Human vs Video-LMMs' performance on only SORA data.
The dataset is licensed under the Creative Commons Attribution Non Commercial Share Alike 4.0 License.
For any questions or issues, please reach out to the dataset maintainers: rohit.bharadwaj@mbzuai.ac.ae or hanan.ghani@mbzuai.ac.ae
@misc{bharadwaj2024vanebench,
title={VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs},
author={Rohit Bharadwaj and Hanan Gani and Muzammal Naseer and Fahad Shahbaz Khan and Salman Khan},
year={2024},
eprint={2406.10326},
archivePrefix={arXiv},
primaryClass={id='cs.CV' full_name='Computer Vision and Pattern Recognition' is_active=True alt_name=None in_archive='cs' is_general=False description='Covers image processing, computer vision, pattern recognition, and scene understanding. Roughly includes material in ACM Subject Classes I.2.10, I.4, and I.5.'}
}