Paper: https://huggingface.co/papers/2406.16338

This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary VideoQA method for comprehensive evaluation, where pairs of basic and hallucinated questions are crafted strategically.
| Object-Relation Hallucination | Temporal Hallucination | Semantic Detail Hallucination | External Factual Hallucination | External Nonfactual Hallucination | |
|---|---|---|---|---|---|
| Questions | 400 | 400 | 400 | 400 | 400 |
| Videos | 183 | 165 | 400 | 200 | 200 |
We provide VideoHallucerKit for evaluation
See our page
Paper: https://huggingface.co/papers/2406.16338

This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary VideoQA method for comprehensive evaluation, where pairs of basic and hallucinated questions are crafted strategically.
| Object-Relation Hallucination | Temporal Hallucination | Semantic Detail Hallucination | External Factual Hallucination | External Nonfactual Hallucination | |
|---|---|---|---|---|---|
| Questions | 400 | 400 | 400 | 400 | 400 |
| Videos | 183 | 165 | 400 | 200 | 200 |
We provide VideoHallucerKit for evaluation
See our page