hukcc/Awesome-Video-Hallucination

[ACL 2026] Curated papers on video LLM hallucination, with benchmarks, mitigation methods, and an interactive browser. Updated monthly.

Python

44

23 commits

updated Oct 1, 2026

See the code

README

Awesome-Video-Hallucination Awesome

arXiv ACL 2026 Findings Entries Auto arXiv Update License: MIT Last Commit

A curated paper list on hallucination in Video Large Language Models (Vid-LLMs), covering 42 benchmarks, 52 mitigation methods, and 1 evaluation analysis. The 95 entries represent 78 distinct papers; a paper may contribute both a benchmark and a method. Updated monthly via arXiv search and manual review.

📄 Survey Paper: Distorted or Fabricated? A Survey on Hallucination in Video LLMs

🔎 Interactive Browser: Browse 78 papers with source-backed summaries, task tags, combined filters, and table or card views.

Framework overview

Table of Contents


Latest Updates

  • [2026/10] Added 7 papers, introduced evaluation analyses, and updated the taxonomy. Review log.
  • [2026/09] Added 2 hallucination mitigation papers.
  • [2026/08] Added 6 papers and updated the taxonomy.
  • [2026/04] Our survey was accepted to ACL 2026 Findings.
Recently added · 2026-10-01 · 7 papers

FlexBench / GM-DPO · Beneath the Scores · VidOmni-Bench · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning)

Repository additions, not publication dates. All recent additions


Find Your Papers

Find Benchmarks · Training-Free Methods · With Code · Latest Papers · Recently Added

Start here: Reading guide, from the survey to benchmarks, mitigation, and evaluation limits.

Scope: Relevance review distinguishes direct hallucination research, broader evaluations, and related-work candidates. All entries remain available pending manual decisions.

Browse by Task

Task index · 9 topics · 78 papers

Paper-level task tags. Newest first; links lead to individual contributions below.

Audio-Visual Understanding (10 papers)

Video-HolmesV2 · TraceAV-Bench · Audio Hallucination QA · EGOILLUSION · AVCD · EmotionHallucer / PEP-MEK · AVHBench / AVHModel-Align-FT · CMM · mrDPO · AVHalluBench

Hallucination Detection (24 papers)

Beneath the Scores · VidOmni-Bench · VHD / TRACE-RC · VidHalLoc · GroundedVQA · MoHallBench · DualFact · Audio Hallucination QA · INFACT · VideoHEDGE · SmartSight · EGOILLUSION · MESH · ELV-Halluc / ELV-Halluc-DPO · ARGUS · EmotionHallucer / PEP-MEK · HAVEN / Video-thinking (TDPO) · VidHal · AVHBench / AVHModel-Align-FT · CMM · VideoHallucer · Vript / Vriptor · AVHalluBench · FactVC

Long-Video Understanding (10 papers)

VidOmni-Bench · Video-HolmesV2 · Reflect-R1 · DistractionBench · TraceAV-Bench · VideoTIR · Video-TwG · VideoTemp-o3 · ELV-Halluc / ELV-Halluc-DPO · Vript / Vriptor

Motion Understanding (9 papers)

FlexBench / GM-DPO · MoHallBench · MotionHalluc / PPV · KPM-Bench / MoPE · MixDPO · SANTA · MHBench · MASH-VLM · VidHalluc / DINO-HEAL

Temporal Grounding (7 papers)

Video-TwG · GraphThinker · STVG-R1 · VideoTemp-o3 · CoE · VTG-LLM · Temporal Insight

Video Captioning (15 papers)

FlexBench / GM-DPO · VidOmni-Bench · VidHalLoc · ProCap · DualFact · Structured Rewards · KPM-Bench / MoPE · SANTA · NOAH · ARGUS · VistaDPO · VidHal · mrDPO · Vript / Vriptor · FactVC

Video QA (41 papers)

Beneath the Scores · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning) · GroundedVQA · MoHallBench · VidPair-Halluc · Reflect-R1 · MotionHalluc / PPV · DistractionBench · TOC-Bench · TraceAV-Bench · Audio Hallucination QA · CCTVBench / C-TCD · VisualTextTrap / VTHM-MoE · Structured Rewards · VideoTIR · GameplayQA · FrameRepeat · ClueNet · INFACT · Video-TwG · KPM-Bench / MoPE · VideoTemp-o3 · ViSSRes · VideoHEDGE · Video-DPL · NOAH · EGOILLUSION · MESH · VideoHallu / VideoHallu-GRPO · VistaDPO · MHBench · RoadSocial · HAVEN / Video-thinking (TDPO) · MASH-VLM · OVBench / VideoChat-Online · VidHalluc / DINO-HEAL · VideoHallucer · Vista-LLaMA

Video Reasoning (26 papers)

Beneath the Scores · STRAND / STRAND (Trajectory Reasoning) · VADER · VidPair-Halluc · TOC-Bench · CCTVBench / C-TCD · Video-ToC · GasVideo-1000 · STEAR · Structured Rewards · GameplayQA · FrameRepeat · ClueNet · INFACT · GraphThinker · OmniVCHall / TriCD · CoE · MixDPO · SEASON · Video-DPL · NOAH · VideoHallu / VideoHallu-GRPO · HAVEN / Video-thinking (TDPO) · VidHalluc / DINO-HEAL · VidHal · EventHallusion / TCD

Video Understanding (21 papers)

VADER · MultiToP · SToP · Video-ToC · GasVideo-1000 · VisualTextTrap / VTHM-MoE · DTR · STEAR · STVG-R1 · MACD · OmniVCHall / TriCD · ViSSRes · SmartSight · SEASON · MMA · TAAE · PaMi-VDPO · RoadSocial · OVBench / VideoChat-Online · EventHallusion / TCD · Vista-LLaMA


Taxonomy of Video Hallucinations

View the full taxonomy tree (95 contributions)

Mechanism-driven taxonomy of Vid-LLM hallucinations
Mechanism-driven taxonomy of Vid-LLM hallucinations. Solid fill = benchmarks; striped fill = mitigation methods; dashed outline = evaluation analyses. Placement indicates a primary indexing category, not exclusive coverage.
Generated from paper data using the LaTeX tree source.


Evaluation Benchmarks

[!NOTE] Newest first within each subtype. Date = first arXiv submission, or publisher issue date when no arXiv record is used; venue years may differ. Sources and review notes.

Resource badge legend

page = Project Page
code = GitHub Repository
dataset = Hugging Face Dataset
dataset = Kaggle Dataset
leaderboard = Leaderboard
- = No verified resource link

🔵 Spatiotemporal Dynamics Benchmarks (Dynamic Distortion)

Event Misordering (7 entries)
PaperBenchmarkVenue / DateResources
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
Probes motion hallucinations caused by prior knowledge, sequential inference, and visual similarity through several question formats.
MoHallBencharXiv 2026
07/2026
-
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
Diagnoses fine-grained motion hallucinations involving direction, attribution, and temporal kinematics, separating distinct failures in interpreting physical movement.
Related method: PPV
MotionHallucarXiv 2026
06/2026
dataset
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
Evaluates fine-grained limb motion through video captioning and question answering, using kinematic parsing to assess motion descriptions.
Related method: MoPE
KPM-BencharXiv 2026
02/2026
-
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
Measures both fabricated content and missing information in free-form video captions against human-written reference descriptions.
ARGUSICCV 2025
06/2025
page code
MHBench: Demystifying Motion Hallucination in VideoLLMs
Tests motion hallucinations using original actions, actions with reversed meanings, and incomplete actions that challenge static visual cues.
MHBenchAAAI 2025
04/2025
code
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
Diagnoses video hallucinations by crossing hallucination causes, object-scene-event aspects, and question formats to expose distinct failures in video understanding.
Related method: Video-thinking (TDPO)
HAVENarXiv 2025
03/2025
code
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
Evaluates hallucinations about actions, temporal order, and scene transitions through video questions, caption generation, and event sorting.
Related method: DINO-HEAL
VidHallucCVPR 2025
12/2024
page code
Duration Distortion (2 entries)
PaperBenchmarkVenue / DateResources
Online Video Understanding: OVBench and VideoChat-Online
Evaluates streaming-video question answering across online perception, memory, and reasoning tasks; its scope extends beyond hallucination-specific evaluation.
Related method: VideoChat-Online
OVBenchCVPR 2025
12/2024
page code
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Uses adversarial question pairs to distinguish hallucinations about visible objects and temporal relations from unsupported external information.
VideoHallucerarXiv 2024
06/2024
code
Frequency Confusion (2 entries)
PaperBenchmarkVenue / DateResources
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
Evaluates temporal hallucinations by asking models to distinguish and rank video captions with different degrees of factual distortion.
VidHalTMLR 2026
11/2024
code
Vript: A Video Is Worth Thousands of Words
Provides dense video descriptions and Vript-Hard evaluations targeting hallucinated captions, information retrieval, and temporal ordering of video events.
Related method: Vriptor
VriptNeurIPS 2024
06/2024
code

🟢 Referential Inconsistency Benchmarks (Dynamic Distortion)

Character Conflation (4 entries)
PaperBenchmarkVenue / DateResources
STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models
Evaluates object-state changes, identity persistence, and relational reasoning, requiring answers to satisfy jointly grounded spatiotemporal prerequisites.
Related method: STRAND (Trajectory Reasoning)
STRANDarXiv 2026
08/2026
page code
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
Tests object identity, state, and continuity with trajectory-grounded questions designed to require temporally ordered visual evidence.
TOC-BencharXiv 2026
05/2026
code
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
Tests hallucinations in first-person videos through human-annotated open and closed questions about visual and auditory evidence.
EGOILLUSIONEMNLP 2025
11/2025
page
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
Uses hierarchical questions and plausible distractors to expose hallucinations about objects, attributes, and subject-action relations across video segments.
MESHACM MM 2025
09/2025
code
Scene Conflation (2 entries)
PaperBenchmarkVenue / DateResources
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
Inserts advertising distractors into long videos to test whether models conflate subjects and events from unrelated segments.
DistractionBencharXiv 2026
05/2026
code dataset
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
Tests hallucinations from semantic aggregation in long videos, where models can incorrectly combine information across separate video segments.
Related method: ELV-Halluc-DPO
ELV-HallucarXiv 2025
08/2025
code

🟠 Context-Driven Fabrication Benchmarks (Content Fabrication)

Object-Action Hallucination (4 entries)
PaperBenchmarkVenue / DateResources
Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
Evaluates person-specific limb-motion captions across shots, distinguishing omissions from fabricated or incorrect actions through graded physical alignment.
Related method: GM-DPO
FlexBencharXiv 2026
09/2026
-
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
Uses synchronized multiplayer gameplay views and structured distractors to test agent identity, role attribution, and grounded reasoning.
GameplayQAACL 2026
03/2026
page code dataset
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Tests prior-driven hallucinations using synthetic videos that violate physical or logical expectations, challenging models to follow observed evidence.
Related method: VideoHallu-GRPO
VideoHalluNeurIPS 2025
05/2025
code
Models See Hallucinations: Evaluating the Factuality in Video Captioning
Studies factual errors in video captions and introduces a weakly supervised factuality metric with human-annotated evaluation data.
FactVCEMNLP 2023
03/2023
code
Scene-Event Hallucination (5 entries)
PaperBenchmarkVenue / DateResources
VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
Tests whether models can verify individual events in dense video captions, using human-checked incorrect descriptions as hard negatives.
VidOmni-BencharXiv 2026
09/2026
-
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Tests consistent traffic-hazard judgments using real accident videos paired with counterfactual counterparts that alter the evidence for an accident.
Related method: C-TCD
CCTVBencharXiv 2026
04/2026
-
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
Inserts unrelated clips into videos to measure narrative-prior hallucinations and omissions through captioning and question-answering tasks.
NOAHarXiv 2025
11/2025
page code
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
Provides diverse social-media road videos and question-answer pairs for evaluating road-event understanding across viewpoints and geographic settings.
RoadSocialCVPR 2025
03/2025
page code
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Diagnoses event hallucinations driven by language priors and visual biases, testing whether answers reflect events actually present in videos.
Related method: TCD
EventHallusionarXiv 2024
09/2024
code
Compositional and Factuality Hallucination (9 entries)
PaperBenchmarkVenue / DateResources
Target-Checked Reliability Score Refinement for Video Question Answering
Provides controlled video-question examples of confident but incorrect answers for diagnosing hallucinations and evaluating the reliability of confidence estimates.
Related method: TRACE-RC
VHDarXiv 2026
09/2026
code dataset
Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators
Compares hallucination detectors on adversarial video questions and captions, with a shared protocol spanning ontology and dynamic errors.
VidHalLocarXiv 2026
09/2026
-
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
Uses adversarial video pairs with similar backgrounds but different foreground events to isolate spatial and temporal hallucinations.
VidPair-HallucECCV 2026
06/2026
page
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding
Evaluates procedural-caption factuality at conceptual and grounded argument levels, using either textual references or direct video evidence.
DualFactACL 2026 Findings
04/2026
-
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
Tests whether misleading conversational feedback causes video models to abandon correct judgments and invent unsupported spatiotemporal explanations.
GasVideo-1000arXiv 2026
04/2026
page
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
Tests whether misleading text overlays cause video models to contradict visual evidence when answering questions about the depicted content.
Related method: VTHM-MoE
VisualTextTraparXiv 2026
04/2026
-
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Separates faithfulness from factuality errors in video answers and probes robustness to degraded visuals, corrupted evidence, and temporal interventions.
INFACTarXiv 2026
03/2026
-
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
Benchmarks isolated and compositional hallucinations across spatial and temporal dimensions, probing failures involving multiple interacting types of video evidence.
Related method: TriCD
OmniVCHallarXiv 2026
01/2026
code
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
Estimates answer reliability by clustering responses to clean and perturbed videos and measuring semantic uncertainty across those responses.
VideoHEDGEarXiv 2026
01/2026
code

🟣 Audio-Visual Conflict Benchmarks (Content Fabrication)

Action Attribution (6 entries)
PaperBenchmarkVenue / DateResources
Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?
Requires long-video answers to cite precise audio-visual evidence, using evidence-aware scoring to penalize guessing and fabricated support.
Video-HolmesV2ECCV 2026
09/2026
-
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
Tests multi-hop reasoning and hallucination robustness using explicit evidence trajectories distributed across long audio-visual recordings.
TraceAV-BencharXiv 2026
05/2026
-
Exploring Audio Hallucination in Egocentric Video Understanding
Probes imagined foreground and background sounds in egocentric videos through questions targeting visible but inaudible events.
Audio Hallucination QAICASSP 2026
04/2026
-
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Evaluates audio-visual hallucinations, cross-modal matching, and reasoning through tasks that distinguish auditory evidence from visible video content.
Related method: AVHModel-Align-FT
AVHBenchICLR 2025
10/2024
code
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
Investigates how unimodal priors and misleading correlations between language, vision, and audio cause multimodal hallucinations.
CMMarXiv 2024
10/2024
code
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
Provides an audio-visual hallucination benchmark with human judgments for evaluating whether generated descriptions remain consistent with multimodal evidence.
AVHalluBencharXiv 2024
05/2024
dataset leaderboard
Emotion Inference (1 entry)
PaperBenchmarkVenue / DateResources
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Evaluates emotion hallucinations in psychological knowledge and multimodal perception through adversarial questions about emotional cues and their interpretation.
Related method: PEP-MEK
EmotionHallucerarXiv 2025
05/2025
code

Back to task index


Mitigation Strategies

[!NOTE] Training-free: ✔︎ No additional parameter learning; ✘ Training required, including auxiliary modules with a frozen backbone. Dates use first publication; newest first within each subtype.

🔵 Spatiotemporal Dynamics Mitigation (Dynamic Distortion)

Event Misordering (6 entries)
PaperMethodVenue / DateTraining-FreeResources
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
Injects measured physical-motion evidence into video reasoning to check kinematic claims, without training or modifying the underlying video model.
Related benchmark: MotionHalluc
PPVarXiv 2026
06/2026
✔︎dataset
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
Jointly trains temporal localization and video answering so an agent can inspect relevant clips and revise inaccurate grounding.
VideoTemp-o3ICML 2026
02/2026
✘page code
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
Creates counterfactual action videos and jointly optimizes visual and textual preferences to reduce action and temporal-order hallucinations.
MixDPOEMNLP 2026
01/2026
✘-
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
Selects among sampled responses using temporal attention collapse and terminates unreliable generations when visual attention vanishes.
SmartSightAAAI 2026
12/2025
✔︎-
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
Diagnoses hallucination tendencies token by token and adaptively contrasts temporal and spatial negatives without additional training.
SEASONarXiv 2025
12/2025
✔︎-
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
Combines supervised reasoning fine-tuning with thinking-based direct preference optimization, giving fabricated reasoning stronger feedback to improve factual grounding.
Related benchmark: HAVEN
Video-thinking (TDPO)arXiv 2025
03/2025
✘code
Duration Distortion (9 entries)
PaperMethodVenue / DateTraining-FreeResources
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Retrieves visual evidence to verify and arbitrate long-video answers, with separate reinforcement-learning objectives for each reflection stage.
Reflect-R1ECCV 2026
06/2026
✘code dataset
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
Rebalances decoder attention toward under-attended frames without training, changing visual encoding, or introducing auxiliary models.
DTRarXiv 2026
04/2026
✔︎-
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
Uses reinforcement learning to coordinate retrieval of video clips, images, and regions for efficient, evidence-grounded long-video answers.
VideoTIRarXiv 2026
03/2026
✘-
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Trains a lightweight scoring module to repeat useful frames during reasoning and counter drift away from visual evidence.
FrameRepeatarXiv 2026
03/2026
✘-
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
Trains long-video models to interleave reasoning with on-demand temporal grounding through a staged reinforcement-learning curriculum.
Video-TwGarXiv 2026
02/2026
✘-
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
Extracts compact, question-relevant visual evidence and uses reinforcement learning to anchor reasoning to the selected temporal evidence.
CoEICME 2026
01/2026
✘-
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
Uses temporal variation to identify and intervene in hallucination-sensitive model activations without further fine-tuning the language model.
TAAEarXiv 2025
05/2025
✘-
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
Uses DINOv2 saliency to reweight visual features during inference, emphasizing informative regions without additional training of the video model.
Related benchmark: VidHalluc
DINO-HEALCVPR 2025
12/2024
✔︎page code
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
Decomposes event queries into characteristic actions and uses visual-language models to estimate timestamps for temporally grounded responses.
Temporal InsightICPR 2024
01/2024
✔︎-
Frequency Confusion (3 entries)
PaperMethodVenue / DateTraining-FreeResources
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
Uses kinematic parsing of motion descriptions as a reinforcement-learning reward to improve fine-grained video captioning and reduce motion hallucinations.
Related benchmark: KPM-Bench
MoPEarXiv 2026
02/2026
✘-
Vript: A Video Is Worth Thousands of Words
Trains a video captioning model on densely annotated videos to generate detailed descriptions of visual content and unfolding events.
Related benchmark: Vript
VriptorNeurIPS 2024
06/2024
✘code
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Improves event timestamp localization by adding explicit time information to video tokens and using slot-based visual compression.
VTG-LLMAAAI 2025
05/2024
✘code

🟢 Referential Inconsistency Mitigation (Dynamic Distortion)

Character Conflation (3 entries)
PaperMethodVenue / DateTraining-FreeResources
STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models
Combines structured visual trajectories with symbolic aggregation to answer object-centric spatiotemporal questions using a fixed video-language model.
Related benchmark: STRAND
STRAND (Trajectory Reasoning)arXiv 2026
08/2026
✔︎page code
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
Separates timestamped perceptual evidence from reasoning and uses factuality-aware process rewards to train hallucination-resistant video models.
Video-DPLarXiv 2025
11/2025
✘code
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
Changes visual-text attention and sequential frame projection to preserve the influence of video evidence during open-ended answer generation.
Vista-LLaMACVPR 2024
12/2023
✘page code
Scene Conflation (2 entries)
PaperMethodVenue / DateTraining-FreeResources
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
Trains on adversarial preference pairs to reduce long-video hallucinations caused by incorrectly combining semantic information across video segments.
Related benchmark: ELV-Halluc
ELV-Halluc-DPOarXiv 2025
08/2025
✘code
Online Video Understanding: OVBench and VideoChat-Online
Combines pyramid memory with offline-to-online instruction training to support streaming-video perception, memory, and reasoning over continuously arriving frames.
Related benchmark: OVBench
VideoChat-OnlineCVPR 2025
12/2024
✘page code

🟠 Context-Driven Fabrication Mitigation (Content Fabrication)

Object-Action Hallucination (5 entries)
PaperMethodVenue / DateTraining-FreeResources
Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
Weights direct preference optimization by the severity of fabricated or incorrect limb actions, promoting physically faithful, person-specific video captions.
Related benchmark: FlexBench
GM-DPOarXiv 2026
09/2026
✘-
ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning
Ranks detected objects by prominence and iteratively revises captions to reduce omissions and hallucinations without retraining the captioning model.
ProCaparXiv 2026
07/2026
✔︎code
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
Replaces coordinate prediction with visually prompted object identities and reinforcement learning for spatially and temporally consistent video grounding.
STVG-R1arXiv 2026
02/2026
✘-
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
Combines hallucination-based negative captions with tracklet-phrase contrastive alignment to improve object and action faithfulness in video descriptions.
SANTAWACV 2026
12/2025
✘page
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Contrasts predictions from original and temporally disrupted video sequences during decoding to reduce event hallucinations without additional model training.
Related benchmark: EventHallusion
TCDarXiv 2024
09/2024
✔︎code
Scene-Event Hallucination (10 entries)
PaperMethodVenue / DateTraining-FreeResources
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
Adaptively reweights visual attention and contrasts selectively erased evidence to suppress prior-driven video hallucinations without training.
VADERarXiv 2026
08/2026
✔︎-
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Uses counterfactual traffic-video counterparts during contrastive decoding to improve the consistency of accident-related judgments without additional model training.
Related benchmark: CCTVBench
C-TCDarXiv 2026
04/2026
✔︎-
Video-ToC: Video Tree-of-Cue Reasoning
Localizes visual evidence through a tree of cues and trains reasoning with rewards adjusted to the video's reasoning demands.
Video-ToCarXiv 2026
04/2026
✘code
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
Separately supervises visual clue extraction and answer reasoning, then filters clues for faithful and interpretable video question answering.
ClueNetarXiv 2026
03/2026
✘-
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Builds event-based video scene graphs and applies visual-attention rewards during reinforcement fine-tuning to ground temporal reasoning.
GraphThinkerarXiv 2026
02/2026
✘-
MACD: Model-Aware Contrastive Decoding via Counterfactual Data
Uses model feedback to construct object-level counterfactual video inputs for evidence-grounded contrastive decoding.
MACDarXiv 2026
02/2026
✔︎-
Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models
Trains a lightweight residual aligner on frozen video representations to improve spatiotemporal consistency and alignment with response semantics.
ViSSResarXiv 2026
01/2026
✘-
Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency
Combines multi-level semantic alignment with progressive training to reduce hallucinations caused by weak discrimination between video and language concepts.
MMAIJCAI 2025
08/2025
✘-
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
Learns video preferences online using prompt-aware augmented clips as rejected inputs, reducing incorrect rejections during hallucination mitigation training.
PaMi-VDPOarXiv 2025
04/2025
✘-
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
Disentangles spatial and temporal attention and introduces Harmonic RoPE to reduce confusion between depicted actions and surrounding scene context.
MASH-VLMCVPR 2025
03/2025
✘-
Compositional and Factuality Hallucination (2 entries)
PaperMethodVenue / DateTraining-FreeResources
Target-Checked Reliability Score Refinement for Video Question Answering
Refines learned reliability scores through target-checked evidence across video samplings, improving answer selection while leaving generated answers unchanged.
Related benchmark: VHD
TRACE-RCarXiv 2026
09/2026
✘code dataset
Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study
Verifies timestamped video-answer claims by independently re-captioning cited frames and checking textual entailment with a separate model.
GroundedVQAarXiv 2026
08/2026
✔︎code
Both Object-Action & Scene-Event (8 entries)
PaperMethodVenue / DateTraining-FreeResources
MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models
Trains a lightweight patcher to selectively replace unreliable visual tokens before generation while leaving the original video model unchanged.
MultiToParXiv 2026
06/2026
✘-
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
Identifies attention sink tokens and suppresses them during visual token pruning to preserve fine-grained grounding and hallucination robustness.
SToPECCV 2026
04/2026
✔︎-
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
Trains specialized experts with dual encoders and adaptive token routing to disentangle misleading text overlays from underlying visual evidence.
Related benchmark: VisualTextTrap
VTHM-MoEarXiv 2026
04/2026
✘-
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
Targets risky decoding steps with layer-specific visual evidence, combining grounding restoration with temporal counterfactual checks.
STEARarXiv 2026
04/2026
✔︎-
Reinforcing Consistency in Video MLLMs with Structured Rewards
Audits captions as factual and temporal claims, then trains with scene-graph, temporal, and video-grounded self-verification rewards.
Structured RewardsCOLM 2026
04/2026
✘-
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
Learns an adaptive controller for triple-path contrastive decoding, combining video perturbations and saliency enhancement to address compositional hallucinations.
Related benchmark: OmniVCHall
TriCDarXiv 2026
01/2026
✘code
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Applies group relative policy optimization to help video models recognize physical and logical violations instead of substituting familiar prior expectations.
Related benchmark: VideoHallu
VideoHallu-GRPONeurIPS 2025
05/2025
✘code
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Optimizes preferences at whole-video, temporal, and object levels using spatially and temporally annotated response pairs.
VistaDPOICML 2025
04/2025
✘code

🟣 Audio-Visual Conflict Mitigation (Content Fabrication)

Action Attribution (3 entries)
PaperMethodVenue / DateTraining-FreeResources
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
Uses modality-aware masking and confidence-guided contrastive decoding to suppress hallucinations across audio, video, and language without training.
AVCDNeurIPS 2025
05/2025
✔︎code
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Fine-tunes audio-visual models with benchmark-derived training data to improve cross-modal alignment and robustness against auditory and visual hallucinations.
Related benchmark: AVHBench
AVHModel-Align-FTICLR 2025
10/2024
✘code
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
Improves audio-visual captions through repeated preference optimization and rebirth tuning designed to retain non-captioning abilities.
mrDPOarXiv 2024
10/2024
✘page code
Emotion Inference (1 entry)
PaperMethodVenue / DateTraining-FreeResources
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Combines perception-enhanced prompting with memory-based emotion knowledge to reduce emotion hallucinations without additional training of the multimodal model.
Related benchmark: EmotionHallucer
PEP-MEKarXiv 2025
05/2025
✔︎code

Back to task index


Evaluation Analyses

Studies of evaluation validity and hallucination mechanisms without a standalone benchmark or mitigation method. Cross-category studies use their closest primary category. Dates use first publication.

Context-Driven Fabrication Analyses (Content Fabrication)

Compositional and Factuality Hallucination (1 entry)
PaperAnalysisVenue / DateResources
Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Intervenes separately on video-agent grounding, observation, and reasoning to test whether benchmark scores predict downstream hallucination risk.
Beneath the ScoresNeurIPS 2026 TAE Workshop
09/2026
-

Back to task index


Citation

If this repository or survey helps your work, please cite:

@article{huang2026distorted,
  title={Distorted or Fabricated? A Survey on Hallucination in Video LLMs},
  author={Huang, Yiyang and Zhang, Yitian and Wang, Yizhou and Zhang, Mingyuan and Shi, Liang and Zeng, Huimin and Fu, Yun},
  journal={arXiv preprint arXiv:2604.12944},
  year={2026}
}

Contributing

[!TIP] Contributions are welcome:

🔀 Pull Request — Add new papers, update resource links, or correct errors
🐛 Open an Issue — Report mistakes, suggest missing papers, or request features

See the curation guide for summary sources, task-tag definitions, publication versus addition dates, and consistency checks. The browser groups contributions by paper; the taxonomy tables retain each benchmark, method, and analysis separately.

Resource gaps tracked in data/papers.json:

  • Add official code links for 47 entries. Browse: missing code
  • Add official project pages for 78 entries. Browse: missing project pages
  • Add official dataset or leaderboard links when available.
📝 PR Format Guide

Edit the shared records in data/papers.json and data/paper_details.json, then regenerate the task index and paper tables with python3 scripts/generate_readme.py. Do not edit generated table rows directly.

Include a contribution-specific English description, a reviewed primary source, verified dates, and official resource links. Preserve stable entry IDs and existing contributions. See the update checklist for all validation steps.


If this repository helps, please consider giving it a ⭐

Maintained by the SmileLab team at Northeastern University.

acl2026
awesome-list
hallucination
hallucination-evaluation
survey
video-llm
video-understanding

hukcc/Awesome-Video-Hallucination

[ACL 2026] Curated papers on video LLM hallucination, with benchmarks, mitigation methods, and an interactive browser. Updated monthly.

Python

44

23 commits

updated Oct 1, 2026

See the code

README

Awesome-Video-Hallucination Awesome

arXiv ACL 2026 Findings Entries Auto arXiv Update License: MIT Last Commit

A curated paper list on hallucination in Video Large Language Models (Vid-LLMs), covering 42 benchmarks, 52 mitigation methods, and 1 evaluation analysis. The 95 entries represent 78 distinct papers; a paper may contribute both a benchmark and a method. Updated monthly via arXiv search and manual review.

📄 Survey Paper: Distorted or Fabricated? A Survey on Hallucination in Video LLMs

🔎 Interactive Browser: Browse 78 papers with source-backed summaries, task tags, combined filters, and table or card views.

Framework overview

Table of Contents


Latest Updates

  • [2026/10] Added 7 papers, introduced evaluation analyses, and updated the taxonomy. Review log.
  • [2026/09] Added 2 hallucination mitigation papers.
  • [2026/08] Added 6 papers and updated the taxonomy.
  • [2026/04] Our survey was accepted to ACL 2026 Findings.
Recently added · 2026-10-01 · 7 papers

FlexBench / GM-DPO · Beneath the Scores · VidOmni-Bench · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning)

Repository additions, not publication dates. All recent additions


Find Your Papers

Find Benchmarks · Training-Free Methods · With Code · Latest Papers · Recently Added

Start here: Reading guide, from the survey to benchmarks, mitigation, and evaluation limits.

Scope: Relevance review distinguishes direct hallucination research, broader evaluations, and related-work candidates. All entries remain available pending manual decisions.

Browse by Task

Task index · 9 topics · 78 papers

Paper-level task tags. Newest first; links lead to individual contributions below.

Audio-Visual Understanding (10 papers)

Video-HolmesV2 · TraceAV-Bench · Audio Hallucination QA · EGOILLUSION · AVCD · EmotionHallucer / PEP-MEK · AVHBench / AVHModel-Align-FT · CMM · mrDPO · AVHalluBench

Hallucination Detection (24 papers)

Beneath the Scores · VidOmni-Bench · VHD / TRACE-RC · VidHalLoc · GroundedVQA · MoHallBench · DualFact · Audio Hallucination QA · INFACT · VideoHEDGE · SmartSight · EGOILLUSION · MESH · ELV-Halluc / ELV-Halluc-DPO · ARGUS · EmotionHallucer / PEP-MEK · HAVEN / Video-thinking (TDPO) · VidHal · AVHBench / AVHModel-Align-FT · CMM · VideoHallucer · Vript / Vriptor · AVHalluBench · FactVC

Long-Video Understanding (10 papers)

VidOmni-Bench · Video-HolmesV2 · Reflect-R1 · DistractionBench · TraceAV-Bench · VideoTIR · Video-TwG · VideoTemp-o3 · ELV-Halluc / ELV-Halluc-DPO · Vript / Vriptor

Motion Understanding (9 papers)

FlexBench / GM-DPO · MoHallBench · MotionHalluc / PPV · KPM-Bench / MoPE · MixDPO · SANTA · MHBench · MASH-VLM · VidHalluc / DINO-HEAL

Temporal Grounding (7 papers)

Video-TwG · GraphThinker · STVG-R1 · VideoTemp-o3 · CoE · VTG-LLM · Temporal Insight

Video Captioning (15 papers)

FlexBench / GM-DPO · VidOmni-Bench · VidHalLoc · ProCap · DualFact · Structured Rewards · KPM-Bench / MoPE · SANTA · NOAH · ARGUS · VistaDPO · VidHal · mrDPO · Vript / Vriptor · FactVC

Video QA (41 papers)

Beneath the Scores · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning) · GroundedVQA · MoHallBench · VidPair-Halluc · Reflect-R1 · MotionHalluc / PPV · DistractionBench · TOC-Bench · TraceAV-Bench · Audio Hallucination QA · CCTVBench / C-TCD · VisualTextTrap / VTHM-MoE · Structured Rewards · VideoTIR · GameplayQA · FrameRepeat · ClueNet · INFACT · Video-TwG · KPM-Bench / MoPE · VideoTemp-o3 · ViSSRes · VideoHEDGE · Video-DPL · NOAH · EGOILLUSION · MESH · VideoHallu / VideoHallu-GRPO · VistaDPO · MHBench · RoadSocial · HAVEN / Video-thinking (TDPO) · MASH-VLM · OVBench / VideoChat-Online · VidHalluc / DINO-HEAL · VideoHallucer · Vista-LLaMA

Video Reasoning (26 papers)

Beneath the Scores · STRAND / STRAND (Trajectory Reasoning) · VADER · VidPair-Halluc · TOC-Bench · CCTVBench / C-TCD · Video-ToC · GasVideo-1000 · STEAR · Structured Rewards · GameplayQA · FrameRepeat · ClueNet · INFACT · GraphThinker · OmniVCHall / TriCD · CoE · MixDPO · SEASON · Video-DPL · NOAH · VideoHallu / VideoHallu-GRPO · HAVEN / Video-thinking (TDPO) · VidHalluc / DINO-HEAL · VidHal · EventHallusion / TCD

Video Understanding (21 papers)

VADER · MultiToP · SToP · Video-ToC · GasVideo-1000 · VisualTextTrap / VTHM-MoE · DTR · STEAR · STVG-R1 · MACD · OmniVCHall / TriCD · ViSSRes · SmartSight · SEASON · MMA · TAAE · PaMi-VDPO · RoadSocial · OVBench / VideoChat-Online · EventHallusion / TCD · Vista-LLaMA


Taxonomy of Video Hallucinations

View the full taxonomy tree (95 contributions)

Mechanism-driven taxonomy of Vid-LLM hallucinations
Mechanism-driven taxonomy of Vid-LLM hallucinations. Solid fill = benchmarks; striped fill = mitigation methods; dashed outline = evaluation analyses. Placement indicates a primary indexing category, not exclusive coverage.
Generated from paper data using the LaTeX tree source.


Evaluation Benchmarks

[!NOTE] Newest first within each subtype. Date = first arXiv submission, or publisher issue date when no arXiv record is used; venue years may differ. Sources and review notes.

Resource badge legend

page = Project Page
code = GitHub Repository
dataset = Hugging Face Dataset
dataset = Kaggle Dataset
leaderboard = Leaderboard
- = No verified resource link

🔵 Spatiotemporal Dynamics Benchmarks (Dynamic Distortion)

Event Misordering (7 entries)
PaperBenchmarkVenue / DateResources
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
Probes motion hallucinations caused by prior knowledge, sequential inference, and visual similarity through several question formats.
MoHallBencharXiv 2026
07/2026
-
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
Diagnoses fine-grained motion hallucinations involving direction, attribution, and temporal kinematics, separating distinct failures in interpreting physical movement.
Related method: PPV
MotionHallucarXiv 2026
06/2026
dataset
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
Evaluates fine-grained limb motion through video captioning and question answering, using kinematic parsing to assess motion descriptions.
Related method: MoPE
KPM-BencharXiv 2026
02/2026
-
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
Measures both fabricated content and missing information in free-form video captions against human-written reference descriptions.
ARGUSICCV 2025
06/2025
page code
MHBench: Demystifying Motion Hallucination in VideoLLMs
Tests motion hallucinations using original actions, actions with reversed meanings, and incomplete actions that challenge static visual cues.
MHBenchAAAI 2025
04/2025
code
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
Diagnoses video hallucinations by crossing hallucination causes, object-scene-event aspects, and question formats to expose distinct failures in video understanding.
Related method: Video-thinking (TDPO)
HAVENarXiv 2025
03/2025
code
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
Evaluates hallucinations about actions, temporal order, and scene transitions through video questions, caption generation, and event sorting.
Related method: DINO-HEAL
VidHallucCVPR 2025
12/2024
page code
Duration Distortion (2 entries)
PaperBenchmarkVenue / DateResources
Online Video Understanding: OVBench and VideoChat-Online
Evaluates streaming-video question answering across online perception, memory, and reasoning tasks; its scope extends beyond hallucination-specific evaluation.
Related method: VideoChat-Online
OVBenchCVPR 2025
12/2024
page code
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
Uses adversarial question pairs to distinguish hallucinations about visible objects and temporal relations from unsupported external information.
VideoHallucerarXiv 2024
06/2024
code
Frequency Confusion (2 entries)
PaperBenchmarkVenue / DateResources
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
Evaluates temporal hallucinations by asking models to distinguish and rank video captions with different degrees of factual distortion.
VidHalTMLR 2026
11/2024
code
Vript: A Video Is Worth Thousands of Words
Provides dense video descriptions and Vript-Hard evaluations targeting hallucinated captions, information retrieval, and temporal ordering of video events.
Related method: Vriptor
VriptNeurIPS 2024
06/2024
code

🟢 Referential Inconsistency Benchmarks (Dynamic Distortion)

Character Conflation (4 entries)
PaperBenchmarkVenue / DateResources
STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models
Evaluates object-state changes, identity persistence, and relational reasoning, requiring answers to satisfy jointly grounded spatiotemporal prerequisites.
Related method: STRAND (Trajectory Reasoning)
STRANDarXiv 2026
08/2026
page code
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
Tests object identity, state, and continuity with trajectory-grounded questions designed to require temporally ordered visual evidence.
TOC-BencharXiv 2026
05/2026
code
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
Tests hallucinations in first-person videos through human-annotated open and closed questions about visual and auditory evidence.
EGOILLUSIONEMNLP 2025
11/2025
page
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
Uses hierarchical questions and plausible distractors to expose hallucinations about objects, attributes, and subject-action relations across video segments.
MESHACM MM 2025
09/2025
code
Scene Conflation (2 entries)
PaperBenchmarkVenue / DateResources
Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models
Inserts advertising distractors into long videos to test whether models conflate subjects and events from unrelated segments.
DistractionBencharXiv 2026
05/2026
code dataset
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
Tests hallucinations from semantic aggregation in long videos, where models can incorrectly combine information across separate video segments.
Related method: ELV-Halluc-DPO
ELV-HallucarXiv 2025
08/2025
code

🟠 Context-Driven Fabrication Benchmarks (Content Fabrication)

Object-Action Hallucination (4 entries)
PaperBenchmarkVenue / DateResources
Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
Evaluates person-specific limb-motion captions across shots, distinguishing omissions from fabricated or incorrect actions through graded physical alignment.
Related method: GM-DPO
FlexBencharXiv 2026
09/2026
-
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
Uses synchronized multiplayer gameplay views and structured distractors to test agent identity, role attribution, and grounded reasoning.
GameplayQAACL 2026
03/2026
page code dataset
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Tests prior-driven hallucinations using synthetic videos that violate physical or logical expectations, challenging models to follow observed evidence.
Related method: VideoHallu-GRPO
VideoHalluNeurIPS 2025
05/2025
code
Models See Hallucinations: Evaluating the Factuality in Video Captioning
Studies factual errors in video captions and introduces a weakly supervised factuality metric with human-annotated evaluation data.
FactVCEMNLP 2023
03/2023
code
Scene-Event Hallucination (5 entries)
PaperBenchmarkVenue / DateResources
VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
Tests whether models can verify individual events in dense video captions, using human-checked incorrect descriptions as hard negatives.
VidOmni-BencharXiv 2026
09/2026
-
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Tests consistent traffic-hazard judgments using real accident videos paired with counterfactual counterparts that alter the evidence for an accident.
Related method: C-TCD
CCTVBencharXiv 2026
04/2026
-
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
Inserts unrelated clips into videos to measure narrative-prior hallucinations and omissions through captioning and question-answering tasks.
NOAHarXiv 2025
11/2025
page code
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
Provides diverse social-media road videos and question-answer pairs for evaluating road-event understanding across viewpoints and geographic settings.
RoadSocialCVPR 2025
03/2025
page code
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Diagnoses event hallucinations driven by language priors and visual biases, testing whether answers reflect events actually present in videos.
Related method: TCD
EventHallusionarXiv 2024
09/2024
code
Compositional and Factuality Hallucination (9 entries)
PaperBenchmarkVenue / DateResources
Target-Checked Reliability Score Refinement for Video Question Answering
Provides controlled video-question examples of confident but incorrect answers for diagnosing hallucinations and evaluating the reliability of confidence estimates.
Related method: TRACE-RC
VHDarXiv 2026
09/2026
code dataset
Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators
Compares hallucination detectors on adversarial video questions and captions, with a shared protocol spanning ontology and dynamic errors.
VidHalLocarXiv 2026
09/2026
-
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
Uses adversarial video pairs with similar backgrounds but different foreground events to isolate spatial and temporal hallucinations.
VidPair-HallucECCV 2026
06/2026
page
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding
Evaluates procedural-caption factuality at conceptual and grounded argument levels, using either textual references or direct video evidence.
DualFactACL 2026 Findings
04/2026
-
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
Tests whether misleading conversational feedback causes video models to abandon correct judgments and invent unsupported spatiotemporal explanations.
GasVideo-1000arXiv 2026
04/2026
page
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
Tests whether misleading text overlays cause video models to contradict visual evidence when answering questions about the depicted content.
Related method: VTHM-MoE
VisualTextTraparXiv 2026
04/2026
-
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Separates faithfulness from factuality errors in video answers and probes robustness to degraded visuals, corrupted evidence, and temporal interventions.
INFACTarXiv 2026
03/2026
-
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
Benchmarks isolated and compositional hallucinations across spatial and temporal dimensions, probing failures involving multiple interacting types of video evidence.
Related method: TriCD
OmniVCHallarXiv 2026
01/2026
code
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
Estimates answer reliability by clustering responses to clean and perturbed videos and measuring semantic uncertainty across those responses.
VideoHEDGEarXiv 2026
01/2026
code

🟣 Audio-Visual Conflict Benchmarks (Content Fabrication)

Action Attribution (6 entries)
PaperBenchmarkVenue / DateResources
Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?
Requires long-video answers to cite precise audio-visual evidence, using evidence-aware scoring to penalize guessing and fabricated support.
Video-HolmesV2ECCV 2026
09/2026
-
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
Tests multi-hop reasoning and hallucination robustness using explicit evidence trajectories distributed across long audio-visual recordings.
TraceAV-BencharXiv 2026
05/2026
-
Exploring Audio Hallucination in Egocentric Video Understanding
Probes imagined foreground and background sounds in egocentric videos through questions targeting visible but inaudible events.
Audio Hallucination QAICASSP 2026
04/2026
-
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Evaluates audio-visual hallucinations, cross-modal matching, and reasoning through tasks that distinguish auditory evidence from visible video content.
Related method: AVHModel-Align-FT
AVHBenchICLR 2025
10/2024
code
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
Investigates how unimodal priors and misleading correlations between language, vision, and audio cause multimodal hallucinations.
CMMarXiv 2024
10/2024
code
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
Provides an audio-visual hallucination benchmark with human judgments for evaluating whether generated descriptions remain consistent with multimodal evidence.
AVHalluBencharXiv 2024
05/2024
dataset leaderboard
Emotion Inference (1 entry)
PaperBenchmarkVenue / DateResources
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Evaluates emotion hallucinations in psychological knowledge and multimodal perception through adversarial questions about emotional cues and their interpretation.
Related method: PEP-MEK
EmotionHallucerarXiv 2025
05/2025
code

Back to task index


Mitigation Strategies

[!NOTE] Training-free: ✔︎ No additional parameter learning; ✘ Training required, including auxiliary modules with a frozen backbone. Dates use first publication; newest first within each subtype.

🔵 Spatiotemporal Dynamics Mitigation (Dynamic Distortion)

Event Misordering (6 entries)
PaperMethodVenue / DateTraining-FreeResources
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
Injects measured physical-motion evidence into video reasoning to check kinematic claims, without training or modifying the underlying video model.
Related benchmark: MotionHalluc
PPVarXiv 2026
06/2026
✔︎dataset
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
Jointly trains temporal localization and video answering so an agent can inspect relevant clips and revise inaccurate grounding.
VideoTemp-o3ICML 2026
02/2026
✘page code
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
Creates counterfactual action videos and jointly optimizes visual and textual preferences to reduce action and temporal-order hallucinations.
MixDPOEMNLP 2026
01/2026
✘-
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
Selects among sampled responses using temporal attention collapse and terminates unreliable generations when visual attention vanishes.
SmartSightAAAI 2026
12/2025
✔︎-
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
Diagnoses hallucination tendencies token by token and adaptively contrasts temporal and spatial negatives without additional training.
SEASONarXiv 2025
12/2025
✔︎-
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
Combines supervised reasoning fine-tuning with thinking-based direct preference optimization, giving fabricated reasoning stronger feedback to improve factual grounding.
Related benchmark: HAVEN
Video-thinking (TDPO)arXiv 2025
03/2025
✘code
Duration Distortion (9 entries)
PaperMethodVenue / DateTraining-FreeResources
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Retrieves visual evidence to verify and arbitrate long-video answers, with separate reinforcement-learning objectives for each reflection stage.
Reflect-R1ECCV 2026
06/2026
✘code dataset
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
Rebalances decoder attention toward under-attended frames without training, changing visual encoding, or introducing auxiliary models.
DTRarXiv 2026
04/2026
✔︎-
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
Uses reinforcement learning to coordinate retrieval of video clips, images, and regions for efficient, evidence-grounded long-video answers.
VideoTIRarXiv 2026
03/2026
✘-
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Trains a lightweight scoring module to repeat useful frames during reasoning and counter drift away from visual evidence.
FrameRepeatarXiv 2026
03/2026
✘-
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
Trains long-video models to interleave reasoning with on-demand temporal grounding through a staged reinforcement-learning curriculum.
Video-TwGarXiv 2026
02/2026
✘-
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
Extracts compact, question-relevant visual evidence and uses reinforcement learning to anchor reasoning to the selected temporal evidence.
CoEICME 2026
01/2026
✘-
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
Uses temporal variation to identify and intervene in hallucination-sensitive model activations without further fine-tuning the language model.
TAAEarXiv 2025
05/2025
✘-
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
Uses DINOv2 saliency to reweight visual features during inference, emphasizing informative regions without additional training of the video model.
Related benchmark: VidHalluc
DINO-HEALCVPR 2025
12/2024
✔︎page code
Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models
Decomposes event queries into characteristic actions and uses visual-language models to estimate timestamps for temporally grounded responses.
Temporal InsightICPR 2024
01/2024
✔︎-
Frequency Confusion (3 entries)
PaperMethodVenue / DateTraining-FreeResources
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
Uses kinematic parsing of motion descriptions as a reinforcement-learning reward to improve fine-grained video captioning and reduce motion hallucinations.
Related benchmark: KPM-Bench
MoPEarXiv 2026
02/2026
✘-
Vript: A Video Is Worth Thousands of Words
Trains a video captioning model on densely annotated videos to generate detailed descriptions of visual content and unfolding events.
Related benchmark: Vript
VriptorNeurIPS 2024
06/2024
✘code
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Improves event timestamp localization by adding explicit time information to video tokens and using slot-based visual compression.
VTG-LLMAAAI 2025
05/2024
✘code

🟢 Referential Inconsistency Mitigation (Dynamic Distortion)

Character Conflation (3 entries)
PaperMethodVenue / DateTraining-FreeResources
STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models
Combines structured visual trajectories with symbolic aggregation to answer object-centric spatiotemporal questions using a fixed video-language model.
Related benchmark: STRAND
STRAND (Trajectory Reasoning)arXiv 2026
08/2026
✔︎page code
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
Separates timestamped perceptual evidence from reasoning and uses factuality-aware process rewards to train hallucination-resistant video models.
Video-DPLarXiv 2025
11/2025
✘code
Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
Changes visual-text attention and sequential frame projection to preserve the influence of video evidence during open-ended answer generation.
Vista-LLaMACVPR 2024
12/2023
✘page code
Scene Conflation (2 entries)
PaperMethodVenue / DateTraining-FreeResources
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
Trains on adversarial preference pairs to reduce long-video hallucinations caused by incorrectly combining semantic information across video segments.
Related benchmark: ELV-Halluc
ELV-Halluc-DPOarXiv 2025
08/2025
✘code
Online Video Understanding: OVBench and VideoChat-Online
Combines pyramid memory with offline-to-online instruction training to support streaming-video perception, memory, and reasoning over continuously arriving frames.
Related benchmark: OVBench
VideoChat-OnlineCVPR 2025
12/2024
✘page code

🟠 Context-Driven Fabrication Mitigation (Content Fabrication)

Object-Action Hallucination (5 entries)
PaperMethodVenue / DateTraining-FreeResources
Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
Weights direct preference optimization by the severity of fabricated or incorrect limb actions, promoting physically faithful, person-specific video captions.
Related benchmark: FlexBench
GM-DPOarXiv 2026
09/2026
✘-
ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning
Ranks detected objects by prominence and iteratively revises captions to reduce omissions and hallucinations without retraining the captioning model.
ProCaparXiv 2026
07/2026
✔︎code
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
Replaces coordinate prediction with visually prompted object identities and reinforcement learning for spatially and temporally consistent video grounding.
STVG-R1arXiv 2026
02/2026
✘-
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
Combines hallucination-based negative captions with tracklet-phrase contrastive alignment to improve object and action faithfulness in video descriptions.
SANTAWACV 2026
12/2025
✘page
EventHallusion: Diagnosing Event Hallucinations in Video LLMs
Contrasts predictions from original and temporally disrupted video sequences during decoding to reduce event hallucinations without additional model training.
Related benchmark: EventHallusion
TCDarXiv 2024
09/2024
✔︎code
Scene-Event Hallucination (10 entries)
PaperMethodVenue / DateTraining-FreeResources
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
Adaptively reweights visual attention and contrasts selectively erased evidence to suppress prior-driven video hallucinations without training.
VADERarXiv 2026
08/2026
✔︎-
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Uses counterfactual traffic-video counterparts during contrastive decoding to improve the consistency of accident-related judgments without additional model training.
Related benchmark: CCTVBench
C-TCDarXiv 2026
04/2026
✔︎-
Video-ToC: Video Tree-of-Cue Reasoning
Localizes visual evidence through a tree of cues and trains reasoning with rewards adjusted to the video's reasoning demands.
Video-ToCarXiv 2026
04/2026
✘code
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
Separately supervises visual clue extraction and answer reasoning, then filters clues for faithful and interpretable video question answering.
ClueNetarXiv 2026
03/2026
✘-
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Builds event-based video scene graphs and applies visual-attention rewards during reinforcement fine-tuning to ground temporal reasoning.
GraphThinkerarXiv 2026
02/2026
✘-
MACD: Model-Aware Contrastive Decoding via Counterfactual Data
Uses model feedback to construct object-level counterfactual video inputs for evidence-grounded contrastive decoding.
MACDarXiv 2026
02/2026
✔︎-
Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models
Trains a lightweight residual aligner on frozen video representations to improve spatiotemporal consistency and alignment with response semantics.
ViSSResarXiv 2026
01/2026
✘-
Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency
Combines multi-level semantic alignment with progressive training to reduce hallucinations caused by weak discrimination between video and language concepts.
MMAIJCAI 2025
08/2025
✘-
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
Learns video preferences online using prompt-aware augmented clips as rejected inputs, reducing incorrect rejections during hallucination mitigation training.
PaMi-VDPOarXiv 2025
04/2025
✘-
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
Disentangles spatial and temporal attention and introduces Harmonic RoPE to reduce confusion between depicted actions and surrounding scene context.
MASH-VLMCVPR 2025
03/2025
✘-
Compositional and Factuality Hallucination (2 entries)
PaperMethodVenue / DateTraining-FreeResources
Target-Checked Reliability Score Refinement for Video Question Answering
Refines learned reliability scores through target-checked evidence across video samplings, improving answer selection while leaving generated answers unchanged.
Related benchmark: VHD
TRACE-RCarXiv 2026
09/2026
✘code dataset
Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study
Verifies timestamped video-answer claims by independently re-captioning cited frames and checking textual entailment with a separate model.
GroundedVQAarXiv 2026
08/2026
✔︎code
Both Object-Action & Scene-Event (8 entries)
PaperMethodVenue / DateTraining-FreeResources
MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models
Trains a lightweight patcher to selectively replace unreliable visual tokens before generation while leaving the original video model unchanged.
MultiToParXiv 2026
06/2026
✘-
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
Identifies attention sink tokens and suppresses them during visual token pruning to preserve fine-grained grounding and hallucination robustness.
SToPECCV 2026
04/2026
✔︎-
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
Trains specialized experts with dual encoders and adaptive token routing to disentangle misleading text overlays from underlying visual evidence.
Related benchmark: VisualTextTrap
VTHM-MoEarXiv 2026
04/2026
✘-
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
Targets risky decoding steps with layer-specific visual evidence, combining grounding restoration with temporal counterfactual checks.
STEARarXiv 2026
04/2026
✔︎-
Reinforcing Consistency in Video MLLMs with Structured Rewards
Audits captions as factual and temporal claims, then trains with scene-graph, temporal, and video-grounded self-verification rewards.
Structured RewardsCOLM 2026
04/2026
✘-
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
Learns an adaptive controller for triple-path contrastive decoding, combining video perturbations and saliency enhancement to address compositional hallucinations.
Related benchmark: OmniVCHall
TriCDarXiv 2026
01/2026
✘code
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
Applies group relative policy optimization to help video models recognize physical and logical violations instead of substituting familiar prior expectations.
Related benchmark: VideoHallu
VideoHallu-GRPONeurIPS 2025
05/2025
✘code
VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Optimizes preferences at whole-video, temporal, and object levels using spatially and temporally annotated response pairs.
VistaDPOICML 2025
04/2025
✘code

🟣 Audio-Visual Conflict Mitigation (Content Fabrication)

Action Attribution (3 entries)
PaperMethodVenue / DateTraining-FreeResources
AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
Uses modality-aware masking and confidence-guided contrastive decoding to suppress hallucinations across audio, video, and language without training.
AVCDNeurIPS 2025
05/2025
✔︎code
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
Fine-tunes audio-visual models with benchmark-derived training data to improve cross-modal alignment and robustness against auditory and visual hallucinations.
Related benchmark: AVHBench
AVHModel-Align-FTICLR 2025
10/2024
✘code
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
Improves audio-visual captions through repeated preference optimization and rebirth tuning designed to retain non-captioning abilities.
mrDPOarXiv 2024
10/2024
✘page code
Emotion Inference (1 entry)
PaperMethodVenue / DateTraining-FreeResources
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Combines perception-enhanced prompting with memory-based emotion knowledge to reduce emotion hallucinations without additional training of the multimodal model.
Related benchmark: EmotionHallucer
PEP-MEKarXiv 2025
05/2025
✔︎code

Back to task index


Evaluation Analyses

Studies of evaluation validity and hallucination mechanisms without a standalone benchmark or mitigation method. Cross-category studies use their closest primary category. Dates use first publication.

Context-Driven Fabrication Analyses (Content Fabrication)

Compositional and Factuality Hallucination (1 entry)
PaperAnalysisVenue / DateResources
Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Intervenes separately on video-agent grounding, observation, and reasoning to test whether benchmark scores predict downstream hallucination risk.
Beneath the ScoresNeurIPS 2026 TAE Workshop
09/2026
-

Back to task index


Citation

If this repository or survey helps your work, please cite:

@article{huang2026distorted,
  title={Distorted or Fabricated? A Survey on Hallucination in Video LLMs},
  author={Huang, Yiyang and Zhang, Yitian and Wang, Yizhou and Zhang, Mingyuan and Shi, Liang and Zeng, Huimin and Fu, Yun},
  journal={arXiv preprint arXiv:2604.12944},
  year={2026}
}

Contributing

[!TIP] Contributions are welcome:

🔀 Pull Request — Add new papers, update resource links, or correct errors
🐛 Open an Issue — Report mistakes, suggest missing papers, or request features

See the curation guide for summary sources, task-tag definitions, publication versus addition dates, and consistency checks. The browser groups contributions by paper; the taxonomy tables retain each benchmark, method, and analysis separately.

Resource gaps tracked in data/papers.json:

  • Add official code links for 47 entries. Browse: missing code
  • Add official project pages for 78 entries. Browse: missing project pages
  • Add official dataset or leaderboard links when available.
📝 PR Format Guide

Edit the shared records in data/papers.json and data/paper_details.json, then regenerate the task index and paper tables with python3 scripts/generate_readme.py. Do not edit generated table rows directly.

Include a contribution-specific English description, a reviewed primary source, verified dates, and official resource links. Preserve stable entry IDs and existing contributions. See the update checklist for all validation steps.


If this repository helps, please consider giving it a ⭐

Maintained by the SmileLab team at Northeastern University.

acl2026
awesome-list
hallucination
hallucination-evaluation
survey
video-llm
video-understanding