[ACL 2026] Curated papers on video LLM hallucination, with benchmarks, mitigation methods, and an interactive browser. Updated monthly.
See the codeA curated paper list on hallucination in Video Large Language Models (Vid-LLMs), covering 42 benchmarks, 52 mitigation methods, and 1 evaluation analysis. The 95 entries represent 78 distinct papers; a paper may contribute both a benchmark and a method. Updated monthly via arXiv search and manual review.
📄 Survey Paper: Distorted or Fabricated? A Survey on Hallucination in Video LLMs
🔎 Interactive Browser: Browse 78 papers with source-backed summaries, task tags, combined filters, and table or card views.

FlexBench / GM-DPO · Beneath the Scores · VidOmni-Bench · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning)
Repository additions, not publication dates. All recent additions
Find Benchmarks · Training-Free Methods · With Code · Latest Papers · Recently Added
Start here: Reading guide, from the survey to benchmarks, mitigation, and evaluation limits.
Scope: Relevance review distinguishes direct hallucination research, broader evaluations, and related-work candidates. All entries remain available pending manual decisions.
Paper-level task tags. Newest first; links lead to individual contributions below.
Video-HolmesV2 · TraceAV-Bench · Audio Hallucination QA · EGOILLUSION · AVCD · EmotionHallucer / PEP-MEK · AVHBench / AVHModel-Align-FT · CMM · mrDPO · AVHalluBench
Beneath the Scores · VidOmni-Bench · VHD / TRACE-RC · VidHalLoc · GroundedVQA · MoHallBench · DualFact · Audio Hallucination QA · INFACT · VideoHEDGE · SmartSight · EGOILLUSION · MESH · ELV-Halluc / ELV-Halluc-DPO · ARGUS · EmotionHallucer / PEP-MEK · HAVEN / Video-thinking (TDPO) · VidHal · AVHBench / AVHModel-Align-FT · CMM · VideoHallucer · Vript / Vriptor · AVHalluBench · FactVC
VidOmni-Bench · Video-HolmesV2 · Reflect-R1 · DistractionBench · TraceAV-Bench · VideoTIR · Video-TwG · VideoTemp-o3 · ELV-Halluc / ELV-Halluc-DPO · Vript / Vriptor
FlexBench / GM-DPO · MoHallBench · MotionHalluc / PPV · KPM-Bench / MoPE · MixDPO · SANTA · MHBench · MASH-VLM · VidHalluc / DINO-HEAL
Video-TwG · GraphThinker · STVG-R1 · VideoTemp-o3 · CoE · VTG-LLM · Temporal Insight
FlexBench / GM-DPO · VidOmni-Bench · VidHalLoc · ProCap · DualFact · Structured Rewards · KPM-Bench / MoPE · SANTA · NOAH · ARGUS · VistaDPO · VidHal · mrDPO · Vript / Vriptor · FactVC
Beneath the Scores · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning) · GroundedVQA · MoHallBench · VidPair-Halluc · Reflect-R1 · MotionHalluc / PPV · DistractionBench · TOC-Bench · TraceAV-Bench · Audio Hallucination QA · CCTVBench / C-TCD · VisualTextTrap / VTHM-MoE · Structured Rewards · VideoTIR · GameplayQA · FrameRepeat · ClueNet · INFACT · Video-TwG · KPM-Bench / MoPE · VideoTemp-o3 · ViSSRes · VideoHEDGE · Video-DPL · NOAH · EGOILLUSION · MESH · VideoHallu / VideoHallu-GRPO · VistaDPO · MHBench · RoadSocial · HAVEN / Video-thinking (TDPO) · MASH-VLM · OVBench / VideoChat-Online · VidHalluc / DINO-HEAL · VideoHallucer · Vista-LLaMA
Beneath the Scores · STRAND / STRAND (Trajectory Reasoning) · VADER · VidPair-Halluc · TOC-Bench · CCTVBench / C-TCD · Video-ToC · GasVideo-1000 · STEAR · Structured Rewards · GameplayQA · FrameRepeat · ClueNet · INFACT · GraphThinker · OmniVCHall / TriCD · CoE · MixDPO · SEASON · Video-DPL · NOAH · VideoHallu / VideoHallu-GRPO · HAVEN / Video-thinking (TDPO) · VidHalluc / DINO-HEAL · VidHal · EventHallusion / TCD
VADER · MultiToP · SToP · Video-ToC · GasVideo-1000 · VisualTextTrap / VTHM-MoE · DTR · STEAR · STVG-R1 · MACD · OmniVCHall / TriCD · ViSSRes · SmartSight · SEASON · MMA · TAAE · PaMi-VDPO · RoadSocial · OVBench / VideoChat-Online · EventHallusion / TCD · Vista-LLaMA
Mechanism-driven taxonomy of Vid-LLM hallucinations. Solid fill = benchmarks; striped fill = mitigation methods; dashed outline = evaluation analyses. Placement indicates a primary indexing category, not exclusive coverage.
Generated from paper data using the LaTeX tree source.
[!NOTE] Newest first within each subtype. Date = first arXiv submission, or publisher issue date when no arXiv record is used; venue years may differ. Sources and review notes.
= Project Page
= GitHub Repository
= Hugging Face Dataset
= Kaggle Dataset
= Leaderboard
- = No verified resource link
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Probes motion hallucinations caused by prior knowledge, sequential inference, and visual similarity through several question formats. | MoHallBench | arXiv 2026 07/2026 | - |
| MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning Diagnoses fine-grained motion hallucinations involving direction, attribution, and temporal kinematics, separating distinct failures in interpreting physical movement. Related method: PPV | MotionHalluc | arXiv 2026 06/2026 | |
| KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding Evaluates fine-grained limb motion through video captioning and question answering, using kinematic parsing to assess motion descriptions. Related method: MoPE | KPM-Bench | arXiv 2026 02/2026 | - |
| ARGUS: Hallucination and Omission Evaluation in Video-LLMs Measures both fabricated content and missing information in free-form video captions against human-written reference descriptions. | ARGUS | ICCV 2025 06/2025 | |
| MHBench: Demystifying Motion Hallucination in VideoLLMs Tests motion hallucinations using original actions, actions with reversed meanings, and incomplete actions that challenge static visual cues. | MHBench | AAAI 2025 04/2025 | |
| Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation Diagnoses video hallucinations by crossing hallucination causes, object-scene-event aspects, and question formats to expose distinct failures in video understanding. Related method: Video-thinking (TDPO) | HAVEN | arXiv 2025 03/2025 | |
| VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding Evaluates hallucinations about actions, temporal order, and scene transitions through video questions, caption generation, and event sorting. Related method: DINO-HEAL | VidHalluc | CVPR 2025 12/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Online Video Understanding: OVBench and VideoChat-Online Evaluates streaming-video question answering across online perception, memory, and reasoning tasks; its scope extends beyond hallucination-specific evaluation. Related method: VideoChat-Online | OVBench | CVPR 2025 12/2024 | |
| VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models Uses adversarial question pairs to distinguish hallucinations about visible objects and temporal relations from unsupported external information. | VideoHallucer | arXiv 2024 06/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| VidHal: Benchmarking Temporal Hallucinations in Vision LLMs Evaluates temporal hallucinations by asking models to distinguish and rank video captions with different degrees of factual distortion. | VidHal | TMLR 2026 11/2024 | |
| Vript: A Video Is Worth Thousands of Words Provides dense video descriptions and Vript-Hard evaluations targeting hallucinated captions, information retrieval, and temporal ordering of video events. Related method: Vriptor | Vript | NeurIPS 2024 06/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models Evaluates object-state changes, identity persistence, and relational reasoning, requiring answers to satisfy jointly grounded spatiotemporal prerequisites. Related method: STRAND (Trajectory Reasoning) | STRAND | arXiv 2026 08/2026 | |
| TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Tests object identity, state, and continuity with trajectory-grounded questions designed to require temporally ordered visual evidence. | TOC-Bench | arXiv 2026 05/2026 | |
| EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding Tests hallucinations in first-person videos through human-annotated open and closed questions about visual and auditory evidence. | EGOILLUSION | EMNLP 2025 11/2025 | |
| MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Uses hierarchical questions and plausible distractors to expose hallucinations about objects, attributes, and subject-action relations across video segments. | MESH | ACM MM 2025 09/2025 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models Inserts advertising distractors into long videos to test whether models conflate subjects and events from unrelated segments. | DistractionBench | arXiv 2026 05/2026 | |
| ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding Tests hallucinations from semantic aggregation in long videos, where models can incorrectly combine information across separate video segments. Related method: ELV-Halluc-DPO | ELV-Halluc | arXiv 2025 08/2025 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning Evaluates person-specific limb-motion captions across shots, distinguishing omissions from fabricated or incorrect actions through graded physical alignment. Related method: GM-DPO | FlexBench | arXiv 2026 09/2026 | - |
| GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents Uses synchronized multiplayer gameplay views and structured distractors to test agent identity, role attribution, and grounded reasoning. | GameplayQA | ACL 2026 03/2026 | |
| VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding Tests prior-driven hallucinations using synthetic videos that violate physical or logical expectations, challenging models to follow observed evidence. Related method: VideoHallu-GRPO | VideoHallu | NeurIPS 2025 05/2025 | |
| Models See Hallucinations: Evaluating the Factuality in Video Captioning Studies factual errors in video captions and introduces a weakly supervised factuality metric with human-annotated evaluation data. | FactVC | EMNLP 2023 03/2023 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration Tests whether models can verify individual events in dense video captions, using human-checked incorrect descriptions as hard negatives. | VidOmni-Bench | arXiv 2026 09/2026 | - |
| CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs Tests consistent traffic-hazard judgments using real accident videos paired with counterfactual counterparts that alter the evidence for an accident. Related method: C-TCD | CCTVBench | arXiv 2026 04/2026 | - |
| NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models Inserts unrelated clips into videos to measure narrative-prior hallucinations and omissions through captioning and question-answering tasks. | NOAH | arXiv 2025 11/2025 | |
| RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives Provides diverse social-media road videos and question-answer pairs for evaluating road-event understanding across viewpoints and geographic settings. | RoadSocial | CVPR 2025 03/2025 | |
| EventHallusion: Diagnosing Event Hallucinations in Video LLMs Diagnoses event hallucinations driven by language priors and visual biases, testing whether answers reflect events actually present in videos. Related method: TCD | EventHallusion | arXiv 2024 09/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Target-Checked Reliability Score Refinement for Video Question Answering Provides controlled video-question examples of confident but incorrect answers for diagnosing hallucinations and evaluating the reliability of confidence estimates. Related method: TRACE-RC | VHD | arXiv 2026 09/2026 | |
| Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators Compares hallucination detectors on adversarial video questions and captions, with a shared protocol spanning ontology and dynamic errors. | VidHalLoc | arXiv 2026 09/2026 | - |
| No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs Uses adversarial video pairs with similar backgrounds but different foreground events to isolate spatial and temporal hallucinations. | VidPair-Halluc | ECCV 2026 06/2026 | |
| DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding Evaluates procedural-caption factuality at conceptual and grounded argument levels, using either textual references or direct video evidence. | DualFact | ACL 2026 Findings 04/2026 | - |
| Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models Tests whether misleading conversational feedback causes video models to abandon correct judgments and invent unsupported spatiotemporal explanations. | GasVideo-1000 | arXiv 2026 04/2026 | |
| When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Tests whether misleading text overlays cause video models to contradict visual evidence when answering questions about the depicted content. Related method: VTHM-MoE | VisualTextTrap | arXiv 2026 04/2026 | - |
| INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs Separates faithfulness from factuality errors in video answers and probes robustness to degraded visuals, corrupted evidence, and temporal interventions. | INFACT | arXiv 2026 03/2026 | - |
| Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models Benchmarks isolated and compositional hallucinations across spatial and temporal dimensions, probing failures involving multiple interacting types of video evidence. Related method: TriCD | OmniVCHall | arXiv 2026 01/2026 | |
| VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations Estimates answer reliability by clustering responses to clean and perturbed videos and measuring semantic uncertainty across those responses. | VideoHEDGE | arXiv 2026 01/2026 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos? Requires long-video answers to cite precise audio-visual evidence, using evidence-aware scoring to penalize guessing and fabricated support. | Video-HolmesV2 | ECCV 2026 09/2026 | - |
| TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Tests multi-hop reasoning and hallucination robustness using explicit evidence trajectories distributed across long audio-visual recordings. | TraceAV-Bench | arXiv 2026 05/2026 | - |
| Exploring Audio Hallucination in Egocentric Video Understanding Probes imagined foreground and background sounds in egocentric videos through questions targeting visible but inaudible events. | Audio Hallucination QA | ICASSP 2026 04/2026 | - |
| AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models Evaluates audio-visual hallucinations, cross-modal matching, and reasoning through tasks that distinguish auditory evidence from visible video content. Related method: AVHModel-Align-FT | AVHBench | ICLR 2025 10/2024 | |
| The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio Investigates how unimodal priors and misleading correlations between language, vision, and audio cause multimodal hallucinations. | CMM | arXiv 2024 10/2024 | |
| CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models Provides an audio-visual hallucination benchmark with human judgments for evaluating whether generated descriptions remain consistent with multimodal evidence. | AVHalluBench | arXiv 2024 05/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models Evaluates emotion hallucinations in psychological knowledge and multimodal perception through adversarial questions about emotional cues and their interpretation. Related method: PEP-MEK | EmotionHallucer | arXiv 2025 05/2025 |
[!NOTE] Training-free: ✔︎ No additional parameter learning; ✘ Training required, including auxiliary modules with a frozen backbone. Dates use first publication; newest first within each subtype.
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning Injects measured physical-motion evidence into video reasoning to check kinematic claims, without training or modifying the underlying video model. Related benchmark: MotionHalluc | PPV | arXiv 2026 06/2026 | ✔︎ | |
| VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos Jointly trains temporal localization and video answering so an agent can inspect relevant clips and revise inaccurate grounding. | VideoTemp-o3 | ICML 2026 02/2026 | ✘ | |
| CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models Creates counterfactual action videos and jointly optimizes visual and textual preferences to reduce action and temporal-order hallucinations. | MixDPO | EMNLP 2026 01/2026 | ✘ | - |
| SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse Selects among sampled responses using temporal attention collapse and terminates unreliable generations when visual attention vanishes. | SmartSight | AAAI 2026 12/2025 | ✔︎ | - |
| SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding Diagnoses hallucination tendencies token by token and adaptively contrasts temporal and spatial negatives without additional training. | SEASON | arXiv 2025 12/2025 | ✔︎ | - |
| Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation Combines supervised reasoning fine-tuning with thinking-based direct preference optimization, giving fabricated reasoning stronger feedback to improve factual grounding. Related benchmark: HAVEN | Video-thinking (TDPO) | arXiv 2025 03/2025 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding Retrieves visual evidence to verify and arbitrate long-video answers, with separate reinforcement-learning objectives for each reflection stage. | Reflect-R1 | ECCV 2026 06/2026 | ✘ | |
| Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models Rebalances decoder attention toward under-attended frames without training, changing visual encoding, or introducing auxiliary models. | DTR | arXiv 2026 04/2026 | ✔︎ | - |
| VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning Uses reinforcement learning to coordinate retrieval of video clips, images, and regions for efficient, evidence-grounded long-video answers. | VideoTIR | arXiv 2026 03/2026 | ✘ | - |
| When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition Trains a lightweight scoring module to repeat useful frames during reasoning and counter drift away from visual evidence. | FrameRepeat | arXiv 2026 03/2026 | ✘ | - |
| Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding Trains long-video models to interleave reasoning with on-demand temporal grounding through a staged reinforcement-learning curriculum. | Video-TwG | arXiv 2026 02/2026 | ✘ | - |
| Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding Extracts compact, question-relevant visual evidence and uses reinforcement learning to anchor reasoning to the selected temporal evidence. | CoE | ICME 2026 01/2026 | ✘ | - |
| Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering Uses temporal variation to identify and intervene in hallucination-sensitive model activations without further fine-tuning the language model. | TAAE | arXiv 2025 05/2025 | ✘ | - |
| VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding Uses DINOv2 saliency to reweight visual features during inference, emphasizing informative regions without additional training of the video model. Related benchmark: VidHalluc | DINO-HEAL | CVPR 2025 12/2024 | ✔︎ | |
| Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models Decomposes event queries into characteristic actions and uses visual-language models to estimate timestamps for temporally grounded responses. | Temporal Insight | ICPR 2024 01/2024 | ✔︎ | - |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding Uses kinematic parsing of motion descriptions as a reinforcement-learning reward to improve fine-grained video captioning and reduce motion hallucinations. Related benchmark: KPM-Bench | MoPE | arXiv 2026 02/2026 | ✘ | - |
| Vript: A Video Is Worth Thousands of Words Trains a video captioning model on densely annotated videos to generate detailed descriptions of visual content and unfolding events. Related benchmark: Vript | Vriptor | NeurIPS 2024 06/2024 | ✘ | |
| VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding Improves event timestamp localization by adding explicit time information to video tokens and using slot-based visual compression. | VTG-LLM | AAAI 2025 05/2024 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models Combines structured visual trajectories with symbolic aggregation to answer object-centric spatiotemporal questions using a fixed video-language model. Related benchmark: STRAND | STRAND (Trajectory Reasoning) | arXiv 2026 08/2026 | ✔︎ | |
| Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding Separates timestamped perceptual evidence from reasoning and uses factuality-aware process rewards to train hallucination-resistant video models. | Video-DPL | arXiv 2025 11/2025 | ✘ | |
| Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens Changes visual-text attention and sequential frame projection to preserve the influence of video evidence during open-ended answer generation. | Vista-LLaMA | CVPR 2024 12/2023 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding Trains on adversarial preference pairs to reduce long-video hallucinations caused by incorrectly combining semantic information across video segments. Related benchmark: ELV-Halluc | ELV-Halluc-DPO | arXiv 2025 08/2025 | ✘ | |
| Online Video Understanding: OVBench and VideoChat-Online Combines pyramid memory with offline-to-online instruction training to support streaming-video perception, memory, and reasoning over continuously arriving frames. Related benchmark: OVBench | VideoChat-Online | CVPR 2025 12/2024 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning Weights direct preference optimization by the severity of fabricated or incorrect limb actions, promoting physically faithful, person-specific video captions. Related benchmark: FlexBench | GM-DPO | arXiv 2026 09/2026 | ✘ | - |
| ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning Ranks detected objects by prominence and iteratively revises captions to reduce omissions and hallucinations without retraining the captioning model. | ProCap | arXiv 2026 07/2026 | ✔︎ | |
| STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning Replaces coordinate prediction with visually prompted object identities and reinforcement learning for spatially and temporally consistent video grounding. | STVG-R1 | arXiv 2026 02/2026 | ✘ | - |
| Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment Combines hallucination-based negative captions with tracklet-phrase contrastive alignment to improve object and action faithfulness in video descriptions. | SANTA | WACV 2026 12/2025 | ✘ | |
| EventHallusion: Diagnosing Event Hallucinations in Video LLMs Contrasts predictions from original and temporally disrupted video sequences during decoding to reduce event hallucinations without additional model training. Related benchmark: EventHallusion | TCD | arXiv 2024 09/2024 | ✔︎ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Adaptively reweights visual attention and contrasts selectively erased evidence to suppress prior-driven video hallucinations without training. | VADER | arXiv 2026 08/2026 | ✔︎ | - |
| CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs Uses counterfactual traffic-video counterparts during contrastive decoding to improve the consistency of accident-related judgments without additional model training. Related benchmark: CCTVBench | C-TCD | arXiv 2026 04/2026 | ✔︎ | - |
| Video-ToC: Video Tree-of-Cue Reasoning Localizes visual evidence through a tree of cues and trains reasoning with rewards adjusted to the video's reasoning demands. | Video-ToC | arXiv 2026 04/2026 | ✘ | |
| Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning Separately supervises visual clue extraction and answer reasoning, then filters clues for faithful and interpretable video question answering. | ClueNet | arXiv 2026 03/2026 | ✘ | - |
| GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Builds event-based video scene graphs and applies visual-attention rewards during reinforcement fine-tuning to ground temporal reasoning. | GraphThinker | arXiv 2026 02/2026 | ✘ | - |
| MACD: Model-Aware Contrastive Decoding via Counterfactual Data Uses model feedback to construct object-level counterfactual video inputs for evidence-grounded contrastive decoding. | MACD | arXiv 2026 02/2026 | ✔︎ | - |
| Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Trains a lightweight residual aligner on frozen video representations to improve spatiotemporal consistency and alignment with response semantics. | ViSSRes | arXiv 2026 01/2026 | ✘ | - |
| Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency Combines multi-level semantic alignment with progressive training to reduce hallucinations caused by weak discrimination between video and language concepts. | MMA | IJCAI 2025 08/2025 | ✘ | - |
| PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning Learns video preferences online using prompt-aware augmented clips as rejected inputs, reducing incorrect rejections during hallucination mitigation training. | PaMi-VDPO | arXiv 2025 04/2025 | ✘ | - |
| MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations Disentangles spatial and temporal attention and introduces Harmonic RoPE to reduce confusion between depicted actions and surrounding scene context. | MASH-VLM | CVPR 2025 03/2025 | ✘ | - |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| Target-Checked Reliability Score Refinement for Video Question Answering Refines learned reliability scores through target-checked evidence across video samplings, improving answer selection while leaving generated answers unchanged. Related benchmark: VHD | TRACE-RC | arXiv 2026 09/2026 | ✘ | |
| Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study Verifies timestamped video-answer claims by independently re-captioning cited frames and checking textual entailment with a separate model. | GroundedVQA | arXiv 2026 08/2026 | ✔︎ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models Trains a lightweight patcher to selectively replace unreliable visual tokens before generation while leaving the original video model unchanged. | MultiToP | arXiv 2026 06/2026 | ✘ | - |
| Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Identifies attention sink tokens and suppresses them during visual token pruning to preserve fine-grained grounding and hallucination robustness. | SToP | ECCV 2026 04/2026 | ✔︎ | - |
| When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Trains specialized experts with dual encoders and adaptive token routing to disentangle misleading text overlays from underlying visual evidence. Related benchmark: VisualTextTrap | VTHM-MoE | arXiv 2026 04/2026 | ✘ | - |
| STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models Targets risky decoding steps with layer-specific visual evidence, combining grounding restoration with temporal counterfactual checks. | STEAR | arXiv 2026 04/2026 | ✔︎ | - |
| Reinforcing Consistency in Video MLLMs with Structured Rewards Audits captions as factual and temporal claims, then trains with scene-graph, temporal, and video-grounded self-verification rewards. | Structured Rewards | COLM 2026 04/2026 | ✘ | - |
| Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models Learns an adaptive controller for triple-path contrastive decoding, combining video perturbations and saliency enhancement to address compositional hallucinations. Related benchmark: OmniVCHall | TriCD | arXiv 2026 01/2026 | ✘ | |
| VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding Applies group relative policy optimization to help video models recognize physical and logical violations instead of substituting familiar prior expectations. Related benchmark: VideoHallu | VideoHallu-GRPO | NeurIPS 2025 05/2025 | ✘ | |
| VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Optimizes preferences at whole-video, temporal, and object levels using spatially and temporally annotated response pairs. | VistaDPO | ICML 2025 04/2025 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding Uses modality-aware masking and confidence-guided contrastive decoding to suppress hallucinations across audio, video, and language without training. | AVCD | NeurIPS 2025 05/2025 | ✔︎ | |
| AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models Fine-tunes audio-visual models with benchmark-derived training data to improve cross-modal alignment and robustness against auditory and visual hallucinations. Related benchmark: AVHBench | AVHModel-Align-FT | ICLR 2025 10/2024 | ✘ | |
| Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization Improves audio-visual captions through repeated preference optimization and rebirth tuning designed to retain non-captioning abilities. | mrDPO | arXiv 2024 10/2024 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models Combines perception-enhanced prompting with memory-based emotion knowledge to reduce emotion hallucinations without additional training of the multimodal model. Related benchmark: EmotionHallucer | PEP-MEK | arXiv 2025 05/2025 | ✔︎ |
Studies of evaluation validity and hallucination mechanisms without a standalone benchmark or mitigation method. Cross-category studies use their closest primary category. Dates use first publication.
| Paper | Analysis | Venue / Date | Resources |
|---|---|---|---|
| Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models Intervenes separately on video-agent grounding, observation, and reasoning to test whether benchmark scores predict downstream hallucination risk. | Beneath the Scores | NeurIPS 2026 TAE Workshop 09/2026 | - |
If this repository or survey helps your work, please cite:
@article{huang2026distorted,
title={Distorted or Fabricated? A Survey on Hallucination in Video LLMs},
author={Huang, Yiyang and Zhang, Yitian and Wang, Yizhou and Zhang, Mingyuan and Shi, Liang and Zeng, Huimin and Fu, Yun},
journal={arXiv preprint arXiv:2604.12944},
year={2026}
}
[!TIP] Contributions are welcome:
🔀 Pull Request — Add new papers, update resource links, or correct errors
🐛 Open an Issue — Report mistakes, suggest missing papers, or request features
See the curation guide for summary sources, task-tag definitions, publication versus addition dates, and consistency checks. The browser groups contributions by paper; the taxonomy tables retain each benchmark, method, and analysis separately.
Resource gaps tracked in data/papers.json:
Edit the shared records in data/papers.json and data/paper_details.json, then regenerate the task index and paper tables with python3 scripts/generate_readme.py. Do not edit generated table rows directly.
Include a contribution-specific English description, a reviewed primary source, verified dates, and official resource links. Preserve stable entry IDs and existing contributions. See the update checklist for all validation steps.
If this repository helps, please consider giving it a ⭐
Maintained by the SmileLab team at Northeastern University.
[ACL 2026] Curated papers on video LLM hallucination, with benchmarks, mitigation methods, and an interactive browser. Updated monthly.
See the codeA curated paper list on hallucination in Video Large Language Models (Vid-LLMs), covering 42 benchmarks, 52 mitigation methods, and 1 evaluation analysis. The 95 entries represent 78 distinct papers; a paper may contribute both a benchmark and a method. Updated monthly via arXiv search and manual review.
📄 Survey Paper: Distorted or Fabricated? A Survey on Hallucination in Video LLMs
🔎 Interactive Browser: Browse 78 papers with source-backed summaries, task tags, combined filters, and table or card views.

FlexBench / GM-DPO · Beneath the Scores · VidOmni-Bench · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning)
Repository additions, not publication dates. All recent additions
Find Benchmarks · Training-Free Methods · With Code · Latest Papers · Recently Added
Start here: Reading guide, from the survey to benchmarks, mitigation, and evaluation limits.
Scope: Relevance review distinguishes direct hallucination research, broader evaluations, and related-work candidates. All entries remain available pending manual decisions.
Paper-level task tags. Newest first; links lead to individual contributions below.
Video-HolmesV2 · TraceAV-Bench · Audio Hallucination QA · EGOILLUSION · AVCD · EmotionHallucer / PEP-MEK · AVHBench / AVHModel-Align-FT · CMM · mrDPO · AVHalluBench
Beneath the Scores · VidOmni-Bench · VHD / TRACE-RC · VidHalLoc · GroundedVQA · MoHallBench · DualFact · Audio Hallucination QA · INFACT · VideoHEDGE · SmartSight · EGOILLUSION · MESH · ELV-Halluc / ELV-Halluc-DPO · ARGUS · EmotionHallucer / PEP-MEK · HAVEN / Video-thinking (TDPO) · VidHal · AVHBench / AVHModel-Align-FT · CMM · VideoHallucer · Vript / Vriptor · AVHalluBench · FactVC
VidOmni-Bench · Video-HolmesV2 · Reflect-R1 · DistractionBench · TraceAV-Bench · VideoTIR · Video-TwG · VideoTemp-o3 · ELV-Halluc / ELV-Halluc-DPO · Vript / Vriptor
FlexBench / GM-DPO · MoHallBench · MotionHalluc / PPV · KPM-Bench / MoPE · MixDPO · SANTA · MHBench · MASH-VLM · VidHalluc / DINO-HEAL
Video-TwG · GraphThinker · STVG-R1 · VideoTemp-o3 · CoE · VTG-LLM · Temporal Insight
FlexBench / GM-DPO · VidOmni-Bench · VidHalLoc · ProCap · DualFact · Structured Rewards · KPM-Bench / MoPE · SANTA · NOAH · ARGUS · VistaDPO · VidHal · mrDPO · Vript / Vriptor · FactVC
Beneath the Scores · Video-HolmesV2 · VHD / TRACE-RC · VidHalLoc · STRAND / STRAND (Trajectory Reasoning) · GroundedVQA · MoHallBench · VidPair-Halluc · Reflect-R1 · MotionHalluc / PPV · DistractionBench · TOC-Bench · TraceAV-Bench · Audio Hallucination QA · CCTVBench / C-TCD · VisualTextTrap / VTHM-MoE · Structured Rewards · VideoTIR · GameplayQA · FrameRepeat · ClueNet · INFACT · Video-TwG · KPM-Bench / MoPE · VideoTemp-o3 · ViSSRes · VideoHEDGE · Video-DPL · NOAH · EGOILLUSION · MESH · VideoHallu / VideoHallu-GRPO · VistaDPO · MHBench · RoadSocial · HAVEN / Video-thinking (TDPO) · MASH-VLM · OVBench / VideoChat-Online · VidHalluc / DINO-HEAL · VideoHallucer · Vista-LLaMA
Beneath the Scores · STRAND / STRAND (Trajectory Reasoning) · VADER · VidPair-Halluc · TOC-Bench · CCTVBench / C-TCD · Video-ToC · GasVideo-1000 · STEAR · Structured Rewards · GameplayQA · FrameRepeat · ClueNet · INFACT · GraphThinker · OmniVCHall / TriCD · CoE · MixDPO · SEASON · Video-DPL · NOAH · VideoHallu / VideoHallu-GRPO · HAVEN / Video-thinking (TDPO) · VidHalluc / DINO-HEAL · VidHal · EventHallusion / TCD
VADER · MultiToP · SToP · Video-ToC · GasVideo-1000 · VisualTextTrap / VTHM-MoE · DTR · STEAR · STVG-R1 · MACD · OmniVCHall / TriCD · ViSSRes · SmartSight · SEASON · MMA · TAAE · PaMi-VDPO · RoadSocial · OVBench / VideoChat-Online · EventHallusion / TCD · Vista-LLaMA
Mechanism-driven taxonomy of Vid-LLM hallucinations. Solid fill = benchmarks; striped fill = mitigation methods; dashed outline = evaluation analyses. Placement indicates a primary indexing category, not exclusive coverage.
Generated from paper data using the LaTeX tree source.
[!NOTE] Newest first within each subtype. Date = first arXiv submission, or publisher issue date when no arXiv record is used; venue years may differ. Sources and review notes.
= Project Page
= GitHub Repository
= Hugging Face Dataset
= Kaggle Dataset
= Leaderboard
- = No verified resource link
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Probes motion hallucinations caused by prior knowledge, sequential inference, and visual similarity through several question formats. | MoHallBench | arXiv 2026 07/2026 | - |
| MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning Diagnoses fine-grained motion hallucinations involving direction, attribution, and temporal kinematics, separating distinct failures in interpreting physical movement. Related method: PPV | MotionHalluc | arXiv 2026 06/2026 | |
| KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding Evaluates fine-grained limb motion through video captioning and question answering, using kinematic parsing to assess motion descriptions. Related method: MoPE | KPM-Bench | arXiv 2026 02/2026 | - |
| ARGUS: Hallucination and Omission Evaluation in Video-LLMs Measures both fabricated content and missing information in free-form video captions against human-written reference descriptions. | ARGUS | ICCV 2025 06/2025 | |
| MHBench: Demystifying Motion Hallucination in VideoLLMs Tests motion hallucinations using original actions, actions with reversed meanings, and incomplete actions that challenge static visual cues. | MHBench | AAAI 2025 04/2025 | |
| Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation Diagnoses video hallucinations by crossing hallucination causes, object-scene-event aspects, and question formats to expose distinct failures in video understanding. Related method: Video-thinking (TDPO) | HAVEN | arXiv 2025 03/2025 | |
| VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding Evaluates hallucinations about actions, temporal order, and scene transitions through video questions, caption generation, and event sorting. Related method: DINO-HEAL | VidHalluc | CVPR 2025 12/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Online Video Understanding: OVBench and VideoChat-Online Evaluates streaming-video question answering across online perception, memory, and reasoning tasks; its scope extends beyond hallucination-specific evaluation. Related method: VideoChat-Online | OVBench | CVPR 2025 12/2024 | |
| VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models Uses adversarial question pairs to distinguish hallucinations about visible objects and temporal relations from unsupported external information. | VideoHallucer | arXiv 2024 06/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| VidHal: Benchmarking Temporal Hallucinations in Vision LLMs Evaluates temporal hallucinations by asking models to distinguish and rank video captions with different degrees of factual distortion. | VidHal | TMLR 2026 11/2024 | |
| Vript: A Video Is Worth Thousands of Words Provides dense video descriptions and Vript-Hard evaluations targeting hallucinated captions, information retrieval, and temporal ordering of video events. Related method: Vriptor | Vript | NeurIPS 2024 06/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models Evaluates object-state changes, identity persistence, and relational reasoning, requiring answers to satisfy jointly grounded spatiotemporal prerequisites. Related method: STRAND (Trajectory Reasoning) | STRAND | arXiv 2026 08/2026 | |
| TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models Tests object identity, state, and continuity with trajectory-grounded questions designed to require temporally ordered visual evidence. | TOC-Bench | arXiv 2026 05/2026 | |
| EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding Tests hallucinations in first-person videos through human-annotated open and closed questions about visual and auditory evidence. | EGOILLUSION | EMNLP 2025 11/2025 | |
| MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Uses hierarchical questions and plausible distractors to expose hallucinations about objects, attributes, and subject-action relations across video segments. | MESH | ACM MM 2025 09/2025 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models Inserts advertising distractors into long videos to test whether models conflate subjects and events from unrelated segments. | DistractionBench | arXiv 2026 05/2026 | |
| ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding Tests hallucinations from semantic aggregation in long videos, where models can incorrectly combine information across separate video segments. Related method: ELV-Halluc-DPO | ELV-Halluc | arXiv 2025 08/2025 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning Evaluates person-specific limb-motion captions across shots, distinguishing omissions from fabricated or incorrect actions through graded physical alignment. Related method: GM-DPO | FlexBench | arXiv 2026 09/2026 | - |
| GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents Uses synchronized multiplayer gameplay views and structured distractors to test agent identity, role attribution, and grounded reasoning. | GameplayQA | ACL 2026 03/2026 | |
| VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding Tests prior-driven hallucinations using synthetic videos that violate physical or logical expectations, challenging models to follow observed evidence. Related method: VideoHallu-GRPO | VideoHallu | NeurIPS 2025 05/2025 | |
| Models See Hallucinations: Evaluating the Factuality in Video Captioning Studies factual errors in video captions and introduces a weakly supervised factuality metric with human-annotated evaluation data. | FactVC | EMNLP 2023 03/2023 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration Tests whether models can verify individual events in dense video captions, using human-checked incorrect descriptions as hard negatives. | VidOmni-Bench | arXiv 2026 09/2026 | - |
| CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs Tests consistent traffic-hazard judgments using real accident videos paired with counterfactual counterparts that alter the evidence for an accident. Related method: C-TCD | CCTVBench | arXiv 2026 04/2026 | - |
| NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models Inserts unrelated clips into videos to measure narrative-prior hallucinations and omissions through captioning and question-answering tasks. | NOAH | arXiv 2025 11/2025 | |
| RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives Provides diverse social-media road videos and question-answer pairs for evaluating road-event understanding across viewpoints and geographic settings. | RoadSocial | CVPR 2025 03/2025 | |
| EventHallusion: Diagnosing Event Hallucinations in Video LLMs Diagnoses event hallucinations driven by language priors and visual biases, testing whether answers reflect events actually present in videos. Related method: TCD | EventHallusion | arXiv 2024 09/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Target-Checked Reliability Score Refinement for Video Question Answering Provides controlled video-question examples of confident but incorrect answers for diagnosing hallucinations and evaluating the reliability of confidence estimates. Related method: TRACE-RC | VHD | arXiv 2026 09/2026 | |
| Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators Compares hallucination detectors on adversarial video questions and captions, with a shared protocol spanning ontology and dynamic errors. | VidHalLoc | arXiv 2026 09/2026 | - |
| No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs Uses adversarial video pairs with similar backgrounds but different foreground events to isolate spatial and temporal hallucinations. | VidPair-Halluc | ECCV 2026 06/2026 | |
| DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding Evaluates procedural-caption factuality at conceptual and grounded argument levels, using either textual references or direct video evidence. | DualFact | ACL 2026 Findings 04/2026 | - |
| Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models Tests whether misleading conversational feedback causes video models to abandon correct judgments and invent unsupported spatiotemporal explanations. | GasVideo-1000 | arXiv 2026 04/2026 | |
| When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Tests whether misleading text overlays cause video models to contradict visual evidence when answering questions about the depicted content. Related method: VTHM-MoE | VisualTextTrap | arXiv 2026 04/2026 | - |
| INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs Separates faithfulness from factuality errors in video answers and probes robustness to degraded visuals, corrupted evidence, and temporal interventions. | INFACT | arXiv 2026 03/2026 | - |
| Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models Benchmarks isolated and compositional hallucinations across spatial and temporal dimensions, probing failures involving multiple interacting types of video evidence. Related method: TriCD | OmniVCHall | arXiv 2026 01/2026 | |
| VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations Estimates answer reliability by clustering responses to clean and perturbed videos and measuring semantic uncertainty across those responses. | VideoHEDGE | arXiv 2026 01/2026 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos? Requires long-video answers to cite precise audio-visual evidence, using evidence-aware scoring to penalize guessing and fabricated support. | Video-HolmesV2 | ECCV 2026 09/2026 | - |
| TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Tests multi-hop reasoning and hallucination robustness using explicit evidence trajectories distributed across long audio-visual recordings. | TraceAV-Bench | arXiv 2026 05/2026 | - |
| Exploring Audio Hallucination in Egocentric Video Understanding Probes imagined foreground and background sounds in egocentric videos through questions targeting visible but inaudible events. | Audio Hallucination QA | ICASSP 2026 04/2026 | - |
| AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models Evaluates audio-visual hallucinations, cross-modal matching, and reasoning through tasks that distinguish auditory evidence from visible video content. Related method: AVHModel-Align-FT | AVHBench | ICLR 2025 10/2024 | |
| The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio Investigates how unimodal priors and misleading correlations between language, vision, and audio cause multimodal hallucinations. | CMM | arXiv 2024 10/2024 | |
| CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models Provides an audio-visual hallucination benchmark with human judgments for evaluating whether generated descriptions remain consistent with multimodal evidence. | AVHalluBench | arXiv 2024 05/2024 |
| Paper | Benchmark | Venue / Date | Resources |
|---|---|---|---|
| EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models Evaluates emotion hallucinations in psychological knowledge and multimodal perception through adversarial questions about emotional cues and their interpretation. Related method: PEP-MEK | EmotionHallucer | arXiv 2025 05/2025 |
[!NOTE] Training-free: ✔︎ No additional parameter learning; ✘ Training required, including auxiliary modules with a frozen backbone. Dates use first publication; newest first within each subtype.
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning Injects measured physical-motion evidence into video reasoning to check kinematic claims, without training or modifying the underlying video model. Related benchmark: MotionHalluc | PPV | arXiv 2026 06/2026 | ✔︎ | |
| VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos Jointly trains temporal localization and video answering so an agent can inspect relevant clips and revise inaccurate grounding. | VideoTemp-o3 | ICML 2026 02/2026 | ✘ | |
| CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models Creates counterfactual action videos and jointly optimizes visual and textual preferences to reduce action and temporal-order hallucinations. | MixDPO | EMNLP 2026 01/2026 | ✘ | - |
| SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse Selects among sampled responses using temporal attention collapse and terminates unreliable generations when visual attention vanishes. | SmartSight | AAAI 2026 12/2025 | ✔︎ | - |
| SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding Diagnoses hallucination tendencies token by token and adaptively contrasts temporal and spatial negatives without additional training. | SEASON | arXiv 2025 12/2025 | ✔︎ | - |
| Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation Combines supervised reasoning fine-tuning with thinking-based direct preference optimization, giving fabricated reasoning stronger feedback to improve factual grounding. Related benchmark: HAVEN | Video-thinking (TDPO) | arXiv 2025 03/2025 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding Retrieves visual evidence to verify and arbitrate long-video answers, with separate reinforcement-learning objectives for each reflection stage. | Reflect-R1 | ECCV 2026 06/2026 | ✘ | |
| Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models Rebalances decoder attention toward under-attended frames without training, changing visual encoding, or introducing auxiliary models. | DTR | arXiv 2026 04/2026 | ✔︎ | - |
| VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning Uses reinforcement learning to coordinate retrieval of video clips, images, and regions for efficient, evidence-grounded long-video answers. | VideoTIR | arXiv 2026 03/2026 | ✘ | - |
| When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition Trains a lightweight scoring module to repeat useful frames during reasoning and counter drift away from visual evidence. | FrameRepeat | arXiv 2026 03/2026 | ✘ | - |
| Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding Trains long-video models to interleave reasoning with on-demand temporal grounding through a staged reinforcement-learning curriculum. | Video-TwG | arXiv 2026 02/2026 | ✘ | - |
| Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding Extracts compact, question-relevant visual evidence and uses reinforcement learning to anchor reasoning to the selected temporal evidence. | CoE | ICME 2026 01/2026 | ✘ | - |
| Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering Uses temporal variation to identify and intervene in hallucination-sensitive model activations without further fine-tuning the language model. | TAAE | arXiv 2025 05/2025 | ✘ | - |
| VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding Uses DINOv2 saliency to reweight visual features during inference, emphasizing informative regions without additional training of the video model. Related benchmark: VidHalluc | DINO-HEAL | CVPR 2025 12/2024 | ✔︎ | |
| Temporal Insight Enhancement: Mitigating Temporal Hallucination in Multimodal Large Language Models Decomposes event queries into characteristic actions and uses visual-language models to estimate timestamps for temporally grounded responses. | Temporal Insight | ICPR 2024 01/2024 | ✔︎ | - |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding Uses kinematic parsing of motion descriptions as a reinforcement-learning reward to improve fine-grained video captioning and reduce motion hallucinations. Related benchmark: KPM-Bench | MoPE | arXiv 2026 02/2026 | ✘ | - |
| Vript: A Video Is Worth Thousands of Words Trains a video captioning model on densely annotated videos to generate detailed descriptions of visual content and unfolding events. Related benchmark: Vript | Vriptor | NeurIPS 2024 06/2024 | ✘ | |
| VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding Improves event timestamp localization by adding explicit time information to video tokens and using slot-based visual compression. | VTG-LLM | AAAI 2025 05/2024 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models Combines structured visual trajectories with symbolic aggregation to answer object-centric spatiotemporal questions using a fixed video-language model. Related benchmark: STRAND | STRAND (Trajectory Reasoning) | arXiv 2026 08/2026 | ✔︎ | |
| Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding Separates timestamped perceptual evidence from reasoning and uses factuality-aware process rewards to train hallucination-resistant video models. | Video-DPL | arXiv 2025 11/2025 | ✘ | |
| Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens Changes visual-text attention and sequential frame projection to preserve the influence of video evidence during open-ended answer generation. | Vista-LLaMA | CVPR 2024 12/2023 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding Trains on adversarial preference pairs to reduce long-video hallucinations caused by incorrectly combining semantic information across video segments. Related benchmark: ELV-Halluc | ELV-Halluc-DPO | arXiv 2025 08/2025 | ✘ | |
| Online Video Understanding: OVBench and VideoChat-Online Combines pyramid memory with offline-to-online instruction training to support streaming-video perception, memory, and reasoning over continuously arriving frames. Related benchmark: OVBench | VideoChat-Online | CVPR 2025 12/2024 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning Weights direct preference optimization by the severity of fabricated or incorrect limb actions, promoting physically faithful, person-specific video captions. Related benchmark: FlexBench | GM-DPO | arXiv 2026 09/2026 | ✘ | - |
| ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning Ranks detected objects by prominence and iteratively revises captions to reduce omissions and hallucinations without retraining the captioning model. | ProCap | arXiv 2026 07/2026 | ✔︎ | |
| STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning Replaces coordinate prediction with visually prompted object identities and reinforcement learning for spatially and temporally consistent video grounding. | STVG-R1 | arXiv 2026 02/2026 | ✘ | - |
| Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment Combines hallucination-based negative captions with tracklet-phrase contrastive alignment to improve object and action faithfulness in video descriptions. | SANTA | WACV 2026 12/2025 | ✘ | |
| EventHallusion: Diagnosing Event Hallucinations in Video LLMs Contrasts predictions from original and temporally disrupted video sequences during decoding to reduce event hallucinations without additional model training. Related benchmark: EventHallusion | TCD | arXiv 2024 09/2024 | ✔︎ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models Adaptively reweights visual attention and contrasts selectively erased evidence to suppress prior-driven video hallucinations without training. | VADER | arXiv 2026 08/2026 | ✔︎ | - |
| CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs Uses counterfactual traffic-video counterparts during contrastive decoding to improve the consistency of accident-related judgments without additional model training. Related benchmark: CCTVBench | C-TCD | arXiv 2026 04/2026 | ✔︎ | - |
| Video-ToC: Video Tree-of-Cue Reasoning Localizes visual evidence through a tree of cues and trains reasoning with rewards adjusted to the video's reasoning demands. | Video-ToC | arXiv 2026 04/2026 | ✘ | |
| Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning Separately supervises visual clue extraction and answer reasoning, then filters clues for faithful and interpretable video question answering. | ClueNet | arXiv 2026 03/2026 | ✘ | - |
| GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Builds event-based video scene graphs and applies visual-attention rewards during reinforcement fine-tuning to ground temporal reasoning. | GraphThinker | arXiv 2026 02/2026 | ✘ | - |
| MACD: Model-Aware Contrastive Decoding via Counterfactual Data Uses model feedback to construct object-level counterfactual video inputs for evidence-grounded contrastive decoding. | MACD | arXiv 2026 02/2026 | ✔︎ | - |
| Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Trains a lightweight residual aligner on frozen video representations to improve spatiotemporal consistency and alignment with response semantics. | ViSSRes | arXiv 2026 01/2026 | ✘ | - |
| Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency Combines multi-level semantic alignment with progressive training to reduce hallucinations caused by weak discrimination between video and language concepts. | MMA | IJCAI 2025 08/2025 | ✘ | - |
| PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning Learns video preferences online using prompt-aware augmented clips as rejected inputs, reducing incorrect rejections during hallucination mitigation training. | PaMi-VDPO | arXiv 2025 04/2025 | ✘ | - |
| MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations Disentangles spatial and temporal attention and introduces Harmonic RoPE to reduce confusion between depicted actions and surrounding scene context. | MASH-VLM | CVPR 2025 03/2025 | ✘ | - |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| Target-Checked Reliability Score Refinement for Video Question Answering Refines learned reliability scores through target-checked evidence across video samplings, improving answer selection while leaving generated answers unchanged. Related benchmark: VHD | TRACE-RC | arXiv 2026 09/2026 | ✘ | |
| Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study Verifies timestamped video-answer claims by independently re-captioning cited frames and checking textual entailment with a separate model. | GroundedVQA | arXiv 2026 08/2026 | ✔︎ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| MultiToP: Learning to Patch Visual Tokens to Mitigate Hallucinations in Video Large Multimodal Models Trains a lightweight patcher to selectively replace unreliable visual tokens before generation while leaving the original video model unchanged. | MultiToP | arXiv 2026 06/2026 | ✘ | - |
| Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Identifies attention sink tokens and suppresses them during visual token pruning to preserve fine-grained grounding and hallucination robustness. | SToP | ECCV 2026 04/2026 | ✔︎ | - |
| When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models Trains specialized experts with dual encoders and adaptive token routing to disentangle misleading text overlays from underlying visual evidence. Related benchmark: VisualTextTrap | VTHM-MoE | arXiv 2026 04/2026 | ✘ | - |
| STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models Targets risky decoding steps with layer-specific visual evidence, combining grounding restoration with temporal counterfactual checks. | STEAR | arXiv 2026 04/2026 | ✔︎ | - |
| Reinforcing Consistency in Video MLLMs with Structured Rewards Audits captions as factual and temporal claims, then trains with scene-graph, temporal, and video-grounded self-verification rewards. | Structured Rewards | COLM 2026 04/2026 | ✘ | - |
| Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models Learns an adaptive controller for triple-path contrastive decoding, combining video perturbations and saliency enhancement to address compositional hallucinations. Related benchmark: OmniVCHall | TriCD | arXiv 2026 01/2026 | ✘ | |
| VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding Applies group relative policy optimization to help video models recognize physical and logical violations instead of substituting familiar prior expectations. Related benchmark: VideoHallu | VideoHallu-GRPO | NeurIPS 2025 05/2025 | ✘ | |
| VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models Optimizes preferences at whole-video, temporal, and object levels using spatially and temporally annotated response pairs. | VistaDPO | ICML 2025 04/2025 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding Uses modality-aware masking and confidence-guided contrastive decoding to suppress hallucinations across audio, video, and language without training. | AVCD | NeurIPS 2025 05/2025 | ✔︎ | |
| AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models Fine-tunes audio-visual models with benchmark-derived training data to improve cross-modal alignment and robustness against auditory and visual hallucinations. Related benchmark: AVHBench | AVHModel-Align-FT | ICLR 2025 10/2024 | ✘ | |
| Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization Improves audio-visual captions through repeated preference optimization and rebirth tuning designed to retain non-captioning abilities. | mrDPO | arXiv 2024 10/2024 | ✘ |
| Paper | Method | Venue / Date | Training-Free | Resources |
|---|---|---|---|---|
| EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models Combines perception-enhanced prompting with memory-based emotion knowledge to reduce emotion hallucinations without additional training of the multimodal model. Related benchmark: EmotionHallucer | PEP-MEK | arXiv 2025 05/2025 | ✔︎ |
Studies of evaluation validity and hallucination mechanisms without a standalone benchmark or mitigation method. Cross-category studies use their closest primary category. Dates use first publication.
| Paper | Analysis | Venue / Date | Resources |
|---|---|---|---|
| Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models Intervenes separately on video-agent grounding, observation, and reasoning to test whether benchmark scores predict downstream hallucination risk. | Beneath the Scores | NeurIPS 2026 TAE Workshop 09/2026 | - |
If this repository or survey helps your work, please cite:
@article{huang2026distorted,
title={Distorted or Fabricated? A Survey on Hallucination in Video LLMs},
author={Huang, Yiyang and Zhang, Yitian and Wang, Yizhou and Zhang, Mingyuan and Shi, Liang and Zeng, Huimin and Fu, Yun},
journal={arXiv preprint arXiv:2604.12944},
year={2026}
}
[!TIP] Contributions are welcome:
🔀 Pull Request — Add new papers, update resource links, or correct errors
🐛 Open an Issue — Report mistakes, suggest missing papers, or request features
See the curation guide for summary sources, task-tag definitions, publication versus addition dates, and consistency checks. The browser groups contributions by paper; the taxonomy tables retain each benchmark, method, and analysis separately.
Resource gaps tracked in data/papers.json:
Edit the shared records in data/papers.json and data/paper_details.json, then regenerate the task index and paper tables with python3 scripts/generate_readme.py. Do not edit generated table rows directly.
Include a contribution-specific English description, a reviewed primary source, verified dates, and official resource links. Preserve stable entry IDs and existing contributions. See the update checklist for all validation steps.
If this repository helps, please consider giving it a ⭐
Maintained by the SmileLab team at Northeastern University.