| What Do Hallucinations Reveal About Multimodal Reasoning? Diagnosing Visual Grounding Failures via Contrastive Decoding Probes | ![resource:code][res-github-code]  | ![Text][mod-text] | 2026-09 | arXiv |
| Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-09 | arXiv |
| Long-to-Short Video Evidence Reasoning for Grounded Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-09 | arXiv |
| BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-09 | arXiv |
| Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-09 | arXiv |
| Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Self-Reflective Multi-modal Reasoning for Short-Video Fake News Detection | N/A | ![Text][mod-text] | 2026-08 | arXiv |
| Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Video-FLAIR: Not Whether to Reason, But How | N/A | ![Text][mod-text] | 2026-08 | arXiv |
| Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos | ![resource:project][res-project] | ![Text][mod-text] | 2026-08 | arXiv |
| AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-08 | arXiv |
| Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition | N/A | ![Text][mod-text] | 2026-07 | arXiv |
| CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| Evidence-Backed Video Question Answering | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| TimeThink: Reasoning with Time for Video LLMs | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization | N/A | ![Text][mod-text] | 2026-07 | arXiv |
| STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| Linguistic Relative Policy Optimization for Video Anomaly Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-07 | arXiv |
| Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| APT: Atomic Physical Transitions for Causal Video-Language Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| Training LLMs with Reinforcement Learning over Digital Twin Representations for Reasoning-Intensive Surgical VideoQA | N/A | ![Text][mod-text] | 2026-06 | arXiv |
| Temporal-Aware Reasoning Optimization for Video Temporal Grounding | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA | N/A | ![Text][mod-text] | 2026-06 | arXiv |
| See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA | N/A | ![Text][mod-text] | 2026-06 | arXiv |
| CACR:Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation | ![resource:code][res-github-code]  | ![Text][mod-text] | 2026-06 | arXiv |
| VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| Question-Aware Evidence Ledgers for Video Relational Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-06 | arXiv |
| TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] ![State][mod-state] | 2026-06 | arXiv |
| Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Disagreement-Based Cross-Model Routing for Implicit Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Rethinking Video-Language Model from the Language Input Perspective | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Reflective Dialogue between Teacher and Solver Agents for Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis | N/A | ![Text][mod-text] | 2026-05 | CVPR 2026 |
| CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| SteerSeg: Attention Steering for Reasoning Video Segmentation | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Video-Zero: Self-Evolution Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought | N/A | ![Text][mod-text] ![Audio][mod-audio] | 2026-05 | arXiv |
| RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| VISD: Enhancing Video Reasoning via Structured Self-Distillation | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling | N/A | ![Text][mod-text] | 2026-05 | arXiv |
| From Priors to Perception: Grounding Video-LLMs in Physical Reality | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-05 | arXiv |
| Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-05 | CVPR 2026 |
| Co-Evolving Policy Distillation | N/A | ![Text][mod-text] | 2026-04 | arXiv |
| StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning | N/A | ![Text][mod-text] | 2026-04 | arXiv |
| UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward | N/A | ![Text][mod-text] | 2026-04 | arXiv |
| Seeing Fast and Slow: Learning the Flow of Time in Videos | ![resource:project][res-project] | ![Text][mod-text] | 2026-04 | arXiv |
| EasyVideoR1: Easier RL for Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| Find, Fix, Reason: Context Repair for Video Reasoning | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| Lost in Adaptation: Layer-Selective Recovery of Temporal Reasoning in Video-Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding | N/A | ![Text][mod-text] | 2026-04 | arXiv |
| OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering | ![resource:project][res-project] | ![Text][mod-text] ![Audio][mod-audio] | 2026-04 | arXiv |
| Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| STEER: Structured Event Evidence for Video Reasoning via Multi-Objective Reinforcement Learning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | CVPR 2026 |
| STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| Reinforcing Consistency in Video MLLMs with Structured Rewards | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-04 | arXiv |
| SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Incentivizing Temporal-Awareness in Egocentric Video Understanding Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-03 | CVPR 2026 |
| Reinforcing Structured Chain-of-Thought for Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | CVPR 2026 |
| GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models | ![resource:project][res-project] | ![Text][mod-text] | 2026-03 | CVPR 2026 |
| Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| CoVR-R:Reason-Aware Composed Video Retrieval | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Narrative Aligned Long Form Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | CVPR 2026 |
| When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Video-CoE: Reinforcing Video Event Prediction via Chain of Events | N/A | ![Text][mod-text] | 2026-03 | CVPR 2026 |
| VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs | ![resource:project][res-project] ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-03 | CVPR 2026 |
| Beyond Single-Sample: Reliable Multi-Sample Distillation for Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Are Video Reasoning Models Ready to Go Outside? | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| Video-Based Reward Modeling for Computer-Use Agents | N/A | ![Text][mod-text] | 2026-03 | arXiv |
| Geometry-Aware Semantic Reasoning for Training Free Video Anomaly Detection | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-03 | arXiv |
| SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning | N/A | ![Text][mod-text] ![Audio][mod-audio] | 2026-03 | arXiv |
| Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | CVPR 2026 |
| APPO: Attention-guided Perception Policy Optimization for Video Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | CVPR 2026 |
| Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | arXiv |
| STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning | N/A | ![Text][mod-text] | 2026-02 | arXiv |
| Credit Where It is Due: Cross-Modality Connectivity Drives Precise Reinforcement Learning for MLLM Reasoning | N/A | ![Text][mod-text] | 2026-02 | arXiv |
| VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning | ![resource:code][res-github-code]  | ![Text][mod-text] | 2026-02 | arXiv |
| Process-of-Thought Reasoning for Videos | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | arXiv |
| OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention | N/A | ![Text][mod-text] ![Audio][mod-audio] | 2026-02 | arXiv |
| AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization | ![resource:project][res-project] | ![Text][mod-text] ![Audio][mod-audio] | 2026-02 | arXiv |
| GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio | N/A | ![Text][mod-text] | 2026-02 | arXiv |
| Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | arXiv |
| RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | arXiv |
| LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | arXiv |
| SRVAU-R1: Enhancing Video Anomaly Understanding via Reflection-Aware Learning | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-02 | arXiv |
| Structure Over Scale: Learning Visual Reasoning from Pedagogical Video | N/A | ![Text][mod-text] | 2026-01 | arXiv |
| Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning | N/A | ![Text][mod-text] | 2026-01 | arXiv |
| Video-KTR: Reinforcing Video Reasoning via Key Token Attribution | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Audio][mod-audio] | 2026-01 | arXiv |
| Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos | N/A | ![Text][mod-text] | 2026-01 | arXiv |
| Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] | 2026-01 | arXiv |
| Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| VERHallu: Evaluating and Mitigating Event Relation Hallucination in Video Large Language Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-01 | CVPR 2026 |
| CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| V-CORE: Temporally Consistent Video Understanding for Video-LLM | N/A | ![Text][mod-text] ![Video][mod-video] | 2026-01 | arXiv |
| Explicit Abstention Knobs for Predictable Reliability in Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-12 | arXiv |
| VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-12 | ACL 2026 |
| Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-12 | arXiv |
| Factorized Learning for Temporally Grounded Video-Language Models | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-12 | arXiv |
| SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-12 | arXiv |
| Xiaomi MiMo-VL-Miloco Technical Report | ![resource:code][res-github-code]  | ![Text][mod-text] | 2025-12 | arXiv |
| Rethinking Chain-of-Thought Reasoning for Videos | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-12 | arXiv |
| 1+1 > 2 : Detector-Empowered Video Large Language Model for Spatio-Temporal Grounding and Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-12 | arXiv |
| TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-12 | CVPR 2026 |
| OneThinker: All-in-one Reasoning Model for Image and Video | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-12 | CVPR 2026 |
| WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-12 | CVPR 2026 |
| Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-11 | CVPR 2026 |
| Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-11 | arXiv |
| Video-CoM: Interactive Video Reasoning via Chain of Manipulations | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-11 | arXiv |
| VideoSeg-R1: Reasoning Video Object Segmentation via Reinforcement Learning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-11 | arXiv |
| AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-11 | arXiv |
| Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-11 | arXiv |
| Video Spatial Reasoning with Object-Centric 3D Rollout | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-11 | arXiv |
| ViSS-R1: Self-Supervised Reinforcement Video Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-11 | arXiv |
| Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-10 | arXiv |
| Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-10 | arXiv |
| MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-09 | arXiv |
| VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-09 | arXiv |
| Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data | ![resource:code][res-github-code]  weights: Google_Drive | ![Text][mod-text] ![Video][mod-video] | 2025-09 | arXiv |
| Kwai Keye-VL 1.5 Technical Report | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-09 | arXiv |
| Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-08 | ICML 2026 |
| Ovis2.5 Technical Report | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-08 | arXiv |
| Reinforcing Video Reasoning Segmentation to Think Before It Segments | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-08 | CVPR 2026 |
| TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-08 | ECCV 2026 |
| ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-08 | arXiv |
| Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2025-08 | arXiv |
| AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-08 | CVPR 2026 |
| VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-08 | ACM-MM 2025 |
| ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-08 | arXiv |
| ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-07 | arXiv |
| METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark | N/A | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-07 | arXiv |
| CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-07 | arXiv |
| EmbRACE-3K: Embodied Reasoning and Action in Complex Environments | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-07 | arXiv |
| ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2025-07 | ACM-MM 2025 |
| Scaling RL to Long Videos | ![resource:code][res-github-code]  ![resource:data][res-hf-data] ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-07 | NeurIPS 2025 |
| Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-07 | EMNLP 2025 |
| Kwai Keye-VL Technical Report | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-07 | arXiv |
| Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-07 | arXiv |
| DIVE: Deep-search Iterative Video Exploration | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-06 | CVPR 2025 |
| HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-06 | arXiv |
| VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| VideoDeepResearch: Long Video Understanding With Agentic Tool Using | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-06 | NeurIPS 2025 |
| VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| MiMo-VL Technical Report | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-06 | EMNLP 2025 (Findings) |
| EgoVLM: Policy Optimization for Egocentric Video Understanding | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-06 | NeurIPS 2025 |
| ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought | ![resource:project][res-project] | ![Text][mod-text] ![Video][mod-video] | 2025-06 | arXiv |
| Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | CVPR 2026 |
| SiLVR: A Simple Language-based Video Reasoning Framework | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-05 | TMLR 2026 |
| Reinforcing Video Reasoning with Focused Thinking | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| Fostering Video Reasoning via Next-Event Prediction | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-05 | NeurIPS 2025 |
| Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-05 | NeurIPS 2025 |
| VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Guided Iterative Policy Optimization | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-05 | NeurIPS 2025 |
| Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | NeurIPS 2025 |
| UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | NeurIPS 2025 |
| VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| Seed1.5-VL Technical Report | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-05 | arXiv |
| AVA: Towards Agentic Video Analytics with Vision Language Models | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-05 | NSDI 2026 |
| MR. Video: "MapReduce" is the Principle for Long Video Understanding | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| Multimodal Long Video Modeling Based on Temporal Dynamic Context | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-04 | arXiv |
| TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| Spatial-R1: Enhancing MLLMs in Video Spatial Reasoning | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| WikiVideo: Article Generation from Multiple Videos | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| Improved Visual-Spatial Reasoning via R1-Zero-Like Training | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-04 | arXiv |
| Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-03 | ICCV 2025 |
| Video-R1: Reinforcing Video Reasoning in MLLMs | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2025-03 | NeurIPS 2025 |
| VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-03 | ICLR 2026 |
| Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-03 | NeurIPS 2025 |
| ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-03 | NeurIPS 2025 |
| TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-02 | ACL 2025 (Oral) |
| video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] ![Audio][mod-audio] | 2025-02 | arXiv |
| CoS: Chain-of-Shot Prompting for Long Video Understanding | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2025-02 | arXiv |
| Temporal Preference Optimization for Long-Form Video Understanding | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-01 | arXiv |
| InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model | ![resource:code][res-github-code]  ![resource:weights][res-hf-weights] | ![Text][mod-text] ![Video][mod-video] | 2025-01 | ACL 2025 (Findings) |
| MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning | ![resource:code][res-github-code]  ![resource:weights][res-github-weights]  | ![Text][mod-text] ![Video][mod-video] | 2025-01 | IEEE TPAMI |
| Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs | N/A | ![Text][mod-text] ![Video][mod-video] | 2025-01 | arXiv |
| Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling | ![resource:code][res-github-code]  ![resource:weights][res-github-weights]  | ![Text][mod-text] ![Video][mod-video] | 2024-12 | arXiv |
| STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training | N/A | ![Text][mod-text] ![Video][mod-video] | 2024-11 | CVPR 2025 |
| VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection | ![resource:code][res-github-code]  ![resource:data][res-hf-data] | ![Text][mod-text] ![Video][mod-video] | 2024-11 | CVPR 2025 |
| Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning | N/A | ![Text][mod-text] ![Video][mod-video] | 2024-10 | NeurIPS 2024 (Workshop) |
| VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2024-09 | EMNLP 2024 (Findings) |
| MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning | ![resource:code][res-github-code]  ![resource:weights][res-github-weights]  | ![Text][mod-text] ![Video][mod-video] | 2024-09 | NeurIPS 2024 (Spotlight) |
| Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition | ![resource:code][res-github-code]  | ![Text][mod-text] ![Video][mod-video] | 2024-05 | ICML 2024 |