iLearn-Lab/TPAMI26-Awesome-MLLMs-for-Video-Temporal-Grounding

Latest Papers, Codes and Datasets on VTG-LLMs.

101

27 commits

updated Jul 12, 2026

See the code

README

Awesome-MLLMs-for-Video-Temporal-Grounding Awesome

🔥 A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu1, Wei Liu1, Ye Liu4, Meng Liu2,
Liqiang Nie1, Zhouchen Lin3, Chang Wen Chen4,
1Harbin Institute of Technology, Shenzhen, 2Shandong Jianzhu University,
3Peking University, 4The Hong Kong Polytechnic University

VTG Task

Video Temporal Grounding (VTG) focuses on locating and understanding temporal segments in untrimmed videos based on multimodal queries. Core tasks include video moment retrieval, dense video captioning, video highlight detection, and temporally grounded video QA, all requiring fine-grained temporal reasoning.

With the rise of Multimodal Large Language Models (MLLMs), VTG has seen transformative progress. These models bring powerful cross-modal alignment and semantic reasoning abilities, enabling flexible, generalizable solutions across VTG tasks.

This repository aims to serve as a curated reference point for researchers and practitioners interested in advancing the field of video temporal grounding through the lens of large multimodal models.

News

  • [2025.09.23] The survey has been accepted to IEEE TPAMI.

Table of Contents


🧠 Functional Roles of MLLMs in VTG-MLLMs

Facilitator

MLLMs generate structured textual representations from video content to support downstream modules.

TitleModelDateLinkVenue
Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long VideosGrounding-Prompter12/2023-arXiv
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPTVTG-GPT03/2024projectApplied Sciences
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity FeaturesGPTSee03/2024-IEEE SPL
Grounded Question-Answering in Long Egocentric VideosGroundVQA04/2024projectCVPR
Context-Enhanced Video Moment Retrieval with Large Language ModelsLMR05/2024-arXiv
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment RetrievalTEA06/2024-arXiv
Training-free Video Temporal Grounding using Large-scale Pre-trained ModelsTFVTG08/2024projectECCV
Infusing Environmental Captions for Long-Form Video Language GroundingEI-VLG08/2024-arXiv
Question-Answering Dense Video EventsDeVi09/2024projectSIGIR
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language ModelsChatVTG10/2024-CVPR
TimeCraft: Navigate Weakly-Supervised Temporal Grounded Video Question Answering via Bi-directional ReasoningTimeCraft10/2024-ECCV
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video UnderstandingVERIFIED10/2024projectarXiv
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment RetrievalVideoLights12/2024projectarXiv
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the WildReCorrect12/2024projectarXiv
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language ModelsMoment-GPT01/2025-AAAI

Executor

MLLMs directly perform temporal boundary prediction via integrated multimodal reasoning.

TitleModelDateLinkVenue
Self-Chained Image-Language Model for Video Localization and Question AnsweringSeViLA05/2023projectNeurIPS
LLaViLo: Boosting Video Moment Retrieval via Adapter-Based Multimodal ModelingLLaViLo10/2023-ICCV
VTimeLLM: Empower LLM to Grasp Video MomentsVTimeLLM11/2023projectCVPR
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingTimeChat12/2023projectCVPR
GroundingGPT:Language Enhanced Multi-modal Grounding ModelGroundingGPT03/2024projectACL
HawkEye: Training Video-Text LLMs for Grounding Text in VideosHawkEye03/2024projectarXiv
LITA: Language Instructed Temporal-Localization AssistantLITA03/2024projectECCV
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingVTG-LLM05/2024projectAAAI
The Surprising Effectiveness of Multimodal Large Language Models for Video Moment RetrievalMr.BLIP06/2024projectarXiv
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningMomentor06/2024projectICML
Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosGeLM08/2024projectAAAI
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLMSlowFocus09/2024projectNeurIPS
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingE.T.Chat09/2024projectNeurIPS
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language ModelsGrounded-VideoLLM10/2024projectarXiv
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding BridgeTGB10/2024projectEMNLP
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningTimeSuite10/2024projectICLR
TRACE: Temporal Grounding Video LLM via Causal Event ModelingTRACE11/2024projectICLR
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosReVisionLLM11/2024projectCVPR
Number it: Temporal Grounding Videos like Flipping MangaNumPro11/2024projectCVPR
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment RetrievalLLaVA-MR11/2024-arXiv
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization AbilityTimeMarker11/2024projectarXiv
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal GroundingSeq2Time11/2024-arXiv
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task AlignmentVideoChat-TPO12/2024projectarXiv
MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video GroundingMLLM-TA12/2024-IEEE SPL
Video LLMs for Temporal Reasoning in Long VideosTemporalVLM12/2024projectarXiv
TimeRefine: Temporal Grounding with Time Refining Video LLMTimeRefine12/2024projectarXiv
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal UnderstandingLLaVA-ST01/2025projectarXiv
Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLMTPE-VLLM01/2025projectCOLING
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal LocalizationMeCo03/2025projectarXiv
VideoMind: A Chain-of-LoRA Agent for Long Video ReasoningVideoMind04/2025projectarXiv
VideoExpert: Augmented LLM for Temporal-Sensitive Video UnderstandingVideoExpert04/2025-arXiv
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-TuningVideoChat-R104/2025projectarXiv
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilitySpaceVLLM04/2025projectarXiv
Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingTime-R105/2025projectarXiv
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingMUSEG05/2025projectarXiv
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsGameplayQA03/2026projectACL

🛠️ Training Paradigms of VTG-MLLMs

Pretraining

Pretraining in VTG-MLLMs aims to establish strong temporal reasoning capabilities through large-scale supervised learning.

TitleModelDateLinkVenue
Self-Chained Image-Language Model for Video Localization and Question AnsweringSeViLA05/2023projectNeurIPS
VTimeLLM: Empower LLM to Grasp Video MomentsVTimeLLM11/2023projectCVPR
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingTimeChat12/2023projectCVPR
GroundingGPT:Language Enhanced Multi-modal Grounding ModelGroundingGPT03/2024projectACL
HawkEye: Training Video-Text LLMs for Grounding Text in VideosHawkEye03/2024projectarXiv
LITA: Language Instructed Temporal-Localization AssistantLITA03/2024projectECCV
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingVTG-LLM05/2024projectAAAI
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningMomentor06/2024projectICML
Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosGeLM08/2024projectAAAI
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLMSlowFocus09/2024projectNeurIPS
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingE.T.Chat09/2024projectNeurIPS
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language ModelsGrounded-VideoLLM10/2024projectarXiv
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningTimeSuite10/2024projectICLR
TRACE: Temporal Grounding Video LLM via Causal Event ModelingTRACE11/2024projectICLR
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosReVisionLLM11/2024projectCVPR
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization AbilityTimeMarker11/2024projectarXiv
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task AlignmentVideoChat-TPO12/2024projectarXiv
MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video GroundingMLLM-TA12/2024-IEEE SPL
Video LLMs for Temporal Reasoning in Long VideosTemporalVLM12/2024projectarXiv
TimeRefine: Temporal Grounding with Time Refining Video LLMTimeRefine12/2024projectarXiv
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal UnderstandingLLaVA-ST01/2025projectarXiv
Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLMTPE-VLLM01/2025projectCOLING
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal LocalizationMeCo03/2025projectarXiv
VideoMind: A Chain-of-LoRA Agent for Long Video ReasoningVideoMind04/2025projectarXiv
VideoExpert: Augmented LLM for Temporal-Sensitive Video UnderstandingVideoExpert04/2025-arXiv
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-TuningVideoChat-R104/2025projectarXiv
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilitySpaceVLLM04/2025projectarXiv
Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingTime-R105/2025projectarXiv
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingMUSEG05/2025projectarXiv

Fine-Tuning

Adapts general-purpose MLLMs to downstream VTG tasks through supervised fine-tuning on temporally annotated training datasets.

Training-Free

Training-free approaches integrate pre-trained foundation models with specialized expert tools through the carefully designed pipeline architecture.


🎞️ Video Feature Processing in VTG-MLLMs

Visual Feature

Efficient visual feature handling is essential for capturing more fine-grained temporal cues without overwhelming the model.

Compression

Directly compress visual features from densely sampled frames within budget constraints.

Refinement

Gradually refine predictions to maintain performance with the input token limitations.

Temporal Feature

Precise temporal feature representation and modeling is crucial for aligning visual content with fine-grained timestamp intervals, enabling accurate temporal reasoning in VTG tasks.

Explicit

Explicit modeling strategies directly furnish MLLMs with unambiguous temporal information, offering direct control and interpretability over temporal cues.

Implicit

Implicit modeling strategies are divided into two main types: Feature Infusion, which subtly integrates temporal context during feature extraction, and Intrinsic Reasoning, which leverages the inherent sequential processing of LLMs. By default, methods under this category are considered part of Intrinsic Reasoning, unless they explicitly incorporate external temporal features during encoding—in which case they are classified as Feature Infusion. The table below highlights representative works employing the Feature Infusion approach.


Contact

If you find our survey is useful in your research, please consider giving us a star 🌟 and cite the following paper:

@article{wu2025surveyvideotemporalgrounding,
  title={A Survey on Video Temporal Grounding with Multimodal Large Language Model},
  author={Wu Jianlong and Liu Wei and Liu Ye and Liu Meng and Nie Liqiang and Lin Zhouchen and Chen Chang Wen},
  journal = {arXiv preprint arXiv:2508.10922},
  year={2025}
}

If you have any question about this project, do not hesitate to contact me liuwei030224@gmail.com.

iLearn-Lab/TPAMI26-Awesome-MLLMs-for-Video-Temporal-Grounding

Latest Papers, Codes and Datasets on VTG-LLMs.

101

27 commits

updated Jul 12, 2026

See the code

README

Awesome-MLLMs-for-Video-Temporal-Grounding Awesome

🔥 A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu1, Wei Liu1, Ye Liu4, Meng Liu2,
Liqiang Nie1, Zhouchen Lin3, Chang Wen Chen4,
1Harbin Institute of Technology, Shenzhen, 2Shandong Jianzhu University,
3Peking University, 4The Hong Kong Polytechnic University

VTG Task

Video Temporal Grounding (VTG) focuses on locating and understanding temporal segments in untrimmed videos based on multimodal queries. Core tasks include video moment retrieval, dense video captioning, video highlight detection, and temporally grounded video QA, all requiring fine-grained temporal reasoning.

With the rise of Multimodal Large Language Models (MLLMs), VTG has seen transformative progress. These models bring powerful cross-modal alignment and semantic reasoning abilities, enabling flexible, generalizable solutions across VTG tasks.

This repository aims to serve as a curated reference point for researchers and practitioners interested in advancing the field of video temporal grounding through the lens of large multimodal models.

News

  • [2025.09.23] The survey has been accepted to IEEE TPAMI.

Table of Contents


🧠 Functional Roles of MLLMs in VTG-MLLMs

Facilitator

MLLMs generate structured textual representations from video content to support downstream modules.

TitleModelDateLinkVenue
Grounding-Prompter: Prompting LLM with Multimodal Information for Temporal Sentence Grounding in Long VideosGrounding-Prompter12/2023-arXiv
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPTVTG-GPT03/2024projectApplied Sciences
GPTSee: Enhancing Moment Retrieval and Highlight Detection via Description-Based Similarity FeaturesGPTSee03/2024-IEEE SPL
Grounded Question-Answering in Long Egocentric VideosGroundVQA04/2024projectCVPR
Context-Enhanced Video Moment Retrieval with Large Language ModelsLMR05/2024-arXiv
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment RetrievalTEA06/2024-arXiv
Training-free Video Temporal Grounding using Large-scale Pre-trained ModelsTFVTG08/2024projectECCV
Infusing Environmental Captions for Long-Form Video Language GroundingEI-VLG08/2024-arXiv
Question-Answering Dense Video EventsDeVi09/2024projectSIGIR
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language ModelsChatVTG10/2024-CVPR
TimeCraft: Navigate Weakly-Supervised Temporal Grounded Video Question Answering via Bi-directional ReasoningTimeCraft10/2024-ECCV
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video UnderstandingVERIFIED10/2024projectarXiv
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment RetrievalVideoLights12/2024projectarXiv
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the WildReCorrect12/2024projectarXiv
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language ModelsMoment-GPT01/2025-AAAI

Executor

MLLMs directly perform temporal boundary prediction via integrated multimodal reasoning.

TitleModelDateLinkVenue
Self-Chained Image-Language Model for Video Localization and Question AnsweringSeViLA05/2023projectNeurIPS
LLaViLo: Boosting Video Moment Retrieval via Adapter-Based Multimodal ModelingLLaViLo10/2023-ICCV
VTimeLLM: Empower LLM to Grasp Video MomentsVTimeLLM11/2023projectCVPR
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingTimeChat12/2023projectCVPR
GroundingGPT:Language Enhanced Multi-modal Grounding ModelGroundingGPT03/2024projectACL
HawkEye: Training Video-Text LLMs for Grounding Text in VideosHawkEye03/2024projectarXiv
LITA: Language Instructed Temporal-Localization AssistantLITA03/2024projectECCV
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingVTG-LLM05/2024projectAAAI
The Surprising Effectiveness of Multimodal Large Language Models for Video Moment RetrievalMr.BLIP06/2024projectarXiv
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningMomentor06/2024projectICML
Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosGeLM08/2024projectAAAI
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLMSlowFocus09/2024projectNeurIPS
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingE.T.Chat09/2024projectNeurIPS
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language ModelsGrounded-VideoLLM10/2024projectarXiv
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding BridgeTGB10/2024projectEMNLP
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningTimeSuite10/2024projectICLR
TRACE: Temporal Grounding Video LLM via Causal Event ModelingTRACE11/2024projectICLR
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosReVisionLLM11/2024projectCVPR
Number it: Temporal Grounding Videos like Flipping MangaNumPro11/2024projectCVPR
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment RetrievalLLaVA-MR11/2024-arXiv
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization AbilityTimeMarker11/2024projectarXiv
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal GroundingSeq2Time11/2024-arXiv
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task AlignmentVideoChat-TPO12/2024projectarXiv
MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video GroundingMLLM-TA12/2024-IEEE SPL
Video LLMs for Temporal Reasoning in Long VideosTemporalVLM12/2024projectarXiv
TimeRefine: Temporal Grounding with Time Refining Video LLMTimeRefine12/2024projectarXiv
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal UnderstandingLLaVA-ST01/2025projectarXiv
Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLMTPE-VLLM01/2025projectCOLING
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal LocalizationMeCo03/2025projectarXiv
VideoMind: A Chain-of-LoRA Agent for Long Video ReasoningVideoMind04/2025projectarXiv
VideoExpert: Augmented LLM for Temporal-Sensitive Video UnderstandingVideoExpert04/2025-arXiv
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-TuningVideoChat-R104/2025projectarXiv
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilitySpaceVLLM04/2025projectarXiv
Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingTime-R105/2025projectarXiv
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingMUSEG05/2025projectarXiv
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsGameplayQA03/2026projectACL

🛠️ Training Paradigms of VTG-MLLMs

Pretraining

Pretraining in VTG-MLLMs aims to establish strong temporal reasoning capabilities through large-scale supervised learning.

TitleModelDateLinkVenue
Self-Chained Image-Language Model for Video Localization and Question AnsweringSeViLA05/2023projectNeurIPS
VTimeLLM: Empower LLM to Grasp Video MomentsVTimeLLM11/2023projectCVPR
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingTimeChat12/2023projectCVPR
GroundingGPT:Language Enhanced Multi-modal Grounding ModelGroundingGPT03/2024projectACL
HawkEye: Training Video-Text LLMs for Grounding Text in VideosHawkEye03/2024projectarXiv
LITA: Language Instructed Temporal-Localization AssistantLITA03/2024projectECCV
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingVTG-LLM05/2024projectAAAI
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningMomentor06/2024projectICML
Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosGeLM08/2024projectAAAI
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLMSlowFocus09/2024projectNeurIPS
E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingE.T.Chat09/2024projectNeurIPS
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language ModelsGrounded-VideoLLM10/2024projectarXiv
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded TuningTimeSuite10/2024projectICLR
TRACE: Temporal Grounding Video LLM via Causal Event ModelingTRACE11/2024projectICLR
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosReVisionLLM11/2024projectCVPR
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization AbilityTimeMarker11/2024projectarXiv
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task AlignmentVideoChat-TPO12/2024projectarXiv
MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video GroundingMLLM-TA12/2024-IEEE SPL
Video LLMs for Temporal Reasoning in Long VideosTemporalVLM12/2024projectarXiv
TimeRefine: Temporal Grounding with Time Refining Video LLMTimeRefine12/2024projectarXiv
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal UnderstandingLLaVA-ST01/2025projectarXiv
Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLMTPE-VLLM01/2025projectCOLING
Measure Twice, Cut Once: Grasping Video Structures and Event Semantics with LLMs for Video Temporal LocalizationMeCo03/2025projectarXiv
VideoMind: A Chain-of-LoRA Agent for Long Video ReasoningVideoMind04/2025projectarXiv
VideoExpert: Augmented LLM for Temporal-Sensitive Video UnderstandingVideoExpert04/2025-arXiv
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-TuningVideoChat-R104/2025projectarXiv
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilitySpaceVLLM04/2025projectarXiv
Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingTime-R105/2025projectarXiv
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingMUSEG05/2025projectarXiv

Fine-Tuning

Adapts general-purpose MLLMs to downstream VTG tasks through supervised fine-tuning on temporally annotated training datasets.

Training-Free

Training-free approaches integrate pre-trained foundation models with specialized expert tools through the carefully designed pipeline architecture.


🎞️ Video Feature Processing in VTG-MLLMs

Visual Feature

Efficient visual feature handling is essential for capturing more fine-grained temporal cues without overwhelming the model.

Compression

Directly compress visual features from densely sampled frames within budget constraints.

Refinement

Gradually refine predictions to maintain performance with the input token limitations.

Temporal Feature

Precise temporal feature representation and modeling is crucial for aligning visual content with fine-grained timestamp intervals, enabling accurate temporal reasoning in VTG tasks.

Explicit

Explicit modeling strategies directly furnish MLLMs with unambiguous temporal information, offering direct control and interpretability over temporal cues.

Implicit

Implicit modeling strategies are divided into two main types: Feature Infusion, which subtly integrates temporal context during feature extraction, and Intrinsic Reasoning, which leverages the inherent sequential processing of LLMs. By default, methods under this category are considered part of Intrinsic Reasoning, unless they explicitly incorporate external temporal features during encoding—in which case they are classified as Feature Infusion. The table below highlights representative works employing the Feature Infusion approach.


Contact

If you find our survey is useful in your research, please consider giving us a star 🌟 and cite the following paper:

@article{wu2025surveyvideotemporalgrounding,
  title={A Survey on Video Temporal Grounding with Multimodal Large Language Model},
  author={Wu Jianlong and Liu Wei and Liu Ye and Liu Meng and Nie Liqiang and Lin Zhouchen and Chen Chang Wen},
  journal = {arXiv preprint arXiv:2508.10922},
  year={2025}
}

If you have any question about this project, do not hesitate to contact me liuwei030224@gmail.com.