yunlong10/Awesome-LLMs-for-Video-Understanding

πŸ”₯πŸ”₯πŸ”₯ [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.

3,289

212 commits

updated Sep 24, 2026

See the code

README

Awesome-LLMs-for-Video-Understanding Awesome

πŸ”₯πŸ”₯πŸ”₯ Video Understanding with Large Language Models: A Survey

Yolo Y. Tang1, Jing Bi1, Siting Xu2, Luchuan Song1, Susan Liang1 , Teng Wang2,3 , Daoan Zhang1 , Jie An1 , Jingyang Lin1 , Rongyi Zhu1 , Ali Vosoughi1 , Chao Huang1 , Zeliang Zhang1 , Pinxin Liu1 , Mingqian Feng1 , Feng Zheng2 , Jianguo Zhang2 , Ping Luo3 , Jiebo Luo1, Chenliang Xu1.

1University of Rochester, 2Southern University of Science and Technology, 3The University of Hong Kong

Paper | arXiv | Project Page

image

πŸ“’ News

[10/06/2025]

πŸ”₯ Our follow-up workβ€”Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Modelsβ€”is now available on arXiv and Hugging Face Papers!

[05/04/2025]

🌟 Our Vid-LLM survey has been accepted to the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)! πŸ‘‰ IEEE Xplore | GitHub

[07/23/2024]

πŸ“’ We've recently updated our survey: β€œVideo Understanding with Large Language Models: A Survey”!

✨ This comprehensive survey covers video understanding techniques powered by large language models (Vid-LLMs), training strategies, relevant tasks, datasets, benchmarks, and evaluation methods, and discusses the applications of Vid-LLMs across various domains.

πŸš€ What's New in This Update:
βœ… Updated to include around 100 additional Vid-LLMs and 15 new benchmarks as of June 2024.
βœ… Introduced a novel taxonomy for Vid-LLMs based on video representation and LLM functionality.
βœ… Added a Preliminary chapter, reclassifying video understanding tasks from the perspectives of granularity and language involvement, and enhanced the LLM Background section.
βœ… Added a new Training Strategies chapter, removing adapters as a factor for model classification.
βœ… All figures and tables have been redesigned.

Multiple minor updates will follow this major update. And the GitHub repository will be gradually updated soon. We welcome your reading and feedback ❀️

Table of Contents

Why we need Vid-LLMs?

image

😎 Vid-LLMs: Models

image

πŸ“‘ Citation

If you find our survey useful for your research, please cite the following paper:

@article{vidllmsurvey,
  author={Tang, Yunlong and Bi, Jing and Xu, Siting and Song, Luchuan and Liang, Susan and Wang, Teng and Zhang, Daoan and An, Jie and Lin, Jingyang and Zhu, Rongyi and Vosoughi, Ali and Huang, Chao and Zhang, Zeliang and Liu, Pinxin and Feng, Mingqian and Zheng, Feng and Zhang, Jianguo and Luo, Ping and Luo, Jiebo and Xu, Chenliang},
  journal={IEEE Transactions on Circuits and Systems for Video Technology}, 
  title={Video Understanding with Large Language Models: A Survey}, 
  year={2025},
  doi={10.1109/TCSVT.2025.3566695}
}

πŸ—’οΈ Taxonomy 1

πŸ•ΉοΈ Video Analyzer Γ— LLM

LLM as Summarizer
LLM as Manager
TitleModelDateCodeVenue
DrVideo: Document Retrieval Based Long Video UnderstandingDrVideo06/2024codearXiv
OmAgent a multi-modal agent framework for complex video understanding with task divide-and-conquerOmAgent06/2024codearXiv
Too Many Frames, not all Useful: Efficient Strategies for Long-Form Video QALVNet06/2024codearXiv
VideoTree adaptive tree-based video representation for LLM reasoning on long videosVideoTree05/2024codearXiv
Harnessing Large Language Models for Training-free Video Anomaly DetectionLAVAD04/2024codeCVPR
TraveLER a multi-LMM agent framework for video question-answeringTraveLER04/2024codearXiv
GPTSee enhancing moment retrieval and highlight detection via description-based similarity featuresGPTSee03/2024codearXiv
Reframe anything LLM agent for open world video reframingRAVA03/2024codearXiv
SCHEMA state CHangEs MAtter for procedure planning in instructional videosSCHEMA03/2024codeICLR
TV-TREES multimodal entailment trees for neuro-symbolic video reasoningTV-TREES02/2024codearXiv
VideoAgent: A Memory-augmented Multimodal Agent for Video UnderstandingVideoAgent03/2024project pagearXiv
VideoAgent long-form video understanding with large language model as agentVideoAgent03/2024codearXiv
VURF a general-purpose reasoning and self-refinement framework for video understandingVURF03/2024codearXiv
Why not use your textbook knowledge-enhanced procedure planning of instructional videosKEPP03/2024codeCVPR
DoraemonGPT toward understanding dynamic scenes with large language modelsDoraemonGPT01/2024codearXiv
LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric VideosLifelongMemory12/2023codearXiv
Zero-Shot Video Question Answering with Procedural ProgramsProViQ12/2023codearXiv
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and LearnAssistGPT06/2023codearXiv
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding SystemChatVideo04/2023project pagearXiv
Video ChatCaptioner: Towards Enriched Spatiotemporal DescriptionsStarVideo ChatCaptioner04/2023codearXiv
ViperGPT: Visual Inference via Python Execution for ReasoningViperGPT03/2023codearXiv
Hawk: Learning to Understand Open-World Video AnomaliesHawk05/2024codearXiv

πŸ‘Ύ Video Embedder Γ— LLM

LLM as Text Decoder
TitleModelDateCodeVenue
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live StreamingTLive-Omni08/2026codearXiv
AuroraCap: Efficient, Performant Video Detailed Captioning and a New BenchmarkAuroraCap10/2024project pagearXiv
Artemis towards referential understanding in complex videosArtemis06/2024codearXiv
EmoLLM multimodal emotional understanding meets large language modelsEmoLLM06/2024codearXiv
Fewer tokens and fewer videos extending video understanding abilities in large vision-language modelsFTFV-LLM06/2024-arXiv
Flash-VStream: Memory-Based Real-Time Understanding for Long Video StreamsFlash-VStream06/2024codearXiv
LLAVIDAL benchmarking large language vision models for daily activities of livingLLAVIDAL06/2024codearXiv
Long context transfer from language to visionLongVA06/2024codearXiv
ShareGPT4Video improving video understanding and generation with better captionsShareGPT4Video06/2024codearXiv
Towards event-oriented long video understandingVIM06/2024codearXiv
Video-SALMONN speech-enhanced audio-visual large language modelsVideo-SALMONN06/2024codeICML
VideoGPT+ integrating image and video encoders for enhanced video understandingVideoGPT+06/2024codearXiv
VideoLLaMA 2 advancing spatial-temporal modeling and audio understanding in video-LLMsVideoLLaMA 206/2024codearXiv
MotionLLM: Understanding Human Behaviors from Human Motions and VideosMotionLLM05/2024project pagearXiv
MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkVideoChat211/2023codeCVPR
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and SummarizationShotluck Holmes05/2024-arXiv
Streaming long video understanding with large language modelsVideoStreaming05/2024-arXiv
Synchronized Video Storytelling: Generating Video Narrations with Structured StorylineVideoNarrator05/2024-arXiv
TOPA extend large language models for video understanding via text-only pre-alignmentTOPA05/2024codeNeurIPS
MovieChat+: Question-aware Sparse Memory for Long Video Question AnsweringMovieChat+04/2024codearXiv
AutoAD III: The Prequel – Back to the PixelsAutoAD III04/2024project pageCVPR
Direct Preference Optimization of Video Large Multimodal Models from Language Model RewardLLaVA-Hound-DPO04/2024codearXiv
From image to video, what do we need in multimodal LLMsRED-VILLM04/2024-arXiv
Koala key frame-conditioned long video-LLMKoala04/2024project pageCVPR
LongVLM efficient long video understanding via large language modelsLongVLM04/2024codeECCV
MA-LMM memory-augmented large multimodal model for long-term video understandingMA-LMM04/2024codeCVPR
MiniGPT4-video advancing multimodal LLMs for video understanding with interleaved visual-textual tokensMiniGPT4-Video04/2024codearXiv
Pegasus-v1 technical reportPegasus-v104/2024codearXiv
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense CaptioningPLLaVA04/2024codearXiv
ST-LLM: Large Language Models Are Effective Temporal LearnersST-LLM04/2024codearXiv
Tarsier recipes for training and evaluating large video description modelsTarsier07/2024codearXiv
X-VARS introducing explainability in football refereeing with multi-modal large language modelX-VARS04/2024codearXiv
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual ScenariosCAT03/2024codearXiv
InternVideo2 scaling video foundation models for multimodal video understandingInternVideo203/2024codeECCV
MovieLLM enhancing long video understanding with AI-generated moviesMovieLLM03/2024codearXiv
LLMs meet long video advancing long video comprehension with an interactive visual adapter in LLMsIVAwithLLM02/2024codearXiv
LSTP language-guided spatial-temporal prompt learning for long-form video-text understandingLSTP02/2024codeEMNLP
LVCHAT facilitating long video comprehensionLVCHAT02/2024codearXiv
OSCaR: Object State Captioning and State Change RepresentationOSCaR02/2024codeNAACL
Slot-VLM SlowFast slots for video-language modelingSlot-VLM02/2024codearXiv
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-TrainingCOSMO01/2024codearXiv
Weakly supervised gaussian contrastive grounding with large multimodal models for video question answeringGCG01/2024codeACMMM
Audio-Visual LLM for Video UnderstandingAV-LLM12/2023codearXiv
Generative Multimodal Models are In-Context LearnersEmu212/2023project pageCVPR
MMICT: Boosting Multi-Modal Fine-Tuning with In-Context ExamplesMMICT12/2023codeTOMM
VaQuitA : Enhancing Alignment in LLM-Assisted Video UnderstandingVaQuitA12/2023codearXiv
VILA: On Pre-training for Visual Language ModelsVILA12/2023codeCVPR
Vista-LLaMA reliable video narrator via equal distance to visual tokensVista-LLaMA12/2023project pagearXiv
Chat-UniVi unified visual representation empowers large language models with image and video understandingChat-UniVi11/2023codeCVPR
LLaMA-VID: An Image is Worth 2 Tokens in Large Language ModelsLLaMA-VID11/2023codearXiv
Video-LLaVA learning united visual representation by alignment before projectionVideo-LLaVA11/2023codearXiv
Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringLLaMA-VQA10/2023codeEMNLP
MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingMovieChat07/2023codeCVPR
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary CaptioningLLMVA-GEBC06/2023codeCVPR
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text IntegrationMacaw-LLM06/2023project pagearXiv
Valley: Video Assistant with Large Language model Enhanced abilitYVALLEY06/2023codearXiv
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsVideo-ChatGPT06/2023codeACL
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video UnderstandingVideo-LLaMA06/2023codeEMNLP
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and BenchmarksmPLUG-video06/2023codearXiv
ChatBridge: Bridging Modalities with Large Language Model as a Language CatalystChatBridge05/2023codearXiv
Otter: A Multi-Modal Model with In-Context Instruction TuningOtter05/2023codearXiv
VideoLLM: Modeling Video Sequence with Large Language ModelsVideoLLM05/2023codearXiv
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory-05/2025codeICCV 2025
LLM as Regressor
LLM as Hidden Layer

🧭 (Analyzer + Embedder) Γ— LLM

LLM as Manager
TitleModelDateCodeVenue
MM-VID: Advancing Video Understanding with GPT-4V(ision)MM-VID10/2023-arXiv
LLM as Summarizer
LLM as Regressor
LLM as Text Decoder
LLM as Hidden Layer
TitleModelDateCodeVenue
PG-Video-LLaVA: Pixel Grounding Large Video-Language ModelsStarPG-Video-LLaVA11/2023codearXiv

πŸ—’οΈ Taxonomy 2

πŸ€– LLM-based Video Agents

πŸŽ₯ Vid-LLM Pretraining

πŸ‘€ Vid-LLM Instruction Tuning

Fine-tuning with Connective Adapters
TitleModelDateCodeVenue
Video-LLaMA: An Instruction-Finetuned Visual Language Model for Video Understanding StarVideo-LLaMA06/2023codearXiv
VALLEY: Video Assistant with Large Language model Enhanced abilitYStarVALLEY06/2023code-
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsStarVideo-ChatGPT06/2023codearXiv
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text IntegrationStarMacaw-LLM06/2023codearXiv
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning StarLLMVA-GEBC06/2023codeCVPR
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks StarmPLUG-video06/2023codearXiv
MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingStarMovieChat07/2023codearXiv
Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringStarLLaMA-VQA10/2023codeEMNLP
Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionStarVideo-LLaVA11/2023codearXiv
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video UnderstandingStarChat-UniVi11/2023codearXiv
LLaMA-VID: An Image is Worth 2 Tokens in Large Language ModelsStarLLaMA-VID11/2023codearXiv
VISTA-LLAMA: Reliable Video Narrator via Equal Distance to Visual TokensVISTA-LLAMA12/2023-arXiv
Audio-Visual LLM for Video Understanding-12/2023-arXiv
AutoAD: Movie Description in ContextAutoAD06/2023codeCVPR
AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionAutoAD II10/2023-ICCV
AutoAD III: The Prequel -- Back to the PixelsAutoAD III04/2024-CVPR
Fine-grained Audio-Visual Joint Representations for Multimodal Large Language ModelsStarFAVOR10/2023codearXiv
VideoLLaMA2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMsStarVideoLLaMA206/2024codearXiv
PAVE: Patching and Adapting Video Large Language ModelsPAVE03/2025codeCVPR
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video UnderstandingTemporal Recipe05/2025codearXiv
Watch Before You Answer: Learning from Visually Grounded Post-TrainingVidGround04/2026codearXiv
Fine-tuning with Insertive Adapters
Fine-tuning with Hybrid Adapters

🦾 Hybrid Methods

πŸ’Ž Training-free Methods


Tasks, Datasets, and Benchmarks

Recognition and Anticipation

Captioning and Description

NamePaperDateLinkVenue
Microsoft Research Video Description Corpus (MSVD)Collecting Highly Parallel Data for Paraphrase Evaluation2011LinkACL
Microsoft Research Video-to-Text (MSR-VTT)MSR-VTT: A Large Video Description Dataset for Bridging Video and Language2016LinkCVPR
Tumblr GIF (TGIF)TGIF: A New Dataset and Benchmark on Animated GIF Description2016LinkCVPR
CharadesHollywood in Homes: Crowdsourcing Data Collection for Activity Understanding2016LinkECCV
Charades-EgoActor and Observer: Joint Modeling of First and Third-Person Videos2018LinkCVPR
ActivityNet CaptionsDense-Captioning Events in Videos2017LinkICCV
HowTo100mHowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips2019LinkICCV
Movie Audio Descriptions (MAD)MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions2021LinkCVPR
YouCook2Towards Automatic Learning of Procedures from Web Instructional Videos2017LinkAAAI
MovieNetMovieNet: A Holistic Dataset for Movie Understanding2020LinkECCV
Youku-mPLUGYouku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks2023LinkarXiv
Video Timeline Tags (ViTT)Multimodal Pretraining for Dense Video Captioning2020LinkAACL-IJCNLP
TVSumTVSum: Summarizing web videos using titles2015LinkCVPR
SumMeCreating Summaries from User Videos2014LinkECCV
VideoXumVideoXum: Cross-modal Visual and Textural Summarization of Videos2023LinkIEEE Trans Multimedia
Multi-Source Video Captioning (MSVC)VideoLLaMA2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs2024LinkarXiv

Grounding and Retrieval

Question Answering

Video Instruction Tuning

Pretraining Dataset
Fine-tuning Dataset

Video-based Large Language Models Benchmark

TitleDateCodeVenue
LVBench: An Extreme Long Video Understanding Benchmark06/2024code-
Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models11/2023code-
Perception Test: A Diagnostic Benchmark for Multimodal Video Models05/2023codeNeurIPS 2023, ICCV 2023 Workshop
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks Star07/2023code-
FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation Star11/2023codeNeurIPS 2023
MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding12/2023code-
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark12/2023code-
TempCompass: Do Video LLMs Really Understand Videos? Star03/2024codeACL 2024
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis Star06/2024code-
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models Star06/2024code-
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events Star06/2025codeCVPR 2025
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment08/2025--
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models11/2025codeAAAI 2026
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMsMVU-Eval11/2025code
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsOmniVideoBench10/2025code
IF-VidCap: Can Video Caption Models Follow Instructions?IF-VidCap10/2025code
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsGameplayQA03/2026code project page

Video Dataset Tools

  • Thordata Video Dataset Toolkit: Documentation and planned tooling for loading, validating, documenting, and preparing video datasets, including manifest and metadata workflows.

Contributing

We welcome everyone to contribute to this repository and help improve it. You can submit pull requests to add new papers, projects, and helpful materials, or to correct any errors that you may find. Please make sure that your pull requests follow the "Title|Model|Date|Code|Venue" format. Thank you for your valuable contributions!

🌟 Star History

Star History Chart

β™₯️ Contributors

Our project wouldn't be possible without the contributions of these amazing people! Thank you all for making this project better.

Yolo Y. Tang @ University of Rochester
Jing Bi @ University of Rochester
Siting Xu @ Southern University of Science and Technology
Luchuan Song @ University of Rochester
Susan Liang @ University of Rochester
Teng Wang @ The University of Hong Kong
Daoan Zhang @ University of Rochester
Jie An @ University of Rochester
Jingyang Lin @ University of Rochester
Rongyi Zhu @ University of Rochester
Ali Vosoughi @ University of Rochester
Chao Huang @ University of Rochester
Zeliang Zhang @ University of Rochester
Pinxin Liu @ University of Rochester
Mingqian Feng @ University of Rochester
Feng Zheng @ Southern University of Science and Technology
Jianguo Zhang @ Southern University of Science and Technology
Ping Luo @ University of Hong Kong
Jiebo Luo @ University of Rochester
Chenliang Xu @ University of Rochester

Contributors

yunlong10

101 commits

sai-01

67 commits

ali-vosoughi

6 commits

inFaaa

6 commits

yunlong10/Awesome-LLMs-for-Video-Understanding

πŸ”₯πŸ”₯πŸ”₯ [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.

3,289

212 commits

updated Sep 24, 2026

See the code

README

Awesome-LLMs-for-Video-Understanding Awesome

πŸ”₯πŸ”₯πŸ”₯ Video Understanding with Large Language Models: A Survey

Yolo Y. Tang1, Jing Bi1, Siting Xu2, Luchuan Song1, Susan Liang1 , Teng Wang2,3 , Daoan Zhang1 , Jie An1 , Jingyang Lin1 , Rongyi Zhu1 , Ali Vosoughi1 , Chao Huang1 , Zeliang Zhang1 , Pinxin Liu1 , Mingqian Feng1 , Feng Zheng2 , Jianguo Zhang2 , Ping Luo3 , Jiebo Luo1, Chenliang Xu1.

1University of Rochester, 2Southern University of Science and Technology, 3The University of Hong Kong

Paper | arXiv | Project Page

image

πŸ“’ News

[10/06/2025]

πŸ”₯ Our follow-up workβ€”Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Modelsβ€”is now available on arXiv and Hugging Face Papers!

[05/04/2025]

🌟 Our Vid-LLM survey has been accepted to the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)! πŸ‘‰ IEEE Xplore | GitHub

[07/23/2024]

πŸ“’ We've recently updated our survey: β€œVideo Understanding with Large Language Models: A Survey”!

✨ This comprehensive survey covers video understanding techniques powered by large language models (Vid-LLMs), training strategies, relevant tasks, datasets, benchmarks, and evaluation methods, and discusses the applications of Vid-LLMs across various domains.

πŸš€ What's New in This Update:
βœ… Updated to include around 100 additional Vid-LLMs and 15 new benchmarks as of June 2024.
βœ… Introduced a novel taxonomy for Vid-LLMs based on video representation and LLM functionality.
βœ… Added a Preliminary chapter, reclassifying video understanding tasks from the perspectives of granularity and language involvement, and enhanced the LLM Background section.
βœ… Added a new Training Strategies chapter, removing adapters as a factor for model classification.
βœ… All figures and tables have been redesigned.

Multiple minor updates will follow this major update. And the GitHub repository will be gradually updated soon. We welcome your reading and feedback ❀️

Table of Contents

Why we need Vid-LLMs?

image

😎 Vid-LLMs: Models

image

πŸ“‘ Citation

If you find our survey useful for your research, please cite the following paper:

@article{vidllmsurvey,
  author={Tang, Yunlong and Bi, Jing and Xu, Siting and Song, Luchuan and Liang, Susan and Wang, Teng and Zhang, Daoan and An, Jie and Lin, Jingyang and Zhu, Rongyi and Vosoughi, Ali and Huang, Chao and Zhang, Zeliang and Liu, Pinxin and Feng, Mingqian and Zheng, Feng and Zhang, Jianguo and Luo, Ping and Luo, Jiebo and Xu, Chenliang},
  journal={IEEE Transactions on Circuits and Systems for Video Technology}, 
  title={Video Understanding with Large Language Models: A Survey}, 
  year={2025},
  doi={10.1109/TCSVT.2025.3566695}
}

πŸ—’οΈ Taxonomy 1

πŸ•ΉοΈ Video Analyzer Γ— LLM

LLM as Summarizer
LLM as Manager
TitleModelDateCodeVenue
DrVideo: Document Retrieval Based Long Video UnderstandingDrVideo06/2024codearXiv
OmAgent a multi-modal agent framework for complex video understanding with task divide-and-conquerOmAgent06/2024codearXiv
Too Many Frames, not all Useful: Efficient Strategies for Long-Form Video QALVNet06/2024codearXiv
VideoTree adaptive tree-based video representation for LLM reasoning on long videosVideoTree05/2024codearXiv
Harnessing Large Language Models for Training-free Video Anomaly DetectionLAVAD04/2024codeCVPR
TraveLER a multi-LMM agent framework for video question-answeringTraveLER04/2024codearXiv
GPTSee enhancing moment retrieval and highlight detection via description-based similarity featuresGPTSee03/2024codearXiv
Reframe anything LLM agent for open world video reframingRAVA03/2024codearXiv
SCHEMA state CHangEs MAtter for procedure planning in instructional videosSCHEMA03/2024codeICLR
TV-TREES multimodal entailment trees for neuro-symbolic video reasoningTV-TREES02/2024codearXiv
VideoAgent: A Memory-augmented Multimodal Agent for Video UnderstandingVideoAgent03/2024project pagearXiv
VideoAgent long-form video understanding with large language model as agentVideoAgent03/2024codearXiv
VURF a general-purpose reasoning and self-refinement framework for video understandingVURF03/2024codearXiv
Why not use your textbook knowledge-enhanced procedure planning of instructional videosKEPP03/2024codeCVPR
DoraemonGPT toward understanding dynamic scenes with large language modelsDoraemonGPT01/2024codearXiv
LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric VideosLifelongMemory12/2023codearXiv
Zero-Shot Video Question Answering with Procedural ProgramsProViQ12/2023codearXiv
AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and LearnAssistGPT06/2023codearXiv
ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding SystemChatVideo04/2023project pagearXiv
Video ChatCaptioner: Towards Enriched Spatiotemporal DescriptionsStarVideo ChatCaptioner04/2023codearXiv
ViperGPT: Visual Inference via Python Execution for ReasoningViperGPT03/2023codearXiv
Hawk: Learning to Understand Open-World Video AnomaliesHawk05/2024codearXiv

πŸ‘Ύ Video Embedder Γ— LLM

LLM as Text Decoder
TitleModelDateCodeVenue
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live StreamingTLive-Omni08/2026codearXiv
AuroraCap: Efficient, Performant Video Detailed Captioning and a New BenchmarkAuroraCap10/2024project pagearXiv
Artemis towards referential understanding in complex videosArtemis06/2024codearXiv
EmoLLM multimodal emotional understanding meets large language modelsEmoLLM06/2024codearXiv
Fewer tokens and fewer videos extending video understanding abilities in large vision-language modelsFTFV-LLM06/2024-arXiv
Flash-VStream: Memory-Based Real-Time Understanding for Long Video StreamsFlash-VStream06/2024codearXiv
LLAVIDAL benchmarking large language vision models for daily activities of livingLLAVIDAL06/2024codearXiv
Long context transfer from language to visionLongVA06/2024codearXiv
ShareGPT4Video improving video understanding and generation with better captionsShareGPT4Video06/2024codearXiv
Towards event-oriented long video understandingVIM06/2024codearXiv
Video-SALMONN speech-enhanced audio-visual large language modelsVideo-SALMONN06/2024codeICML
VideoGPT+ integrating image and video encoders for enhanced video understandingVideoGPT+06/2024codearXiv
VideoLLaMA 2 advancing spatial-temporal modeling and audio understanding in video-LLMsVideoLLaMA 206/2024codearXiv
MotionLLM: Understanding Human Behaviors from Human Motions and VideosMotionLLM05/2024project pagearXiv
MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkVideoChat211/2023codeCVPR
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and SummarizationShotluck Holmes05/2024-arXiv
Streaming long video understanding with large language modelsVideoStreaming05/2024-arXiv
Synchronized Video Storytelling: Generating Video Narrations with Structured StorylineVideoNarrator05/2024-arXiv
TOPA extend large language models for video understanding via text-only pre-alignmentTOPA05/2024codeNeurIPS
MovieChat+: Question-aware Sparse Memory for Long Video Question AnsweringMovieChat+04/2024codearXiv
AutoAD III: The Prequel – Back to the PixelsAutoAD III04/2024project pageCVPR
Direct Preference Optimization of Video Large Multimodal Models from Language Model RewardLLaVA-Hound-DPO04/2024codearXiv
From image to video, what do we need in multimodal LLMsRED-VILLM04/2024-arXiv
Koala key frame-conditioned long video-LLMKoala04/2024project pageCVPR
LongVLM efficient long video understanding via large language modelsLongVLM04/2024codeECCV
MA-LMM memory-augmented large multimodal model for long-term video understandingMA-LMM04/2024codeCVPR
MiniGPT4-video advancing multimodal LLMs for video understanding with interleaved visual-textual tokensMiniGPT4-Video04/2024codearXiv
Pegasus-v1 technical reportPegasus-v104/2024codearXiv
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense CaptioningPLLaVA04/2024codearXiv
ST-LLM: Large Language Models Are Effective Temporal LearnersST-LLM04/2024codearXiv
Tarsier recipes for training and evaluating large video description modelsTarsier07/2024codearXiv
X-VARS introducing explainability in football refereeing with multi-modal large language modelX-VARS04/2024codearXiv
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual ScenariosCAT03/2024codearXiv
InternVideo2 scaling video foundation models for multimodal video understandingInternVideo203/2024codeECCV
MovieLLM enhancing long video understanding with AI-generated moviesMovieLLM03/2024codearXiv
LLMs meet long video advancing long video comprehension with an interactive visual adapter in LLMsIVAwithLLM02/2024codearXiv
LSTP language-guided spatial-temporal prompt learning for long-form video-text understandingLSTP02/2024codeEMNLP
LVCHAT facilitating long video comprehensionLVCHAT02/2024codearXiv
OSCaR: Object State Captioning and State Change RepresentationOSCaR02/2024codeNAACL
Slot-VLM SlowFast slots for video-language modelingSlot-VLM02/2024codearXiv
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-TrainingCOSMO01/2024codearXiv
Weakly supervised gaussian contrastive grounding with large multimodal models for video question answeringGCG01/2024codeACMMM
Audio-Visual LLM for Video UnderstandingAV-LLM12/2023codearXiv
Generative Multimodal Models are In-Context LearnersEmu212/2023project pageCVPR
MMICT: Boosting Multi-Modal Fine-Tuning with In-Context ExamplesMMICT12/2023codeTOMM
VaQuitA : Enhancing Alignment in LLM-Assisted Video UnderstandingVaQuitA12/2023codearXiv
VILA: On Pre-training for Visual Language ModelsVILA12/2023codeCVPR
Vista-LLaMA reliable video narrator via equal distance to visual tokensVista-LLaMA12/2023project pagearXiv
Chat-UniVi unified visual representation empowers large language models with image and video understandingChat-UniVi11/2023codeCVPR
LLaMA-VID: An Image is Worth 2 Tokens in Large Language ModelsLLaMA-VID11/2023codearXiv
Video-LLaVA learning united visual representation by alignment before projectionVideo-LLaVA11/2023codearXiv
Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringLLaMA-VQA10/2023codeEMNLP
MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingMovieChat07/2023codeCVPR
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary CaptioningLLMVA-GEBC06/2023codeCVPR
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text IntegrationMacaw-LLM06/2023project pagearXiv
Valley: Video Assistant with Large Language model Enhanced abilitYVALLEY06/2023codearXiv
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsVideo-ChatGPT06/2023codeACL
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video UnderstandingVideo-LLaMA06/2023codeEMNLP
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and BenchmarksmPLUG-video06/2023codearXiv
ChatBridge: Bridging Modalities with Large Language Model as a Language CatalystChatBridge05/2023codearXiv
Otter: A Multi-Modal Model with In-Context Instruction TuningOtter05/2023codearXiv
VideoLLM: Modeling Video Sequence with Large Language ModelsVideoLLM05/2023codearXiv
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory-05/2025codeICCV 2025
LLM as Regressor
LLM as Hidden Layer

🧭 (Analyzer + Embedder) Γ— LLM

LLM as Manager
TitleModelDateCodeVenue
MM-VID: Advancing Video Understanding with GPT-4V(ision)MM-VID10/2023-arXiv
LLM as Summarizer
LLM as Regressor
LLM as Text Decoder
LLM as Hidden Layer
TitleModelDateCodeVenue
PG-Video-LLaVA: Pixel Grounding Large Video-Language ModelsStarPG-Video-LLaVA11/2023codearXiv

πŸ—’οΈ Taxonomy 2

πŸ€– LLM-based Video Agents

πŸŽ₯ Vid-LLM Pretraining

πŸ‘€ Vid-LLM Instruction Tuning

Fine-tuning with Connective Adapters
TitleModelDateCodeVenue
Video-LLaMA: An Instruction-Finetuned Visual Language Model for Video Understanding StarVideo-LLaMA06/2023codearXiv
VALLEY: Video Assistant with Large Language model Enhanced abilitYStarVALLEY06/2023code-
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsStarVideo-ChatGPT06/2023codearXiv
Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text IntegrationStarMacaw-LLM06/2023codearXiv
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning StarLLMVA-GEBC06/2023codeCVPR
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks StarmPLUG-video06/2023codearXiv
MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingStarMovieChat07/2023codearXiv
Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringStarLLaMA-VQA10/2023codeEMNLP
Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionStarVideo-LLaVA11/2023codearXiv
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video UnderstandingStarChat-UniVi11/2023codearXiv
LLaMA-VID: An Image is Worth 2 Tokens in Large Language ModelsStarLLaMA-VID11/2023codearXiv
VISTA-LLAMA: Reliable Video Narrator via Equal Distance to Visual TokensVISTA-LLAMA12/2023-arXiv
Audio-Visual LLM for Video Understanding-12/2023-arXiv
AutoAD: Movie Description in ContextAutoAD06/2023codeCVPR
AutoAD II: The Sequel - Who, When, and What in Movie Audio DescriptionAutoAD II10/2023-ICCV
AutoAD III: The Prequel -- Back to the PixelsAutoAD III04/2024-CVPR
Fine-grained Audio-Visual Joint Representations for Multimodal Large Language ModelsStarFAVOR10/2023codearXiv
VideoLLaMA2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMsStarVideoLLaMA206/2024codearXiv
PAVE: Patching and Adapting Video Large Language ModelsPAVE03/2025codeCVPR
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video UnderstandingTemporal Recipe05/2025codearXiv
Watch Before You Answer: Learning from Visually Grounded Post-TrainingVidGround04/2026codearXiv
Fine-tuning with Insertive Adapters
Fine-tuning with Hybrid Adapters

🦾 Hybrid Methods

πŸ’Ž Training-free Methods


Tasks, Datasets, and Benchmarks

Recognition and Anticipation

Captioning and Description

NamePaperDateLinkVenue
Microsoft Research Video Description Corpus (MSVD)Collecting Highly Parallel Data for Paraphrase Evaluation2011LinkACL
Microsoft Research Video-to-Text (MSR-VTT)MSR-VTT: A Large Video Description Dataset for Bridging Video and Language2016LinkCVPR
Tumblr GIF (TGIF)TGIF: A New Dataset and Benchmark on Animated GIF Description2016LinkCVPR
CharadesHollywood in Homes: Crowdsourcing Data Collection for Activity Understanding2016LinkECCV
Charades-EgoActor and Observer: Joint Modeling of First and Third-Person Videos2018LinkCVPR
ActivityNet CaptionsDense-Captioning Events in Videos2017LinkICCV
HowTo100mHowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips2019LinkICCV
Movie Audio Descriptions (MAD)MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions2021LinkCVPR
YouCook2Towards Automatic Learning of Procedures from Web Instructional Videos2017LinkAAAI
MovieNetMovieNet: A Holistic Dataset for Movie Understanding2020LinkECCV
Youku-mPLUGYouku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks2023LinkarXiv
Video Timeline Tags (ViTT)Multimodal Pretraining for Dense Video Captioning2020LinkAACL-IJCNLP
TVSumTVSum: Summarizing web videos using titles2015LinkCVPR
SumMeCreating Summaries from User Videos2014LinkECCV
VideoXumVideoXum: Cross-modal Visual and Textural Summarization of Videos2023LinkIEEE Trans Multimedia
Multi-Source Video Captioning (MSVC)VideoLLaMA2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs2024LinkarXiv

Grounding and Retrieval

Question Answering

Video Instruction Tuning

Pretraining Dataset
Fine-tuning Dataset

Video-based Large Language Models Benchmark

TitleDateCodeVenue
LVBench: An Extreme Long Video Understanding Benchmark06/2024code-
Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models11/2023code-
Perception Test: A Diagnostic Benchmark for Multimodal Video Models05/2023codeNeurIPS 2023, ICCV 2023 Workshop
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks Star07/2023code-
FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation Star11/2023codeNeurIPS 2023
MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding12/2023code-
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark12/2023code-
TempCompass: Do Video LLMs Really Understand Videos? Star03/2024codeACL 2024
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis Star06/2024code-
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models Star06/2024code-
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events Star06/2025codeCVPR 2025
Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment08/2025--
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models11/2025codeAAAI 2026
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMsMVU-Eval11/2025code
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsOmniVideoBench10/2025code
IF-VidCap: Can Video Caption Models Follow Instructions?IF-VidCap10/2025code
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual AgentsGameplayQA03/2026code project page

Video Dataset Tools

  • Thordata Video Dataset Toolkit: Documentation and planned tooling for loading, validating, documenting, and preparing video datasets, including manifest and metadata workflows.

Contributing

We welcome everyone to contribute to this repository and help improve it. You can submit pull requests to add new papers, projects, and helpful materials, or to correct any errors that you may find. Please make sure that your pull requests follow the "Title|Model|Date|Code|Venue" format. Thank you for your valuable contributions!

🌟 Star History

Star History Chart

β™₯️ Contributors

Our project wouldn't be possible without the contributions of these amazing people! Thank you all for making this project better.

Yolo Y. Tang @ University of Rochester
Jing Bi @ University of Rochester
Siting Xu @ Southern University of Science and Technology
Luchuan Song @ University of Rochester
Susan Liang @ University of Rochester
Teng Wang @ The University of Hong Kong
Daoan Zhang @ University of Rochester
Jie An @ University of Rochester
Jingyang Lin @ University of Rochester
Rongyi Zhu @ University of Rochester
Ali Vosoughi @ University of Rochester
Chao Huang @ University of Rochester
Zeliang Zhang @ University of Rochester
Pinxin Liu @ University of Rochester
Mingqian Feng @ University of Rochester
Feng Zheng @ Southern University of Science and Technology
Jianguo Zhang @ Southern University of Science and Technology
Ping Luo @ University of Hong Kong
Jiebo Luo @ University of Rochester
Chenliang Xu @ University of Rochester

Contributors

yunlong10

101 commits

sai-01

67 commits

ali-vosoughi

6 commits

inFaaa

6 commits