This repository accompanies the survey Video Understanding in the Era of Multimodal Large Language Models by Yiming Zhong, Chang Nie, Yan Yang, Xiaoyu Liu, and Caifeng Shan. It favors representative, reproducible work over an exhaustive paper dump.
| If you want to… | Recommended path | Good first reads |
|---|---|---|
| Understand the field | Milestones → architecture → evaluation | Video-LLaMA, VideoChat, LLaVA-Video |
| Process long videos | Long-context encoding | MovieChat, LLaMA-VID, LongVU, LongVILA |
| Build a live assistant | Streaming encoding | StreamingVLM, Flash-VStream, StreamChat |
| Train a model | Pre-training → instruction tuning → alignment | InternVideo, ShareGPT4Video, VideoChat-R1 |
| Evaluate a model | General → long-video → temporal → streaming | Video-MME, MLVU, TempCompass, StreamingBench |
Yiming Zhong · Chang Nie · Yan Yang · Xiaoyu Liu · Caifeng Shan
The survey organizes modern Video-LLMs through three architectural decisions—frame encoding, multimodal alignment, and LLM selection—then connects them to training strategies, datasets, evaluation protocols, and open challenges.
📌 The paper link and BibTeX will be added here when the publisher activates the public article page and DOI.
Sorted by first public release, newest first. Star badges are live and update automatically from GitHub. A dash means that no official repository was available when the entry was added.
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2026-06 | Harnessing Streaming Video in the Wild | Streaming system, proactive interaction, 12-hour memory | 📄 | — |
| 2026-05 | Linear Scaling Video VLMs for Long Video Understanding | Linear-time video encoding | 📄 | — |
| 2026-05 | CRPO: Learning Spatiotemporal Sensitivity via Counterfactual RL | Counterfactual reinforcement learning | 📄 | — |
| 2026-05 | EvoVid: Temporal-Centric Self-Evolution for Video LLMs | Temporal self-improvement | 📄 | — |
| 2026-05 | VSTAT: Benchmarking Visual State Tracking in Multimodal Video Understanding | Long-form state tracking | 💻 | |
| 2026-05 | VideoOdyssey: Ultra-Long-Context and Omni-Modal Video Understanding | Ultra-long visual and audio-visual evaluation | 📄 💻 | |
| 2026-05 | VideoZeroBench: Spatio-Temporal Evidence Verification | Hierarchical evidence verification | 💻 | |
| 2026-03 | FlexMem: Scaling Long Video Understanding via Visual Memory | Training-free visual memory | 📄 💻 | |
| 2026-03 | RIVER: A Real-Time Interaction Benchmark for Video LLMs | Streaming perception, memory, proactive response | 📄 💻 | |
| 2026-01 | Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams | Event-aware streaming and persistent memory | 📄 | — |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2025-12 | MMSI-Video-Bench: Video-Based Spatial Intelligence | Perception, planning, prediction, cross-video reasoning | 💻 | |
| 2025-10 | EgoThinker: Egocentric Reasoning with Spatio-Temporal CoT | Egocentric reasoning | 📄 | — |
| 2025-10 | DSI-Bench: A Benchmark for Dynamic Spatial Intelligence | Dynamic spatial reasoning | 📄 | — |
| 2025-10 | FineVision: Open Data Is All You Need | Open multimodal training data | 📄 | — |
| 2025-10 | K-Frames: Scene-Driven Any-k Keyframe Selection | Adaptive keyframe selection | 📄 | — |
| 2025-09 | LLaVA-OneVision-1.5 | Fully open multimodal training framework | 📄 💻 | |
| 2025-09 | Kwai Keye-VL 1.5 | Video-native multimodal foundation model | 📄 | — |
| 2025-08 | InternVL3.5 | Cascade RL and visual resolution routing | 📄 💻 | |
| 2025-08 | Thinking with Videos | Tool-augmented RL for long-video reasoning | 📄 | — |
| 2025-06 | Video-XL-2 | Task-aware KV sparsification for very long videos | 📄 | — |
| 2025-06 | MiniMax-M1 | Efficient test-time scaling with lightning attention | 📄 | — |
| 2025-06 | DeepVideo-R1 | Difficulty-aware regressive GRPO | 📄 | — |
| 2025-06 | Reinforcement Learning Tuning for VideoLLMs | Reward design and data efficiency | 📄 | — |
| 2025-06 | Video-SALMONN 2: Captioning-Enhanced Audio-Visual LLMs | Fine-grained audio-visual understanding | 📄 | — |
| 2025-05 | CrossLMM: Dual Cross-Attention for Long Video Sequences | Decoupled long-video alignment | 📄 | — |
| 2025-05 | VideoEval-Pro | Robust real-world long-video evaluation | 📄 | — |
| 2025-05 | RTV-Bench | Continuous real-time perception and reasoning | 📄 | — |
| 2025-05 | Video-Holmes: Complex Video Reasoning | Multi-clue causal reasoning | 💻 | |
| 2025-04 | Video-MMLU | Multi-discipline lecture understanding | 📄 | — |
| 2025-04 | InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source MLLMs | General-purpose multimodal foundation model | 📄 💻 | |
| 2025-04 | VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning | Reinforcement learning for video reasoning | 📄 💻 | |
| 2025-04 | VILAMP: Scaling Video-Language Models to 10K Frames | Hierarchical differential distillation | 📄 | — |
| 2025-04 | SpaceR: Reinforcing MLLMs in Video Spatial Reasoning | Spatial reasoning with RL | 📄 | — |
| 2025-03 | Exploring Hallucination in Video Understanding | Benchmark, analysis, and mitigation | 📄 | — |
| 2025-03 | FAVOR-Bench | Fine-grained video motion understanding | 📄 | — |
| 2025-03 | AdaReTAKE | Adaptive redundancy reduction | 📄 | — |
| 2025-03 | Agentic Keyframe Search for Video QA | Agentic frame retrieval | 📄 | — |
| 2025-03 | Video-R1: Reinforcing Video Reasoning in MLLMs | Rule-based reinforcement learning | 📄 💻 | |
| 2025-02 | EgoNormia | Physical and social norm understanding | 📄 | — |
| 2025-02 | video-SALMONN-o1 | Reasoning-enhanced audio-visual LLM | 📄 | — |
| 2025-02 | SVBench | Temporal multi-turn streaming dialogue | 📄 | — |
| 2025-02 | MM-RLHF | Multimodal preference alignment | 📄 | — |
| 2025-02 | Qwen2.5-VL | Dynamic-resolution perception and long-video comprehension | 📄 💻 | |
| 2025-01 | VideoChat-Flash | Hierarchical compression for long-context video modeling | 📄 | — |
| 2025-01 | VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding | Vision-centric alignment and adaptive tokenization | 📄 💻 | |
| 2025-01 | Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding | Dense description and general video understanding | 📄 💻 | |
| 2025-01 | Apollo: An Exploration of Video Understanding in Large Multimodal Models | Data, architecture, and scaling study | 📄 🌐 | — |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2024-12 | StreamChat: Chatting with Streaming Video | Online video dialogue | 📄 | — |
| 2024-12 | Inst-IT | Explicit visual-prompt instruction tuning | 📄 | — |
| 2024-11 | StreamingBench: Assessing the Gap for Streaming Video Understanding | Streaming evaluation | 📄 💻 | |
| 2024-11 | Video-RAG | Retrieval-augmented long-video comprehension | 📄 | — |
| 2024-11 | Mixed Preference Optimization for MLLMs | Multimodal preference optimization | 📄 | — |
| 2024-10 | LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding | Query-aware frame and token compression | 📄 💻 | |
| 2024-10 | TemporalBench | Fine-grained temporal understanding | 📄 | — |
| 2024-10 | LLaVA-Video: Video Instruction Tuning with Synthetic Data | Large-scale synthetic instruction tuning | 📄 💻 | |
| 2024-09 | Q-Bench-Video: Video Quality Understanding of LMMs | Perceptual video quality evaluation | 💻 | |
| 2024-08 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos | Multimodal sequence parallelism | 📄 💻 | |
| 2024-08 | Kangaroo: A Powerful Long-Context Video-Language Model | Long-context video input | 📄 | — |
| 2024-07 | LongVideoBench: Long-Context Interleaved Video-Language Understanding | Long-video benchmark | 📄 💻 | |
| 2024-06 | VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding | Audio-visual spatial-temporal modeling | 📄 💻 | |
| 2024-06 | Long Context Transfer from Language to Vision | Cross-modal context transfer | 📄 | — |
| 2024-06 | OmAgent | Divide-and-conquer video agent | 📄 | — |
| 2024-06 | VideoVista | Versatile video understanding and reasoning benchmark | 📄 | — |
| 2024-06 | VideoGPT+ | Joint image and video encoders | 📄 | — |
| 2024-06 | MMWorld | Multi-discipline world-model evaluation in videos | 📄 | — |
| 2024-06 | LVBench | Extreme long-video understanding | 📄 💻 | |
| 2024-06 | Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams | Streaming visual memory | 📄 💻 | |
| 2024-06 | MLVU: Multi-Task Long Video Understanding | Long-video evaluation | 📄 💻 | |
| 2024-06 | LongVA: Long Context Transfer from Language to Vision | Text-to-vision context transfer | 📄 💻 | |
| 2024-05 | Video-MME: Comprehensive Evaluation of MLLMs in Video Analysis | Duration- and modality-balanced evaluation | 📄 💻 | |
| 2024-05 | RLAIF-V | Open-source AI feedback for trustworthy MLLMs | 📄 | — |
| 2024-05 | CinePile | Long-video question answering dataset | 📄 | — |
| 2024-04 | PLLaVA: Parameter-Free LLaVA Extension for Video Dense Captioning | Adaptive spatial pooling | 📄 💻 | |
| 2024-04 | MiniGPT4-Video | Interleaved visual-textual tokens | 📄 | — |
| 2024-03 | TempCompass: Do Video LLMs Really Understand Videos? | Temporal perception diagnosis | 📄 💻 | |
| 2024-02 | ALLAVA | GPT-4V-synthesized data for lightweight VLMs | 📄 | — |
| 2024-02 | Momentor | Fine-grained temporal reasoning | 📄 | — |
| 2024-02 | RLAIF for Video MLLMs | Reinforcement learning from AI feedback | 📄 | — |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2023-12 | Silkie | Preference distillation for visual language models | 📄 | — |
| 2023-12 | TimeChat: A Time-Sensitive Multimodal Large Language Model for Long Video Understanding | Timestamp-aware video dialogue | 📄 💻 | |
| 2023-11 | LLaMA-VID: An Image Is Worth 2 Tokens in Large Language Models | Two-token frame representation | 📄 💻 | |
| 2023-11 | Video-Bench | Video-LLM benchmark and toolkit | 📄 | — |
| 2023-11 | GPT-4V for Visual Instruction Tuning | Synthetic visual instruction generation | 📄 | — |
| 2023-11 | Video-LLaVA: Learning United Visual Representation by Alignment Before Projection | Unified image-video representation | 📄 💻 | |
| 2023-10 | UltraFeedback | High-quality preference feedback | 📄 | — |
| 2023-09 | Factually Augmented RLHF for MLLMs | Factual alignment | 📄 | — |
| 2023-08 | EgoSchema | Very long-form video-language diagnosis | 📄 | — |
| 2023-07 | MovieChat: From Dense Token to Sparse Memory for Long Video Understanding | Long-video memory | 📄 💻 | |
| 2023-07 | InternVid | Large-scale video-text pre-training data | 📄 💻 | |
| 2023-06 | Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models | Video dialogue and instruction data | 📄 💻 | |
| 2023-06 | Video-LLaMA: An Instruction-Tuned Audio-Visual Language Model | Audio-visual instruction tuning | 📄 💻 | |
| 2023-05 | VideoChat: Chat-Centric Video Understanding | Video-centric dialogue | 📄 💻 | |
| 2023-03 | EVA-CLIP | Scaling vision-language pre-training | 📄 | — |
| 2023-02 | Vid2Seq | Dense video captioning with sequence prediction | 📄 | — |
| 2023-01 | BLIP-2 | Query-based visual-language alignment | 📄 💻 |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2022-12 | InternVideo: General Video Foundation Models | Generative and discriminative video pre-training | 📄 💻 | |
| 2022-04 | Flamingo: A Visual Language Model for Few-Shot Learning | Interleaved visual-language learning | 📄 | — |
| 2022-01 | MERLOT Reserve | Script knowledge from video and audio | 📄 | — |
| 2021-04 | Frozen in Time | End-to-end video-text representation learning | 📄 💻 | |
| 2021-03 | CLIP: Learning Transferable Visual Models from Natural Language Supervision | Image-text contrastive pre-training | 📄 💻 | |
| 2020-05 | HERO: Hierarchical Encoder for Video+Language Omni-Representation | Hierarchical video-language pre-training | 📄 | — |
| 2019-06 | HowTo100M | Large-scale narrated instructional video data | 📄 🌐 | — |
| 2019-04 | VideoBERT | Joint video-language representation learning | 📄 | — |
| Year | Work | Why it matters | Resources |
|---|---|---|---|
| 2025 | Video-SALMONN 2 | Strengthens fine-grained audio-visual understanding using caption-enhanced training. | 📄 |
| 2024 | LLaVA-Video | Scales video instruction tuning with synthetic video data. | 📄 💻 |
| 2023 | LLaMA-VID | Compresses each frame into two tokens for efficient long-video processing. | 📄 💻 |
| 2023 | Video-LLaMA | Aligns visual and audio streams with an instruction-tuned language model. | 📄 💻 |
| 2023 | VideoChat | Introduces a chat-centric framework for video understanding. | 📄 💻 |
| 2022 | InternVideo | Establishes a general video foundation model through generative and discriminative learning. | 📄 💻 |
Encode sampled frames independently with an image encoder, or model motion directly with a video-native encoder.
| Strategy | Representative work | Core idea | Resources |
|---|---|---|---|
| Token pooling | PLLaVA | Adaptive spatial pooling preserves temporal order while reducing visual tokens. | 📄 💻 |
| Token abstraction | LLaMA-VID | Represents one frame with a context token and a content token. | 📄 💻 |
| Memory | MovieChat | Uses short- and long-term memory to process videos beyond the context window. | 📄 💻 |
| Adaptive compression | LongVU | Removes redundant frames and preserves query-relevant detail at higher fidelity. | 📄 💻 |
| Context transfer | LongVA | Transfers text-only long-context ability to video without long-video training. | 📄 💻 |
| Sequence parallelism | LongVILA | Extends VILA to thousands of frames through staged training and distributed systems. | 📄 💻 |
| Family | Representative systems | Strength | Limitation |
|---|---|---|---|
| Projection-based | LLaVA-NeXT-Video, Qwen2-VL, InternVL | Simple, stable, parameter-efficient | Token count grows with frames |
| Query-based | BLIP-2, Video-LLaMA, VideoChat, TimeChat | Fixed or adaptive visual abstraction | May lose spatial and fine-grained detail |
| Audio-visual | video-SALMONN, Video-LLaMA 2 | Adds speech, music, and environmental cues | Fine-grained temporal alignment remains difficult |
Selected resources:
| Dataset | Modality | Scale / focus | Resources |
|---|---|---|---|
| WebVid-2M | Video–text | Large-scale web video captions | 📄 🤗 |
| HowTo100M | Video–speech | Instructional videos with narrated text | 📄 🌐 |
| InternVid | Video–text | Large, diverse video-text pre-training corpus | 📄 💻 |
| Panda-70M | Video–text | High-quality captions generated at web scale | 📄 🌐 |
| Benchmark | Primary focus | Format | Resources |
|---|---|---|---|
| MVBench | 20 temporal and multimodal video tasks | Multiple choice | 📄 💻 |
| Video-MME | Duration-, domain-, and modality-balanced evaluation | Multiple choice | 📄 💻 |
| MLVU | Multi-task long-video understanding | Mixed | 📄 💻 |
| LongVideoBench | Interleaved video-language reasoning over long contexts | Multiple choice | 📄 💻 |
| LVBench | Hour-long real-world video understanding | Multiple choice | 📄 💻 |
Video-LLMs
├── Perception → fine detail, motion, audio, OCR, spatial grounding
├── Temporal memory → long context, streaming, event boundaries
├── Reasoning → causality, compositionality, multi-step inference
├── Reliability → hallucination, calibration, evidence attribution
├── Efficiency → token compression, adaptive compute, edge inference
└── Responsibility → privacy, bias, safety, provenance
Contributions are welcome. Please read CONTRIBUTING.md before opening a pull request. A good addition includes a stable paper link, an official code or project link when available, and one sentence explaining why the work belongs in its category.
The presentation follows the curation principles of the Awesome manifesto and draws organizational inspiration from established multimodal research lists such as Awesome Multimodal Large Language Models.
To the extent possible under law, the maintainers have waived copyright and related rights to this curated list under CC0 1.0. Individual papers, codebases, models, and datasets retain their original licenses.
6 commits
Ruby
100.0%
This repository accompanies the survey Video Understanding in the Era of Multimodal Large Language Models by Yiming Zhong, Chang Nie, Yan Yang, Xiaoyu Liu, and Caifeng Shan. It favors representative, reproducible work over an exhaustive paper dump.
| If you want to… | Recommended path | Good first reads |
|---|---|---|
| Understand the field | Milestones → architecture → evaluation | Video-LLaMA, VideoChat, LLaVA-Video |
| Process long videos | Long-context encoding | MovieChat, LLaMA-VID, LongVU, LongVILA |
| Build a live assistant | Streaming encoding | StreamingVLM, Flash-VStream, StreamChat |
| Train a model | Pre-training → instruction tuning → alignment | InternVideo, ShareGPT4Video, VideoChat-R1 |
| Evaluate a model | General → long-video → temporal → streaming | Video-MME, MLVU, TempCompass, StreamingBench |
Yiming Zhong · Chang Nie · Yan Yang · Xiaoyu Liu · Caifeng Shan
The survey organizes modern Video-LLMs through three architectural decisions—frame encoding, multimodal alignment, and LLM selection—then connects them to training strategies, datasets, evaluation protocols, and open challenges.
📌 The paper link and BibTeX will be added here when the publisher activates the public article page and DOI.
Sorted by first public release, newest first. Star badges are live and update automatically from GitHub. A dash means that no official repository was available when the entry was added.
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2026-06 | Harnessing Streaming Video in the Wild | Streaming system, proactive interaction, 12-hour memory | 📄 | — |
| 2026-05 | Linear Scaling Video VLMs for Long Video Understanding | Linear-time video encoding | 📄 | — |
| 2026-05 | CRPO: Learning Spatiotemporal Sensitivity via Counterfactual RL | Counterfactual reinforcement learning | 📄 | — |
| 2026-05 | EvoVid: Temporal-Centric Self-Evolution for Video LLMs | Temporal self-improvement | 📄 | — |
| 2026-05 | VSTAT: Benchmarking Visual State Tracking in Multimodal Video Understanding | Long-form state tracking | 💻 | |
| 2026-05 | VideoOdyssey: Ultra-Long-Context and Omni-Modal Video Understanding | Ultra-long visual and audio-visual evaluation | 📄 💻 | |
| 2026-05 | VideoZeroBench: Spatio-Temporal Evidence Verification | Hierarchical evidence verification | 💻 | |
| 2026-03 | FlexMem: Scaling Long Video Understanding via Visual Memory | Training-free visual memory | 📄 💻 | |
| 2026-03 | RIVER: A Real-Time Interaction Benchmark for Video LLMs | Streaming perception, memory, proactive response | 📄 💻 | |
| 2026-01 | Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams | Event-aware streaming and persistent memory | 📄 | — |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2025-12 | MMSI-Video-Bench: Video-Based Spatial Intelligence | Perception, planning, prediction, cross-video reasoning | 💻 | |
| 2025-10 | EgoThinker: Egocentric Reasoning with Spatio-Temporal CoT | Egocentric reasoning | 📄 | — |
| 2025-10 | DSI-Bench: A Benchmark for Dynamic Spatial Intelligence | Dynamic spatial reasoning | 📄 | — |
| 2025-10 | FineVision: Open Data Is All You Need | Open multimodal training data | 📄 | — |
| 2025-10 | K-Frames: Scene-Driven Any-k Keyframe Selection | Adaptive keyframe selection | 📄 | — |
| 2025-09 | LLaVA-OneVision-1.5 | Fully open multimodal training framework | 📄 💻 | |
| 2025-09 | Kwai Keye-VL 1.5 | Video-native multimodal foundation model | 📄 | — |
| 2025-08 | InternVL3.5 | Cascade RL and visual resolution routing | 📄 💻 | |
| 2025-08 | Thinking with Videos | Tool-augmented RL for long-video reasoning | 📄 | — |
| 2025-06 | Video-XL-2 | Task-aware KV sparsification for very long videos | 📄 | — |
| 2025-06 | MiniMax-M1 | Efficient test-time scaling with lightning attention | 📄 | — |
| 2025-06 | DeepVideo-R1 | Difficulty-aware regressive GRPO | 📄 | — |
| 2025-06 | Reinforcement Learning Tuning for VideoLLMs | Reward design and data efficiency | 📄 | — |
| 2025-06 | Video-SALMONN 2: Captioning-Enhanced Audio-Visual LLMs | Fine-grained audio-visual understanding | 📄 | — |
| 2025-05 | CrossLMM: Dual Cross-Attention for Long Video Sequences | Decoupled long-video alignment | 📄 | — |
| 2025-05 | VideoEval-Pro | Robust real-world long-video evaluation | 📄 | — |
| 2025-05 | RTV-Bench | Continuous real-time perception and reasoning | 📄 | — |
| 2025-05 | Video-Holmes: Complex Video Reasoning | Multi-clue causal reasoning | 💻 | |
| 2025-04 | Video-MMLU | Multi-discipline lecture understanding | 📄 | — |
| 2025-04 | InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source MLLMs | General-purpose multimodal foundation model | 📄 💻 | |
| 2025-04 | VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning | Reinforcement learning for video reasoning | 📄 💻 | |
| 2025-04 | VILAMP: Scaling Video-Language Models to 10K Frames | Hierarchical differential distillation | 📄 | — |
| 2025-04 | SpaceR: Reinforcing MLLMs in Video Spatial Reasoning | Spatial reasoning with RL | 📄 | — |
| 2025-03 | Exploring Hallucination in Video Understanding | Benchmark, analysis, and mitigation | 📄 | — |
| 2025-03 | FAVOR-Bench | Fine-grained video motion understanding | 📄 | — |
| 2025-03 | AdaReTAKE | Adaptive redundancy reduction | 📄 | — |
| 2025-03 | Agentic Keyframe Search for Video QA | Agentic frame retrieval | 📄 | — |
| 2025-03 | Video-R1: Reinforcing Video Reasoning in MLLMs | Rule-based reinforcement learning | 📄 💻 | |
| 2025-02 | EgoNormia | Physical and social norm understanding | 📄 | — |
| 2025-02 | video-SALMONN-o1 | Reasoning-enhanced audio-visual LLM | 📄 | — |
| 2025-02 | SVBench | Temporal multi-turn streaming dialogue | 📄 | — |
| 2025-02 | MM-RLHF | Multimodal preference alignment | 📄 | — |
| 2025-02 | Qwen2.5-VL | Dynamic-resolution perception and long-video comprehension | 📄 💻 | |
| 2025-01 | VideoChat-Flash | Hierarchical compression for long-context video modeling | 📄 | — |
| 2025-01 | VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding | Vision-centric alignment and adaptive tokenization | 📄 💻 | |
| 2025-01 | Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding | Dense description and general video understanding | 📄 💻 | |
| 2025-01 | Apollo: An Exploration of Video Understanding in Large Multimodal Models | Data, architecture, and scaling study | 📄 🌐 | — |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2024-12 | StreamChat: Chatting with Streaming Video | Online video dialogue | 📄 | — |
| 2024-12 | Inst-IT | Explicit visual-prompt instruction tuning | 📄 | — |
| 2024-11 | StreamingBench: Assessing the Gap for Streaming Video Understanding | Streaming evaluation | 📄 💻 | |
| 2024-11 | Video-RAG | Retrieval-augmented long-video comprehension | 📄 | — |
| 2024-11 | Mixed Preference Optimization for MLLMs | Multimodal preference optimization | 📄 | — |
| 2024-10 | LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding | Query-aware frame and token compression | 📄 💻 | |
| 2024-10 | TemporalBench | Fine-grained temporal understanding | 📄 | — |
| 2024-10 | LLaVA-Video: Video Instruction Tuning with Synthetic Data | Large-scale synthetic instruction tuning | 📄 💻 | |
| 2024-09 | Q-Bench-Video: Video Quality Understanding of LMMs | Perceptual video quality evaluation | 💻 | |
| 2024-08 | LongVILA: Scaling Long-Context Visual Language Models for Long Videos | Multimodal sequence parallelism | 📄 💻 | |
| 2024-08 | Kangaroo: A Powerful Long-Context Video-Language Model | Long-context video input | 📄 | — |
| 2024-07 | LongVideoBench: Long-Context Interleaved Video-Language Understanding | Long-video benchmark | 📄 💻 | |
| 2024-06 | VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding | Audio-visual spatial-temporal modeling | 📄 💻 | |
| 2024-06 | Long Context Transfer from Language to Vision | Cross-modal context transfer | 📄 | — |
| 2024-06 | OmAgent | Divide-and-conquer video agent | 📄 | — |
| 2024-06 | VideoVista | Versatile video understanding and reasoning benchmark | 📄 | — |
| 2024-06 | VideoGPT+ | Joint image and video encoders | 📄 | — |
| 2024-06 | MMWorld | Multi-discipline world-model evaluation in videos | 📄 | — |
| 2024-06 | LVBench | Extreme long-video understanding | 📄 💻 | |
| 2024-06 | Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams | Streaming visual memory | 📄 💻 | |
| 2024-06 | MLVU: Multi-Task Long Video Understanding | Long-video evaluation | 📄 💻 | |
| 2024-06 | LongVA: Long Context Transfer from Language to Vision | Text-to-vision context transfer | 📄 💻 | |
| 2024-05 | Video-MME: Comprehensive Evaluation of MLLMs in Video Analysis | Duration- and modality-balanced evaluation | 📄 💻 | |
| 2024-05 | RLAIF-V | Open-source AI feedback for trustworthy MLLMs | 📄 | — |
| 2024-05 | CinePile | Long-video question answering dataset | 📄 | — |
| 2024-04 | PLLaVA: Parameter-Free LLaVA Extension for Video Dense Captioning | Adaptive spatial pooling | 📄 💻 | |
| 2024-04 | MiniGPT4-Video | Interleaved visual-textual tokens | 📄 | — |
| 2024-03 | TempCompass: Do Video LLMs Really Understand Videos? | Temporal perception diagnosis | 📄 💻 | |
| 2024-02 | ALLAVA | GPT-4V-synthesized data for lightweight VLMs | 📄 | — |
| 2024-02 | Momentor | Fine-grained temporal reasoning | 📄 | — |
| 2024-02 | RLAIF for Video MLLMs | Reinforcement learning from AI feedback | 📄 | — |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2023-12 | Silkie | Preference distillation for visual language models | 📄 | — |
| 2023-12 | TimeChat: A Time-Sensitive Multimodal Large Language Model for Long Video Understanding | Timestamp-aware video dialogue | 📄 💻 | |
| 2023-11 | LLaMA-VID: An Image Is Worth 2 Tokens in Large Language Models | Two-token frame representation | 📄 💻 | |
| 2023-11 | Video-Bench | Video-LLM benchmark and toolkit | 📄 | — |
| 2023-11 | GPT-4V for Visual Instruction Tuning | Synthetic visual instruction generation | 📄 | — |
| 2023-11 | Video-LLaVA: Learning United Visual Representation by Alignment Before Projection | Unified image-video representation | 📄 💻 | |
| 2023-10 | UltraFeedback | High-quality preference feedback | 📄 | — |
| 2023-09 | Factually Augmented RLHF for MLLMs | Factual alignment | 📄 | — |
| 2023-08 | EgoSchema | Very long-form video-language diagnosis | 📄 | — |
| 2023-07 | MovieChat: From Dense Token to Sparse Memory for Long Video Understanding | Long-video memory | 📄 💻 | |
| 2023-07 | InternVid | Large-scale video-text pre-training data | 📄 💻 | |
| 2023-06 | Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models | Video dialogue and instruction data | 📄 💻 | |
| 2023-06 | Video-LLaMA: An Instruction-Tuned Audio-Visual Language Model | Audio-visual instruction tuning | 📄 💻 | |
| 2023-05 | VideoChat: Chat-Centric Video Understanding | Video-centric dialogue | 📄 💻 | |
| 2023-03 | EVA-CLIP | Scaling vision-language pre-training | 📄 | — |
| 2023-02 | Vid2Seq | Dense video captioning with sequence prediction | 📄 | — |
| 2023-01 | BLIP-2 | Query-based visual-language alignment | 📄 💻 |
| Date | Title | Topic | Paper / Code | Stars |
|---|---|---|---|---|
| 2022-12 | InternVideo: General Video Foundation Models | Generative and discriminative video pre-training | 📄 💻 | |
| 2022-04 | Flamingo: A Visual Language Model for Few-Shot Learning | Interleaved visual-language learning | 📄 | — |
| 2022-01 | MERLOT Reserve | Script knowledge from video and audio | 📄 | — |
| 2021-04 | Frozen in Time | End-to-end video-text representation learning | 📄 💻 | |
| 2021-03 | CLIP: Learning Transferable Visual Models from Natural Language Supervision | Image-text contrastive pre-training | 📄 💻 | |
| 2020-05 | HERO: Hierarchical Encoder for Video+Language Omni-Representation | Hierarchical video-language pre-training | 📄 | — |
| 2019-06 | HowTo100M | Large-scale narrated instructional video data | 📄 🌐 | — |
| 2019-04 | VideoBERT | Joint video-language representation learning | 📄 | — |
| Year | Work | Why it matters | Resources |
|---|---|---|---|
| 2025 | Video-SALMONN 2 | Strengthens fine-grained audio-visual understanding using caption-enhanced training. | 📄 |
| 2024 | LLaVA-Video | Scales video instruction tuning with synthetic video data. | 📄 💻 |
| 2023 | LLaMA-VID | Compresses each frame into two tokens for efficient long-video processing. | 📄 💻 |
| 2023 | Video-LLaMA | Aligns visual and audio streams with an instruction-tuned language model. | 📄 💻 |
| 2023 | VideoChat | Introduces a chat-centric framework for video understanding. | 📄 💻 |
| 2022 | InternVideo | Establishes a general video foundation model through generative and discriminative learning. | 📄 💻 |
Encode sampled frames independently with an image encoder, or model motion directly with a video-native encoder.
| Strategy | Representative work | Core idea | Resources |
|---|---|---|---|
| Token pooling | PLLaVA | Adaptive spatial pooling preserves temporal order while reducing visual tokens. | 📄 💻 |
| Token abstraction | LLaMA-VID | Represents one frame with a context token and a content token. | 📄 💻 |
| Memory | MovieChat | Uses short- and long-term memory to process videos beyond the context window. | 📄 💻 |
| Adaptive compression | LongVU | Removes redundant frames and preserves query-relevant detail at higher fidelity. | 📄 💻 |
| Context transfer | LongVA | Transfers text-only long-context ability to video without long-video training. | 📄 💻 |
| Sequence parallelism | LongVILA | Extends VILA to thousands of frames through staged training and distributed systems. | 📄 💻 |
| Family | Representative systems | Strength | Limitation |
|---|---|---|---|
| Projection-based | LLaVA-NeXT-Video, Qwen2-VL, InternVL | Simple, stable, parameter-efficient | Token count grows with frames |
| Query-based | BLIP-2, Video-LLaMA, VideoChat, TimeChat | Fixed or adaptive visual abstraction | May lose spatial and fine-grained detail |
| Audio-visual | video-SALMONN, Video-LLaMA 2 | Adds speech, music, and environmental cues | Fine-grained temporal alignment remains difficult |
Selected resources:
| Dataset | Modality | Scale / focus | Resources |
|---|---|---|---|
| WebVid-2M | Video–text | Large-scale web video captions | 📄 🤗 |
| HowTo100M | Video–speech | Instructional videos with narrated text | 📄 🌐 |
| InternVid | Video–text | Large, diverse video-text pre-training corpus | 📄 💻 |
| Panda-70M | Video–text | High-quality captions generated at web scale | 📄 🌐 |
| Benchmark | Primary focus | Format | Resources |
|---|---|---|---|
| MVBench | 20 temporal and multimodal video tasks | Multiple choice | 📄 💻 |
| Video-MME | Duration-, domain-, and modality-balanced evaluation | Multiple choice | 📄 💻 |
| MLVU | Multi-task long-video understanding | Mixed | 📄 💻 |
| LongVideoBench | Interleaved video-language reasoning over long contexts | Multiple choice | 📄 💻 |
| LVBench | Hour-long real-world video understanding | Multiple choice | 📄 💻 |
Video-LLMs
├── Perception → fine detail, motion, audio, OCR, spatial grounding
├── Temporal memory → long context, streaming, event boundaries
├── Reasoning → causality, compositionality, multi-step inference
├── Reliability → hallucination, calibration, evidence attribution
├── Efficiency → token compression, adaptive compute, edge inference
└── Responsibility → privacy, bias, safety, provenance
Contributions are welcome. Please read CONTRIBUTING.md before opening a pull request. A good addition includes a stable paper link, an official code or project link when available, and one sentence explaining why the work belongs in its category.
The presentation follows the curation principles of the Awesome manifesto and draws organizational inspiration from established multimodal research lists such as Awesome Multimodal Large Language Models.
To the extent possible under law, the maintainers have waived copyright and related rights to this curated list under CC0 1.0. Individual papers, codebases, models, and datasets retain their original licenses.
6 commits
Ruby
100.0%