A survey on MM-LLMs for long video understanding: From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
25
10 commits
updated Sep 12, 2025
๐ A Comprehensive Survey on MultiModal Large Language Models for Long Video Understanding
This repository contains the most comprehensive, up-to-date, and innovative survey on MultiModal Large Language Models (MM-LLMs) for Long Video Understanding. As video content continues to grow exponentially, understanding videos that span from seconds to hours becomes increasingly crucial for various applications including video analysis, content moderation, educational technology, and entertainment.
Updated: January 15, 2025
graph TD
A[Long Video Understanding Tasks] --> B[Video QA]
A --> C[Temporal Localization]
A --> D[Video Summarization]
A --> E[Multi-hour Analysis]
B --> B1[Question Answering]
B --> B2[Content Understanding]
C --> C1[Event Detection]
C --> C2[Temporal Grounding]
D --> D1[Key Moment Extraction]
D --> D2[Narrative Summary]
E --> E1[Long-term Dependencies]
E --> E2[Cross-temporal Relations]
The integration of Large Language Models (LLMs) with visual encoders has recently shown promising performance in visual understanding tasks, leveraging their inherent capability to comprehend and generate human-like text for visual reasoning. This paper reviews the advancements in MultiModal Large Language Models (MM-LLMs) for long video understanding.
We highlight the unique challenges posed by long videos, including fine-grained spatiotemporal details, dynamic events, and long-term dependencies. We summarize the progress in model design and training methodologies for MM-LLMs understanding long videos and compare their performance on various long video understanding benchmarks. Finally, we discuss future directions for MM-LLMs in long video understanding.
This survey provides a comprehensive review of MultiModal Large Language Models (MM-LLMs) for long video understanding, covering:
timeline
title Evolution of Long Video Understanding Models
2023 Q2 : InstructBLIP (23.05)
: VideoChat (23.05)
: Video-LLaMA (23.06)
: Video-ChatGPT (23.06)
: Valley (23.06)
2023 Q3 : MovieChat (23.07)
2023 Q4 : LLaMA-VID (23.11)
: VideoChat2 (23.11)
: TimeChat (23.12)
2024 Q1 : LongVLM (23.04)
: Momentor (24.02)
: MovieLLM (24.03)
: MA-LMM (24.04)
: ST-LLM (24.04)
2024 Q3 : LONGVILA (24.08)
: Qwen2-VL (24.09)
: Oryx-1.5 (24.10)
2024 Q4 : TimeMarker (24.11)
: NVILA (24.12)
2025 Q1 : VideoChat-Flash (25.01)
: R1-VL (25.03)
| Model | Year | Backbone | Connector | Frame | Token | Training | Long | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Visual Encoder | LLMs | Image-level | Video-level | Long-video-level | Hardware | PreT | IT | ||||
| InstructBLIP | 23.05 | EVA-CLIP-ViT-G/14 | FlanT5, Vicuna-7B/13B | Q-Former | -- | -- | 4 | 32/128 | 16 A100-40G | Y-N-N | Y-N-N | No |
| VideoChat | 23.05 | EVA-CLIP-ViT-G/14 | StableVicuna-13B | Q-Former | Global multi-head relation aggregator | -- | 8 | /32 | 1 A10 | Y-Y-N | Y-Y-N | No |
| MovieChat | 23.07 | EVA-CLIP-ViT-G/14 | LLama-7B | Q-Former | Frame merging, Q-Former | Merging adjacent frames | 2048 | 32/32 | - | E2E | E2E | โ Yes |
| TimeChat | 23.12 | EVA-CLIP-ViT-G/14 | LLaMA2-7B | Q-Former | Sliding window Q-Former | Time-aware encoding | 96 | /96 | 8 V100-32G | Y-Y-N | N-N-Y | โ Yes |
| LONGVILA | 24.08 | SigLIP-SO400M | Qwen2-1.5B/7B | Multi-Modal Sequence Parallelism | 1024 | 256/ | 256 A100 80G | Y-Y-N | Y-Y-Y | โ Yes | ||
| NVILA | 24.12 | SigLIP-SO400M | Qwen2-7B/14B | Spatial-to-Channel Reshaping | Temporal Averaging | 256 | /8192 | 128 H100-80G | Y-Y-N | Y-Y-Y | โ Yes | |
Note: This is a condensed view. The full table contains 50+ models with detailed specifications.
| Benchmark | Videos | Annotations | Avg Duration | Focus |
|---|---|---|---|---|
| Video-MME | 900 | 2,700 | 17.0 min | Multi-scale evaluation |
| VideoVista | - | - | - | Long video understanding |
| EgoSchema | - | - | 180 sec | Egocentric video reasoning |
| LongVideoBench | - | - | - | Reference-based evaluation |
| MLVU | - | - | - | Multi-task long video understanding |
| HourVideo | 500 | 12,976 | 45.7 min | Hour-level understanding |
| HLV-1K | 1,009 | 14,847 | 55.0 min | Comprehensive evaluation |
| LVBench | 103 | 1,549 | 68.4 min | Long-form analysis |
This survey analyzes how multimodal large language models process long videos through different architectural components:
graph LR
A[Video Input] --> B[Visual Encoder]
A --> C[Temporal Modeling]
A --> D[Language Integration]
B --> B1[Frame Features]
B --> B2[Spatial Attention]
C --> C1[Temporal Attention]
C --> C2[Memory Mechanisms]
D --> D1[Cross-modal Fusion]
D --> D2[Language Generation]
๐ Key Insights:
| Reasoning Type | Complexity | Representative Models | Performance Range |
|---|---|---|---|
| Frame-level Events | Low | Most MM-LLMs | 85-95% |
| Short-term Patterns | Medium | Video-LLaVA, TimeChat | 75-85% |
| Long-term Dependencies | High | MovieChat, LongVA | 65-80% |
| Cross-temporal Relations | Very High | LONGVILA, NVILA | 60-75% |
flowchart TD
A[Multimodal Input] --> B{Fusion Strategy}
B --> C[Early Fusion]
B --> D[Late Fusion]
B --> E[Hierarchical Fusion]
C --> C1[Feature Concatenation]
C --> C2[Cross-modal Attention]
D --> D1[Independent Processing]
D --> D2[Decision Combination]
E --> E1[Multi-level Integration]
E --> E2[Adaptive Weighting]
Key Findings: Hierarchical fusion strategies show better performance for long video understanding tasks.
๐ Memory-Augmented Models (15+ models)
โโโ ๐ฌ Sparse Memory (MovieChat, MA-LMM)
โโโ ๐ Sliding Windows (TimeChat, LLaMA-VID)
โโโ ๐ Dynamic Compression (Video-XL, Oryx-1.5)
๐ Efficiency Techniques
โโโ ๐ Token Merging (LongVLM, Video-LLaVA)
โโโ ๐ Hierarchical Processing (SlowFast-LLaVA)
โโโ ๐ Parallel Processing (LONGVILA)
โโโ ๐ Adaptive Pooling (PLLaVA, VideoGPT+)
๐ง Connector Types
โโโ ๐ค Q-Former Based (MovieChat, TimeChat)
โโโ ๐ Cross-Attention (Qwen-VL, EVLM)
โโโ ๐ MLP Projectors (VITA, LLaVA-OneVision)
โโโ ๐ง Advanced Fusion (Kangaroo, NVILA)
| Strategy | Models | Advantages | Challenges |
|---|---|---|---|
| End-to-End | MovieChat, MA-LMM | Optimal performance | High computational cost |
| Stage-wise | Video-LLaVA, TimeChat | Stable training | Suboptimal alignment |
| Hybrid | LongVA, LONGVILA | Balanced approach | Complex implementation |
Based on emerging trends from recent research, the following developments are expected:
Based on current challenges and limitations in long video understanding, several key research directions emerge:
If you find our survey useful in your research, please consider citing:
@article{zou2024seconds,
title={From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding},
author={Zou, Heqing and Luo, Tianze and Xie, Guiyang and Lv, Fengmao and Wang, Guangcong and Chen, Juanyang and Wang, Zhuochen and Zhang, Hansheng and Zhang, Huaijian and others},
journal={arXiv preprint arXiv:2409.18938},
year={2024}
}
We welcome contributions to this survey! Here's how you can help:
git checkout -b feature/new-model)git commit -am 'Add new model: ModelName')git push origin feature/new-model)This project is licensed under the MIT License - see the LICENSE file for details.
A survey on MM-LLMs for long video understanding: From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
25
10 commits
updated Sep 12, 2025
๐ A Comprehensive Survey on MultiModal Large Language Models for Long Video Understanding
This repository contains the most comprehensive, up-to-date, and innovative survey on MultiModal Large Language Models (MM-LLMs) for Long Video Understanding. As video content continues to grow exponentially, understanding videos that span from seconds to hours becomes increasingly crucial for various applications including video analysis, content moderation, educational technology, and entertainment.
Updated: January 15, 2025
graph TD
A[Long Video Understanding Tasks] --> B[Video QA]
A --> C[Temporal Localization]
A --> D[Video Summarization]
A --> E[Multi-hour Analysis]
B --> B1[Question Answering]
B --> B2[Content Understanding]
C --> C1[Event Detection]
C --> C2[Temporal Grounding]
D --> D1[Key Moment Extraction]
D --> D2[Narrative Summary]
E --> E1[Long-term Dependencies]
E --> E2[Cross-temporal Relations]
The integration of Large Language Models (LLMs) with visual encoders has recently shown promising performance in visual understanding tasks, leveraging their inherent capability to comprehend and generate human-like text for visual reasoning. This paper reviews the advancements in MultiModal Large Language Models (MM-LLMs) for long video understanding.
We highlight the unique challenges posed by long videos, including fine-grained spatiotemporal details, dynamic events, and long-term dependencies. We summarize the progress in model design and training methodologies for MM-LLMs understanding long videos and compare their performance on various long video understanding benchmarks. Finally, we discuss future directions for MM-LLMs in long video understanding.
This survey provides a comprehensive review of MultiModal Large Language Models (MM-LLMs) for long video understanding, covering:
timeline
title Evolution of Long Video Understanding Models
2023 Q2 : InstructBLIP (23.05)
: VideoChat (23.05)
: Video-LLaMA (23.06)
: Video-ChatGPT (23.06)
: Valley (23.06)
2023 Q3 : MovieChat (23.07)
2023 Q4 : LLaMA-VID (23.11)
: VideoChat2 (23.11)
: TimeChat (23.12)
2024 Q1 : LongVLM (23.04)
: Momentor (24.02)
: MovieLLM (24.03)
: MA-LMM (24.04)
: ST-LLM (24.04)
2024 Q3 : LONGVILA (24.08)
: Qwen2-VL (24.09)
: Oryx-1.5 (24.10)
2024 Q4 : TimeMarker (24.11)
: NVILA (24.12)
2025 Q1 : VideoChat-Flash (25.01)
: R1-VL (25.03)
| Model | Year | Backbone | Connector | Frame | Token | Training | Long | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Visual Encoder | LLMs | Image-level | Video-level | Long-video-level | Hardware | PreT | IT | ||||
| InstructBLIP | 23.05 | EVA-CLIP-ViT-G/14 | FlanT5, Vicuna-7B/13B | Q-Former | -- | -- | 4 | 32/128 | 16 A100-40G | Y-N-N | Y-N-N | No |
| VideoChat | 23.05 | EVA-CLIP-ViT-G/14 | StableVicuna-13B | Q-Former | Global multi-head relation aggregator | -- | 8 | /32 | 1 A10 | Y-Y-N | Y-Y-N | No |
| MovieChat | 23.07 | EVA-CLIP-ViT-G/14 | LLama-7B | Q-Former | Frame merging, Q-Former | Merging adjacent frames | 2048 | 32/32 | - | E2E | E2E | โ Yes |
| TimeChat | 23.12 | EVA-CLIP-ViT-G/14 | LLaMA2-7B | Q-Former | Sliding window Q-Former | Time-aware encoding | 96 | /96 | 8 V100-32G | Y-Y-N | N-N-Y | โ Yes |
| LONGVILA | 24.08 | SigLIP-SO400M | Qwen2-1.5B/7B | Multi-Modal Sequence Parallelism | 1024 | 256/ | 256 A100 80G | Y-Y-N | Y-Y-Y | โ Yes | ||
| NVILA | 24.12 | SigLIP-SO400M | Qwen2-7B/14B | Spatial-to-Channel Reshaping | Temporal Averaging | 256 | /8192 | 128 H100-80G | Y-Y-N | Y-Y-Y | โ Yes | |
Note: This is a condensed view. The full table contains 50+ models with detailed specifications.
| Benchmark | Videos | Annotations | Avg Duration | Focus |
|---|---|---|---|---|
| Video-MME | 900 | 2,700 | 17.0 min | Multi-scale evaluation |
| VideoVista | - | - | - | Long video understanding |
| EgoSchema | - | - | 180 sec | Egocentric video reasoning |
| LongVideoBench | - | - | - | Reference-based evaluation |
| MLVU | - | - | - | Multi-task long video understanding |
| HourVideo | 500 | 12,976 | 45.7 min | Hour-level understanding |
| HLV-1K | 1,009 | 14,847 | 55.0 min | Comprehensive evaluation |
| LVBench | 103 | 1,549 | 68.4 min | Long-form analysis |
This survey analyzes how multimodal large language models process long videos through different architectural components:
graph LR
A[Video Input] --> B[Visual Encoder]
A --> C[Temporal Modeling]
A --> D[Language Integration]
B --> B1[Frame Features]
B --> B2[Spatial Attention]
C --> C1[Temporal Attention]
C --> C2[Memory Mechanisms]
D --> D1[Cross-modal Fusion]
D --> D2[Language Generation]
๐ Key Insights:
| Reasoning Type | Complexity | Representative Models | Performance Range |
|---|---|---|---|
| Frame-level Events | Low | Most MM-LLMs | 85-95% |
| Short-term Patterns | Medium | Video-LLaVA, TimeChat | 75-85% |
| Long-term Dependencies | High | MovieChat, LongVA | 65-80% |
| Cross-temporal Relations | Very High | LONGVILA, NVILA | 60-75% |
flowchart TD
A[Multimodal Input] --> B{Fusion Strategy}
B --> C[Early Fusion]
B --> D[Late Fusion]
B --> E[Hierarchical Fusion]
C --> C1[Feature Concatenation]
C --> C2[Cross-modal Attention]
D --> D1[Independent Processing]
D --> D2[Decision Combination]
E --> E1[Multi-level Integration]
E --> E2[Adaptive Weighting]
Key Findings: Hierarchical fusion strategies show better performance for long video understanding tasks.
๐ Memory-Augmented Models (15+ models)
โโโ ๐ฌ Sparse Memory (MovieChat, MA-LMM)
โโโ ๐ Sliding Windows (TimeChat, LLaMA-VID)
โโโ ๐ Dynamic Compression (Video-XL, Oryx-1.5)
๐ Efficiency Techniques
โโโ ๐ Token Merging (LongVLM, Video-LLaVA)
โโโ ๐ Hierarchical Processing (SlowFast-LLaVA)
โโโ ๐ Parallel Processing (LONGVILA)
โโโ ๐ Adaptive Pooling (PLLaVA, VideoGPT+)
๐ง Connector Types
โโโ ๐ค Q-Former Based (MovieChat, TimeChat)
โโโ ๐ Cross-Attention (Qwen-VL, EVLM)
โโโ ๐ MLP Projectors (VITA, LLaVA-OneVision)
โโโ ๐ง Advanced Fusion (Kangaroo, NVILA)
| Strategy | Models | Advantages | Challenges |
|---|---|---|---|
| End-to-End | MovieChat, MA-LMM | Optimal performance | High computational cost |
| Stage-wise | Video-LLaVA, TimeChat | Stable training | Suboptimal alignment |
| Hybrid | LongVA, LONGVILA | Balanced approach | Complex implementation |
Based on emerging trends from recent research, the following developments are expected:
Based on current challenges and limitations in long video understanding, several key research directions emerge:
If you find our survey useful in your research, please consider citing:
@article{zou2024seconds,
title={From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding},
author={Zou, Heqing and Luo, Tianze and Xie, Guiyang and Lv, Fengmao and Wang, Guangcong and Chen, Juanyang and Wang, Zhuochen and Zhang, Hansheng and Zhang, Huaijian and others},
journal={arXiv preprint arXiv:2409.18938},
year={2024}
}
We welcome contributions to this survey! Here's how you can help:
git checkout -b feature/new-model)git commit -am 'Add new model: ModelName')git push origin feature/new-model)This project is licensed under the MIT License - see the LICENSE file for details.