pipixin321/Awesome-Video-MLLMs

:fire: :fire: :fire: Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding :video_camera:

74

15 commits

updated Sep 1, 2025

See the code

README

Awesome-Video-Multimodal-Large-Language-Models Awesome Awesome MLLM GitHub last commit

SVG Banners

🔥🔥🔥 Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding

Welcome to stars ⭐ & comments 😀 & sharing :chart_with_upwards_trend: !!

📖 Contents


General Works

TitleVenueDateCodeFrames/FPS
Star
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
arXiv2025-08Github-
MiniCPM-V 4.5: A GPT-4o Level MLLM for Single Image, Multi Image and Video Understanding on Your Phone-2025-08Github10fps
GLM-4.5VZhipu AI2025-08API-
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGoogle2025-06--
ERNIE 4.5Baidu2025-06Github-
Seed1.5-VL Technical ReportarXiv2025-05--
Star
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
arXiv2025-04Github-
Star
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
arXiv2025-04Github-
Star
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuray
arXiv2025-02Github-
Star
Qwen2.5-VL Technical Report
arXiv2025-02Github-
Apollo: An Exploration of Video Understanding in Large Multimodal ModelsarXiv2024-12--
Star
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
arXiv2024-11Github128
Star
ARIA : An Open Multimodal Native Mixture-of-Experts Model
arXiv2024-10Github256
Star
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
arXiv2024-10Github768(2fps)
Star
LLaVA-Video: Video Instruction Tuning With Synthetic Data
arXiv2024-10Github64
Star
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
arXiv2024-10Github1FPS
Star
MiniCPM-V 2.6: A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone
arXiv2023-08Github64
Star
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
arXiv2024-08Github128
Star
InternVL2: Better than the Best—Expanding Performance Boundaries of Open-Source Multimodal Models with the Progressive Scaling Strategy
blog2024-07Github16
Star
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
arXiv2024-06Github32
Star
ShareGPT4Video: Improving video understanding and generation with better captions
arXiv2024-06Github16
Star
LongVA: Long context transfer from language to vision
arXiv2024-06Github1FPS
Star
LongVLM: Efficient long video understanding via large language models
ECCV2024-04Github100
Star
VILA: On Pre-training for Visual Language Models
CVPR2023-12Github8
Star
TimeChat: A time-sensitive multimodal large language model for long video understanding
CVPR2023-12Github96
Star
Chat-UniVi unified visual representation empowers large language models with image and video understanding
CVPR2023-11Github64
Star
VTimeLLM: Empower LLM to Grasp Video Moments
CVPR2023-11Github100
Star
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
ECCV2023-11Github1FPS
Star
Video-LLaVA: Learning united visual representation by alignment before projection
arXiv2023-11Github8
Star
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
arXiv2023-07Github2048
Star
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
ACL2023-06Github100
Star
VALLEY: Video Assistant with Large Language model Enhanced ability
arXiv2023-06Github0.5FPS
Star
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
EMNLP2023-06Github8
Star
VideoChat: Chat-Centric Video Understanding
arXiv2023-05Github4~32
Star
LLaMA-Adapter: Efficient Fine-tuning of LLaMA
ICLR2023-03Github-

Streaming Videos

Interesting Works

Benchmarks for Evaluation

General

:bar_chart: Opencomprass Leaderboard
awesome-list
benchmarks
large-language-models
mllm
multimodal-instruction-tuning
multimodal-video-understanding
video-understanding

Contributors

pipixin321

15 commits

pipixin321/Awesome-Video-MLLMs

:fire: :fire: :fire: Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding :video_camera:

74

15 commits

updated Sep 1, 2025

See the code

README

Awesome-Video-Multimodal-Large-Language-Models Awesome Awesome MLLM GitHub last commit

SVG Banners

🔥🔥🔥 Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding

Welcome to stars ⭐ & comments 😀 & sharing :chart_with_upwards_trend: !!

📖 Contents


General Works

TitleVenueDateCodeFrames/FPS
Star
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
arXiv2025-08Github-
MiniCPM-V 4.5: A GPT-4o Level MLLM for Single Image, Multi Image and Video Understanding on Your Phone-2025-08Github10fps
GLM-4.5VZhipu AI2025-08API-
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic CapabilitiesGoogle2025-06--
ERNIE 4.5Baidu2025-06Github-
Seed1.5-VL Technical ReportarXiv2025-05--
Star
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
arXiv2025-04Github-
Star
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
arXiv2025-04Github-
Star
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuray
arXiv2025-02Github-
Star
Qwen2.5-VL Technical Report
arXiv2025-02Github-
Apollo: An Exploration of Video Understanding in Large Multimodal ModelsarXiv2024-12--
Star
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
arXiv2024-11Github128
Star
ARIA : An Open Multimodal Native Mixture-of-Experts Model
arXiv2024-10Github256
Star
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
arXiv2024-10Github768(2fps)
Star
LLaVA-Video: Video Instruction Tuning With Synthetic Data
arXiv2024-10Github64
Star
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
arXiv2024-10Github1FPS
Star
MiniCPM-V 2.6: A GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone
arXiv2023-08Github64
Star
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
arXiv2024-08Github128
Star
InternVL2: Better than the Best—Expanding Performance Boundaries of Open-Source Multimodal Models with the Progressive Scaling Strategy
blog2024-07Github16
Star
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
arXiv2024-06Github32
Star
ShareGPT4Video: Improving video understanding and generation with better captions
arXiv2024-06Github16
Star
LongVA: Long context transfer from language to vision
arXiv2024-06Github1FPS
Star
LongVLM: Efficient long video understanding via large language models
ECCV2024-04Github100
Star
VILA: On Pre-training for Visual Language Models
CVPR2023-12Github8
Star
TimeChat: A time-sensitive multimodal large language model for long video understanding
CVPR2023-12Github96
Star
Chat-UniVi unified visual representation empowers large language models with image and video understanding
CVPR2023-11Github64
Star
VTimeLLM: Empower LLM to Grasp Video Moments
CVPR2023-11Github100
Star
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
ECCV2023-11Github1FPS
Star
Video-LLaVA: Learning united visual representation by alignment before projection
arXiv2023-11Github8
Star
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
arXiv2023-07Github2048
Star
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
ACL2023-06Github100
Star
VALLEY: Video Assistant with Large Language model Enhanced ability
arXiv2023-06Github0.5FPS
Star
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
EMNLP2023-06Github8
Star
VideoChat: Chat-Centric Video Understanding
arXiv2023-05Github4~32
Star
LLaMA-Adapter: Efficient Fine-tuning of LLaMA
ICLR2023-03Github-

Streaming Videos

Interesting Works

Benchmarks for Evaluation

General

:bar_chart: Opencomprass Leaderboard
awesome-list
benchmarks
large-language-models
mllm
multimodal-instruction-tuning
multimodal-video-understanding
video-understanding

Contributors

pipixin321

15 commits