:fire: :fire: :fire: Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding :video_camera:
74
15 commits
updated Sep 1, 2025
🔥🔥🔥 Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding
| Title | Venue | Date | Code | Frames |
|---|---|---|---|---|
[Benchmark] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? | arXiv | 2025-01 | Github | Streaming |
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction | arXiv | 2025-01 | Github | Streaming |
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction | arXiv | 2025-01 | Github | Streaming |
| Streaming long video understanding with large language models | arXiv | 2024-05 | - | 16(Streaming) |
| Title | Venue | Date | Code | Frames |
|---|---|---|---|---|
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation | arXiv | 2025-03 | Github | - |
T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs | arXiv | 2024-12 | Github | - |
15 commits
:fire: :fire: :fire: Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding :video_camera:
74
15 commits
updated Sep 1, 2025
🔥🔥🔥 Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding
| Title | Venue | Date | Code | Frames |
|---|---|---|---|---|
[Benchmark] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? | arXiv | 2025-01 | Github | Streaming |
Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction | arXiv | 2025-01 | Github | Streaming |
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction | arXiv | 2025-01 | Github | Streaming |
| Streaming long video understanding with large language models | arXiv | 2024-05 | - | 16(Streaming) |
| Title | Venue | Date | Code | Frames |
|---|---|---|---|---|
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation | arXiv | 2025-03 | Github | - |
T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs | arXiv | 2024-12 | Github | - |
15 commits