StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. π
[NEW! 2025.05.15] π₯: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] β: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the Omni-Source Understanding.
[NEW! 2025.01.14] π: MiniCPM-o 2.6 achieved Streaming SOTA with a score of 66.01 on the Overall benchmark.
[NEW! 2025.01.06] π: Dispider achieved Streaming SOTA with a score of 53.12 on the Overall benchmark.
[NEW! 2024.12.09] π: InternLM-XComposer2.5-OmniLive achieved 73.79 on Real-Time Visual Understanding.
As MLLMs continue to advance, they remain largely focused on offline video comprehension, where all frames are pre-loaded before making queries. However, this is far from the human ability to process and respond to video streams in real-time, capturing the dynamic nature of multimedia content. To bridge this gap, StreamingBench introduces the first comprehensive benchmark for streaming video understanding in MLLMs.
"β€ xs" means that the answer is considered correct if the actual output time is within x seconds of the ground truth.
@article{lin2024streaming,
title={StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding},
author={Junming Lin and Zheng Fang and Chi Chen and Zihao Wan and Fuwen Luo and Peng Li and Yang Liu and Maosong Sun},
journal={arXiv preprint arXiv:2411.03628},
year={2024}
}
122 commits
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. π
[NEW! 2025.05.15] π₯: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] β: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the Omni-Source Understanding.
[NEW! 2025.01.14] π: MiniCPM-o 2.6 achieved Streaming SOTA with a score of 66.01 on the Overall benchmark.
[NEW! 2025.01.06] π: Dispider achieved Streaming SOTA with a score of 53.12 on the Overall benchmark.
[NEW! 2024.12.09] π: InternLM-XComposer2.5-OmniLive achieved 73.79 on Real-Time Visual Understanding.
As MLLMs continue to advance, they remain largely focused on offline video comprehension, where all frames are pre-loaded before making queries. However, this is far from the human ability to process and respond to video streams in real-time, capturing the dynamic nature of multimedia content. To bridge this gap, StreamingBench introduces the first comprehensive benchmark for streaming video understanding in MLLMs.
"β€ xs" means that the answer is considered correct if the actual output time is within x seconds of the ground truth.
@article{lin2024streaming,
title={StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding},
author={Junming Lin and Zheng Fang and Chi Chen and Zihao Wan and Fuwen Luo and Peng Li and Yang Liu and Maosong Sun},
journal={arXiv preprint arXiv:2411.03628},
year={2024}
}
122 commits