[ECCV 26] Video Streaming Thinking
122
stars
20
commits
Python
primary language
Jul 28, 2026
updated
π VST has been accepted to ECCV 2026!
Video Streaming Thinking introduces a new paradigm for streaming video understanding that interleaves active reasoning with continuous video consumption, enabling amortized test-time scaling with real-time responsiveness.
Existing online VideoLLMs focus on efficient streaming perception but lack explicit analytical reasoning. Offline VideoLLMs with Chain-of-Thought (CoT) can reason deeply, but incur high query-answer (QA) latency that violates real-time constraints. VST bridges this gap by shifting the LLM backend from passive waiting to active, intermittent reasoning during video consumption, implementing a thinking-while-watching mechanism inspired by human neural coupling.
https://github.com/user-attachments/assets/49846db5-bf76-4cf8-b923-4b9b88117482
Instead of deferring all reasoning until a user query arrives, VST continuously processes incoming video clips and produces intermediate streaming thoughts in real time. This front-loads and amortizes the reasoning cost, so the final response is both deeply grounded and instantly available.
| Model | HuggingFace | OVO-Bench | StreamingBench | VideoMME | LongVideoBench | VideoHolmes |
|---|---|---|---|---|---|---|
| VST-3B | π€ Link | 56.2 | 75.5 | 59.5 | 54.1 | 36.1 |
| VST-7B | π€ Link | 59.3 | 79.5 | 64.9 | 58.0 | 41.9 |
| VST-32B | π€ Link | 63.5 | 80.7 | 67.2 | 60.7 | 45.1 |
We release the full training data used for both SFT and RL stages on HuggingFace and ModelScope:
| Dataset | HuggingFace | ModelScope | Description |
|---|---|---|---|
| vst_sft_data | π€ Link | π€ Link | SFT data including video-text pairs from multiple sources |
| vst_rl_data | π€ Link | π€ Link | RL data for reinforcement learning stage |
We thank the following great works and open-source repositories:
@inproceedings{guan2026videostreamingthinking,
title={Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously},
author={Yiran Guan and Liang Yin and Dingkang Liang and Jianzhong Ju and Zhenbo Luo and Jian Luan and Yuliang Liu and Xiang Bai},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026},
}
Python
98.3%
Jupyter Notebook
1.1%
[ECCV 26] Video Streaming Thinking
122
stars
20
commits
Python
primary language
Jul 28, 2026
updated
π VST has been accepted to ECCV 2026!
Video Streaming Thinking introduces a new paradigm for streaming video understanding that interleaves active reasoning with continuous video consumption, enabling amortized test-time scaling with real-time responsiveness.
Existing online VideoLLMs focus on efficient streaming perception but lack explicit analytical reasoning. Offline VideoLLMs with Chain-of-Thought (CoT) can reason deeply, but incur high query-answer (QA) latency that violates real-time constraints. VST bridges this gap by shifting the LLM backend from passive waiting to active, intermittent reasoning during video consumption, implementing a thinking-while-watching mechanism inspired by human neural coupling.
https://github.com/user-attachments/assets/49846db5-bf76-4cf8-b923-4b9b88117482
Instead of deferring all reasoning until a user query arrives, VST continuously processes incoming video clips and produces intermediate streaming thoughts in real time. This front-loads and amortizes the reasoning cost, so the final response is both deeply grounded and instantly available.
| Model | HuggingFace | OVO-Bench | StreamingBench | VideoMME | LongVideoBench | VideoHolmes |
|---|---|---|---|---|---|---|
| VST-3B | π€ Link | 56.2 | 75.5 | 59.5 | 54.1 | 36.1 |
| VST-7B | π€ Link | 59.3 | 79.5 | 64.9 | 58.0 | 41.9 |
| VST-32B | π€ Link | 63.5 | 80.7 | 67.2 | 60.7 | 45.1 |
We release the full training data used for both SFT and RL stages on HuggingFace and ModelScope:
| Dataset | HuggingFace | ModelScope | Description |
|---|---|---|---|
| vst_sft_data | π€ Link | π€ Link | SFT data including video-text pairs from multiple sources |
| vst_rl_data | π€ Link | π€ Link | RL data for reinforcement learning stage |
We thank the following great works and open-source repositories:
@inproceedings{guan2026videostreamingthinking,
title={Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously},
author={Yiran Guan and Liang Yin and Dingkang Liang and Jianzhong Ju and Zhenbo Luo and Jian Luan and Yuliang Liu and Xiang Bai},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026},
}
Python
98.3%
Jupyter Notebook
1.1%