Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
📄 Paper | 🌐 Project Page | 💻 Code | 🤗 Models
This dataset contains the full training data used for Video Streaming Thinking (VST), including both supervised fine-tuning (SFT) and reinforcement learning (RL) stages.
| Subset | Description |
|---|---|
vst_sft_data | SFT data including video-text pairs from multiple sources |
vst_rl_data | RL data for reinforcement learning stage |
Also available on ModelScope.
| Model | Link |
|---|---|
| VST-3B | 🤗 Catalan258/VST-3B |
| VST-7B | 🤗 Catalan258/VST-7B |
| VST-32B | 🤗 Catalan258/VST-32B |
@article{guan2026videostreamingthinking,
title={Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously},
author={Yiran Guan and Liang Yin and Dingkang Liang and Jianzhong Ju and Zhenbo Luo and Jian Luan and Yuliang Liu and Xiang Bai},
journal={arXiv preprint arXiv:2603.12262},
year={2026},
}
39 commits
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
📄 Paper | 🌐 Project Page | 💻 Code | 🤗 Models
This dataset contains the full training data used for Video Streaming Thinking (VST), including both supervised fine-tuning (SFT) and reinforcement learning (RL) stages.
| Subset | Description |
|---|---|
vst_sft_data | SFT data including video-text pairs from multiple sources |
vst_rl_data | RL data for reinforcement learning stage |
Also available on ModelScope.
| Model | Link |
|---|---|
| VST-3B | 🤗 Catalan258/VST-3B |
| VST-7B | 🤗 Catalan258/VST-7B |
| VST-32B | 🤗 Catalan258/VST-32B |
@article{guan2026videostreamingthinking,
title={Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously},
author={Yiran Guan and Liang Yin and Dingkang Liang and Jianzhong Ju and Zhenbo Luo and Jian Luan and Yuliang Liu and Xiang Bai},
journal={arXiv preprint arXiv:2603.12262},
year={2026},
}
39 commits