This is the official repository for the paper "Thinking in Streaming Video".
Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios.
To address this, we introduce ThinkStream, a framework for streaming video reasoning based on a Watch-Think-Speak paradigm that enables models to incrementally update their understanding as new video observations arrive.
Experiments on multiple streaming video benchmarks show that ThinkStream significantly outperforms existing online video models while maintaining low latency and memory usage.
ThinkStream/
βββ scripts/ # Scripts for training, evaluation, and inference demos
β βββ eval/ # Evaluation scripts (OVO-Bench, StreamingBench)
β βββ demo.py # Inference demo code
β βββ rl.sh # Reinforcement Learning (RL) training script
β βββ sft.sh # Supervised Fine-Tuning (SFT) training script
βββ thinkstream/ # Core codebase
β βββ data/ # Data processing logic and dataset path configuration
β βββ eval/ # Evaluation code and format conversion scripts
β βββ model/ # Core model architecture, streaming attention mechanism, and inference engine
β βββ trainer/ # Training logic for SFT and RL
βββ train.py # Main training entry point
βββ requirements.txt # Python dependencies
βββ README.md
First, install the required dependencies:
pip install -r requirements.txt
Data Preparation:
Note: The dataset path configurations are located in thinkstream/data/__init__.py, which follows a similar logic to qwen-vl-finetune.
Run Training: Simply run the corresponding training scripts (please note you need to modify the model paths inside the scripts):
# Supervised Fine-Tuning (SFT)
./scripts/sft.sh
# Reinforcement Learning (RL)
./scripts/rl.sh
First, prepare the official datasets for OVO-Bench and StreamingBench.
Run the respective transfer_annotation_format.py scripts under the thinkstream/eval folder to convert the format:
thinkstream/eval/ovo_bench/transfer_annotation_format.pythinkstream/eval/rtvu/transfer_annotation_format.pyAfter conversion, start the evaluation script:
bash ./scripts/eval/eval.sh
Note: You need to change the model checkpoint (ckpt) path to your own path.
Use Python to run scripts/demo.py for inference testing.
Before running, please change MODEL_ID and VIDEO_PATH in the code to your own paths. Meanwhile, you need to manually fill in the content (question/instruction) and timestamp in the queries list.
Then, simply run:
python scripts/demo.py
You will see the output results in the command line.
We would like to thank the following open-source projects for their valuable contributions:
If you find this work helpful, you can cite the following papers:
@misc{liu2026thinkingstreamingvideo,
title={Thinking in Streaming Video},
author={Zikang Liu and Longteng Guo and Handong Li and Ru Zhen and Xingjian He and Ruyi Ji and Xiaoming Ren and Yanhao Zhang and Haonan Lu and Jing Liu},
year={2026},
eprint={2603.12938},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.12938},
}
This is the official repository for the paper "Thinking in Streaming Video".
Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios.
To address this, we introduce ThinkStream, a framework for streaming video reasoning based on a Watch-Think-Speak paradigm that enables models to incrementally update their understanding as new video observations arrive.
Experiments on multiple streaming video benchmarks show that ThinkStream significantly outperforms existing online video models while maintaining low latency and memory usage.
ThinkStream/
βββ scripts/ # Scripts for training, evaluation, and inference demos
β βββ eval/ # Evaluation scripts (OVO-Bench, StreamingBench)
β βββ demo.py # Inference demo code
β βββ rl.sh # Reinforcement Learning (RL) training script
β βββ sft.sh # Supervised Fine-Tuning (SFT) training script
βββ thinkstream/ # Core codebase
β βββ data/ # Data processing logic and dataset path configuration
β βββ eval/ # Evaluation code and format conversion scripts
β βββ model/ # Core model architecture, streaming attention mechanism, and inference engine
β βββ trainer/ # Training logic for SFT and RL
βββ train.py # Main training entry point
βββ requirements.txt # Python dependencies
βββ README.md
First, install the required dependencies:
pip install -r requirements.txt
Data Preparation:
Note: The dataset path configurations are located in thinkstream/data/__init__.py, which follows a similar logic to qwen-vl-finetune.
Run Training: Simply run the corresponding training scripts (please note you need to modify the model paths inside the scripts):
# Supervised Fine-Tuning (SFT)
./scripts/sft.sh
# Reinforcement Learning (RL)
./scripts/rl.sh
First, prepare the official datasets for OVO-Bench and StreamingBench.
Run the respective transfer_annotation_format.py scripts under the thinkstream/eval folder to convert the format:
thinkstream/eval/ovo_bench/transfer_annotation_format.pythinkstream/eval/rtvu/transfer_annotation_format.pyAfter conversion, start the evaluation script:
bash ./scripts/eval/eval.sh
Note: You need to change the model checkpoint (ckpt) path to your own path.
Use Python to run scripts/demo.py for inference testing.
Before running, please change MODEL_ID and VIDEO_PATH in the code to your own paths. Meanwhile, you need to manually fill in the content (question/instruction) and timestamp in the queries list.
Then, simply run:
python scripts/demo.py
You will see the output results in the command line.
We would like to thank the following open-source projects for their valuable contributions:
If you find this work helpful, you can cite the following papers:
@misc{liu2026thinkingstreamingvideo,
title={Thinking in Streaming Video},
author={Zikang Liu and Longteng Guo and Handong Li and Ru Zhen and Xingjian He and Ruyi Ji and Xiaoming Ren and Yanhao Zhang and Haonan Lu and Jing Liu},
year={2026},
eprint={2603.12938},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.12938},
}