Project Page | Paper | GitHub | Model Checkpoint
β οΈ Important Notice: Research-Only Use
Before downloading or using this dataset, you must agree to the LICENSE terms.
This dataset contains video data that may be copyright-sensitive.
It is provided solely for non-commercial research and educational purposes.
By accessing this dataset, you confirm that you understand and agree to the terms in the LICENSE file.
For flexible real-time interaction, we introduce a comprehensive streaming video dataset TimeChat-Online-139K with backward-tracing, real-time visual perception, and future-responding scenarios.
.tar.gz archives of raw video frames extracted from diverse video sources. Each archive corresponds to a dataset.Statistics
We now provide video frame caption annotations in annotations_caption_flt.jsonl. This file contains detailed frame-level captions for video key frames.
Each line in the JSONL file represents a single frame annotation with the following structure:
{
"frame_id": 142,
"segment_id": 146,
"timestamp": 393.3,
"caption": "This frame presents a stark contrast to the preceding one...",
"video": "Youcook2/gZuDMKXWU_E"
}
frame_id (int): Sequential key frame identifier within the videosegment_id (int): Key segment identifiertimestamp (float): Video timestamp in seconds of the key framecaption (str): Detailed textual description of the visual content in the frame via GPT-4ovideo (str): Video identifier in the format {dataset}/{video_id}The annotation file provides comprehensive coverage with 876,398 frame-level detailed captions across approximately 10,949 videos from multiple datasets. Each frame caption averages 176 words in length, with an average of 87.8 key frames per video. This extensive collection offers rich visual descriptions ideal for training and research in video understanding tasks.
The dataset consists of 11,043 videos sampled from the following 13 public video datasets:
| Dataset | #Videos | Dataset | #Videos | Dataset | #Videos |
|---|---|---|---|---|---|
| COIN [52] | 151 | QV-Highlights [23] | 1778 | ActivityNet [15] | 12 |
| HD-VILA [64] | 695 | YouCook2 [83] | 710 | TVSum [50] | 10 |
| ViTT [20] | 2000 | QuerYD [40] | 566 | YouMakeup [56] | 1801 |
| VideoIC [55] | 2649 | Movie101 [71] | 202 | HiREST [72] | 469 |
Total videos: 11,043
Please refer to the original dataset papers and licenses:
This dataset is released under a custom research-only license. By accessing it, you agree:
For full terms, see LICENSE. Contact us if you're unsure about permitted uses.
If you use this dataset in your research, please cite:
@misc{timechatonline,
title={TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos},
author={Linli Yao and Yicheng Li and Yuancheng Wei and Lei Li and Shuhuai Ren and Yuanxin Liu and Kun Ouyang and Lean Wang and Shicheng Li and Sida Li and Lingpeng Kong and Qi Liu and Yuanxing Zhang and Xu Sun},
year={2025},
eprint={2504.17343},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2504.17343},
}
Project Page | Paper | GitHub | Model Checkpoint
β οΈ Important Notice: Research-Only Use
Before downloading or using this dataset, you must agree to the LICENSE terms.
This dataset contains video data that may be copyright-sensitive.
It is provided solely for non-commercial research and educational purposes.
By accessing this dataset, you confirm that you understand and agree to the terms in the LICENSE file.
For flexible real-time interaction, we introduce a comprehensive streaming video dataset TimeChat-Online-139K with backward-tracing, real-time visual perception, and future-responding scenarios.
.tar.gz archives of raw video frames extracted from diverse video sources. Each archive corresponds to a dataset.Statistics
We now provide video frame caption annotations in annotations_caption_flt.jsonl. This file contains detailed frame-level captions for video key frames.
Each line in the JSONL file represents a single frame annotation with the following structure:
{
"frame_id": 142,
"segment_id": 146,
"timestamp": 393.3,
"caption": "This frame presents a stark contrast to the preceding one...",
"video": "Youcook2/gZuDMKXWU_E"
}
frame_id (int): Sequential key frame identifier within the videosegment_id (int): Key segment identifiertimestamp (float): Video timestamp in seconds of the key framecaption (str): Detailed textual description of the visual content in the frame via GPT-4ovideo (str): Video identifier in the format {dataset}/{video_id}The annotation file provides comprehensive coverage with 876,398 frame-level detailed captions across approximately 10,949 videos from multiple datasets. Each frame caption averages 176 words in length, with an average of 87.8 key frames per video. This extensive collection offers rich visual descriptions ideal for training and research in video understanding tasks.
The dataset consists of 11,043 videos sampled from the following 13 public video datasets:
| Dataset | #Videos | Dataset | #Videos | Dataset | #Videos |
|---|---|---|---|---|---|
| COIN [52] | 151 | QV-Highlights [23] | 1778 | ActivityNet [15] | 12 |
| HD-VILA [64] | 695 | YouCook2 [83] | 710 | TVSum [50] | 10 |
| ViTT [20] | 2000 | QuerYD [40] | 566 | YouMakeup [56] | 1801 |
| VideoIC [55] | 2649 | Movie101 [71] | 202 | HiREST [72] | 469 |
Total videos: 11,043
Please refer to the original dataset papers and licenses:
This dataset is released under a custom research-only license. By accessing it, you agree:
For full terms, see LICENSE. Contact us if you're unsure about permitted uses.
If you use this dataset in your research, please cite:
@misc{timechatonline,
title={TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos},
author={Linli Yao and Yicheng Li and Yuancheng Wei and Lei Li and Shuhuai Ren and Yuanxin Liu and Kun Ouyang and Lean Wang and Shicheng Li and Sida Li and Lingpeng Kong and Qi Liu and Yuanxing Zhang and Xu Sun},
year={2025},
eprint={2504.17343},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2504.17343},
}