This dataset provides a comprehensive collection for Online Spatial-Temporal Understanding tasks, covering multiple domains including Dense Video Captioning, Video Grounding, Step Localization, Spatial-Temporal Action Localization, and Object Tracking.
Our pipeline begins with 96K high-quality samples curated from 5 tasks across 12 datasets. The conversion process enhances online spatiotemporal understanding through template transformation. We strategically insert queries along the timeline in an organized interleaved format for each video sample to facilitate temporal context differentiation.
| Category | Dataset | Count | Query | Response |
|---|---|---|---|---|
| Temporal Grounding | DiDeMo | 33,002 | Identify whether a specific event is still ongoing at present or has it concluded. Provide the start time of the event and its duration up to the query timestamp. | <start time> - <event duration>: duration up to query timestamp. |
| QuerYD | 14,620 | |||
| HiREST | 459 | |||
| Charades-STA | 12,408 | |||
| Object Tracking | LaSOT | 1,120 | Track the object currently based on a brief description or box. | (1) Past trajectory up to the present with brief descriptions; (2) Track the object sequentially in future frames as they become available. |
| GOT10k | 8,250 | |||
| Step Localization and Captioning | COIN | 9,029 | List steps completed up to the current point, excluding previously reported ones. | <start time> - <end time>, <step description>... |
| HiREST | 459 | |||
| Dense Video Captioning | ActivityNet Captions | 10,009 | Identify and list events up to the current point, excluding previously reported ones. | <start time> - <end time>, <event description>... |
| VITT | 5,141 | |||
| YouCook2 | 1,192 | |||
| Spatial Temporal Action Localization | AVA | 160 | Identify current and past actions of a person at a specific box at present. | List actions for the person over time, with corresponding positions. |
| Total number of datasets: | 96k |
{
"video": "116/NLy71UrHElw.mp4",
"conversations": [
{
"from": "human",
"timestamps": 1026.0, # Video timestamp in seconds
"value": "<video>\nBased on current observation, list events..."
},
{
"from": "gpt",
"value": "21.0s - 22.0s (duration: 1.0s), begin to run up..."
}
]
}
Format 2: Template-based Tracking
{
"video": "GOT-10k_Train_006457",
"fps": 1, # Frame rate
"all_image_files": ["00000001.jpg", ...], # Keyframe paths
"image_bboxes": [ # Temporal object tracking data
{
"timestamp": 0.0,
"bbox": [0.412, 0.517, 0.452, 0.753] # [x1,y1,x2,y2]
},
...
],
"query_template": { # Randomized temporal insertion
"from": "human",
"value": "Track the location of \"person\" at <bbox> over time..."
}
}
| Task | Dataset | Source |
|---|---|---|
| Dense Video Captioning | ActivityNet Captions | Source |
ViTT | Source | |
YouCook2 | Source | |
| Temporal Video Grounding | DiDeMo | Source |
QuerYD | Source | |
HiREST_grounding | Source | |
Charades-STA | Source | |
| Step Localization | COIN | Source |
HiREST_step | Source | |
| Spatial Temporal Action Localization | AVA | Source |
| Object Tracking | GOT 10K | Source |
LaSOT | Source |
If you find this project useful in your research, please consider cite:
@article{huang2024online,
title={Online Video Understanding: A Comprehensive Benchmark and Memory-Augmented Method},
author={Huang, Zhenpeng and Li, Xinhao and Li, Jiaqi and Wang, Jing and Zeng, Xiangyu and Liang, Cheng and Wu, Tao and Chen, Xi and Li, Liang and Wang, Limin},
journal={arXiv preprint arXiv:2501.00584},
year={2024}
}
4 commits
This dataset provides a comprehensive collection for Online Spatial-Temporal Understanding tasks, covering multiple domains including Dense Video Captioning, Video Grounding, Step Localization, Spatial-Temporal Action Localization, and Object Tracking.
Our pipeline begins with 96K high-quality samples curated from 5 tasks across 12 datasets. The conversion process enhances online spatiotemporal understanding through template transformation. We strategically insert queries along the timeline in an organized interleaved format for each video sample to facilitate temporal context differentiation.
| Category | Dataset | Count | Query | Response |
|---|---|---|---|---|
| Temporal Grounding | DiDeMo | 33,002 | Identify whether a specific event is still ongoing at present or has it concluded. Provide the start time of the event and its duration up to the query timestamp. | <start time> - <event duration>: duration up to query timestamp. |
| QuerYD | 14,620 | |||
| HiREST | 459 | |||
| Charades-STA | 12,408 | |||
| Object Tracking | LaSOT | 1,120 | Track the object currently based on a brief description or box. | (1) Past trajectory up to the present with brief descriptions; (2) Track the object sequentially in future frames as they become available. |
| GOT10k | 8,250 | |||
| Step Localization and Captioning | COIN | 9,029 | List steps completed up to the current point, excluding previously reported ones. | <start time> - <end time>, <step description>... |
| HiREST | 459 | |||
| Dense Video Captioning | ActivityNet Captions | 10,009 | Identify and list events up to the current point, excluding previously reported ones. | <start time> - <end time>, <event description>... |
| VITT | 5,141 | |||
| YouCook2 | 1,192 | |||
| Spatial Temporal Action Localization | AVA | 160 | Identify current and past actions of a person at a specific box at present. | List actions for the person over time, with corresponding positions. |
| Total number of datasets: | 96k |
{
"video": "116/NLy71UrHElw.mp4",
"conversations": [
{
"from": "human",
"timestamps": 1026.0, # Video timestamp in seconds
"value": "<video>\nBased on current observation, list events..."
},
{
"from": "gpt",
"value": "21.0s - 22.0s (duration: 1.0s), begin to run up..."
}
]
}
Format 2: Template-based Tracking
{
"video": "GOT-10k_Train_006457",
"fps": 1, # Frame rate
"all_image_files": ["00000001.jpg", ...], # Keyframe paths
"image_bboxes": [ # Temporal object tracking data
{
"timestamp": 0.0,
"bbox": [0.412, 0.517, 0.452, 0.753] # [x1,y1,x2,y2]
},
...
],
"query_template": { # Randomized temporal insertion
"from": "human",
"value": "Track the location of \"person\" at <bbox> over time..."
}
}
| Task | Dataset | Source |
|---|---|---|
| Dense Video Captioning | ActivityNet Captions | Source |
ViTT | Source | |
YouCook2 | Source | |
| Temporal Video Grounding | DiDeMo | Source |
QuerYD | Source | |
HiREST_grounding | Source | |
Charades-STA | Source | |
| Step Localization | COIN | Source |
HiREST_step | Source | |
| Spatial Temporal Action Localization | AVA | Source |
| Object Tracking | GOT 10K | Source |
LaSOT | Source |
If you find this project useful in your research, please consider cite:
@article{huang2024online,
title={Online Video Understanding: A Comprehensive Benchmark and Memory-Augmented Method},
author={Huang, Zhenpeng and Li, Xinhao and Li, Jiaqi and Wang, Jing and Zeng, Xiangyu and Liang, Cheng and Wu, Tao and Chen, Xi and Li, Liang and Wang, Limin},
journal={arXiv preprint arXiv:2501.00584},
year={2024}
}
4 commits