zhang9302002/MultiTaskVideoReasoning

Dataset

Multi Task Video Reasoning Dataset

6

6 commits

6 linked in READMEs

updated Aug 25, 2025

See the code

README

Multi Task Video Reasoning Dataset

This is the official training dataset for Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

[Project] [arXiv] [Code]

Data Structure

└── MultiTaskVideoReasoning
    β”œβ”€β”€ MTVR_CoT
    β”‚   β”œβ”€β”€ actnet.json
    β”‚   β”œβ”€β”€ charades.json
    β”‚   β”œβ”€β”€ longvideo-reason.json
    β”‚   β”œβ”€β”€ nextgqa.json
    β”‚   β”œβ”€β”€ rextime.json
    β”‚   β”œβ”€β”€ vidchapters.json
    β”‚   β”œβ”€β”€ Video-R1-data-image.json
    β”‚   └── Video-R1-data-video.json
    β”œβ”€β”€ MTVR_RL
    β”‚   β”œβ”€β”€ actnet.json
    β”‚   β”œβ”€β”€ charades.json
    β”‚   β”œβ”€β”€ nextgqa.json
    β”‚   β”œβ”€β”€ rextime.json
    β”‚   β”œβ”€β”€ vidchapters.json
    β”‚   β”œβ”€β”€ Video-R1-data-image.json
    β”‚   └── Video-R1-data-video.json
    β”œβ”€β”€ MTVR_Tool_CoT
    β”‚   β”œβ”€β”€ longvideo-reason.json
    β”‚   └── vidchapters.json
    β”œβ”€β”€ MTVR_Tool_RL
    β”‚   β”œβ”€β”€ longvideo-reason.json
    β”‚   └── vidchapters.json
    └── README.md

Data Preperation

Please download original videos from these sources.

Citation

If you use this dataset or find it useful, please cite our paper:

@article{zhang2025thinking,
  title={Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning},
  author={Zhang, Haoji and Gu, Xin and Li, Jiawen and Ma, Chixiang and Bai, Sule and Zhang, Chubin and Zhang, Bowen and Zhou, Zhichao and He, Dongliang and Tang, Yansong},
  journal={arXiv preprint arXiv:2508.04416},
  year={2025}
}

Contributors

ZH
zhanghaoji.1

5 commits

zhang9302002

1 commits

zhang9302002/MultiTaskVideoReasoning

Dataset

Multi Task Video Reasoning Dataset

6

6 commits

6 linked in READMEs

updated Aug 25, 2025

See the code

README

Multi Task Video Reasoning Dataset

This is the official training dataset for Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

[Project] [arXiv] [Code]

Data Structure

└── MultiTaskVideoReasoning
    β”œβ”€β”€ MTVR_CoT
    β”‚   β”œβ”€β”€ actnet.json
    β”‚   β”œβ”€β”€ charades.json
    β”‚   β”œβ”€β”€ longvideo-reason.json
    β”‚   β”œβ”€β”€ nextgqa.json
    β”‚   β”œβ”€β”€ rextime.json
    β”‚   β”œβ”€β”€ vidchapters.json
    β”‚   β”œβ”€β”€ Video-R1-data-image.json
    β”‚   └── Video-R1-data-video.json
    β”œβ”€β”€ MTVR_RL
    β”‚   β”œβ”€β”€ actnet.json
    β”‚   β”œβ”€β”€ charades.json
    β”‚   β”œβ”€β”€ nextgqa.json
    β”‚   β”œβ”€β”€ rextime.json
    β”‚   β”œβ”€β”€ vidchapters.json
    β”‚   β”œβ”€β”€ Video-R1-data-image.json
    β”‚   └── Video-R1-data-video.json
    β”œβ”€β”€ MTVR_Tool_CoT
    β”‚   β”œβ”€β”€ longvideo-reason.json
    β”‚   └── vidchapters.json
    β”œβ”€β”€ MTVR_Tool_RL
    β”‚   β”œβ”€β”€ longvideo-reason.json
    β”‚   └── vidchapters.json
    └── README.md

Data Preperation

Please download original videos from these sources.

Citation

If you use this dataset or find it useful, please cite our paper:

@article{zhang2025thinking,
  title={Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning},
  author={Zhang, Haoji and Gu, Xin and Li, Jiawen and Ma, Chixiang and Bai, Sule and Zhang, Chubin and Zhang, Bowen and Zhou, Zhichao and He, Dongliang and Tang, Yansong},
  journal={arXiv preprint arXiv:2508.04416},
  year={2025}
}

Contributors

ZH
zhanghaoji.1

5 commits

zhang9302002

1 commits