๐ MMR1-SFT: Long Chain-of-Thought Cold-Start Dataset
3
9 commits
1 linked in READMEs
updated Sep 30, 2025
MMR1-SFT is a large-scale, carefully curated visionโlanguage long chain-of-thought (CoT) dataset for cold-start supervised fine-tuning of multimodal reasoning models.
It accompanies our work on Variance-Aware Sampling (VAS) for RL post-training and the MMR1 model family.
๐ Paper: MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
| Domain | #Samples |
|---|---|
| Math | ~1.6M |
| Science | ~13K |
| Chart-Figure | ~8K |
| Doc-Table | ~8K |
| General | ~8K |
The sft.json file contains an array of 1,620,490 training examples in the following structured format:
{
"conversations": [
{
"from": "human",
"value": "<instruction_template_with_image_tag>"
},
{
"from": "gpt",
"value": "<think>detailed_reasoning_process</think><answer>final_answer</answer>"
}
],
"images": [
"path/to/image.png"
]
}
<image> placeholder and question<think>reasoning</think><answer>result</answer> formatAll assistant responses follow a consistent chain-of-thought structure:
<think>...</think>: Detailed step-by-step reasoning process with error checking and corrections<answer>...</answer>: Final concise answer to the questionThis format enables training models for structured reasoning with explicit thought processes before providing final answers.
If you find MMR1 useful for your research and applications, please cite using this BibTeX:
@misc{leng2025mmr1,
title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources},
author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
year={2025},
eprint={2509.21268},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.21268},
}
๐ MMR1-SFT: Long Chain-of-Thought Cold-Start Dataset
3
9 commits
1 linked in READMEs
updated Sep 30, 2025
MMR1-SFT is a large-scale, carefully curated visionโlanguage long chain-of-thought (CoT) dataset for cold-start supervised fine-tuning of multimodal reasoning models.
It accompanies our work on Variance-Aware Sampling (VAS) for RL post-training and the MMR1 model family.
๐ Paper: MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
| Domain | #Samples |
|---|---|
| Math | ~1.6M |
| Science | ~13K |
| Chart-Figure | ~8K |
| Doc-Table | ~8K |
| General | ~8K |
The sft.json file contains an array of 1,620,490 training examples in the following structured format:
{
"conversations": [
{
"from": "human",
"value": "<instruction_template_with_image_tag>"
},
{
"from": "gpt",
"value": "<think>detailed_reasoning_process</think><answer>final_answer</answer>"
}
],
"images": [
"path/to/image.png"
]
}
<image> placeholder and question<think>reasoning</think><answer>result</answer> formatAll assistant responses follow a consistent chain-of-thought structure:
<think>...</think>: Detailed step-by-step reasoning process with error checking and corrections<answer>...</answer>: Final concise answer to the questionThis format enables training models for structured reasoning with explicit thought processes before providing final answers.
If you find MMR1 useful for your research and applications, please cite using this BibTeX:
@misc{leng2025mmr1,
title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources},
author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
year={2025},
eprint={2509.21268},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.21268},
}