MMR1/MMR1-SFT

Dataset

๐Ÿ“˜ MMR1-SFT: Long Chain-of-Thought Cold-Start Dataset

3

9 commits

1 linked in READMEs

updated Sep 30, 2025

See the code

README

๐Ÿ“˜ MMR1-SFT: Long Chain-of-Thought Cold-Start Dataset

arXiv Hugging Face

MMR1-SFT is a large-scale, carefully curated visionโ€“language long chain-of-thought (CoT) dataset for cold-start supervised fine-tuning of multimodal reasoning models.
It accompanies our work on Variance-Aware Sampling (VAS) for RL post-training and the MMR1 model family.

  • Scale: ~1.6M multimodal QA examples with verified long CoT rationales and short answers
  • Quality control: CoTs generated by Gemini-2.5 Pro/Flash, verified by GPT-4o
  • Domains: Math (majority), Science, Chart/Figure, Doc/Table, General

๐Ÿ“„ Paper: MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources


๐Ÿ” Summary

  • Purpose. Provide an open, reproducible cold-start corpus with long CoTs for reasoning-oriented MLLMs.
  • Intended use. SFT for format alignment and long-cot trajectories; warm-up for RL.

๐Ÿ“Š Data Sources & Curation

  • Prompt origins: Public VLM instruction datasets and exam-style corpora
  • Annotation: Long CoTs by Gemini-2.5 Pro/Flash, answers verified by GPT-4o
  • Filtering: Retain only medium/hard samples via pass rate
  • Balancing: Additional sourcing for Science, Chart/Figure, Doc/Table

โšก Domain Coverage

Domain#Samples
Math~1.6M
Science~13K
Chart-Figure~8K
Doc-Table~8K
General~8K

๐Ÿ“‹ Data Structure

The sft.json file contains an array of 1,620,490 training examples in the following structured format:

Schema Overview

{
  "conversations": [
    {
      "from": "human",
      "value": "<instruction_template_with_image_tag>"
    },
    {
      "from": "gpt",
      "value": "<think>detailed_reasoning_process</think><answer>final_answer</answer>"
    }
  ],
  "images": [
    "path/to/image.png"
  ]
}

Field Descriptions

  • conversations: Array containing the conversation turns
    • from: Role identifier ("human" or "gpt")
    • value: The content of the message
      • Human messages: Instruction template with <image> placeholder and question
      • Assistant messages: Structured response with <think>reasoning</think><answer>result</answer> format
  • images: Array of image file paths referenced in the conversation
    • Contains relative paths to associated visual content
    • Supports multimodal reasoning tasks with mathematical figures, charts, and diagrams

Response Format

All assistant responses follow a consistent chain-of-thought structure:

  • <think>...</think>: Detailed step-by-step reasoning process with error checking and corrections
  • <answer>...</answer>: Final concise answer to the question

This format enables training models for structured reasoning with explicit thought processes before providing final answers.



๐Ÿ“Œ Citation

If you find MMR1 useful for your research and applications, please cite using this BibTeX:

@misc{leng2025mmr1,
  title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources}, 
  author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
  year={2025},
  eprint={2509.21268},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2509.21268}, 
}
chain-of-thought
multimodal
reasoning
reinforcement-learning-support
vision-language

MMR1/MMR1-SFT

Dataset

๐Ÿ“˜ MMR1-SFT: Long Chain-of-Thought Cold-Start Dataset

3

9 commits

1 linked in READMEs

updated Sep 30, 2025

See the code

README

๐Ÿ“˜ MMR1-SFT: Long Chain-of-Thought Cold-Start Dataset

arXiv Hugging Face

MMR1-SFT is a large-scale, carefully curated visionโ€“language long chain-of-thought (CoT) dataset for cold-start supervised fine-tuning of multimodal reasoning models.
It accompanies our work on Variance-Aware Sampling (VAS) for RL post-training and the MMR1 model family.

  • Scale: ~1.6M multimodal QA examples with verified long CoT rationales and short answers
  • Quality control: CoTs generated by Gemini-2.5 Pro/Flash, verified by GPT-4o
  • Domains: Math (majority), Science, Chart/Figure, Doc/Table, General

๐Ÿ“„ Paper: MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources


๐Ÿ” Summary

  • Purpose. Provide an open, reproducible cold-start corpus with long CoTs for reasoning-oriented MLLMs.
  • Intended use. SFT for format alignment and long-cot trajectories; warm-up for RL.

๐Ÿ“Š Data Sources & Curation

  • Prompt origins: Public VLM instruction datasets and exam-style corpora
  • Annotation: Long CoTs by Gemini-2.5 Pro/Flash, answers verified by GPT-4o
  • Filtering: Retain only medium/hard samples via pass rate
  • Balancing: Additional sourcing for Science, Chart/Figure, Doc/Table

โšก Domain Coverage

Domain#Samples
Math~1.6M
Science~13K
Chart-Figure~8K
Doc-Table~8K
General~8K

๐Ÿ“‹ Data Structure

The sft.json file contains an array of 1,620,490 training examples in the following structured format:

Schema Overview

{
  "conversations": [
    {
      "from": "human",
      "value": "<instruction_template_with_image_tag>"
    },
    {
      "from": "gpt",
      "value": "<think>detailed_reasoning_process</think><answer>final_answer</answer>"
    }
  ],
  "images": [
    "path/to/image.png"
  ]
}

Field Descriptions

  • conversations: Array containing the conversation turns
    • from: Role identifier ("human" or "gpt")
    • value: The content of the message
      • Human messages: Instruction template with <image> placeholder and question
      • Assistant messages: Structured response with <think>reasoning</think><answer>result</answer> format
  • images: Array of image file paths referenced in the conversation
    • Contains relative paths to associated visual content
    • Supports multimodal reasoning tasks with mathematical figures, charts, and diagrams

Response Format

All assistant responses follow a consistent chain-of-thought structure:

  • <think>...</think>: Detailed step-by-step reasoning process with error checking and corrections
  • <answer>...</answer>: Final concise answer to the question

This format enables training models for structured reasoning with explicit thought processes before providing final answers.



๐Ÿ“Œ Citation

If you find MMR1 useful for your research and applications, please cite using this BibTeX:

@misc{leng2025mmr1,
  title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources}, 
  author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
  year={2025},
  eprint={2509.21268},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2509.21268}, 
}
chain-of-thought
multimodal
reasoning
reinforcement-learning-support
vision-language