MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
0
5 commits
4 linked in READMEs
updated Oct 1, 2025
This repository hosts the MMR1 model, a family of multimodal reasoning models introduced in the paper MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources.
Large multimodal reasoning models have achieved rapid progress, but their advancement is constrained by two major limitations: the absence of open, large-scale, high-quality long chain-of-thought (CoT) data, and the instability of reinforcement learning (RL) algorithms in post-training. Group Relative Policy Optimization (GRPO), the standard framework for RL fine-tuning, is prone to gradient vanishing when reward variance is low, which weakens optimization signals and impairs convergence. This work makes three contributions: (1) We propose Variance-Aware Sampling (VAS), a data selection strategy guided by Variance Promotion Score (VPS) that combines outcome variance and trajectory diversity to promote reward variance and stabilize policy optimization. (2) We release large-scale, carefully curated resources containing ~1.6M long CoT cold-start data and ~15k RL QA pairs, designed to ensure quality, difficulty, and diversity, along with a fully reproducible end-to-end training codebase. (3) We open-source a family of multimodal reasoning models in multiple scales, establishing standardized baselines for the community. Experiments across mathematical reasoning benchmarks demonstrate the effectiveness of both the curated data and the proposed VAS. Comprehensive ablation studies and analyses provide further insight into the contributions of each component. In addition, we theoretically establish that reward variance lower-bounds the expected policy gradient magnitude, with VAS serving as a practical mechanism to realize this guarantee. Our code, data, and checkpoints are available at this https URL.
This repository introduces our work on enhancing multimodal reasoning models. Current progress is limited by:
Variance-Aware Sampling (VAS):
A new data selection strategy guided by the Variance Promotion Score (VPS). VAS combines outcome variance and trajectory diversity to promote reward variance, stabilize policy optimization, and improve convergence.
Large-scale curated resources:
Open-source codebase & models:
Please refer to our TRAIN.md for detailed instructions on training with VAS.
Our method introduces Variance-Aware Sampling (VAS) to address the gradient vanishing problem in reinforcement learning with Group Relative Policy Optimization (GRPO).
As illustrated in Figure 1, training begins with a pool of prompts from the dataset:
This design ensures that training consistently focuses on prompts that provide strong learning signals, while still maintaining sufficient randomness for coverage.
Algorithm 1 provides a step-by-step description of VAS within the GRPO framework:
By adaptively steering training toward prompts with higher reward variance, VAS effectively stabilizes optimization and amplifies gradient signals, enabling more efficient and robust learning.
We release the following resources for the community:
The dataset spans diverse domains—including mathematics, science, charts/figures, document tables, and general understanding—covering ~1.6M math samples and an additional ~37K samples across other domains. It integrates existing public resources (e.g., MathVerse, ScienceQA, ChartQA, DocVQA, GQA) together with newly curated and self-collected data, ensuring quality, difficulty, and diversity. This collection establishes one of the most comprehensive open resources for multimodal reasoning models. We hope these resources can serve as a benchmark for the community and facilitate the research of multimodal reasoning.
We evaluate our models on a suite of mathematics-related multimodal reasoning benchmarks (MathVerse, MathVista, MathVision, LogicVista, and ChartQA).
We further analyze the effectiveness of Variance-Aware Sampling (VAS) through training efficiency and the evolution of Variance Promotion Score (VPS).
Training Efficiency (Fig. 2).
VPS Dynamics (Fig. 3).
👉 Together, these analyses highlight how VAS effectively mitigates gradient vanishing, improves sample efficiency, and adapts dynamically to the evolving training landscape.
To illustrate the reasoning capability of our models, we provide qualitative examples from MathVerse.
The demo showcases how the model carefully analyzes the problem, plans a structured solution, executes step-by-step reasoning, verifies results, and even provides alternative solution paths.
This demonstrates the model’s ability to maintain logical consistency, perform reflective verification, and present human-readable reasoning traces.
You can use the MMR1 model with the Hugging Face transformers library. Ensure you have transformers, torch, and Pillow installed. You may also need requests for downloading images from URLs.
First, install the necessary libraries:
pip install transformers torch Pillow requests
Here's a quick inference code example using the MMR1/MMR1-7B-RL checkpoint:
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoProcessor
from PIL import Image
import torch
import requests # For downloading images from URLs
import io # For handling image bytes
# Load model and tokenizer/processor
# Replace "MMR1/MMR1-7B-RL" with your desired model checkpoint, e.g., MMR1/MMR1-3B-RL, MMR1/MMR1-7B-SFT
model_id = "MMR1/MMR1-7B-RL"
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True # Required for custom Qwen2.5-VL architecture
)
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
# Prepare conversation history (example from Qwen2.5-VL chat template format)
messages = []
text_query = "Generate a comprehensive and detailed description for this image."
image_url = "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pond.jpg" # Example image URL
# Fetch the image
response = requests.get(image_url)
image = Image.open(io.BytesIO(response.content))
# Append user message with image and text
# The chat template expects content as a list of dictionaries for multimodal inputs
messages.append({"role": "user", "content": [{"type": "image", "image": image}, {"type": "text", "text": text_query}]})
# Apply chat template and process inputs
# `apply_chat_template` generates the text portion of the input, `processor` handles images.
text_inputs = processor.apply_chat_template(
messages,
tokenize=False, # We tokenize later with the full processor
add_generation_prompt=True
)
# Note: The processor expects a list of PIL Images for the `images` argument.
# It automatically handles vision token insertion based on the model's `chat_template`.
inputs = processor(text=text_inputs, images=[image], return_tensors="pt")
inputs = inputs.to(model.device)
# Generate response
generated_ids = model.generate(**inputs, max_new_tokens=1024, do_sample=False)
response_text = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(f"User: {text_query}")
print(f"Assistant: {response_text}")
This project is still under active development. Community feedback and contributions are highly appreciated. If you want to contribute, please feel free to make a pull request or create an issue.
Our MMR1 is build on top of Qwen2.5VL, LLaMA-Factory and EasyR1. Besides, our MMR1 benefits from tons of open-source efforts. We sincerely appreciate these efforts and compile a list in ACKNOWLEDGEMENT.md to express our gratitude. If your work is used in MMR1 but not mentioned in either this repo or the technical report, feel free to let us know ❤️.
If you find MMR1 useful for your research and applications, please cite using this BibTeX:
@misc{leng2025mmr1,
title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources},
author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
year={2025},
eprint={2509.21268},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.21268},
}
This project is released under the Apache 2.0 license as found in the LICENSE file. The service is a research preview intended for non-commercial use ONLY, subject to the model Licenses of Qwen, Terms of Use of the data generated by OpenAI and Gemini, and Privacy Practices of ShareGPT. Please get in touch with us if you find any potential violations.
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
0
5 commits
4 linked in READMEs
updated Oct 1, 2025
This repository hosts the MMR1 model, a family of multimodal reasoning models introduced in the paper MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources.
Large multimodal reasoning models have achieved rapid progress, but their advancement is constrained by two major limitations: the absence of open, large-scale, high-quality long chain-of-thought (CoT) data, and the instability of reinforcement learning (RL) algorithms in post-training. Group Relative Policy Optimization (GRPO), the standard framework for RL fine-tuning, is prone to gradient vanishing when reward variance is low, which weakens optimization signals and impairs convergence. This work makes three contributions: (1) We propose Variance-Aware Sampling (VAS), a data selection strategy guided by Variance Promotion Score (VPS) that combines outcome variance and trajectory diversity to promote reward variance and stabilize policy optimization. (2) We release large-scale, carefully curated resources containing ~1.6M long CoT cold-start data and ~15k RL QA pairs, designed to ensure quality, difficulty, and diversity, along with a fully reproducible end-to-end training codebase. (3) We open-source a family of multimodal reasoning models in multiple scales, establishing standardized baselines for the community. Experiments across mathematical reasoning benchmarks demonstrate the effectiveness of both the curated data and the proposed VAS. Comprehensive ablation studies and analyses provide further insight into the contributions of each component. In addition, we theoretically establish that reward variance lower-bounds the expected policy gradient magnitude, with VAS serving as a practical mechanism to realize this guarantee. Our code, data, and checkpoints are available at this https URL.
This repository introduces our work on enhancing multimodal reasoning models. Current progress is limited by:
Variance-Aware Sampling (VAS):
A new data selection strategy guided by the Variance Promotion Score (VPS). VAS combines outcome variance and trajectory diversity to promote reward variance, stabilize policy optimization, and improve convergence.
Large-scale curated resources:
Open-source codebase & models:
Please refer to our TRAIN.md for detailed instructions on training with VAS.
Our method introduces Variance-Aware Sampling (VAS) to address the gradient vanishing problem in reinforcement learning with Group Relative Policy Optimization (GRPO).
As illustrated in Figure 1, training begins with a pool of prompts from the dataset:
This design ensures that training consistently focuses on prompts that provide strong learning signals, while still maintaining sufficient randomness for coverage.
Algorithm 1 provides a step-by-step description of VAS within the GRPO framework:
By adaptively steering training toward prompts with higher reward variance, VAS effectively stabilizes optimization and amplifies gradient signals, enabling more efficient and robust learning.
We release the following resources for the community:
The dataset spans diverse domains—including mathematics, science, charts/figures, document tables, and general understanding—covering ~1.6M math samples and an additional ~37K samples across other domains. It integrates existing public resources (e.g., MathVerse, ScienceQA, ChartQA, DocVQA, GQA) together with newly curated and self-collected data, ensuring quality, difficulty, and diversity. This collection establishes one of the most comprehensive open resources for multimodal reasoning models. We hope these resources can serve as a benchmark for the community and facilitate the research of multimodal reasoning.
We evaluate our models on a suite of mathematics-related multimodal reasoning benchmarks (MathVerse, MathVista, MathVision, LogicVista, and ChartQA).
We further analyze the effectiveness of Variance-Aware Sampling (VAS) through training efficiency and the evolution of Variance Promotion Score (VPS).
Training Efficiency (Fig. 2).
VPS Dynamics (Fig. 3).
👉 Together, these analyses highlight how VAS effectively mitigates gradient vanishing, improves sample efficiency, and adapts dynamically to the evolving training landscape.
To illustrate the reasoning capability of our models, we provide qualitative examples from MathVerse.
The demo showcases how the model carefully analyzes the problem, plans a structured solution, executes step-by-step reasoning, verifies results, and even provides alternative solution paths.
This demonstrates the model’s ability to maintain logical consistency, perform reflective verification, and present human-readable reasoning traces.
You can use the MMR1 model with the Hugging Face transformers library. Ensure you have transformers, torch, and Pillow installed. You may also need requests for downloading images from URLs.
First, install the necessary libraries:
pip install transformers torch Pillow requests
Here's a quick inference code example using the MMR1/MMR1-7B-RL checkpoint:
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoProcessor
from PIL import Image
import torch
import requests # For downloading images from URLs
import io # For handling image bytes
# Load model and tokenizer/processor
# Replace "MMR1/MMR1-7B-RL" with your desired model checkpoint, e.g., MMR1/MMR1-3B-RL, MMR1/MMR1-7B-SFT
model_id = "MMR1/MMR1-7B-RL"
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True # Required for custom Qwen2.5-VL architecture
)
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
# Prepare conversation history (example from Qwen2.5-VL chat template format)
messages = []
text_query = "Generate a comprehensive and detailed description for this image."
image_url = "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pond.jpg" # Example image URL
# Fetch the image
response = requests.get(image_url)
image = Image.open(io.BytesIO(response.content))
# Append user message with image and text
# The chat template expects content as a list of dictionaries for multimodal inputs
messages.append({"role": "user", "content": [{"type": "image", "image": image}, {"type": "text", "text": text_query}]})
# Apply chat template and process inputs
# `apply_chat_template` generates the text portion of the input, `processor` handles images.
text_inputs = processor.apply_chat_template(
messages,
tokenize=False, # We tokenize later with the full processor
add_generation_prompt=True
)
# Note: The processor expects a list of PIL Images for the `images` argument.
# It automatically handles vision token insertion based on the model's `chat_template`.
inputs = processor(text=text_inputs, images=[image], return_tensors="pt")
inputs = inputs.to(model.device)
# Generate response
generated_ids = model.generate(**inputs, max_new_tokens=1024, do_sample=False)
response_text = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(f"User: {text_query}")
print(f"Assistant: {response_text}")
This project is still under active development. Community feedback and contributions are highly appreciated. If you want to contribute, please feel free to make a pull request or create an issue.
Our MMR1 is build on top of Qwen2.5VL, LLaMA-Factory and EasyR1. Besides, our MMR1 benefits from tons of open-source efforts. We sincerely appreciate these efforts and compile a list in ACKNOWLEDGEMENT.md to express our gratitude. If your work is used in MMR1 but not mentioned in either this repo or the technical report, feel free to let us know ❤️.
If you find MMR1 useful for your research and applications, please cite using this BibTeX:
@misc{leng2025mmr1,
title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources},
author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
year={2025},
eprint={2509.21268},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.21268},
}
This project is released under the Apache 2.0 license as found in the LICENSE file. The service is a research preview intended for non-commercial use ONLY, subject to the model Licenses of Qwen, Terms of Use of the data generated by OpenAI and Gemini, and Privacy Practices of ShareGPT. Please get in touch with us if you find any potential violations.