MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
0
5 commits
1 linked in READMEs
updated Oct 1, 2025
This repository introduces the MMR1 family of multimodal reasoning models, presented in the paper "MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources".
MMR1 addresses critical limitations in the advancement of large multimodal reasoning models, specifically the absence of open, large-scale, high-quality long chain-of-thought (CoT) data, and the instability of reinforcement learning (RL) algorithms during post-training.
MMR1 introduces Variance-Aware Sampling (VAS) to mitigate the gradient vanishing problem in reinforcement learning fine-tuning with GRPO. The framework balances exploration and coverage by combining a random sampler with a weighted sampler guided by the Variance Promotion Score (VPS). This ensures that training focuses on prompts providing strong learning signals, with VPS scores periodically re-estimated for dynamic adaptation.
The project open-sources the following resources for the community:
These resources cover diverse domains, including mathematics, science, charts/figures, document tables, and general understanding, integrating existing public resources with newly curated data.
MMR1 models have been evaluated on a suite of mathematics-related multimodal reasoning benchmarks (MathVerse, MathVista, MathVision, LogicVista, and ChartQA).
For detailed instructions on installation, training, and further evaluation, please refer to the GitHub repository.
If you find MMR1 useful for your research and applications, please cite using this BibTeX:
@misc{leng2025mmr1,
title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources},
author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
year={2025},
eprint={2509.21268},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.21268},
}
This project is released under the Apache 2.0 license as found in the LICENSE file. The service is a research preview intended for non-commercial use ONLY, subject to the model Licenses of Qwen, Terms of Use of the data generated by OpenAI and Gemini, and Privacy Practices of ShareGPT. Please get in touch with us if you find any potential violations.
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
0
5 commits
1 linked in READMEs
updated Oct 1, 2025
This repository introduces the MMR1 family of multimodal reasoning models, presented in the paper "MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources".
MMR1 addresses critical limitations in the advancement of large multimodal reasoning models, specifically the absence of open, large-scale, high-quality long chain-of-thought (CoT) data, and the instability of reinforcement learning (RL) algorithms during post-training.
MMR1 introduces Variance-Aware Sampling (VAS) to mitigate the gradient vanishing problem in reinforcement learning fine-tuning with GRPO. The framework balances exploration and coverage by combining a random sampler with a weighted sampler guided by the Variance Promotion Score (VPS). This ensures that training focuses on prompts providing strong learning signals, with VPS scores periodically re-estimated for dynamic adaptation.
The project open-sources the following resources for the community:
These resources cover diverse domains, including mathematics, science, charts/figures, document tables, and general understanding, integrating existing public resources with newly curated data.
MMR1 models have been evaluated on a suite of mathematics-related multimodal reasoning benchmarks (MathVerse, MathVista, MathVision, LogicVista, and ChartQA).
For detailed instructions on installation, training, and further evaluation, please refer to the GitHub repository.
If you find MMR1 useful for your research and applications, please cite using this BibTeX:
@misc{leng2025mmr1,
title={MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources},
author={Sicong Leng and Jing Wang and Jiaxi Li and Hao Zhang and Zhiqiang Hu and Boqiang Zhang and Yuming Jiang and Hang Zhang and Xin Li and Lidong Bing and Deli Zhao and Wei Lu and Yu Rong and Aixin Sun and Shijian Lu},
year={2025},
eprint={2509.21268},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.21268},
}
This project is released under the Apache 2.0 license as found in the LICENSE file. The service is a research preview intended for non-commercial use ONLY, subject to the model Licenses of Qwen, Terms of Use of the data generated by OpenAI and Gemini, and Privacy Practices of ShareGPT. Please get in touch with us if you find any potential violations.