A Survey of Reinforcement Learning for Large Reasoning Models
See the codeWe welcome everyone to open an issue for any related work we haven’t discussed, and we’ll try to address it in the next release!
If you find this survey helpful, please cite our work:
@article{zhang2025survey,
title={A survey of reinforcement learning for large reasoning models},
author={Zhang, Kaiyan and Zuo, Yuxin and He, Bingxiang and Sun, Youbang and Liu, Runze and Jiang, Che and Fan, Yuchen and Tian, Kai and Jia, Guoli and Li, Pengfei and others},
journal={arXiv preprint arXiv:2509.08827},
year={2025}
}
Our survey provides a comprehensive examination of Reinforcement Learning for Large Reasoning Models.
We organize the survey into five main sections:
| Date | Name | Title | Paper | Github |
|---|---|---|---|---|
| 2025-09 | Baichuan-M2 | Baichuan-M2: Scaling Medical Capability with Large Verifier System | - | |
| 2025-08 | CX-Mind | CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning | ||
| 2025-08 | MORE-CLEAR | MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation | - | |
| 2025-08 | ARMed | Breaking Reward Collapse: Adaptive Reinforcement for Open-ended Medical Reasoning with Enhanced Semantic Discrimination | - | |
| 2025-08 | ProMed | ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs | ||
| 2025-08 | OwkinZero | OwkinZero: Accelerating Biological Discovery with AI | - | |
| 2025-08 | MolReasoner | MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs | ||
| 2025-08 | MedGR$^2$ | MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning | - | |
| 2025-07 | MedGround-R1 | MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization | ||
| 2025-07 | MedGemma | MedGemma Technical Report | [![Paper](https://img.shields.io/badge/paper-A42C |
Truncated — view the full README on GitHub.
TeX
100.0%
A Survey of Reinforcement Learning for Large Reasoning Models
See the codeWe welcome everyone to open an issue for any related work we haven’t discussed, and we’ll try to address it in the next release!
If you find this survey helpful, please cite our work:
@article{zhang2025survey,
title={A survey of reinforcement learning for large reasoning models},
author={Zhang, Kaiyan and Zuo, Yuxin and He, Bingxiang and Sun, Youbang and Liu, Runze and Jiang, Che and Fan, Yuchen and Tian, Kai and Jia, Guoli and Li, Pengfei and others},
journal={arXiv preprint arXiv:2509.08827},
year={2025}
}
Our survey provides a comprehensive examination of Reinforcement Learning for Large Reasoning Models.
We organize the survey into five main sections:
| Date | Name | Title | Paper | Github |
|---|---|---|---|---|
| 2025-09 | Baichuan-M2 | Baichuan-M2: Scaling Medical Capability with Large Verifier System | - | |
| 2025-08 | CX-Mind | CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning | ||
| 2025-08 | MORE-CLEAR | MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation | - | |
| 2025-08 | ARMed | Breaking Reward Collapse: Adaptive Reinforcement for Open-ended Medical Reasoning with Enhanced Semantic Discrimination | - | |
| 2025-08 | ProMed | ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs | ||
| 2025-08 | OwkinZero | OwkinZero: Accelerating Biological Discovery with AI | - | |
| 2025-08 | MolReasoner | MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs | ||
| 2025-08 | MedGR$^2$ | MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning | - | |
| 2025-07 | MedGround-R1 | MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization | ||
| 2025-07 | MedGemma | MedGemma Technical Report | [![Paper](https://img.shields.io/badge/paper-A42C |
Truncated — view the full README on GitHub.
TeX
100.0%