arXiv:2505.02387 · 2 repos reference this paper in their README
RM-R1-UIUC/RM-R1
171
·
[ICLR'26] RM-R1: Unleashing the Reasoning Potential of Reward Models
gaotang/RM-R1-Qwen2.5-Instruct-7B
4
No description