4 repos
ModalMinds/MM-EUREKA
MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
770
98 commits
Zkkkai/CPGD-7B
We proposed a novel RL algorithm called Clipped Policy Gradient Optimization with Policy Drift…
1
6 commits
FanqingM/MMK12
MMK12
26
9 commits
ModalMinds/MM-PRM
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
30
5 commits