GD-ML/Open-RS1

Model

GPG: A Simple and Strong Reinforcement Learning

0

8 commits

1 linked in READMEs

updated May 8, 2025

See the code

README

GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning https://arxiv.org/abs/2504.02546

The RL model trained on the Open-r1 dataset based on GPG, using DeepSeek-R1-Distill-Qwen-1.5B as the baseline model.

qwen2
safetensors

Contributors

xiao23451

8 commits

GD-ML/Open-RS1

Model

GPG: A Simple and Strong Reinforcement Learning

0

8 commits

1 linked in READMEs

updated May 8, 2025

See the code

README

GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning https://arxiv.org/abs/2504.02546

The RL model trained on the Open-r1 dataset based on GPG, using DeepSeek-R1-Distill-Qwen-1.5B as the baseline model.

qwen2
safetensors

Contributors

xiao23451

8 commits