3 repos
lsdefine/simple_GRPO
A very simple GRPO implement for reproducing r1-like LLM thinking.
1,708
45 commits
zysNLP/quickllm
A repo for update and debug Mixtral-7x8B、MOE、ChatGLM3、LLaMa2、 BaChuan、Qwen an other LLM models…
47
55 commits
Tim-Siu/reft-exp
A research repo for experiments about Reinforcement Finetuning
56
44 commits