轻量级 LLM Post-training 框架,支持 SFT、RLVR、On-Policy KD、Guide KD 及混合训练;实现单轮/多轮 Guide 蒸馏、多教师蒸馏、Reward 混合训练与自动化数据分流👩🎓👨🎓
Python
939
268 commits
updated Mar 8, 2026
A lightweight playground for
RLHFandSFTexperiments, with support forRLVR,KD, andGuide-KD.轻量级 RLHF/SFT 实验平台,支持
RLVR、KD与Guide-KD。
| 模块 | 说明 |
|---|---|
openrlhf-kd | 当前主线,基于 OpenRLHF 重构,实现 RLVR + KD + Guide-KD |
main_train.py | 简洁 SFT 训练入口 |
openrlhf-kd 是这个仓库当前最核心的部分,基于 OpenRLHF 构建,具体训练使用可参见文档 openrlhf-kd/examples/README.md
主要改动:
RLVR 部分,移除了 critic 等不需要的内容KD、Guide-KD 与 reward 的混合训练,支持按 datasource 路由根目录的 SFT 部分保持了比较简洁的训练入口,适合快速微调实验。
特性:
DeepspeedLoRA、QLoRA、全参微调示例文件可参见 data/sft_data.jsonl
Quick Start:
bash run_example.sh
或:
deepspeed --include localhost:0,1 main_train.py \
--train_data_path /path/to/data.jsonl \
--model_name_or_path /path/to/model \
--task_type sft \
--train_mode qlora \
--output_dir /path/to/output
31 followers · starred May 2025
11 followers · starred Jun 2024
Python
99.3%
轻量级 LLM Post-training 框架,支持 SFT、RLVR、On-Policy KD、Guide KD 及混合训练;实现单轮/多轮 Guide 蒸馏、多教师蒸馏、Reward 混合训练与自动化数据分流👩🎓👨🎓
Python
939
268 commits
updated Mar 8, 2026
A lightweight playground for
RLHFandSFTexperiments, with support forRLVR,KD, andGuide-KD.轻量级 RLHF/SFT 实验平台,支持
RLVR、KD与Guide-KD。
| 模块 | 说明 |
|---|---|
openrlhf-kd | 当前主线,基于 OpenRLHF 重构,实现 RLVR + KD + Guide-KD |
main_train.py | 简洁 SFT 训练入口 |
openrlhf-kd 是这个仓库当前最核心的部分,基于 OpenRLHF 构建,具体训练使用可参见文档 openrlhf-kd/examples/README.md
主要改动:
RLVR 部分,移除了 critic 等不需要的内容KD、Guide-KD 与 reward 的混合训练,支持按 datasource 路由根目录的 SFT 部分保持了比较简洁的训练入口,适合快速微调实验。
特性:
DeepspeedLoRA、QLoRA、全参微调示例文件可参见 data/sft_data.jsonl
Quick Start:
bash run_example.sh
或:
deepspeed --include localhost:0,1 main_train.py \
--train_data_path /path/to/data.jsonl \
--model_name_or_path /path/to/model \
--task_type sft \
--train_mode qlora \
--output_dir /path/to/output
31 followers · starred May 2025
11 followers · starred Jun 2024
Python
99.3%