用强化学习提升LLM性能,ChatGPT
3
stars
commits
Python
primary language
Apr 10, 2023
updated
基于 基于TRL库 构建ChatGPT训练流程
python train_sft.py
python train_reward.py
python train_rl.py
3 commits
ethanyanjiali/minChatGPT
A minimum example of aligning language models with RLHF similar to ChatGPT
226
AI-Study-Han/Zero-Chatgpt
从0开始,将chatgpt的技术路线跑一遍。
283
huggingface/trl
Train transformer language models with reinforcement learning.
19,286
agentscope-ai/Trinity-RFT
Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement…
700
RLHFlow/Online-RLHF
A recipe for online RLHF and online iterative DPO.
544
cy0307/lm-rlhf-ppo
0
weirayao/Retroformer
40
RLHFlow/Llama3-SFT-v2.0-epoch1
93.5%
Shell
6.5%