LLM Fine-tuning and Reinforcement Learning

11 repos

Models and implementations for fine-tuning large language models using reinforcement learning techniques, particularly policy optimization methods like PPO (Proximal Policy Optimization). The cluster centers on adapted versions of Meta's Llama and Microsoft's Phi models with various RL-based training approaches (token-level, bandit, and segment-based variants), all compatible with the Hugging Face transformers ecosystem and text-generation-inference infrastructure. Someone exploring this area would encounter practical approaches to RLHF (Reinforcement Learning from Human Feedback) and post-training optimization for instruction-tuned language models.

Python · 1
conversational ·1
text-generation-inference ·1
transformers ·1
endpoints_compatible ·1
llama ·1
safetensors ·1
text-generation ·1
custom_code ·0
phi3 ·0