LLM Fine-tuning with Reinforcement Learning

6 repos

This cluster focuses on fine-tuned large language models optimized through reinforcement learning techniques, particularly PPO (Proximal Policy Optimization) and bandit-based approaches. The repositories center on instruction-tuned variants of popular open models like Llama 3.1 and Phi-3, published as Hugging Face model checkpoints ready for deployment. Someone exploring this area would find practical implementations of RLHF workflows, model variants trained with different optimization strategies, and endpoints-compatible inference artifacts.

Python · 1
transformers ·1
endpoints_compatible ·1
conversational ·1
mistral ·1
safetensors ·1
text-generation ·1
text-generation-inference ·1
llama ·0