6 repos
This cluster focuses on fine-tuned large language models optimized through reinforcement learning techniques, particularly PPO (Proximal Policy Optimization) and bandit-based approaches. The repositories center on instruction-tuned variants of popular open models like Llama 3.1 and Phi-3, published as Hugging Face model checkpoints ready for deployment. Someone exploring this area would find practical implementations of RLHF workflows, model variants trained with different optimization strategies, and endpoints-compatible inference artifacts.