10 repos
Techniques and implementations for fine-tuning large language models through reinforcement learning from human feedback (RLHF) and self-play preference optimization. The cluster centers on iterative model improvement methods applied to models like Llama, Gemma, and Mistral, focusing on alignment and performance optimization. These repos represent both reference implementations and specialized variants of preference-based training approaches for adapting pretrained models to specific tasks and behavioral objectives.