13 repos
Techniques and datasets for training language models through preference-based learning, including Direct Preference Optimization (DPO), Reinforcement Learning from Human Feedback (RLHF), and related alignment methods. The cluster covers preference datasets for multiple languages, training pipelines, and approaches to fine-tuning models using human or model-generated preference signals rather than supervised labels. Central repos demonstrate binarized feedback mechanisms and multimodal extensions, with Python as the primary implementation language.