training Contains training pipelines for Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). Note: The Self-Play training pipeline is still under development.inference_and_scoring contains the modified implementaion from https://github.com/hkust-nlp/deita to score samples on quality and complexity. Added multi-GPU inference support using Hugging Face’s accelerate library.environment.yml file5 commits
Python
92.8%
Shell
7.2%
training Contains training pipelines for Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). Note: The Self-Play training pipeline is still under development.inference_and_scoring contains the modified implementaion from https://github.com/hkust-nlp/deita to score samples on quality and complexity. Added multi-GPU inference support using Hugging Face’s accelerate library.environment.yml file5 commits
Python
92.8%
Shell
7.2%