5 repos
CatoG/LLM-DPO
LLM Direct Preference Optimization: Huggingface Space
0
12 commits
CatoG/DPO_Demo
No description
Airmomo/tpo-llm-webui
TPO 是一个优化 LLM 输出文本的框架,通过迭代反馈和优化提示的方式来“微调模型”,而非直接调整模型的参数,使模型在推理过程中与人类偏好对齐以生成更好的结果。本项目提供了一个友好的 WebUI…
13
11 commits
deep-diver/LLM-Pref-Mark-UI
37
7 commits
Phylliida/ModelPreferences
Learning Preferences of LLMs
2
76 commits