Learning Preferences of LLMs
2
stars
76
commits
Jupyter Notebook
primary language
Sep 17, 2025
updated
See Llama 3.2 1B has non-transitive preferences
76 commits
CatoG/LLM-DPO
LLM Direct Preference Optimization: Huggingface Space
0
adityaemmanuel/explainable-rl-alignment
This repository contains all the work as part of our research work exploring breaking down…
1
xu1998hz/llm_self_bias
This is the project to quantify the issues with LLM's self evaluation
9
CatoG/DPO_Demo
littlemesie/llm-learning
大模型学习
ramoneirao/llm-projects
LoselSpt/LLM
muthuka/sample-llm-finetuning
61.4%
Python
37.7%