0
stars
7
commits
1
linked in READMEs
Oct 16, 2025
updated
GitHub repository: https://github.com/mzhaoshuai/RefAlign
This is the model aligned with SimPO described in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.
This model is a fine-tuned version of meta-llama/Meta-Llama-3-8B-Instruct on the mzhaoshuai/llama3-ultrafeedback-bertscore-bart-large-mnli dataset. It achieves the following results on the evaluation set:
The following hyperparameters were used during training:
6 commits
1 commits
0
stars
7
commits
1
linked in READMEs
Oct 16, 2025
updated
GitHub repository: https://github.com/mzhaoshuai/RefAlign
This is the model aligned with SimPO described in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.
This model is a fine-tuned version of meta-llama/Meta-Llama-3-8B-Instruct on the mzhaoshuai/llama3-ultrafeedback-bertscore-bart-large-mnli dataset. It achieves the following results on the evaluation set:
The following hyperparameters were used during training:
6 commits
1 commits