mzhaoshuai/Llama-3-8B-Instruct-refalign

Model

0

stars

7

commits

1

linked in READMEs

Oct 16, 2025

updated

conversational
endpoints_compatible
llama
safetensors
text-generation
text-generation-inference
transformers
Browse cluster: Large Language Model Fine-tuning & Quantization

README

RefAlign: RL with Similarity-based Rewards

GitHub repository: https://github.com/mzhaoshuai/RefAlign

Paper: Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.

The training data is mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.

When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is BERTScore.

Hyper-ParametersValue
LR2.5e-6
Batch Size512
Epoch1
Prompt Length600
Generation Length1200
Advantage CLIP0.08
Sampled Generations (K)2
BertScore Modelbart-large-mnli

Contributors

mzhaoshuai

6 commits

nielsr

1 commits

mzhaoshuai/Llama-3-8B-Instruct-refalign

Model

0

stars

7

commits

1

linked in READMEs

Oct 16, 2025

updated

conversational
endpoints_compatible
llama
safetensors
text-generation
text-generation-inference
transformers
Browse cluster: Large Language Model Fine-tuning & Quantization

README

RefAlign: RL with Similarity-based Rewards

GitHub repository: https://github.com/mzhaoshuai/RefAlign

Paper: Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.

The training data is mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.

When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is BERTScore.

Hyper-ParametersValue
LR2.5e-6
Batch Size512
Epoch1
Prompt Length600
Generation Length1200
Advantage CLIP0.08
Sampled Generations (K)2
BertScore Modelbart-large-mnli

Contributors

mzhaoshuai

6 commits

nielsr

1 commits