mzhaoshuai/alpaca-7b-ref-meteor

Model

0

stars

10

commits

1

linked in READMEs

Oct 16, 2025

updated

endpoints_compatible
llama
safetensors
text-generation
text-generation-inference
transformers
Browse cluster: Large Language Model Fine-tuning & Quantization

README

RefAlign: RL with Similarity-based Rewards

GitHub repository: https://github.com/mzhaoshuai/RefAlign

Paper: Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.

This is the model aligned with RefAlign described in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.

It is primarily aligned for safety.

The training data is https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF.

For the project code, please refer to the GitHub repository.

When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is Meteor.

Hyper-ParametersValue
LR2e-6
Batch Size512
Epoch2
Prompt Length192
Generation Length384
Sampled Generations (K)2
Reward functionMeteor
harmless advantage weight4.0

Contributors

mzhaoshuai

9 commits

nielsr

1 commits

mzhaoshuai/alpaca-7b-ref-meteor

Model

0

stars

10

commits

1

linked in READMEs

Oct 16, 2025

updated

endpoints_compatible
llama
safetensors
text-generation
text-generation-inference
transformers
Browse cluster: Large Language Model Fine-tuning & Quantization

README

RefAlign: RL with Similarity-based Rewards

GitHub repository: https://github.com/mzhaoshuai/RefAlign

Paper: Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.

This is the model aligned with RefAlign described in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data.

It is primarily aligned for safety.

The training data is https://huggingface.co/datasets/mzhaoshuai/Llama-3.3-70B-Inst-awq_SafeRLHF.

For the project code, please refer to the GitHub repository.

When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is Meteor.

Hyper-ParametersValue
LR2e-6
Batch Size512
Epoch2
Prompt Length192
Generation Length384
Sampled Generations (K)2
Reward functionMeteor
harmless advantage weight4.0

Contributors

mzhaoshuai

9 commits

nielsr

1 commits