0
stars
7
commits
1
linked in READMEs
Oct 16, 2025
updated
GitHub repository: https://github.com/mzhaoshuai/RefAlign
The training data is mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.
When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is BERTScore.
| Hyper-Parameters | Value |
|---|---|
| LR | 2.5e-6 |
| Batch Size | 512 |
| Epoch | 1 |
| Prompt Length | 600 |
| Generation Length | 1200 |
| Advantage CLIP | 0.08 |
| Sampled Generations (K) | 2 |
| BertScore Model | bart-large-mnli |
6 commits
1 commits
0
stars
7
commits
1
linked in READMEs
Oct 16, 2025
updated
GitHub repository: https://github.com/mzhaoshuai/RefAlign
The training data is mzhaoshuai/Llama-3.3-70B-Inst-awq_ultrafeedback_1in3.
When conducting Reinforcement Learning with Similarity-based Rewards, the reward function is BERTScore.
| Hyper-Parameters | Value |
|---|---|
| LR | 2.5e-6 |
| Batch Size | 512 |
| Epoch | 1 |
| Prompt Length | 600 |
| Generation Length | 1200 |
| Advantage CLIP | 0.08 |
| Sampled Generations (K) | 2 |
| BertScore Model | bart-large-mnli |
6 commits
1 commits