IMPORTANT: This dataset is the decontaminated version of Skywork-Reward-Preference-80K-v0.1. We removed 4,957 pairs from the magpie-ultra-v0.1 subset that have a significant n-gram overlap with the evaluation prompts in RewardBench. You can find the set of removed pairs here. For more information, see this GitHub gist.
If your task involves evaluation on RewardBench, we strongly encourage you to use v0.2 instead of v0.1 of the dataset.
We will soon release our new version of the reward models!
Skywork Reward Preference 80K is a subset of 80K preference pairs, sourced from publicly available data. This subset is used to train Skywork-Reward-Gemma-2-27B-v0.2 and Skywork-Reward-Llama-3.1-8B-v0.2.
We carefully curate the Skywork Reward Data Collection (1) to include high-quality preference pairs and (2) to target specific capability and knowledge domains. The curated training dataset consists of approximately 80K samples, subsampled from multiple publicly available data sources, including
Disclaimer: We made no modifications to the original datasets listed above, other than subsampling the datasets to create the Skywork Reward Data Collection.
During dataset curation, we adopt several tricks to achieve both performance improvement and a balance between each domain, without compromising the overall performance:
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
If you have any questions, please feel free to reach us at yuhao.liuu@kunlun-inc.com or liang.zeng@kunlun-inc.com.
If you find our work helpful, please feel free to cite us using the following BibTeX entry:
@article{liu2024skywork,
title={Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs},
author={Liu, Chris Yuhao and Zeng, Liang and Liu, Jiacai and Yan, Rui and He, Jujie and Wang, Chaojie and Yan, Shuicheng and Liu, Yang and Zhou, Yahui},
journal={arXiv preprint arXiv:2410.18451},
year={2024}
}
10 commits
IMPORTANT: This dataset is the decontaminated version of Skywork-Reward-Preference-80K-v0.1. We removed 4,957 pairs from the magpie-ultra-v0.1 subset that have a significant n-gram overlap with the evaluation prompts in RewardBench. You can find the set of removed pairs here. For more information, see this GitHub gist.
If your task involves evaluation on RewardBench, we strongly encourage you to use v0.2 instead of v0.1 of the dataset.
We will soon release our new version of the reward models!
Skywork Reward Preference 80K is a subset of 80K preference pairs, sourced from publicly available data. This subset is used to train Skywork-Reward-Gemma-2-27B-v0.2 and Skywork-Reward-Llama-3.1-8B-v0.2.
We carefully curate the Skywork Reward Data Collection (1) to include high-quality preference pairs and (2) to target specific capability and knowledge domains. The curated training dataset consists of approximately 80K samples, subsampled from multiple publicly available data sources, including
Disclaimer: We made no modifications to the original datasets listed above, other than subsampling the datasets to create the Skywork Reward Data Collection.
During dataset curation, we adopt several tricks to achieve both performance improvement and a balance between each domain, without compromising the overall performance:
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
If you have any questions, please feel free to reach us at yuhao.liuu@kunlun-inc.com or liang.zeng@kunlun-inc.com.
If you find our work helpful, please feel free to cite us using the following BibTeX entry:
@article{liu2024skywork,
title={Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs},
author={Liu, Chris Yuhao and Zeng, Liang and Liu, Jiacai and Yan, Rui and He, Jujie and Wang, Chaojie and Yan, Shuicheng and Liu, Yang and Zhou, Yahui},
journal={arXiv preprint arXiv:2410.18451},
year={2024}
}
10 commits