This contains the MBPP-Plus correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from MBPP-Plus.
This dataset is meant for benchmarking and evaluation, not for training.
User prompts are licensed under Apache-2.0, and model outputs are governed by the terms of use set by the respective model providers.
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for RLHF},
author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
year={2024},
eprint={2410.14872},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2410.14872},
}
This contains the MBPP-Plus correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from MBPP-Plus.
This dataset is meant for benchmarking and evaluation, not for training.
User prompts are licensed under Apache-2.0, and model outputs are governed by the terms of use set by the respective model providers.
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for RLHF},
author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
year={2024},
eprint={2410.14872},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2410.14872},
}