lmarena-ai/PPE-GPQA-Best-of-K

Dataset

1

stars

5

commits

1

linked in READMEs

Oct 22, 2024

updated

README

Overview

This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations.

The prompts are sampled from GPQA.

This dataset is meant for benchmarking and evaluation, not for training.

Paper

Code

License

User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers.

Citation

@misc{frick2024evaluaterewardmodelsrlhf,
      title={How to Evaluate Reward Models for RLHF}, 
      author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
      year={2024},
      eprint={2410.14872},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2410.14872}, 
}

Contributors

Timmli

3 commits

evanfrick

2 commits

lmarena-ai/PPE-GPQA-Best-of-K

Dataset

1

stars

5

commits

1

linked in READMEs

Oct 22, 2024

updated

README

Overview

This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations.

The prompts are sampled from GPQA.

This dataset is meant for benchmarking and evaluation, not for training.

Paper

Code

License

User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers.

Citation

@misc{frick2024evaluaterewardmodelsrlhf,
      title={How to Evaluate Reward Models for RLHF}, 
      author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica},
      year={2024},
      eprint={2410.14872},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2410.14872}, 
}

Linked in READMEs

Contributors

Timmli

3 commits

evanfrick

2 commits