allenai/reward-bench

Dataset

109

stars

113

commits

9

linked in READMEs

Sep 9, 2024

updated

README

RewardBench Logo

Code | Leaderboard | Prior Preference Sets | Results | Paper

Reward Bench Evaluation Dataset Card

The RewardBench evaluation dataset evaluates capabilities of reward models over the following categories:

  1. Chat: Includes the easy chat subsets (alpacaeval-easy, alpacaeval-length, alpacaeval-hard, mt-bench-easy, mt-bench-medium)
  2. Chat Hard: Includes the hard chat subsets (mt-bench-hard, llmbar-natural, llmbar-adver-neighbor, llmbar-adver-GPTInst, llmbar-adver-GPTOut, llmbar-adver-manual)
  3. Safety: Includes the safety subsets (refusals-dangerous, refusals-offensive, xstest-should-refuse, xstest-should-respond, do not answer)
  4. Reasoning: Includes the code and math subsets (math-prm, hep-cpp, hep-go, hep-java, hep-js, hep-python, hep-rust)

The RewardBench leaderboard averages over these subsets and a final category from prior preference data test sets including Anthropic Helpful, Anthropic HHH in BIG-Bench, Stanford Human Preferences (SHP), and OpenAI's Learning to Summarize data.

The scoring for RewardBench compares the score of a prompt-chosen pair to a prompt-rejected pair. Success is when the chosen score is higher than rejected.

RewardBench Scoring

In order to create a representative, single evaluation score, we perform a limited mixture of averaging across results. For all the subsets detailed below except for Reasoning, we perform per-prompt weighted averaging across all the prompts in the subset to get the section score. For example, in Chat we take a weighted average of the AlpacaEval and MT Bench sets based on the number of prompts. For Reasoning, we increase the weight of the PRM-Math subset so code and math abilities are weighed equally in the final number, rather than increasing the relevance of code. Once all subsets weighted averages are achieved, the final RewardBench score is the average across the subset scores (including Prior Sets).

Dataset Details

In order to maintain all the relevant data, the samples in the dataset will have the following items. Note, the dataset is single-turn:

  • prompt (str): the instruction given in the various test sets.
  • chosen (str): the response from the better model or the better rated prompt.
  • chosen_model (str): where applicable
  • rejected (str): the response with the lower score or from word model.
  • rejected_model (str): where applicable
  • subset (str): the subset (e.g. alpacaeval-easy) of the associated prompt as the dataset is all in one split.
  • id (int): an incremented id for every prompt in the benchmark.

To select a specific subset use HuggingFace Datasets .filter functionality.

dataset = dataset.filter(lambda ex: ex["subset"] == "alpacaeval-easy")

This can easily be converted to the standard chosen/rejected list of messages format (see UltraFeedback for an example), for example with our data loading utilities on GitHub.

Subset Summary

Total number of the prompts is: 2985.

SubsetNum. Samples (Pre-filtering, post-filtering)Description
alpacaeval-easy805, 100Great model vs poor model; GPT4-Turbo 97.7% v. Alpaca 7b 26.46% (data here)
alpacaeval-length805, 95Good model vs low model, similar length; Llama2chat 70B 92.66% vs Guanaco 13B 52.61% (data here)
alpacaeval-hard805, 95Great model vs baseline model; Tulu 2 95.0% v. Davinici003 50.0% (data here)
mt-bench-easy28, 28MT Bench 10s vs 1s (source data)
mt-bench-medium45, 40MT Bench 9s vs 2-5s (source data)
mt-bench-hard45, 37MT Bench 7-8 vs 5-6 (source data)
refusals-dangerous505, 100Dangerous rejected response vs polite chosen refusal
refusals-offensive704, 100Offensive rejected response vs polite chosen refusal
llmbar-natural100Manually curated instruction pairs (See paper)
llmbar-adver-neighbor134Adversarial instruction response vs. off-topic prompt response (See paper)
llmbar-adver-GPTInst92Adversarial instruction response vs. GPT4 generated off-topic prompt response (See paper)
llmbar-adver-GPTOut47Adversarial instruction response vs. unhelpful-prompted GPT4 responses (See paper)
llmbar-adver-manual46Challenge set manually designed chosen vs. rejected
xstest-should-refuse450, 154False response dataset (see paper)
xstest-should-respond450, 250False refusal dataset (see paper)
do not answer939, 136Prompts which responsible LLMs do not answer: Refusals are chosen and responses are rejected
hep-cpp164C++ working code vs. buggy code (See dataset or paper)
hep-go164Go working code vs. buggy code
hep-java164Java working code vs. buggy code
hep-js164Javascript working code vs. buggy code
hep-python164Python working code vs. buggy code
hep-rust164Rust working code vs. buggy code
math-prm447Human references vs. model error (see paper)

The length distribution of the subsets with a Llama tokenizer is shown below.

subsetChosen Mean TokensRejected Mean TokensChosen Max TokensRejected Max TokensChosen Min TokensRejected Min TokensChosen Mean Unique TokensRejected Mean Unique TokensChosen Max Unique TokensRejected Max Unique TokensChosen Min Unique TokensRejected Min Unique Tokens
alpacaeval-easy591.26167.33133210434015252.9183.446302903312
alpacaeval-hard411.684136.92611127115712172.53770.9684359297458
alpacaeval-length510.589596.895160422425552192.442188.5474346643038
donotanswer169.61320.57457352020103.743156.9413583371813
hep-cpp261.262259.488833835535799.853799.3722012013740
hep-go266.22264.598732720555799.62299.1892012013637
hep-java263.14260.9397487335554102.311101.9272072063941
hep-js251.165249.695771774535293.274492.92681921923740
hep-python211.988211.146624612534985.646385.30491901903635
hep-rust221.256219.049988993464995.140294.83541921923636
llmbar-adver-GPTInst170.109377.359636959151592.9457179.372874711213
llmbar-adver-GPTOut96.4255101393476182060.042655.04262412281314
llmbar-adver-manual159.804264.37607737233391.9565140.132733851824
llmbar-adver-neighbor70.2239172.50760386591343.313490.932825032489
llmbar-natural139.42129.82907900171874.9970.073543521414
math-prm279.313488.84116081165357783.6264124.5822372572346
mt-bench-easy391.821481.929778112615531169.071121.3212884347419
mt-bench-hard287.784301.64957311766862133.622121.6762613095048
mt-bench-med351.375466.025655129714552159.9140.3252854958241
refusals-dangerous208.4458.6138080487103128.532112003657155
refusals-offensive139.82298.632781117752695.98134.021704776019
xstest-should-refuse129.227217.019402549181580.5519116.1491942451613
xstest-should-respond188.708107.3565154652016103.78867.3282312021516

Filtering Summary

The RewardBench dataset is manually filtered from 5123 source prompts to manually verify the chosen-rejected ranking of prompts.

  • The categories of AlpacaEval and MT Bench are manually filtered for every prompt.
  • LLMBar, DoNotAnswer, HEP, and Math PRM all contained structured metadata for automatic filtering.
  • XSTest is a hybrid of manual confirmation with metadata from the project.
  • Refusals are automatically generated as a refusal or response (where refusal is preffered) with manual confirmation.

Substantial filtering details are available in the appendix of the papr. If there are any bugs in the data, please reach out!

License information

Licensing an aggregated dataset is a complex task. We release the RewardBench dataset under ODC-BY requiring the user to follow the licenses of the subsequent parts. Licensing LLM datasets is an evolving topic. The licenses primarily apply to the prompts and the completions generated by models are often unlicensed. The details for the datasets used in this work vary in the level of the detail on licenses and method of applying them.

DatasetVariantsData License
AlpacaEval{Easy, Length, Hard}CC By NC 4.0
MT Bench{Easy, Medium, Hard}Apache 2.0
LLMBar{Natural, Neighbor, GPTInst, GPTOut, Manual}MIT License
Do Not AnswerCC BY NC SA 4.0
XSTest{Should Respond, Should Refuse}CC By 4.0
HumanEvalPack{HEP CPP, Go, Javascript, Rust, Python, Rust}MIT License
PRM MathMIT License

Within this dataset are prompts created by AI2 (the refusals data, released as MIT for now, see official release soon) with completions from API and open models. More details will come on this soon.

Development

Requirements

Building the dataset requires datasets. Maintaining the script and notebooks requites notebook.

pip install datasets notebook nbconvert

Convert with:

jupyter nbconvert --to script [YOUR_NOTEBOOK].ipynb

With no changes to the ipynb, the dataset can be re-built and pushed with the following (PLEASE BE CAREFUL):

python build_dataset.py

Git LFS notes

If your uploads fail with:

Git LFS upload failed:  14% (1/7), 4.2 MB | 0 B/s                                                                                                                                                 
  (missing) data/train-00000-of-00001.parquet (425c88744455a9b0e7248cdd81fe4716085aae22849798f653f59fc878117a4d)
hint: Your push was rejected due to missing or corrupt local objects.
hint: You can disable this check with: `git config lfs.allowincompletepush true`

First fetch all lfs objects:

git lfs fetch --all origin main

Filtering script (basic)

To filter data, run the following script:

python scripts/filter.py subset-name 0

with a subset from the dataset and a start index.


Citation

@misc{RewardBench,
    title={RewardBench: Evaluating Reward Models for Language Modeling},
    author={Lambert, Nathan and Pyatkin, Valentina and Morrison, Jacob and Miranda, LJ and Lin, Bill Yuchen and Chandu, Khyathi and Dziri, Nouha and Kumar, Sachin and Zick, Tom and Choi, Yejin and Smith, Noah A. and Hajishirzi, Hannaneh},
    year={2024},
    howpublished={\url{https://huggingface.co/spaces/allenai/reward-bench}
}

Contributors

natolambert

111 commits

khyathi

1 commits

reciprocate

1 commits

allenai/reward-bench

Dataset

109

stars

113

commits

9

linked in READMEs

Sep 9, 2024

updated

README

RewardBench Logo

Code | Leaderboard | Prior Preference Sets | Results | Paper

Reward Bench Evaluation Dataset Card

The RewardBench evaluation dataset evaluates capabilities of reward models over the following categories:

  1. Chat: Includes the easy chat subsets (alpacaeval-easy, alpacaeval-length, alpacaeval-hard, mt-bench-easy, mt-bench-medium)
  2. Chat Hard: Includes the hard chat subsets (mt-bench-hard, llmbar-natural, llmbar-adver-neighbor, llmbar-adver-GPTInst, llmbar-adver-GPTOut, llmbar-adver-manual)
  3. Safety: Includes the safety subsets (refusals-dangerous, refusals-offensive, xstest-should-refuse, xstest-should-respond, do not answer)
  4. Reasoning: Includes the code and math subsets (math-prm, hep-cpp, hep-go, hep-java, hep-js, hep-python, hep-rust)

The RewardBench leaderboard averages over these subsets and a final category from prior preference data test sets including Anthropic Helpful, Anthropic HHH in BIG-Bench, Stanford Human Preferences (SHP), and OpenAI's Learning to Summarize data.

The scoring for RewardBench compares the score of a prompt-chosen pair to a prompt-rejected pair. Success is when the chosen score is higher than rejected.

RewardBench Scoring

In order to create a representative, single evaluation score, we perform a limited mixture of averaging across results. For all the subsets detailed below except for Reasoning, we perform per-prompt weighted averaging across all the prompts in the subset to get the section score. For example, in Chat we take a weighted average of the AlpacaEval and MT Bench sets based on the number of prompts. For Reasoning, we increase the weight of the PRM-Math subset so code and math abilities are weighed equally in the final number, rather than increasing the relevance of code. Once all subsets weighted averages are achieved, the final RewardBench score is the average across the subset scores (including Prior Sets).

Dataset Details

In order to maintain all the relevant data, the samples in the dataset will have the following items. Note, the dataset is single-turn:

  • prompt (str): the instruction given in the various test sets.
  • chosen (str): the response from the better model or the better rated prompt.
  • chosen_model (str): where applicable
  • rejected (str): the response with the lower score or from word model.
  • rejected_model (str): where applicable
  • subset (str): the subset (e.g. alpacaeval-easy) of the associated prompt as the dataset is all in one split.
  • id (int): an incremented id for every prompt in the benchmark.

To select a specific subset use HuggingFace Datasets .filter functionality.

dataset = dataset.filter(lambda ex: ex["subset"] == "alpacaeval-easy")

This can easily be converted to the standard chosen/rejected list of messages format (see UltraFeedback for an example), for example with our data loading utilities on GitHub.

Subset Summary

Total number of the prompts is: 2985.

SubsetNum. Samples (Pre-filtering, post-filtering)Description
alpacaeval-easy805, 100Great model vs poor model; GPT4-Turbo 97.7% v. Alpaca 7b 26.46% (data here)
alpacaeval-length805, 95Good model vs low model, similar length; Llama2chat 70B 92.66% vs Guanaco 13B 52.61% (data here)
alpacaeval-hard805, 95Great model vs baseline model; Tulu 2 95.0% v. Davinici003 50.0% (data here)
mt-bench-easy28, 28MT Bench 10s vs 1s (source data)
mt-bench-medium45, 40MT Bench 9s vs 2-5s (source data)
mt-bench-hard45, 37MT Bench 7-8 vs 5-6 (source data)
refusals-dangerous505, 100Dangerous rejected response vs polite chosen refusal
refusals-offensive704, 100Offensive rejected response vs polite chosen refusal
llmbar-natural100Manually curated instruction pairs (See paper)
llmbar-adver-neighbor134Adversarial instruction response vs. off-topic prompt response (See paper)
llmbar-adver-GPTInst92Adversarial instruction response vs. GPT4 generated off-topic prompt response (See paper)
llmbar-adver-GPTOut47Adversarial instruction response vs. unhelpful-prompted GPT4 responses (See paper)
llmbar-adver-manual46Challenge set manually designed chosen vs. rejected
xstest-should-refuse450, 154False response dataset (see paper)
xstest-should-respond450, 250False refusal dataset (see paper)
do not answer939, 136Prompts which responsible LLMs do not answer: Refusals are chosen and responses are rejected
hep-cpp164C++ working code vs. buggy code (See dataset or paper)
hep-go164Go working code vs. buggy code
hep-java164Java working code vs. buggy code
hep-js164Javascript working code vs. buggy code
hep-python164Python working code vs. buggy code
hep-rust164Rust working code vs. buggy code
math-prm447Human references vs. model error (see paper)

The length distribution of the subsets with a Llama tokenizer is shown below.

subsetChosen Mean TokensRejected Mean TokensChosen Max TokensRejected Max TokensChosen Min TokensRejected Min TokensChosen Mean Unique TokensRejected Mean Unique TokensChosen Max Unique TokensRejected Max Unique TokensChosen Min Unique TokensRejected Min Unique Tokens
alpacaeval-easy591.26167.33133210434015252.9183.446302903312
alpacaeval-hard411.684136.92611127115712172.53770.9684359297458
alpacaeval-length510.589596.895160422425552192.442188.5474346643038
donotanswer169.61320.57457352020103.743156.9413583371813
hep-cpp261.262259.488833835535799.853799.3722012013740
hep-go266.22264.598732720555799.62299.1892012013637
hep-java263.14260.9397487335554102.311101.9272072063941
hep-js251.165249.695771774535293.274492.92681921923740
hep-python211.988211.146624612534985.646385.30491901903635
hep-rust221.256219.049988993464995.140294.83541921923636
llmbar-adver-GPTInst170.109377.359636959151592.9457179.372874711213
llmbar-adver-GPTOut96.4255101393476182060.042655.04262412281314
llmbar-adver-manual159.804264.37607737233391.9565140.132733851824
llmbar-adver-neighbor70.2239172.50760386591343.313490.932825032489
llmbar-natural139.42129.82907900171874.9970.073543521414
math-prm279.313488.84116081165357783.6264124.5822372572346
mt-bench-easy391.821481.929778112615531169.071121.3212884347419
mt-bench-hard287.784301.64957311766862133.622121.6762613095048
mt-bench-med351.375466.025655129714552159.9140.3252854958241
refusals-dangerous208.4458.6138080487103128.532112003657155
refusals-offensive139.82298.632781117752695.98134.021704776019
xstest-should-refuse129.227217.019402549181580.5519116.1491942451613
xstest-should-respond188.708107.3565154652016103.78867.3282312021516

Filtering Summary

The RewardBench dataset is manually filtered from 5123 source prompts to manually verify the chosen-rejected ranking of prompts.

  • The categories of AlpacaEval and MT Bench are manually filtered for every prompt.
  • LLMBar, DoNotAnswer, HEP, and Math PRM all contained structured metadata for automatic filtering.
  • XSTest is a hybrid of manual confirmation with metadata from the project.
  • Refusals are automatically generated as a refusal or response (where refusal is preffered) with manual confirmation.

Substantial filtering details are available in the appendix of the papr. If there are any bugs in the data, please reach out!

License information

Licensing an aggregated dataset is a complex task. We release the RewardBench dataset under ODC-BY requiring the user to follow the licenses of the subsequent parts. Licensing LLM datasets is an evolving topic. The licenses primarily apply to the prompts and the completions generated by models are often unlicensed. The details for the datasets used in this work vary in the level of the detail on licenses and method of applying them.

DatasetVariantsData License
AlpacaEval{Easy, Length, Hard}CC By NC 4.0
MT Bench{Easy, Medium, Hard}Apache 2.0
LLMBar{Natural, Neighbor, GPTInst, GPTOut, Manual}MIT License
Do Not AnswerCC BY NC SA 4.0
XSTest{Should Respond, Should Refuse}CC By 4.0
HumanEvalPack{HEP CPP, Go, Javascript, Rust, Python, Rust}MIT License
PRM MathMIT License

Within this dataset are prompts created by AI2 (the refusals data, released as MIT for now, see official release soon) with completions from API and open models. More details will come on this soon.

Development

Requirements

Building the dataset requires datasets. Maintaining the script and notebooks requites notebook.

pip install datasets notebook nbconvert

Convert with:

jupyter nbconvert --to script [YOUR_NOTEBOOK].ipynb

With no changes to the ipynb, the dataset can be re-built and pushed with the following (PLEASE BE CAREFUL):

python build_dataset.py

Git LFS notes

If your uploads fail with:

Git LFS upload failed:  14% (1/7), 4.2 MB | 0 B/s                                                                                                                                                 
  (missing) data/train-00000-of-00001.parquet (425c88744455a9b0e7248cdd81fe4716085aae22849798f653f59fc878117a4d)
hint: Your push was rejected due to missing or corrupt local objects.
hint: You can disable this check with: `git config lfs.allowincompletepush true`

First fetch all lfs objects:

git lfs fetch --all origin main

Filtering script (basic)

To filter data, run the following script:

python scripts/filter.py subset-name 0

with a subset from the dataset and a start index.


Citation

@misc{RewardBench,
    title={RewardBench: Evaluating Reward Models for Language Modeling},
    author={Lambert, Nathan and Pyatkin, Valentina and Morrison, Jacob and Miranda, LJ and Lin, Bill Yuchen and Chandu, Khyathi and Dziri, Nouha and Kumar, Sachin and Zick, Tom and Choi, Yejin and Smith, Noah A. and Hajishirzi, Hannaneh},
    year={2024},
    howpublished={\url{https://huggingface.co/spaces/allenai/reward-bench}
}

Contributors

natolambert

111 commits

khyathi

1 commits

reciprocate

1 commits