The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.
1
3 commits
1 linked in READMEs
updated Apr 1, 2024
The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.
Steps are split by the newlines in the response. step_labels indicates the logical correctness of steps, defined as "logically correct and it's based on accurate premises, not necessarily helps to solve the problem"; step_labels_progress indicates helpfulness of steps, defined as "logically correct, based on accurate premises, and helps to solve the problem".
3 commits
The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.
1
3 commits
1 linked in READMEs
updated Apr 1, 2024
The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.
Steps are split by the newlines in the response. step_labels indicates the logical correctness of steps, defined as "logically correct and it's based on accurate premises, not necessarily helps to solve the problem"; step_labels_progress indicates helpfulness of steps, defined as "logically correct, based on accurate premises, and helps to solve the problem".
3 commits