FreedomIntelligence/OVM-process

Dataset

The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.

1

3 commits

1 linked in READMEs

updated Apr 1, 2024

See the code

README

The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.

Steps are split by the newlines in the response. step_labels indicates the logical correctness of steps, defined as "logically correct and it's based on accurate premises, not necessarily helps to solve the problem"; step_labels_progress indicates helpfulness of steps, defined as "logically correct, based on accurate premises, and helps to solve the problem".

Contributors

OakYU

3 commits

FreedomIntelligence/OVM-process

Dataset

The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.

1

3 commits

1 linked in READMEs

updated Apr 1, 2024

See the code

README

The training dataset of GSM8K for process reward models in the paper OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning, where the responses were generated by llama2-7b and the labels were annotated by GPT-4.

Steps are split by the newlines in the response. step_labels indicates the logical correctness of steps, defined as "logically correct and it's based on accurate premises, not necessarily helps to solve the problem"; step_labels_progress indicates helpfulness of steps, defined as "logically correct, based on accurate premises, and helps to solve the problem".

Contributors

OakYU

3 commits