The training dataset for verifiers, which is generated by the finetuned models in GSM8K and Game of 24. The models are open-sourced in OVM-llama2-7b and OVM-Mistral-7b.
4
2 commits
1 linked in READMEs
updated Dec 15, 2023
The training dataset for verifiers, which is generated by the finetuned models in GSM8K and Game of 24. The models are open-sourced in OVM-llama2-7b and OVM-Mistral-7b.
See the paper Outcome-supervised Verifiers for Planning in Mathematical Reasoning and the code in github
The training dataset for verifiers, which is generated by the finetuned models in GSM8K and Game of 24. The models are open-sourced in OVM-llama2-7b and OVM-Mistral-7b.
4
2 commits
1 linked in READMEs
updated Dec 15, 2023
The training dataset for verifiers, which is generated by the finetuned models in GSM8K and Game of 24. The models are open-sourced in OVM-llama2-7b and OVM-Mistral-7b.
See the paper Outcome-supervised Verifiers for Planning in Mathematical Reasoning and the code in github