Guided by the observation that evaluation is easier than generation, we enabled large language models to excel on hard math problems beyond human evaluation capabilities through the easy-to-hard generalization of evaluators (e.g., process reward models). For comprehensive details and insights, we kindly direct you to our paper.
We provide model checkpoints for the supervised fine-tuned models and reward models. The current list of models includes:
SFT models:
Reward models:
Please check the examples for the training scripts and data for the data preparation.
14 commits
Python
99.8%
Guided by the observation that evaluation is easier than generation, we enabled large language models to excel on hard math problems beyond human evaluation capabilities through the easy-to-hard generalization of evaluators (e.g., process reward models). For comprehensive details and insights, we kindly direct you to our paper.
We provide model checkpoints for the supervised fine-tuned models and reward models. The current list of models includes:
SFT models:
Reward models:
Please check the examples for the training scripts and data for the data preparation.
14 commits
Python
99.8%