11 repos
UW-Madison-Lee-Lab/VersaPRM
should probably proofread and complete it, then remove this comment. -->
3
7 commits
UW-Madison-Lee-Lab/VersaPRM-Base-8B
0
3 commits
UW-Madison-Lee-Lab/VersaPRM-Aug
ScalableMath/llemma-7b-prm-prm800k-level-1to3-hf
PRMs are trained to predict the correctness of each step on the positions of "\n\n" and "\<eos\>".
6 commits
ScalableMath/llemma-7b-oprm-prm800k-level-1to3-hf
OPRMs (Outcome \& Process Reward Models) are trained to predict the correctness of each step on the…
ScalableMath/llemma-7b-prm-metamath-level-1to3-hf
9 commits
ScalableMath/llemma-7b-orm-prm800k-level-1to3-hf
ORMs are trained to predict the correctness of the whole solution on the position of "\<eos\>".
1
UW-Madison-Lee-Lab/Llama-PRM800K
ylacombe/musicgen-melody-lora-punk
ScalableMath/llemma-7b-sft-prm800k-level-1to3-hf
Usage:
2
4 commits
ScalableMath/llemma-7b-sft-metamath-level-1to3-hf