they say have Q* and such, but seems not very impressive thinking from first impression.
however their Qwen2.5-Math-7B-Instruct based PRM with llama 3.1 8b instruct looks to be second best to Qwen2.5-Math-RM-72B (about as good as Maj@64 with llama) - source
so, we can probably use this PRM along with mathshepherd? however i don't know how this one is trained though
another math shepherd moded, but based on llama 3.1 8b instruct, so it may be better
they say have Q* and such, but seems not very impressive thinking from first impression.
however their Qwen2.5-Math-7B-Instruct based PRM with llama 3.1 8b instruct looks to be second best to Qwen2.5-Math-RM-72B (about as good as Maj@64 with llama) - source
so, we can probably use this PRM along with mathshepherd? however i don't know how this one is trained though
another math shepherd moded, but based on llama 3.1 8b instruct, so it may be better