Process Reward Models for LLM Reasoning

7 repos

Process reward models (PRMs) and related techniques for improving large language model reasoning, particularly in mathematical and complex problem-solving domains. These models train on intermediate reasoning steps rather than just final outputs, enabling better evaluation of chain-of-thought reasoning paths. The cluster includes multiple model variants (ThinkPRM at different scales) and verification/evaluation frameworks that apply process supervision to enhance LLM decision-making and solution quality.

process supervision ·19
chain-of-thought ·19
math reasoning ·19
endpoints_compatible ·11
generative reward model ·11
qwen2 ·11
reward-model ·11
verification ·11
code verification ·11
conversational ·11