7 repos
Process reward models (PRMs) and related techniques for improving large language model reasoning, particularly in mathematical and complex problem-solving domains. These models train on intermediate reasoning steps rather than just final outputs, enabling better evaluation of chain-of-thought reasoning paths. The cluster includes multiple model variants (ThinkPRM at different scales) and verification/evaluation frameworks that apply process supervision to enhance LLM decision-making and solution quality.