7 repos
Reinforcement learning from human feedback (RLHF) and reinforcement learning from verification/reasoning (RLVR) techniques applied to large language models and agentic AI systems. The cluster focuses on training methodologies, inference optimization, and frameworks for building AI agents that can reason, verify outputs, and improve through human and automated feedback signals. Central projects include distributed training frameworks (verl, LongVT), reasoning-based systems (ReasonGen-R1), and multimodal agent platforms (cogflow_code, dart-gui).