RLHF and Agentic AI Systems

7 repos

Reinforcement learning from human feedback (RLHF) and reinforcement learning from verification/reasoning (RLVR) techniques applied to large language models and agentic AI systems. The cluster focuses on training methodologies, inference optimization, and frameworks for building AI agents that can reason, verify outputs, and improve through human and automated feedback signals. Central projects include distributed training frameworks (verl, LongVT), reasoning-based systems (ReasonGen-R1), and multimodal agent platforms (cogflow_code, dart-gui).

Python · 7
vlm ·266
agi ·266
multimodal ·266
multimodal-large-language-models ·266
tool-using-agent ·266
large-multimodal-models ·266
long-video-understanding ·266
mllm ·266
computer-use-agent ·97
gui-agent ·97