An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
3,391
stars
469
commits
Python
primary language
Sep 11, 2026
updated
ROLL is an efficient and user-friendly RL library designed for Large Language Models (LLMs) utilizing Large Scale GPU resources. It significantly enhances LLM performance in key areas such as human preference alignment, complex reasoning, and multi-turn agentic interaction scenarios.
Leveraging a multi-role distributed architecture with Ray for flexible resource allocation and heterogeneous task scheduling, ROLL integrates cutting-edge technologies like Megatron-Core, SGLang and vLLM to accelerate model training and inference.
| 📣 Updates |
|---|
| [16/06/2026] 🎉 Our OSDI’26 RollArt paper is now available on arxiv. |
| [03/06/2026] 🎉 We support Qwen3.5 Dense and MoE series models and on-policy distill. Welcome to use! |
| [02/03/2026] 🎉 We released FSDP2 Strategy, Megatron with LoRA, GPU partial overlapping, Qwen3-Omni supports and other features. For more details, please refer to the release notes. Welcome to use! |
| [01/01/2026] 🎉 Our Let It Flow: Agentic Crafting on Rock and Roll report released! Introducing ALE ecosystem and ROME, an open-source agentic model with novel IPA algorithm. |
| [11/08/2025] 🎉 Our ROCK: Reinforcement Open Construction Kit released, Explore the new capabilities!. |
| [10/23/2025] 🎉 Our Papers released, see Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning and Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization. |
| [10/14/2025] 🎉 Our Paper released, see Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony. |
| [09/28/2025] 🎉 Ascend NPU support — see usage guide. |
| [09/25/2025] 🎉 Our Paper released, see RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training |
| [09/24/2025] 🎉 Support Wan2_2 Reward FL pipeline. Explore the new capabilities! |
| [09/23/2025] 🎉 ROLL aligns with GEM environment definition, providing agentic Tool Use training capabilities, ToolUse docs. |
| [09/16/2025] 🎉 Qwen3-Next model training is supported, refer to configuration. |
| [09/04/2025] 🎉 ROLL supports vLLM dynamic FP8 rollout and remove_padding for acceleration. |
| [08/28/2025] 🎉 ROLL supports SFT pipeline, refer to configuration. |
[08/13/2025] 🎉 ROLL supports AMD GPUs with out-of-box image docker and Dockerfile and specific yamls under examples/ directory. Please refer to Installation. |
| [08/11/2025] 🎉 Our Paper released, see Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning. |
| [08/10/2025] 🎉 Agentic RL supports stepwise learning, like GiGPO; Distill supports VLM. Explore the new capabilities! |
| [08/06/2025] 🎉 ROLL PPT is now available, Slides. |
| [07/31/2025] 🎉 Refactor agentic rl design. Support agentic rl async training. Explore the new capabilities! |
| [07/31/2025] 🎉 Support DistillPipeline/DpoPipeline. Support lora. Support GSPO |
| [06/25/2025] 🎉 Support thread env for env scaling and support qwen2.5 VL agentic pipeline. |
| [06/13/2025] 🎉 Support Qwen2.5 VL rlvr pipeline and upgrade mcore to 0.12 version. |
| [06/09/2025] 🎉 ROLL tech report is now available! Access the report here. |
| [06/08/2025] 🎉Supports Qwen3(8B/14B/32B), Qwen3-MoE(30A3/235A22), Qwen2.5(7B/14B/32B/72B) LLM models. |
| [05/30/2025] 🎉 Training RLVR and Agentic RL with ROLL is now available! Explore the new capabilities. |
Installation
Config System Explanation
Debugging Guide
Trackers and Metrics
Checkpoint Saving and Resuming Guide
Converting MCoreAdapter Models to Hugging Face Format
Quick Start: Single-Node Deployment Guide
Quick Start: Multi-Node Deployment Guide
Quick Start: Using Alibaba Cloud Function Compute DevPod for Rapid Development
Frequently Asked Questions
RLVR Pipeline
Agentic Pipeline
Agentic Comprehensive Guide
Distill Pipeline
Reinforce++
TOPR
GiGPO
PPO
Lite PPO
GRPO
GSPO
RAFT++
StarPO
RewardFL
Asynchronous Parallel Rollout
Asynchronous Training Feature
Resource Config
GPU Time-Division Multiplexing Control
domain_batch_size distribution control.reward_fresh priority) and asynchronous full-buffer refresh, providing fresher and higher-signal off-policy samples for both step- and trajectory-level agentic RL. codeROLL is inspired by the design of OpenRLHF, VeRL, Nemo-Aligner, and RAGEN.
The project is developed by Alibaba TAOBAO & TMALL Group and Alibaba Group. The code is distributed under the Apache License (Version 2.0). This product contains various third-party components under other open-source licenses. See the NOTICE file for more information.
The following repositories have been used in ROLL, either in their close-to-original form or as an inspiration:
If you use ROLL in your research or project, please consider citing us:
@article{wang2025reinforcement,
title={Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library},
author={Wang, Weixun and Xiong, Shaopan and Chen, Gengru and Gao, Wei and Guo, Sheng and He, Yancheng and Huang, Ju and Liu, Jiaheng and Li, Zhendong and Li, Xiaoyang and others},
journal={arXiv preprint arXiv:2506.06122},
year={2025}
}
ROLL is a project jointly developed by Taotian Future Living Lab and Alibaba AI Engine Team, with a strong emphasis on pioneering the future of Reinforcement Learning (RL). Our mission is to explore and shape innovative forms of future living powered by advanced RL technologies. If you are passionate about the future of RL and want to be part of its evolution, we warmly welcome you to join us!👇
We are HIRING!
Python
96.1%
JavaScript
1.9%
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
3,391
stars
469
commits
Python
primary language
Sep 11, 2026
updated
ROLL is an efficient and user-friendly RL library designed for Large Language Models (LLMs) utilizing Large Scale GPU resources. It significantly enhances LLM performance in key areas such as human preference alignment, complex reasoning, and multi-turn agentic interaction scenarios.
Leveraging a multi-role distributed architecture with Ray for flexible resource allocation and heterogeneous task scheduling, ROLL integrates cutting-edge technologies like Megatron-Core, SGLang and vLLM to accelerate model training and inference.
| 📣 Updates |
|---|
| [16/06/2026] 🎉 Our OSDI’26 RollArt paper is now available on arxiv. |
| [03/06/2026] 🎉 We support Qwen3.5 Dense and MoE series models and on-policy distill. Welcome to use! |
| [02/03/2026] 🎉 We released FSDP2 Strategy, Megatron with LoRA, GPU partial overlapping, Qwen3-Omni supports and other features. For more details, please refer to the release notes. Welcome to use! |
| [01/01/2026] 🎉 Our Let It Flow: Agentic Crafting on Rock and Roll report released! Introducing ALE ecosystem and ROME, an open-source agentic model with novel IPA algorithm. |
| [11/08/2025] 🎉 Our ROCK: Reinforcement Open Construction Kit released, Explore the new capabilities!. |
| [10/23/2025] 🎉 Our Papers released, see Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning and Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization. |
| [10/14/2025] 🎉 Our Paper released, see Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony. |
| [09/28/2025] 🎉 Ascend NPU support — see usage guide. |
| [09/25/2025] 🎉 Our Paper released, see RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training |
| [09/24/2025] 🎉 Support Wan2_2 Reward FL pipeline. Explore the new capabilities! |
| [09/23/2025] 🎉 ROLL aligns with GEM environment definition, providing agentic Tool Use training capabilities, ToolUse docs. |
| [09/16/2025] 🎉 Qwen3-Next model training is supported, refer to configuration. |
| [09/04/2025] 🎉 ROLL supports vLLM dynamic FP8 rollout and remove_padding for acceleration. |
| [08/28/2025] 🎉 ROLL supports SFT pipeline, refer to configuration. |
[08/13/2025] 🎉 ROLL supports AMD GPUs with out-of-box image docker and Dockerfile and specific yamls under examples/ directory. Please refer to Installation. |
| [08/11/2025] 🎉 Our Paper released, see Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning. |
| [08/10/2025] 🎉 Agentic RL supports stepwise learning, like GiGPO; Distill supports VLM. Explore the new capabilities! |
| [08/06/2025] 🎉 ROLL PPT is now available, Slides. |
| [07/31/2025] 🎉 Refactor agentic rl design. Support agentic rl async training. Explore the new capabilities! |
| [07/31/2025] 🎉 Support DistillPipeline/DpoPipeline. Support lora. Support GSPO |
| [06/25/2025] 🎉 Support thread env for env scaling and support qwen2.5 VL agentic pipeline. |
| [06/13/2025] 🎉 Support Qwen2.5 VL rlvr pipeline and upgrade mcore to 0.12 version. |
| [06/09/2025] 🎉 ROLL tech report is now available! Access the report here. |
| [06/08/2025] 🎉Supports Qwen3(8B/14B/32B), Qwen3-MoE(30A3/235A22), Qwen2.5(7B/14B/32B/72B) LLM models. |
| [05/30/2025] 🎉 Training RLVR and Agentic RL with ROLL is now available! Explore the new capabilities. |
Installation
Config System Explanation
Debugging Guide
Trackers and Metrics
Checkpoint Saving and Resuming Guide
Converting MCoreAdapter Models to Hugging Face Format
Quick Start: Single-Node Deployment Guide
Quick Start: Multi-Node Deployment Guide
Quick Start: Using Alibaba Cloud Function Compute DevPod for Rapid Development
Frequently Asked Questions
RLVR Pipeline
Agentic Pipeline
Agentic Comprehensive Guide
Distill Pipeline
Reinforce++
TOPR
GiGPO
PPO
Lite PPO
GRPO
GSPO
RAFT++
StarPO
RewardFL
Asynchronous Parallel Rollout
Asynchronous Training Feature
Resource Config
GPU Time-Division Multiplexing Control
domain_batch_size distribution control.reward_fresh priority) and asynchronous full-buffer refresh, providing fresher and higher-signal off-policy samples for both step- and trajectory-level agentic RL. codeROLL is inspired by the design of OpenRLHF, VeRL, Nemo-Aligner, and RAGEN.
The project is developed by Alibaba TAOBAO & TMALL Group and Alibaba Group. The code is distributed under the Apache License (Version 2.0). This product contains various third-party components under other open-source licenses. See the NOTICE file for more information.
The following repositories have been used in ROLL, either in their close-to-original form or as an inspiration:
If you use ROLL in your research or project, please consider citing us:
@article{wang2025reinforcement,
title={Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library},
author={Wang, Weixun and Xiong, Shaopan and Chen, Gengru and Gao, Wei and Guo, Sheng and He, Yancheng and Huang, Ju and Liu, Jiaheng and Li, Zhendong and Li, Xiaoyang and others},
journal={arXiv preprint arXiv:2506.06122},
year={2025}
}
ROLL is a project jointly developed by Taotian Future Living Lab and Alibaba AI Engine Team, with a strong emphasis on pioneering the future of Reinforcement Learning (RL). Our mission is to explore and shape innovative forms of future living powered by advanced RL technologies. If you are passionate about the future of RL and want to be part of its evolution, we warmly welcome you to join us!👇
We are HIRING!
Python
96.1%
JavaScript
1.9%