OpenHelix-Team/Awesome-VLA-RL

This repository summarizes recent advances in the VLA + RL paradigm and provides a taxonomic classification of relevant works.

435

15 commits

updated Oct 10, 2025

See the code

README

Awesome-VLA-RL

With significant advances in Vision-Language-Action (VLA)🍔 models based on large-scale imitation learning, integrating VLA with Reinforcement Learning (RL)🥤 has emerged as a promising paradigm. This paradigm leverages the benefits of trial-and-error interactions with environments or pre-collected sub-optimal data.

This repository summarizes recent advances in the VLA🍔 + RL🥤 paradigm and provides a classification of relevant works (offline RL training(without env.), online RL training(with env.), Model-Based RL (with world model as env.) test-time RL(during deployment), and RL alignment).

Contributions are welcome! Please feel free to submit an issue or reach out via email to add papers!

If you find this repository useful, please giving this list a star ⭐. Feel free to share it with others!

Offline RL

The Offline RL pre-trained VLA models leverage both human demonstrations and autonomously collected data.

MethodTitleVenueDateCode/ProjectKey feature/finding
Q-TransformerQ-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-FunctionsArxiv18/9/2023Github
Detailsoffline Q-learning with Transformer models: 1. Autoregressive Discrete Q-Learning; 2. Conservative Q-Learning; 3. Monte Carlo and n-step Returns
Perceiver-Actor-CriticOffline Actor-Critic Reinforcement Learning Scales to Large ModelsICML202408/2/2024Project
DetailsAn offline actor-critic method that scales to large models of up to 1B parameters and learn a wide variety of 132 control and robotics tasks
GeRMGeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped RobotIROS202420/3/2024Github
DetailsMixtureof-Experts structure; Quadruped robot learning
ReinboTReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement LearningICML202512/5/2025
DetailsMax-Return Sequence Modeling as Reinformer; Reward Densification with heuristic methods
MoREMoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action ModelsICRA202511/3/2025
DetailsIntegrates multiple low-rank adaptation modules as distinct experts within a dense multi-modal large language model (MLLM), forming a sparse-activated mixture-of-experts model
CO-RFTCO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement LearningArxiv04/8/2025
DetailsChunk-level offline RL finetuning. It proposed Chunked RL via n-step TD learning
ARFMBalancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow ModelsArxiv04/9/2025
DetailsBy introducing an adaptively adjusted scaling factor in the VLA flow model loss, we construct a principled bias-variance trade-off objective function to optimally control the impact of RL signal on flow loss. ARFM adaptively balances RL advantage preservation and flow loss gradient variance control, resulting in a more stable and efficient fine-tuning process.

Online RL

With trial-and-error interactions in online environments, VLA models can be further optimized to improve their performance.

in Simulator

MethodTitleVenueDateCode/ProjectKey feature/finding
FLaReFLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-TuningICRA 2025 Best Paper Finalist30/9/2024Code
DetailsFor large-scale fine-tuning in simulation, it performs extensive domain randomization, extract visual features through DinoV2, and utilize the KV-cache technique during inference and a set of algorithmic choices to ensure the stability of RL fine-tuning
PA-RLPolicy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and BackboneArxiv9/12/2024Project
Detailsa single method that fine-tunes multiple policy classes, with varying architectures and sizes. It enables sample-efficient improvement of diffusion and transformer-based autoregressive policies. PA-RL sets a new state of the art for offline to online RL, and it makes it possible, for the first time, to improve OpenVLA
iRe-VLAImproving Vision-Language-Action Model with Online Reinforcement LearningRAL202528/1/2025
DetailsAdopt SFT & RL two-stage iterative optimization to Stabilizing Training Process and Managing the Model Training Burden.
RIPT-VLAInteractive Post-Training for Vision-Language-Action ModelsArxiv22/5/2025Github
DetailsA critic-free optimization framework called Leave-One-Out Proximal Policy Optimization (LOOP); Dynamic rollout sampling
VLA-RLVLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningArxiv24/5/2025Github
DetailsRobotic process reward model and the VLA-RL System with (1) Curriculum Selection Strategy (2) Critic Warmup (3) GPU-balanced Vectorized Environments (4) PPO infrastructure
RLVLAWhat Can RL Bring to VLA Generalization? An Empirical StudyNeurIPS 202526/5/2025Github
DetailsPPO consistently outperforms GRPO and DPO; Shared actor-critic backbone; VLA warm-up
RFTFRFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal FeedbackArxiv26/5/2025
DetailsFor the sparse reward problem, RFTF leverages a value model trained using temporal information to generate dense rewards
SimpleVLA-RLSimpleVLA-RL: Scaling VLA Training via Reinforcement LearningArxiv12/9/2025Github
TGRPOTGRPO: Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy OptimizationArxiv10/6/2025Github
DetailsFrom GRPO in LLM to TGRPO in VLA
OctoNavOctoNav: Towards Generalist Embodied NavigationArxiv11/6/2025Project
DetailsFor Navigation tasks, it proposes a VLA+RL Hybrid Training Paradigm, including SFT, Nav-GRPO, Online RL stages. The VLA model also obtains thinking-before-action ability.
RLRCRLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action ModelsArxiv21/6/2025Project
DetailsA RL-based VLA compression Paradigm. Through a carefully designed three-stage pipeline, structured pruning, performance recovery based on SFT and RL, and 4bit quantization, they significantly reduce model size and boost inference speed while preserving, and in some cases surpassing, the original model’s ability to execute robotic tasks
RLinfRLinf: Reinforcement Learning Infrastructure for Agentic AIArxiv8/2025Project
DetailsRLinf is a flexible and scalable open-source infrastructure designed for post-training foundation models via reinforcement learning. The ‘inf’ in RLinf stands for Infrastructure, highlighting its role as a robust backbone for next-generation training. It also stands for Infinite, symbolizing the system’s support for open-ended learning, continuous generalization, and limitless possibilities in intelligence development.
RLinf-VLARLinf-VLA: A Unified and Efficient Framework for VLA+RL TrainingArxiv10/2025Project
Details

in Real-World

MethodTitleVenueDateCode/ProjectKey feature/finding
RLDGRLDG: Robotic Generalist Policy Distillation via Reinforcement LearningRSS202512/2024Project
DetailsPretrain task-specific RL policies with HIL-SERL; Distill RL policies into VLA for Knowledge Transfer.
PA-RLPolicy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and BackboneArxiv9/12/2024Project
Detailsa single method that fine-tunes multiple policy classes, with varying architectures and sizes. It enables sample-efficient improvement of diffusion and transformer-based autoregressive policies. PA-RL sets a new state of the art for offline to online RL, and it makes it possible, for the first time, to improve OpenVLA
iRe-VLAImproving Vision-Language-Action Model with Online Reinforcement LearningRAL202528/1/2025
DetailsAdopt SFT & RL two-stage iterative optimization to Stabilizing Training Process and Managing the Model Training Burden.
ConRFTConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency PolicyRSS202514/4/2025Github
DetailsOffline fine-tuning(Cal-QL) and online fine-tuning(CPQL+HIL-SERL)
VLACVLAC: A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement LearningGithub16/9/2025Github
Details VLAC is a general-purpose pair-wise critic and manipulation model which designed for real world robot reinforcement learning and data refinement.
GeneralistSelf-Improving Embodied Foundation ModelsNeurIPS 202518/9/2025
Details A two stage paradigm: The first stage, Supervised Fine-Tuning (SFT), fine-tunes pretrained foundation models using both: a) behavioral cloning, and b) steps-to-go prediction objectives. In the second stage, Self-Improvement, steps-to-go prediction enables the extraction of a well-shaped reward function and a robust success detector, enabling a fleet of robots to autonomously practice downstream tasks with minimal human supervision.

World Model (Model-Based RL)

MethodTitleVenueDateCode/ProjectKey feature/finding
World-EnvWorld-Env: Leveraging World Model as a Virtual Environment for VLA Post-TrainingArxiv30/9/2025
Details A world model-based framework that enables low-cost, safe reinforcement learning post-training for VLA policies under extreme data scarcity, eliminating the need for real-world interaction.
VLA-RFTVLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World SimulatorsArxiv01/10/2025Github
Details VLA-RFT, a reinforcement fine-tuning framework that leverages a data-driven world model as a controllable simulator.

Test-Time RL

Leverage a value function pre-trained via offline RL.

MethodTitleVenueDateCode/ProjectKey feature/finding
Bellman-Guided RetrialsTo Err is Robotic: Rapid Value-Based Trial-and-Error during DeploymentArxiv22/6/2024Github
DetailsPre-train a value function to estimate task completion, recover the robot and sample a new strategy if failed
V-GPSSteering Your Generalists: Improving Robotic Foundation Models via Value GuidanceCoRL202417/10/2024Project
DetailsRe-ranking multiple action proposals from a generalist policy using a value function at test-time
HumeHume: Introducing System-2 Thinking in Visual-Language-Action ModelArxiv2/6/2025Github
Details Pre-train a value function, perform best-of-N selection of candidate action chunks with state-action value estimation
VLA-ReasonerVLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree SearchArxiv26/9/2025
Details plug-in framework named VLA-Reasoner that empowers VLAs with test-time MTCS to address their incremental deviations during deployment.

RL Alignment

MethodTitleVenueDateCode/ProjectKey feature/finding
GRAPEGRAPE: Generalizing Robot Policy via Preference AlignmentICLR2025 workshop4/2/2025Github
DetailsTrajectory-wise Preference Optimization aligns VLA policies on a trajectory level
SafeVLASafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained LearningNeurIPS 202531/5/2025Project
DetailsConstraining VLA policies via safe reinforcement learning

Unclassified

MethodTitleVenueDateCode/ProjectKey feature/finding
RPDRefined Policy Distillation: From VLA Generalists to RL ExpertsArxiv6/3/2025
DetailsLeverage VLA model as policy prior to improve sample-efficiency of RL, as Jump-Start RL

Significant stargazers

Charlie Cheng

19 followers · starred Jun 2025

OpenHelix-Team/Awesome-VLA-RL

This repository summarizes recent advances in the VLA + RL paradigm and provides a taxonomic classification of relevant works.

435

15 commits

updated Oct 10, 2025

See the code

README

Awesome-VLA-RL

With significant advances in Vision-Language-Action (VLA)🍔 models based on large-scale imitation learning, integrating VLA with Reinforcement Learning (RL)🥤 has emerged as a promising paradigm. This paradigm leverages the benefits of trial-and-error interactions with environments or pre-collected sub-optimal data.

This repository summarizes recent advances in the VLA🍔 + RL🥤 paradigm and provides a classification of relevant works (offline RL training(without env.), online RL training(with env.), Model-Based RL (with world model as env.) test-time RL(during deployment), and RL alignment).

Contributions are welcome! Please feel free to submit an issue or reach out via email to add papers!

If you find this repository useful, please giving this list a star ⭐. Feel free to share it with others!

Offline RL

The Offline RL pre-trained VLA models leverage both human demonstrations and autonomously collected data.

MethodTitleVenueDateCode/ProjectKey feature/finding
Q-TransformerQ-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-FunctionsArxiv18/9/2023Github
Detailsoffline Q-learning with Transformer models: 1. Autoregressive Discrete Q-Learning; 2. Conservative Q-Learning; 3. Monte Carlo and n-step Returns
Perceiver-Actor-CriticOffline Actor-Critic Reinforcement Learning Scales to Large ModelsICML202408/2/2024Project
DetailsAn offline actor-critic method that scales to large models of up to 1B parameters and learn a wide variety of 132 control and robotics tasks
GeRMGeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped RobotIROS202420/3/2024Github
DetailsMixtureof-Experts structure; Quadruped robot learning
ReinboTReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement LearningICML202512/5/2025
DetailsMax-Return Sequence Modeling as Reinformer; Reward Densification with heuristic methods
MoREMoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action ModelsICRA202511/3/2025
DetailsIntegrates multiple low-rank adaptation modules as distinct experts within a dense multi-modal large language model (MLLM), forming a sparse-activated mixture-of-experts model
CO-RFTCO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement LearningArxiv04/8/2025
DetailsChunk-level offline RL finetuning. It proposed Chunked RL via n-step TD learning
ARFMBalancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow ModelsArxiv04/9/2025
DetailsBy introducing an adaptively adjusted scaling factor in the VLA flow model loss, we construct a principled bias-variance trade-off objective function to optimally control the impact of RL signal on flow loss. ARFM adaptively balances RL advantage preservation and flow loss gradient variance control, resulting in a more stable and efficient fine-tuning process.

Online RL

With trial-and-error interactions in online environments, VLA models can be further optimized to improve their performance.

in Simulator

MethodTitleVenueDateCode/ProjectKey feature/finding
FLaReFLaRe: Achieving Masterful and Adaptive Robot Policies with Large-Scale Reinforcement Learning Fine-TuningICRA 2025 Best Paper Finalist30/9/2024Code
DetailsFor large-scale fine-tuning in simulation, it performs extensive domain randomization, extract visual features through DinoV2, and utilize the KV-cache technique during inference and a set of algorithmic choices to ensure the stability of RL fine-tuning
PA-RLPolicy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and BackboneArxiv9/12/2024Project
Detailsa single method that fine-tunes multiple policy classes, with varying architectures and sizes. It enables sample-efficient improvement of diffusion and transformer-based autoregressive policies. PA-RL sets a new state of the art for offline to online RL, and it makes it possible, for the first time, to improve OpenVLA
iRe-VLAImproving Vision-Language-Action Model with Online Reinforcement LearningRAL202528/1/2025
DetailsAdopt SFT & RL two-stage iterative optimization to Stabilizing Training Process and Managing the Model Training Burden.
RIPT-VLAInteractive Post-Training for Vision-Language-Action ModelsArxiv22/5/2025Github
DetailsA critic-free optimization framework called Leave-One-Out Proximal Policy Optimization (LOOP); Dynamic rollout sampling
VLA-RLVLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement LearningArxiv24/5/2025Github
DetailsRobotic process reward model and the VLA-RL System with (1) Curriculum Selection Strategy (2) Critic Warmup (3) GPU-balanced Vectorized Environments (4) PPO infrastructure
RLVLAWhat Can RL Bring to VLA Generalization? An Empirical StudyNeurIPS 202526/5/2025Github
DetailsPPO consistently outperforms GRPO and DPO; Shared actor-critic backbone; VLA warm-up
RFTFRFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal FeedbackArxiv26/5/2025
DetailsFor the sparse reward problem, RFTF leverages a value model trained using temporal information to generate dense rewards
SimpleVLA-RLSimpleVLA-RL: Scaling VLA Training via Reinforcement LearningArxiv12/9/2025Github
TGRPOTGRPO: Fine-tuning Vision-Language-Action Model via Trajectory-wise Group Relative Policy OptimizationArxiv10/6/2025Github
DetailsFrom GRPO in LLM to TGRPO in VLA
OctoNavOctoNav: Towards Generalist Embodied NavigationArxiv11/6/2025Project
DetailsFor Navigation tasks, it proposes a VLA+RL Hybrid Training Paradigm, including SFT, Nav-GRPO, Online RL stages. The VLA model also obtains thinking-before-action ability.
RLRCRLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action ModelsArxiv21/6/2025Project
DetailsA RL-based VLA compression Paradigm. Through a carefully designed three-stage pipeline, structured pruning, performance recovery based on SFT and RL, and 4bit quantization, they significantly reduce model size and boost inference speed while preserving, and in some cases surpassing, the original model’s ability to execute robotic tasks
RLinfRLinf: Reinforcement Learning Infrastructure for Agentic AIArxiv8/2025Project
DetailsRLinf is a flexible and scalable open-source infrastructure designed for post-training foundation models via reinforcement learning. The ‘inf’ in RLinf stands for Infrastructure, highlighting its role as a robust backbone for next-generation training. It also stands for Infinite, symbolizing the system’s support for open-ended learning, continuous generalization, and limitless possibilities in intelligence development.
RLinf-VLARLinf-VLA: A Unified and Efficient Framework for VLA+RL TrainingArxiv10/2025Project
Details

in Real-World

MethodTitleVenueDateCode/ProjectKey feature/finding
RLDGRLDG: Robotic Generalist Policy Distillation via Reinforcement LearningRSS202512/2024Project
DetailsPretrain task-specific RL policies with HIL-SERL; Distill RL policies into VLA for Knowledge Transfer.
PA-RLPolicy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and BackboneArxiv9/12/2024Project
Detailsa single method that fine-tunes multiple policy classes, with varying architectures and sizes. It enables sample-efficient improvement of diffusion and transformer-based autoregressive policies. PA-RL sets a new state of the art for offline to online RL, and it makes it possible, for the first time, to improve OpenVLA
iRe-VLAImproving Vision-Language-Action Model with Online Reinforcement LearningRAL202528/1/2025
DetailsAdopt SFT & RL two-stage iterative optimization to Stabilizing Training Process and Managing the Model Training Burden.
ConRFTConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency PolicyRSS202514/4/2025Github
DetailsOffline fine-tuning(Cal-QL) and online fine-tuning(CPQL+HIL-SERL)
VLACVLAC: A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement LearningGithub16/9/2025Github
Details VLAC is a general-purpose pair-wise critic and manipulation model which designed for real world robot reinforcement learning and data refinement.
GeneralistSelf-Improving Embodied Foundation ModelsNeurIPS 202518/9/2025
Details A two stage paradigm: The first stage, Supervised Fine-Tuning (SFT), fine-tunes pretrained foundation models using both: a) behavioral cloning, and b) steps-to-go prediction objectives. In the second stage, Self-Improvement, steps-to-go prediction enables the extraction of a well-shaped reward function and a robust success detector, enabling a fleet of robots to autonomously practice downstream tasks with minimal human supervision.

World Model (Model-Based RL)

MethodTitleVenueDateCode/ProjectKey feature/finding
World-EnvWorld-Env: Leveraging World Model as a Virtual Environment for VLA Post-TrainingArxiv30/9/2025
Details A world model-based framework that enables low-cost, safe reinforcement learning post-training for VLA policies under extreme data scarcity, eliminating the need for real-world interaction.
VLA-RFTVLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World SimulatorsArxiv01/10/2025Github
Details VLA-RFT, a reinforcement fine-tuning framework that leverages a data-driven world model as a controllable simulator.

Test-Time RL

Leverage a value function pre-trained via offline RL.

MethodTitleVenueDateCode/ProjectKey feature/finding
Bellman-Guided RetrialsTo Err is Robotic: Rapid Value-Based Trial-and-Error during DeploymentArxiv22/6/2024Github
DetailsPre-train a value function to estimate task completion, recover the robot and sample a new strategy if failed
V-GPSSteering Your Generalists: Improving Robotic Foundation Models via Value GuidanceCoRL202417/10/2024Project
DetailsRe-ranking multiple action proposals from a generalist policy using a value function at test-time
HumeHume: Introducing System-2 Thinking in Visual-Language-Action ModelArxiv2/6/2025Github
Details Pre-train a value function, perform best-of-N selection of candidate action chunks with state-action value estimation
VLA-ReasonerVLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree SearchArxiv26/9/2025
Details plug-in framework named VLA-Reasoner that empowers VLAs with test-time MTCS to address their incremental deviations during deployment.

RL Alignment

MethodTitleVenueDateCode/ProjectKey feature/finding
GRAPEGRAPE: Generalizing Robot Policy via Preference AlignmentICLR2025 workshop4/2/2025Github
DetailsTrajectory-wise Preference Optimization aligns VLA policies on a trajectory level
SafeVLASafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained LearningNeurIPS 202531/5/2025Project
DetailsConstraining VLA policies via safe reinforcement learning

Unclassified

MethodTitleVenueDateCode/ProjectKey feature/finding
RPDRefined Policy Distillation: From VLA Generalists to RL ExpertsArxiv6/3/2025
DetailsLeverage VLA model as policy prior to improve sample-efficiency of RL, as Jump-Start RL

Significant stargazers

Charlie Cheng

19 followers · starred Jun 2025