Tigerdwgth/AwesomeRWRL4M

Awesome Real-World Reinforcement Learning for Manipulation

30

15 commits

updated Sep 23, 2026

See the code

README

Awesome Real-World Reinforcement Learning for Manipulation 🦾

A curated list of research papers on real-world reinforcement learning for robotic manipulation, where policies are trained or fine-tuned through direct interaction with physical manipulators (e.g., robotic arms, dexterous hands), rather than simulation-only or purely offline learning.

This list focuses on end-to-end or policy-level reinforcement learning applied to real-world manipulation tasks, including grasping, assembly, deformable object manipulation, and long-horizon dexterous control. Closely related interactive learning methods that improve policies from on-robot interaction (e.g., DAgger-style human correction) are also tracked.

Papers within each section are roughly ordered by first release date.


Algorithms, Systems & Infrastructure

RL Post-Training of VLAs & Generalist Policies

Residual RL & Policy Steering

Human-in-the-Loop RL

DAgger & Human Correction

World Models & Digital Twins

Reward Models for Real-World RL

Dexterous & Contact-Rich Manipulation


Contributions welcome.

Tigerdwgth/AwesomeRWRL4M

Awesome Real-World Reinforcement Learning for Manipulation

30

15 commits

updated Sep 23, 2026

See the code

README

Awesome Real-World Reinforcement Learning for Manipulation 🦾

A curated list of research papers on real-world reinforcement learning for robotic manipulation, where policies are trained or fine-tuned through direct interaction with physical manipulators (e.g., robotic arms, dexterous hands), rather than simulation-only or purely offline learning.

This list focuses on end-to-end or policy-level reinforcement learning applied to real-world manipulation tasks, including grasping, assembly, deformable object manipulation, and long-horizon dexterous control. Closely related interactive learning methods that improve policies from on-robot interaction (e.g., DAgger-style human correction) are also tracked.

Papers within each section are roughly ordered by first release date.


Algorithms, Systems & Infrastructure

RL Post-Training of VLAs & Generalist Policies

Residual RL & Policy Steering

Human-in-the-Loop RL

DAgger & Human Correction

World Models & Digital Twins

Reward Models for Real-World RL

Dexterous & Contact-Rich Manipulation


Contributions welcome.