[Project Page] [GitHub] [Model] [Paper]
This repository contains the dataset for Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning.
The Robo-ValueRL dataset provides heterogeneous real-robot experience for studying reliable value estimation, value-guided offline policy pretraining, and online residual adaptation.
The Robo-ValueRL dataset contains real-robot trajectories collected on two long-horizon manipulation tasks:
The dataset includes:
The dataset is released together with the Robo-ValueRL model suite:
The associated models include a history-conditioned value estimator, a quality-conditioned VLA policy, and an online residual adaptation module.
The dataset is designed for:
A precision manipulation task where the robot must grasp a PCB, adjust it to a feasible insertion pose, grasp a chip, and insert it into millimeter-scale clearance.
A generalizable manipulation task where the robot must grasp, separate, and classify block components under varied configurations.
This dataset can be used to reproduce the Robo-ValueRL pipeline or to study new methods for:
Please refer to the GitHub repository for data loading, preprocessing, and training scripts.
If you use the Robo-ValueRL dataset in your research, please cite our work. Citation will be updated after the arXiv release.
Please refer to the license file in the GitHub repository.
For questions, please open an issue on our GitHub repository.
[Project Page] [GitHub] [Model] [Paper]
This repository contains the dataset for Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning.
The Robo-ValueRL dataset provides heterogeneous real-robot experience for studying reliable value estimation, value-guided offline policy pretraining, and online residual adaptation.
The Robo-ValueRL dataset contains real-robot trajectories collected on two long-horizon manipulation tasks:
The dataset includes:
The dataset is released together with the Robo-ValueRL model suite:
The associated models include a history-conditioned value estimator, a quality-conditioned VLA policy, and an online residual adaptation module.
The dataset is designed for:
A precision manipulation task where the robot must grasp a PCB, adjust it to a feasible insertion pose, grasp a chip, and insert it into millimeter-scale clearance.
A generalizable manipulation task where the robot must grasp, separate, and classify block components under varied configurations.
This dataset can be used to reproduce the Robo-ValueRL pipeline or to study new methods for:
Please refer to the GitHub repository for data loading, preprocessing, and training scripts.
If you use the Robo-ValueRL dataset in your research, please cite our work. Citation will be updated after the arXiv release.
Please refer to the license file in the GitHub repository.
For questions, please open an issue on our GitHub repository.