I am currently implementing pi07, so many things will be broken for now.
This repo is a research-oriented version of LeRobot. The focus is on having SOTA algorithms (e.g., RECAP) and SOTA models (e.g., MolmoAct2) and tooling for examining the model's internals such as attention maps, clustering of internal representations.
This is an active project and we expect to continually add more features and capabilities. Happy to take requests.
Results obtained using MolmoAct2 on anchor actions.
https://github.com/user-attachments/assets/fa476815-7a97-4b58-a62d-7ee1dcd91d88
A longer video with successes, recoveries, and failures is availabe on Youtube.
MolmoAct2 is fully integrated. We have been doing preliminary tests and runs. It looks very promising, but there is a lot more to test. The next thing is a full training run with RECAP.
Clone this repo
git clone https://github.com/cijerezg/lerobot.git
cd lerobot
and to set up the environment run:
uv sync --extra molmoact2 --extra async --extra training
The quality of the dataset is extremely important. We suggest at least 50 episodes with a consistent strategy for task execution, e.g., try to grasp object in a consistent way as much as possible. MolmoAct2 has shown generalization capabilities, so we suggest to try a diverse dataset where objects are moved throughout the scene.
[!IMPORTANT] This pipeline assumes a LeRobot v3.0 dataset format (introduced in
lerobot 5.0.0). If your dataset is v2.1, see Using a v2.1 Dataset for the one-command migration.
Pre-decode the dataset to a cache. Images are stored at their original resolution, so the offline buffer would otherwise hold every decoded frame in RAM at startup — a 50k-transition dataset with two 480×640 cameras runs ~90 GB. With the cache, the dataset lives on disk and only the data sampled during a training step is loaded into memory; subsequent runs load instantly.
Generate the cache once per dataset:
python -m lerobot.scripts.lerobot_memmap_buffer_cache \
--root /path/to/local/dataset \
--cache-dir outputs/buffer_cache \
--image-storage-dtype uint8 \
--image-storage-size 480 640
You can also pass repo-id instead of root if the dataset is on HF hub.
Then point the YAML at it with buffer_cache_dir: outputs/buffer_cache. Benefits:
--image-storage-size and --image-storage-dtype flags are model-agnostic, so the same pipeline works for any policy that uses a different image resolution (378x378 for MolmoAct2, 224×224 for $\pi_{0.5}, etc.).More on edge cases in advanced usage.
This is the file used for all the scripts in this repo. An example of the config file can be found in the rl/config_rl.yaml.
To get started with training, the key fields to change are:
root: this is the path to your dataset.task: this is the task prompt.base_path: this should point to a local copy of the HF MolmoAct2 model. It can be downloaded locally using: hf download allenai/molmoact2-so100_101.pretrained_path: path to your finetuned model, or null if you don't have one. base_path must still be set either way — it's the upstream base model that supplies the architecture and norm stats.[!NOTE] Naming heads-up: this fork keeps the upstream LeRobot field name
pretrained_path, butbase_pathis the actual pretrained foundation model. Readpretrained_pathas "finetune to load on top of base."
We suggest you take a look at the entire config file before launching a training run.
Before launching any training, decide how actions are represented. The same encoding has to be used end-to-end — switching it later means recomputing statistics and retraining from base. Three options:
absolute — raw joint positions. Simplest, but does not generalize across starting configurations.anchor (recommended) — offsets from the chunk's initial state. Translation-invariant, generalizes well.delta — first-order differences between consecutive actions. Compact, but errors can accumulate.Set the choice via policy.action_encoding in the config. anchor and delta also require precomputed normalization statistics — see the advanced usage guide for details and the script that generates them.
The same scripts work for MolmoAct2 and $\pi_{0.5}$ — the policy is selected by policy.type in your config.
We suggest to start with a round of offline training so that the policy has a better starting point. To run it use:
uv run python -m lerobot.scripts.rl_offline --config path/to/config.yaml
[!NOTE] Offline training runs validation probes (attention maps, action drift, etc.) that render MP4 artifacts. On Linux these can fail with
[Errno 12] Cannot allocate memory— a virtual-memory overcommit accounting quirk with large PyTorch processes, not actual OOM. Persistent fix:echo 'vm.overcommit_memory = 1' | sudo tee /etc/sysctl.d/99-overcommit.conf sudo sysctl --systemBackground and tradeoffs in docs/engineering_notes/runbooks/system_overcommit.md. If you'd rather not touch sysctl, set val_on_start to false and a large number to val_freq.
Once the offline training has run for a while, set pretrained_path in the config to the resulting checkpoint and proceed to online training.
First run the learner script:
uv run python -m lerobot.rl.rl_learner --config path/to/config.yaml
and then on another terminal run the actor script:
uv run python -m lerobot.rl.rl_actor_async --config path/to/config.yaml
The learner will automatically save buffers to disk with the online data. After processing, those can be reused for the next round of offline or online training.
Following the suggestions from the RECAP paper, we suggest to retrain every time from the base model, and just include the additional data to avoid drift.
[!IMPORTANT] While these instructions might get you started, we recommend reading the advanced usage guide for better results. The $\pi_{0.5}$ variant lives at advanced_usage_pi05.md.
Highly recommended before first inference: set
action_clamp_limits. Teleop the arm through its safe range, record min/max per joint, then set the limits in degrees:policy: action_clamp_limits: - [-150, 150] # joint 1 - [-150, 0] # joint 2 - ... # one [min, max] per jointAnything outside isclamped before reaching the servos.
Once your config has a trained model, camera indices, and follower/leader ports, run inference with:
uv run python -m lerobot.rl.inference_async --config path/to/config.yaml
Note: Initial model loading takes 1 to 2 minutes.
Many important details were omitted in this introduction, and as we all know the devil is in the details, especially in a research-oriented repo. We strongly recommend reading the rest of the documentation, which is structured as:
(top 30 of 210)
Python
99.6%
I am currently implementing pi07, so many things will be broken for now.
This repo is a research-oriented version of LeRobot. The focus is on having SOTA algorithms (e.g., RECAP) and SOTA models (e.g., MolmoAct2) and tooling for examining the model's internals such as attention maps, clustering of internal representations.
This is an active project and we expect to continually add more features and capabilities. Happy to take requests.
Results obtained using MolmoAct2 on anchor actions.
https://github.com/user-attachments/assets/fa476815-7a97-4b58-a62d-7ee1dcd91d88
A longer video with successes, recoveries, and failures is availabe on Youtube.
MolmoAct2 is fully integrated. We have been doing preliminary tests and runs. It looks very promising, but there is a lot more to test. The next thing is a full training run with RECAP.
Clone this repo
git clone https://github.com/cijerezg/lerobot.git
cd lerobot
and to set up the environment run:
uv sync --extra molmoact2 --extra async --extra training
The quality of the dataset is extremely important. We suggest at least 50 episodes with a consistent strategy for task execution, e.g., try to grasp object in a consistent way as much as possible. MolmoAct2 has shown generalization capabilities, so we suggest to try a diverse dataset where objects are moved throughout the scene.
[!IMPORTANT] This pipeline assumes a LeRobot v3.0 dataset format (introduced in
lerobot 5.0.0). If your dataset is v2.1, see Using a v2.1 Dataset for the one-command migration.
Pre-decode the dataset to a cache. Images are stored at their original resolution, so the offline buffer would otherwise hold every decoded frame in RAM at startup — a 50k-transition dataset with two 480×640 cameras runs ~90 GB. With the cache, the dataset lives on disk and only the data sampled during a training step is loaded into memory; subsequent runs load instantly.
Generate the cache once per dataset:
python -m lerobot.scripts.lerobot_memmap_buffer_cache \
--root /path/to/local/dataset \
--cache-dir outputs/buffer_cache \
--image-storage-dtype uint8 \
--image-storage-size 480 640
You can also pass repo-id instead of root if the dataset is on HF hub.
Then point the YAML at it with buffer_cache_dir: outputs/buffer_cache. Benefits:
--image-storage-size and --image-storage-dtype flags are model-agnostic, so the same pipeline works for any policy that uses a different image resolution (378x378 for MolmoAct2, 224×224 for $\pi_{0.5}, etc.).More on edge cases in advanced usage.
This is the file used for all the scripts in this repo. An example of the config file can be found in the rl/config_rl.yaml.
To get started with training, the key fields to change are:
root: this is the path to your dataset.task: this is the task prompt.base_path: this should point to a local copy of the HF MolmoAct2 model. It can be downloaded locally using: hf download allenai/molmoact2-so100_101.pretrained_path: path to your finetuned model, or null if you don't have one. base_path must still be set either way — it's the upstream base model that supplies the architecture and norm stats.[!NOTE] Naming heads-up: this fork keeps the upstream LeRobot field name
pretrained_path, butbase_pathis the actual pretrained foundation model. Readpretrained_pathas "finetune to load on top of base."
We suggest you take a look at the entire config file before launching a training run.
Before launching any training, decide how actions are represented. The same encoding has to be used end-to-end — switching it later means recomputing statistics and retraining from base. Three options:
absolute — raw joint positions. Simplest, but does not generalize across starting configurations.anchor (recommended) — offsets from the chunk's initial state. Translation-invariant, generalizes well.delta — first-order differences between consecutive actions. Compact, but errors can accumulate.Set the choice via policy.action_encoding in the config. anchor and delta also require precomputed normalization statistics — see the advanced usage guide for details and the script that generates them.
The same scripts work for MolmoAct2 and $\pi_{0.5}$ — the policy is selected by policy.type in your config.
We suggest to start with a round of offline training so that the policy has a better starting point. To run it use:
uv run python -m lerobot.scripts.rl_offline --config path/to/config.yaml
[!NOTE] Offline training runs validation probes (attention maps, action drift, etc.) that render MP4 artifacts. On Linux these can fail with
[Errno 12] Cannot allocate memory— a virtual-memory overcommit accounting quirk with large PyTorch processes, not actual OOM. Persistent fix:echo 'vm.overcommit_memory = 1' | sudo tee /etc/sysctl.d/99-overcommit.conf sudo sysctl --systemBackground and tradeoffs in docs/engineering_notes/runbooks/system_overcommit.md. If you'd rather not touch sysctl, set val_on_start to false and a large number to val_freq.
Once the offline training has run for a while, set pretrained_path in the config to the resulting checkpoint and proceed to online training.
First run the learner script:
uv run python -m lerobot.rl.rl_learner --config path/to/config.yaml
and then on another terminal run the actor script:
uv run python -m lerobot.rl.rl_actor_async --config path/to/config.yaml
The learner will automatically save buffers to disk with the online data. After processing, those can be reused for the next round of offline or online training.
Following the suggestions from the RECAP paper, we suggest to retrain every time from the base model, and just include the additional data to avoid drift.
[!IMPORTANT] While these instructions might get you started, we recommend reading the advanced usage guide for better results. The $\pi_{0.5}$ variant lives at advanced_usage_pi05.md.
Highly recommended before first inference: set
action_clamp_limits. Teleop the arm through its safe range, record min/max per joint, then set the limits in degrees:policy: action_clamp_limits: - [-150, 150] # joint 1 - [-150, 0] # joint 2 - ... # one [min, max] per jointAnything outside isclamped before reaching the servos.
Once your config has a trained model, camera indices, and follower/leader ports, run inference with:
uv run python -m lerobot.rl.inference_async --config path/to/config.yaml
Note: Initial model loading takes 1 to 2 minutes.
Many important details were omitted in this introduction, and as we all know the devil is in the details, especially in a research-oriented repo. We strongly recommend reading the rest of the documentation, which is structured as:
(top 30 of 210)
Python
99.6%