Boyu Chen1,2,* ·
Yi Chen1,3,* ·
Lu Qiu1,3 ·
Jerry Bai1 ·
Yuying Ge1,† ·
Yixiao Ge1
1XPENG Robotics ·
2Tsinghua University ·
3The University of Hong Kong
*Equal contribution †Corresponding author
Correspondence: yyge13@gmail.com
Project page: https://xpeng-robotics.github.io/unit/
Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that learns a unified physical language for human-to-humanoid manipulation transfer. Grounded in the philosophy that heterogeneous kinematics share consistent visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch integrates these purified modalities into a shared discrete latent space of cross-embodiment physical intents.
We validate UniT across two paradigms: (1) Policy Learning (VLA-UniT): By predicting these unified tokens, VLA-UniT achieves state-of-the-art performance with high data efficiency on the RoboCasa GR1 benchmark. Leveraging diverse human data further improves out-of-distribution (OOD) generalization in simulation and real-world deployment, and enables zero-shot task transfer in the real world. (2) World Modeling (WM-UniT): By aligning cross-embodiment dynamics via unified tokens as conditions, it supports direct human-to-humanoid action-conditioned generation, translating human knowledge into enhanced action controllability for humanoid video generation.
Ultimately, by inducing a more aligned cross-embodiment representation (empirically supported by t-SNE visualizations revealing improved alignment of human and humanoid features), UniT offers a scalable path to distill human priors into humanoid manipulation capabilities.
Inside Fe₀ is our follow-up embodied foundation model that scales UniT to heterogeneous robots and human data (on the order of 10k hours—roughly 1,587× the target teleoperation anchor in duration), unlocking broader generalization across multiple axes. Built on UniT’s unified physical language, Fe₀ is evaluated on IRON-R01-1.11 through an L1–L5 framework: from visual robustness and semantic grounding, through relational reasoning and compositional planning, to novel action execution. The study also maps where transfer is strongest (vision and semantics) versus where it remains fragile (mid-level relational/temporal reasoning and precise physical execution). See the blog post for details.
Requirements: conda; NVIDIA GPU and a PyTorch/CUDA build that matches your driver.
Set eagle_path in the model config JSON to point at the vision-language backbone (e.g. a Qwen2.5-VL checkpoint on Hugging Face if you use the default recipe).
Bootstrap the Python env (optional overrides CONDA_ENV_NAME, PIP_INDEX_URL): examples/environment_setup.sh.
Developed with CUDA 12.x; default hyperparameters assume 120+ GB GPU memory per device.
bash examples/environment_setup.sh
conda activate unit
Requires the same unit base installation as Training environment, plus simulation dependencies for RoboCasa / MuJoCo eval. Rollout was tested on 24 GB GPUs (NVIDIA RTX 4090).
# System dependencies
sudo apt-get install -y libegl1-mesa libegl1-mesa-dev libosmesa6-dev patchelf
third_party/ (local editable installs, not committed — see .gitignore):
mkdir -p third_party
# Install robosuite v1.5.1
git clone https://github.com/ARISE-Initiative/robosuite.git third_party/robosuite
cd third_party/robosuite
git checkout v1.5.1
pip install -e .
cd ../..
# Install robocasa-gr1-tabletop-tasks
git clone https://github.com/robocasa/robocasa-gr1-tabletop-tasks.git third_party/robocasa-gr1-tabletop-tasks
pip install -e third_party/robocasa-gr1-tabletop-tasks
Required patch: Apply the following fix to third_party/robocasa-gr1-tabletop-tasks/robocasa/models/objects/kitchen_object_utils.py.
In the sample_kitchen_object_helper() function, after the line reg_choices = reg_choices[split_th:] under elif split == "B":, add:
if "assets/objects/sketchfab/basket/" in reg_choices[0]:
reg_choices = [c for c in reg_choices if not c.endswith('basket_4/model.xml')]
Download tabletop assets:
python third_party/robocasa-gr1-tabletop-tasks/robocasa/scripts/download_tabletop_assets.py -y
Pipelines and eval env vars: examples/README.md.
Source. nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim on Hugging Face (HDF5 + LeRobot).
Joints-only (default). Main UniT recipes use joint actions only—no EEF. Download the LeRobot splits from that dataset; you do not need HDF5 or the augmentation steps below.
EEF + joints (optional). Download both HDF5 and LeRobot. Install the IK stack used by the EEF extraction scripts:
pip install pin==3.9.0
pip install mink==0.0.5
Raw LeRobot data only includes joint angles. To add end-effector poses (3D position + 6D rotation), run two steps per task:
# Step 1: Extract EEF poses by replaying trajectories in the simulator
# Input: HDF5 file (e.g., HDF5/TaskName.hdf5)
# Output: Replay parquet directory (e.g., Replay-Correct/TaskName/)
python preprocessing/extract_and_visualize_3d-pos_6d-rot_from_gr1.py \
--dataset /path/to/HDF5/TaskName.hdf5 \
--output_dir /path/to/Replay-Correct/TaskName \
--render_image_names egoview \
--verbose \
--render_height 800 \
--render_width 1280 \
--num_parallel_jobs 10
# Step 2: Augment the LeRobot dataset with extracted poses
# Input: LeRobot dataset + Replay parquet from Step 1
# Output: Augmented LeRobot dataset (LeRobot-AugPosRot-Correct/)
python preprocessing/aug_lerobot_data.py \
--lerobot_base_path /path/to/LeRobot/gr1_unified.TaskName \
--replay_base_path /path/to/Replay-Correct/TaskName/parquet/ \
--output_base_path /path/to/LeRobot-AugPosRot-Correct/gr1_unified.TaskName
Repeat for all 24 tasks. The final augmented datasets under LeRobot-AugPosRot-Correct/ are used for training.
GR1_DATASET_DIR: see examples/README.md.
Download raw EgoDex data from apple/ml-egodex.
Install LeRobot (required for data format conversion only):
git clone https://github.com/huggingface/lerobot.git
cd lerobot
git checkout d602e8169cbad9e93a4a3b3ee1dd8b332af7ebf8
pip install -e .
pip install tyro h5py
cd ..
# Step 1: Convert raw EgoDex HDF5 to LeRobot v2.1 format
python preprocessing/convert_egodex_data_to_lerobot.py \
--raw_dir /path/to/egodex/basic_pick_place \
--repo_id /path/to/egodex_lerobot/part2/basic_pick_place
# Step 2: Convert from LeRobot v2.1 to v2.0 (GR00T-compatible format)
python preprocessing/convert_dataset_v21_to_v20_gr00t.py \
--repo-id=part2/basic_pick_place \
--root=/path/to/egodex_lerobot/part2/basic_pick_place
EGODEX_DATASET_DIR: see examples/README.md.
UniT training is staged: tokenizer → dual-system pretrain on the chosen data mix, then dual-system finetune on GR1 only when using the few-shot recipe. Scripts under examples/ run these steps in order; a failure aborts the pipeline and the next step always resumes from the previous checkpoint.
We ship two GR1-Joints entry points:
examples/run_gr1_100_egodex.sh — EgoDex + GR1-100: tokenizer, mixed pretrain (EgoDex + GR1), then finetune on GR1 alone.examples/run_gr1_full.sh — GR1-full: tokenizer, then dual-system pretrain on all GR1-Joints data (no separate finetune stage in this script).examples/README.md lists stage details, defaults, environment overrides, and copy-paste commands.
Pretrained VLA-UniT checkpoints are released on Hugging Face: xpeng-robotics/VLA-UniT-checkpoints.
Before evaluating a released checkpoint: set
unit_cfg.groot_tokenizer_pathin the checkpoint'sconfig.jsonto the corresponding tokenizer.
Closed-loop rollout against a trained checkpoint in the RoboCasa GR1 simulation environment:
bash examples/run_eval.sh <model_path> <eval_type>
Replace the placeholders:
<model_path> — absolute or repo-relative path to your trained checkpoint directory (the same one consumed by scripts/inference_service_unit.py).<eval_type> — one of the five entries below.Evaluation types:
| Type | Description | Tasks |
|---|---|---|
id | In-distribution training tasks | 24 (6 PnPClose + 18 PosttrainPnPNovel SplitA) |
ood_object_appearance | OOD unseen object appearance | 18 (EvalPnPNovel SplitB) |
ood_container_combination | OOD unseen source-target containers | 14 (PretrainPnPNovel SplitA) |
ood_object_type | OOD unseen object types | 32 (PretrainPnPBase SplitA) |
unseen_close | OOD variants of PnP*Close | 9 |
Example:
# In-distribution evaluation
bash examples/run_eval.sh /path/to/checkpoint id
# OOD unseen object appearance
bash examples/run_eval.sh /path/to/checkpoint ood_object_appearance
# Customize port, env count, episodes per task
PORT=8891 N_ENVS=1 N_EPISODES=50 \
bash examples/run_eval.sh /path/to/checkpoint ood_object_type
run_eval.sh launches the inference server, retries each task up to 5 times against scripts/simulation/simulation_service_unit.py, then writes per-task success rates to <model_path>/evaluation_sim_<eval_type>_${N_ENVS}envs${EVAL_TAG}/results.json.
Tunables (env vars): DATA_CONFIG, N_ENVS, N_EPISODES, CUDA_VISIBLE_DEVICES, PORT, EVAL_TAG.
Open-loop action-reconstruction MSE on held-out GR1-Joints trajectories, no simulator required:
bash examples/run_eval_loss.sh <model_path>
<model_path> is the trained checkpoint directory (same meaning as in online evaluation). The script runs scripts/eval_policy_unit.py over the 24 GR1 training datasets, using the tail slice of each dataset (DATA_SPLIT) as held-out trajectories, and writes per-task action plots and metrics to <model_path>/eval_action_dual_system_${DATA_SPLIT}/.
Examples:
# Default settings (last 2 trajectories per task, GR1-Joints data config)
bash examples/run_eval_loss.sh /path/to/checkpoint
# Point at a local GR1 LeRobot root
GR1_DATASET_DIR=/abs/path/to/gr1 \
bash examples/run_eval_loss.sh /path/to/checkpoint
Tunables (env vars): GR1_DATASET_DIR, DATA_CONFIG, DATA_SPLIT (default [-2:]), TRAJS (default 2), CUDA_VISIBLE_DEVICES.
Note: Default
GR1_DATASET_DIRinexamples/run_eval_loss.shpoints at the EEF-augmented layout (LeRobot-AugPosRot-Correct/). For the joints-only main recipe, point it at your plain LeRobot root.
Overall in-distribution success rates from examples/run_eval.sh ... id (24 tasks, 50 episodes each, N_ENVS=1):
| Recipe | Script | Overall SR |
|---|---|---|
| GR1-full | examples/run_gr1_full.sh | 66.4% |
| EgoDex + GR1-100 (few-shot) | examples/run_gr1_100_egodex.sh | 50.9% |
Per-task numbers are in docs/evaluation_id_results.md.
If you find this work useful, please cite:
@article{chen2026unit,
title={{UniT}: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling},
author={Chen, Boyu and Chen, Yi and Qiu, Lu and Bai, Jerry and Ge, Yuying and Ge, Yixiao},
journal={arXiv preprint arXiv:2604.19734},
year={2026},
eprint={2604.19734},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2604.19734}
}
This codebase is built on top of NVIDIA Isaac GR00T N1.5, an open foundation model for generalized humanoid robot policies. We thank NVIDIA for open-sourcing the GR00T N1.5 model, data pipeline, and training infrastructure, which served as the foundation for our work. We also thank the authors of DIAL, EgoDex, RoboCasa, LeRobot, and lucidrains/vector-quantize-pytorch (for the VQ-VAE building blocks) for their open-source contributions.
This project is licensed under the Apache License 2.0.
The code release is in progress. Data preparation, the UniT tokenizer, VLA-UniT (training + RoboCasa GR1 evaluation), and pretrained VLA-UniT checkpoints (Hugging Face) are available now. WM-UniT and the real-world deployment stack will follow.
Planned release order:
For questions about the paper or the upcoming release, please open an issue in this repository, or reach out to:
4 commits
Python
100.0%
Boyu Chen1,2,* ·
Yi Chen1,3,* ·
Lu Qiu1,3 ·
Jerry Bai1 ·
Yuying Ge1,† ·
Yixiao Ge1
1XPENG Robotics ·
2Tsinghua University ·
3The University of Hong Kong
*Equal contribution †Corresponding author
Correspondence: yyge13@gmail.com
Project page: https://xpeng-robotics.github.io/unit/
Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that learns a unified physical language for human-to-humanoid manipulation transfer. Grounded in the philosophy that heterogeneous kinematics share consistent visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch integrates these purified modalities into a shared discrete latent space of cross-embodiment physical intents.
We validate UniT across two paradigms: (1) Policy Learning (VLA-UniT): By predicting these unified tokens, VLA-UniT achieves state-of-the-art performance with high data efficiency on the RoboCasa GR1 benchmark. Leveraging diverse human data further improves out-of-distribution (OOD) generalization in simulation and real-world deployment, and enables zero-shot task transfer in the real world. (2) World Modeling (WM-UniT): By aligning cross-embodiment dynamics via unified tokens as conditions, it supports direct human-to-humanoid action-conditioned generation, translating human knowledge into enhanced action controllability for humanoid video generation.
Ultimately, by inducing a more aligned cross-embodiment representation (empirically supported by t-SNE visualizations revealing improved alignment of human and humanoid features), UniT offers a scalable path to distill human priors into humanoid manipulation capabilities.
Inside Fe₀ is our follow-up embodied foundation model that scales UniT to heterogeneous robots and human data (on the order of 10k hours—roughly 1,587× the target teleoperation anchor in duration), unlocking broader generalization across multiple axes. Built on UniT’s unified physical language, Fe₀ is evaluated on IRON-R01-1.11 through an L1–L5 framework: from visual robustness and semantic grounding, through relational reasoning and compositional planning, to novel action execution. The study also maps where transfer is strongest (vision and semantics) versus where it remains fragile (mid-level relational/temporal reasoning and precise physical execution). See the blog post for details.
Requirements: conda; NVIDIA GPU and a PyTorch/CUDA build that matches your driver.
Set eagle_path in the model config JSON to point at the vision-language backbone (e.g. a Qwen2.5-VL checkpoint on Hugging Face if you use the default recipe).
Bootstrap the Python env (optional overrides CONDA_ENV_NAME, PIP_INDEX_URL): examples/environment_setup.sh.
Developed with CUDA 12.x; default hyperparameters assume 120+ GB GPU memory per device.
bash examples/environment_setup.sh
conda activate unit
Requires the same unit base installation as Training environment, plus simulation dependencies for RoboCasa / MuJoCo eval. Rollout was tested on 24 GB GPUs (NVIDIA RTX 4090).
# System dependencies
sudo apt-get install -y libegl1-mesa libegl1-mesa-dev libosmesa6-dev patchelf
third_party/ (local editable installs, not committed — see .gitignore):
mkdir -p third_party
# Install robosuite v1.5.1
git clone https://github.com/ARISE-Initiative/robosuite.git third_party/robosuite
cd third_party/robosuite
git checkout v1.5.1
pip install -e .
cd ../..
# Install robocasa-gr1-tabletop-tasks
git clone https://github.com/robocasa/robocasa-gr1-tabletop-tasks.git third_party/robocasa-gr1-tabletop-tasks
pip install -e third_party/robocasa-gr1-tabletop-tasks
Required patch: Apply the following fix to third_party/robocasa-gr1-tabletop-tasks/robocasa/models/objects/kitchen_object_utils.py.
In the sample_kitchen_object_helper() function, after the line reg_choices = reg_choices[split_th:] under elif split == "B":, add:
if "assets/objects/sketchfab/basket/" in reg_choices[0]:
reg_choices = [c for c in reg_choices if not c.endswith('basket_4/model.xml')]
Download tabletop assets:
python third_party/robocasa-gr1-tabletop-tasks/robocasa/scripts/download_tabletop_assets.py -y
Pipelines and eval env vars: examples/README.md.
Source. nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim on Hugging Face (HDF5 + LeRobot).
Joints-only (default). Main UniT recipes use joint actions only—no EEF. Download the LeRobot splits from that dataset; you do not need HDF5 or the augmentation steps below.
EEF + joints (optional). Download both HDF5 and LeRobot. Install the IK stack used by the EEF extraction scripts:
pip install pin==3.9.0
pip install mink==0.0.5
Raw LeRobot data only includes joint angles. To add end-effector poses (3D position + 6D rotation), run two steps per task:
# Step 1: Extract EEF poses by replaying trajectories in the simulator
# Input: HDF5 file (e.g., HDF5/TaskName.hdf5)
# Output: Replay parquet directory (e.g., Replay-Correct/TaskName/)
python preprocessing/extract_and_visualize_3d-pos_6d-rot_from_gr1.py \
--dataset /path/to/HDF5/TaskName.hdf5 \
--output_dir /path/to/Replay-Correct/TaskName \
--render_image_names egoview \
--verbose \
--render_height 800 \
--render_width 1280 \
--num_parallel_jobs 10
# Step 2: Augment the LeRobot dataset with extracted poses
# Input: LeRobot dataset + Replay parquet from Step 1
# Output: Augmented LeRobot dataset (LeRobot-AugPosRot-Correct/)
python preprocessing/aug_lerobot_data.py \
--lerobot_base_path /path/to/LeRobot/gr1_unified.TaskName \
--replay_base_path /path/to/Replay-Correct/TaskName/parquet/ \
--output_base_path /path/to/LeRobot-AugPosRot-Correct/gr1_unified.TaskName
Repeat for all 24 tasks. The final augmented datasets under LeRobot-AugPosRot-Correct/ are used for training.
GR1_DATASET_DIR: see examples/README.md.
Download raw EgoDex data from apple/ml-egodex.
Install LeRobot (required for data format conversion only):
git clone https://github.com/huggingface/lerobot.git
cd lerobot
git checkout d602e8169cbad9e93a4a3b3ee1dd8b332af7ebf8
pip install -e .
pip install tyro h5py
cd ..
# Step 1: Convert raw EgoDex HDF5 to LeRobot v2.1 format
python preprocessing/convert_egodex_data_to_lerobot.py \
--raw_dir /path/to/egodex/basic_pick_place \
--repo_id /path/to/egodex_lerobot/part2/basic_pick_place
# Step 2: Convert from LeRobot v2.1 to v2.0 (GR00T-compatible format)
python preprocessing/convert_dataset_v21_to_v20_gr00t.py \
--repo-id=part2/basic_pick_place \
--root=/path/to/egodex_lerobot/part2/basic_pick_place
EGODEX_DATASET_DIR: see examples/README.md.
UniT training is staged: tokenizer → dual-system pretrain on the chosen data mix, then dual-system finetune on GR1 only when using the few-shot recipe. Scripts under examples/ run these steps in order; a failure aborts the pipeline and the next step always resumes from the previous checkpoint.
We ship two GR1-Joints entry points:
examples/run_gr1_100_egodex.sh — EgoDex + GR1-100: tokenizer, mixed pretrain (EgoDex + GR1), then finetune on GR1 alone.examples/run_gr1_full.sh — GR1-full: tokenizer, then dual-system pretrain on all GR1-Joints data (no separate finetune stage in this script).examples/README.md lists stage details, defaults, environment overrides, and copy-paste commands.
Pretrained VLA-UniT checkpoints are released on Hugging Face: xpeng-robotics/VLA-UniT-checkpoints.
Before evaluating a released checkpoint: set
unit_cfg.groot_tokenizer_pathin the checkpoint'sconfig.jsonto the corresponding tokenizer.
Closed-loop rollout against a trained checkpoint in the RoboCasa GR1 simulation environment:
bash examples/run_eval.sh <model_path> <eval_type>
Replace the placeholders:
<model_path> — absolute or repo-relative path to your trained checkpoint directory (the same one consumed by scripts/inference_service_unit.py).<eval_type> — one of the five entries below.Evaluation types:
| Type | Description | Tasks |
|---|---|---|
id | In-distribution training tasks | 24 (6 PnPClose + 18 PosttrainPnPNovel SplitA) |
ood_object_appearance | OOD unseen object appearance | 18 (EvalPnPNovel SplitB) |
ood_container_combination | OOD unseen source-target containers | 14 (PretrainPnPNovel SplitA) |
ood_object_type | OOD unseen object types | 32 (PretrainPnPBase SplitA) |
unseen_close | OOD variants of PnP*Close | 9 |
Example:
# In-distribution evaluation
bash examples/run_eval.sh /path/to/checkpoint id
# OOD unseen object appearance
bash examples/run_eval.sh /path/to/checkpoint ood_object_appearance
# Customize port, env count, episodes per task
PORT=8891 N_ENVS=1 N_EPISODES=50 \
bash examples/run_eval.sh /path/to/checkpoint ood_object_type
run_eval.sh launches the inference server, retries each task up to 5 times against scripts/simulation/simulation_service_unit.py, then writes per-task success rates to <model_path>/evaluation_sim_<eval_type>_${N_ENVS}envs${EVAL_TAG}/results.json.
Tunables (env vars): DATA_CONFIG, N_ENVS, N_EPISODES, CUDA_VISIBLE_DEVICES, PORT, EVAL_TAG.
Open-loop action-reconstruction MSE on held-out GR1-Joints trajectories, no simulator required:
bash examples/run_eval_loss.sh <model_path>
<model_path> is the trained checkpoint directory (same meaning as in online evaluation). The script runs scripts/eval_policy_unit.py over the 24 GR1 training datasets, using the tail slice of each dataset (DATA_SPLIT) as held-out trajectories, and writes per-task action plots and metrics to <model_path>/eval_action_dual_system_${DATA_SPLIT}/.
Examples:
# Default settings (last 2 trajectories per task, GR1-Joints data config)
bash examples/run_eval_loss.sh /path/to/checkpoint
# Point at a local GR1 LeRobot root
GR1_DATASET_DIR=/abs/path/to/gr1 \
bash examples/run_eval_loss.sh /path/to/checkpoint
Tunables (env vars): GR1_DATASET_DIR, DATA_CONFIG, DATA_SPLIT (default [-2:]), TRAJS (default 2), CUDA_VISIBLE_DEVICES.
Note: Default
GR1_DATASET_DIRinexamples/run_eval_loss.shpoints at the EEF-augmented layout (LeRobot-AugPosRot-Correct/). For the joints-only main recipe, point it at your plain LeRobot root.
Overall in-distribution success rates from examples/run_eval.sh ... id (24 tasks, 50 episodes each, N_ENVS=1):
| Recipe | Script | Overall SR |
|---|---|---|
| GR1-full | examples/run_gr1_full.sh | 66.4% |
| EgoDex + GR1-100 (few-shot) | examples/run_gr1_100_egodex.sh | 50.9% |
Per-task numbers are in docs/evaluation_id_results.md.
If you find this work useful, please cite:
@article{chen2026unit,
title={{UniT}: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling},
author={Chen, Boyu and Chen, Yi and Qiu, Lu and Bai, Jerry and Ge, Yuying and Ge, Yixiao},
journal={arXiv preprint arXiv:2604.19734},
year={2026},
eprint={2604.19734},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2604.19734}
}
This codebase is built on top of NVIDIA Isaac GR00T N1.5, an open foundation model for generalized humanoid robot policies. We thank NVIDIA for open-sourcing the GR00T N1.5 model, data pipeline, and training infrastructure, which served as the foundation for our work. We also thank the authors of DIAL, EgoDex, RoboCasa, LeRobot, and lucidrains/vector-quantize-pytorch (for the VQ-VAE building blocks) for their open-source contributions.
This project is licensed under the Apache License 2.0.
The code release is in progress. Data preparation, the UniT tokenizer, VLA-UniT (training + RoboCasa GR1 evaluation), and pretrained VLA-UniT checkpoints (Hugging Face) are available now. WM-UniT and the real-world deployment stack will follow.
Planned release order:
For questions about the paper or the upcoming release, please open an issue in this repository, or reach out to:
4 commits
Python
100.0%