conda create -n a1 python=3.10
conda activate a1
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -e .[all]
pip install --no-deps --force-reinstall git+https://github.com/moojink/dlimp_openvla
pip install -r requirements.txt
๐ cp .env.example .env.personal
Edit .env.personal with your personal settings:
# Example content
CONDA_ROOT=/path/to/conda
CONDA_ENV=a1
WANDB_ENTITY=your_entity
WANDB_PROJECT=your_project
๐ This file is Git-ignored and won't be committed!
source .env.personal
โ ๏ธ Security Note:
.env.personalcontains sensitive information (paths, API keys, etc.). Do NOT commit it to Git.
Start the API server for model inference:
๐ฅ๏ธ bash deploy/deploy.sh --weight /put/checkpoint/here --port <port>
๐ Arguments:
| Argument | Required | Description |
|---|---|---|
--weight | โ | Path to model checkpoint |
--port | โ | Server port (auto-selected if not provided) |
--norm | โ | Enable normalization (0 or 1) |
โจ Example:
bash deploy/deploy.sh --weight ./model/checkpoints/pretrain --port 8000
๐ฅ git submodule update --init robot_experiments/libero/LIBERO
๐ฅ pip install -e robot_experiments/libero/LIBERO
| Mode | Command | Description |
|---|---|---|
| ๐ฏ Standard | bash eval_libero.sh | Standard evaluation |
| โก Early Exit | bash eval_libero_exit.sh | Evaluates all 4 LIBERO task suites |
VLABench evaluation requires running both a server and a client in separate terminals.
๐ Setup Steps:
| Step | Action | Command |
|---|---|---|
| 1๏ธโฃ | Install VLABench | pip install -r ... |
| 2๏ธโฃ | Download Assets | python scripts/download_assets.py |
| 3๏ธโฃ | Start Server ๐ป | bash deploy/deploy.sh ... |
| 4๏ธโฃ | Run Client ๐ฎ | python eval_client.py |
๐ง Step 1: Install VLABench
pip install -r robot_experiments/vlabench/VLABench/requirements.txt
pip install -e robot_experiments/vlabench/VLABench
๐ฅ Step 2: Download Assets (if not already downloaded)
cd robot_experiments/vlabench/VLABench
python scripts/download_assets.py --choice all
๐ฅ๏ธ Step 3: Start the Evaluation Server (Terminal 1)
# Load environment and start the API server
bash deploy/deploy.sh --weight <path_to_checkpoint> --port 8000
๐ฎ Step 4: Run the Evaluation Client (Terminal 2)
cd robot_experiments/vlabench
python eval_client.py
โ ๏ธ Note: The server and client must run in separate terminals. The server loads the model and waits for client connections, while the client sends evaluation requests and receives results.
RoboChallenge evaluation is executed through the run_task.py script, supporting two modes:
| ๐ฎ Mode | ๐ Description | ๐ฏ Purpose |
|---|---|---|
mock | Automatic evaluation using local pre-recorded data | Local testing, debugging |
real | Connect to real robot and submit official evaluation | Official competition evaluation |
๐ฆพ Supported Robot Types:
ALOHA ๐คARX5 ๐งUR5 โกFRANKA ๐ฆฟ# ๐ป Terminal 1: Deploy model
bash deploy/deploy.sh --weight <path_to_checkpoint> --port 8000
# ๐ฎ Terminal 2: Run mock evaluation
cd robot_experiments/RoboChallengeInference
python run_task.py \
--task_name open_the_drawer \
--test_type mock \
--url http://localhost:8000
# ๐ป Terminal 1: Deploy model
bash deploy/deploy.sh --weight <path_to_checkpoint> --port 8000
# ๐ค Terminal 2: Run real robot evaluation
cd robot_experiments/RoboChallengeInference
python run_task.py \
--task_name open_the_drawer \
--test_type real \
--url http://localhost:8000 \
--user_token <your_token> \
--run_id <run_id> \
--action_nums 30
๐ก Tips:
- โ Mock mode automatically starts the mock server, no manual startup required
- ๐ Real mode requires valid
user_tokenandrun_idfor official evaluation- ๐
task_namemust be defined intask_config.ROBO_CHALLENGE_TASKS
Pretraining trains the model from scratch using large-scale VLA datasets, supporting distributed training on Slurm clusters.
๐ Configuration files:
configs/experiments/pretrain.yaml - Pretraining experiment configurationconfigs/datasets/pretrain.yaml - Pretraining dataset configurationscripts/slurms/pretrain.sh - Pretraining script (runs on Slurm cluster)๐ฅ๏ธ Slurm Cluster Training:
1๏ธโฃ Configure Slurm submission script scripts/slurms/submit_job.sh:
nnodes: Number of nodes required (default 8 nodes)gpus_per_node: GPUs per node (default 8)partition and quotatype: Partition name and QOS type2๏ธโฃ Submit pretraining job:
๐ bash scripts/slurms/submit_job.sh
๐ป Single-node multi-GPU training (non-Slurm):
๐ bash scripts/slurms/pretrain.sh
โ ๏ธ Note: Pretraining requires significant computational resources. Distributed training on Slurm clusters is recommended. Global batch size = 128 ร number of nodes.
LIBERO training fine-tunes on simulation data, supporting single-node multi-GPU training.
๐ Configuration files:
configs/experiments/libero_simulation.yaml - LIBERO training configurationconfigs/datasets/libero_4_tasks.yaml - LIBERO 4-task dataset configuration๐ Run training:
bash train_libero.sh
VLAbench training fine-tunes in the VLAbench simulation environment.
๐ Configuration files:
configs/experiments/vlabench.yaml - VLAbench training configurationconfigs/datasets/vlabench.yaml - VLAbench dataset configuration๐ Run training:
bash train_vlabench.sh
RoboChallenge training uses the train_rc.sh script to fine-tune on specific tasks (e.g., open_the_drawer, put_cup_on_coaster).
๐ Configuration files:
| File | Description |
|---|---|
configs/experiments/rc_open_the_drawer.yaml | ๐๏ธ Open the drawer task config |
configs/experiments/rc_put_cup_on_coaster.yaml | โ Put cup on coaster task config |
configs/datasets/rc_*.yaml | ๐ค Dataset configs (ARX5, etc.) |
๐ Run training:
bash train_rc.sh
๐ก Tip: Modify the
vla_config_pathvariable in the script to switch between different task configurations.
| Model | Description | Checkpoint |
|---|---|---|
| pretrain | Pretrained model on large-scale VLA datasets | Link |
| libero | Fine-tuned on LIBERO simulation tasks | Link |
| libero_exit | LIBERO model with early exit mechanism | Link |
| vlabench | Fine-tuned on VLABench simulation tasks | Link |
| rc_put_cup_on_coaster | Fine-tuned on RoboChallenge put cup task | Link |
| rc_open_the_drawer | Fine-tuned on RoboChallenge open drawer task | Link |
| Dataset | Description | Download |
|---|---|---|
| Droid | DROID dataset for robotic manipulation | Link |
| RoboChallenge | RoboChallenge competition data | Link |
| RoboCOIN | RoboCOIN dataset | Link |
| RoboMIND | RoboMIND benchmark dataset | Link |
| AgiBot | AgiBot dataset | Link |
| LIBERO | LIBERO simulation tasks | Link |
| VlaBench | VLABench simulation environment | Link |
๐ Storage Paths:
model/ directorydata/ directoryโ ๏ธ Important Notes:
- RoboMIND dataset preprocessing: Before using RoboMIND dataset, you need to run the indexing script:
bash scripts/robomind_build_index.sh- LeRobot dataset patch for pretraining: Before pretraining, you must replace the LeRobot dataset file:
Replacecp a1/data/vla/lerobot_datasets_replace.py <CONDA_ENV_PATH>/lib/python3.10/site-packages/lerobot/datasets/lerobot_dataset.py<CONDA_ENV_PATH>with your actual conda environment path (e.g.,/path/to/conda/envs/a1)
๐ Note: Please fill in the actual download links for models and datasets in the table above.
This project is built upon the Molmo project. We thank the Allen Institute for AI for their excellent open-source work.
If you find this work useful for your research, please consider citing:
@misc{zhang2026a1fullytransparentopensource,
title={A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model},
author={Kaidong Zhang and Jian Zhang and Rongtao Xu and Yu Sun and Shuoshuo Xue and Youpeng Wen and Xiaoyu Guo and Minghao Guo and Weijia Liufu and Liu Zihou and Kangyi Ji and Yangsong Zhang and Jiarun Zhu and Jingzhi Liu and Zihang Li and Ruiyi Chen and Meng Cao and Jingming Zhang and Shen Zhao and Xiaojun Chang and Feng Zheng and Ivan Laptev and Xiaodan Liang},
year={2026},
eprint={2604.05672},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2604.05672},
}
conda create -n a1 python=3.10
conda activate a1
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -e .[all]
pip install --no-deps --force-reinstall git+https://github.com/moojink/dlimp_openvla
pip install -r requirements.txt
๐ cp .env.example .env.personal
Edit .env.personal with your personal settings:
# Example content
CONDA_ROOT=/path/to/conda
CONDA_ENV=a1
WANDB_ENTITY=your_entity
WANDB_PROJECT=your_project
๐ This file is Git-ignored and won't be committed!
source .env.personal
โ ๏ธ Security Note:
.env.personalcontains sensitive information (paths, API keys, etc.). Do NOT commit it to Git.
Start the API server for model inference:
๐ฅ๏ธ bash deploy/deploy.sh --weight /put/checkpoint/here --port <port>
๐ Arguments:
| Argument | Required | Description |
|---|---|---|
--weight | โ | Path to model checkpoint |
--port | โ | Server port (auto-selected if not provided) |
--norm | โ | Enable normalization (0 or 1) |
โจ Example:
bash deploy/deploy.sh --weight ./model/checkpoints/pretrain --port 8000
๐ฅ git submodule update --init robot_experiments/libero/LIBERO
๐ฅ pip install -e robot_experiments/libero/LIBERO
| Mode | Command | Description |
|---|---|---|
| ๐ฏ Standard | bash eval_libero.sh | Standard evaluation |
| โก Early Exit | bash eval_libero_exit.sh | Evaluates all 4 LIBERO task suites |
VLABench evaluation requires running both a server and a client in separate terminals.
๐ Setup Steps:
| Step | Action | Command |
|---|---|---|
| 1๏ธโฃ | Install VLABench | pip install -r ... |
| 2๏ธโฃ | Download Assets | python scripts/download_assets.py |
| 3๏ธโฃ | Start Server ๐ป | bash deploy/deploy.sh ... |
| 4๏ธโฃ | Run Client ๐ฎ | python eval_client.py |
๐ง Step 1: Install VLABench
pip install -r robot_experiments/vlabench/VLABench/requirements.txt
pip install -e robot_experiments/vlabench/VLABench
๐ฅ Step 2: Download Assets (if not already downloaded)
cd robot_experiments/vlabench/VLABench
python scripts/download_assets.py --choice all
๐ฅ๏ธ Step 3: Start the Evaluation Server (Terminal 1)
# Load environment and start the API server
bash deploy/deploy.sh --weight <path_to_checkpoint> --port 8000
๐ฎ Step 4: Run the Evaluation Client (Terminal 2)
cd robot_experiments/vlabench
python eval_client.py
โ ๏ธ Note: The server and client must run in separate terminals. The server loads the model and waits for client connections, while the client sends evaluation requests and receives results.
RoboChallenge evaluation is executed through the run_task.py script, supporting two modes:
| ๐ฎ Mode | ๐ Description | ๐ฏ Purpose |
|---|---|---|
mock | Automatic evaluation using local pre-recorded data | Local testing, debugging |
real | Connect to real robot and submit official evaluation | Official competition evaluation |
๐ฆพ Supported Robot Types:
ALOHA ๐คARX5 ๐งUR5 โกFRANKA ๐ฆฟ# ๐ป Terminal 1: Deploy model
bash deploy/deploy.sh --weight <path_to_checkpoint> --port 8000
# ๐ฎ Terminal 2: Run mock evaluation
cd robot_experiments/RoboChallengeInference
python run_task.py \
--task_name open_the_drawer \
--test_type mock \
--url http://localhost:8000
# ๐ป Terminal 1: Deploy model
bash deploy/deploy.sh --weight <path_to_checkpoint> --port 8000
# ๐ค Terminal 2: Run real robot evaluation
cd robot_experiments/RoboChallengeInference
python run_task.py \
--task_name open_the_drawer \
--test_type real \
--url http://localhost:8000 \
--user_token <your_token> \
--run_id <run_id> \
--action_nums 30
๐ก Tips:
- โ Mock mode automatically starts the mock server, no manual startup required
- ๐ Real mode requires valid
user_tokenandrun_idfor official evaluation- ๐
task_namemust be defined intask_config.ROBO_CHALLENGE_TASKS
Pretraining trains the model from scratch using large-scale VLA datasets, supporting distributed training on Slurm clusters.
๐ Configuration files:
configs/experiments/pretrain.yaml - Pretraining experiment configurationconfigs/datasets/pretrain.yaml - Pretraining dataset configurationscripts/slurms/pretrain.sh - Pretraining script (runs on Slurm cluster)๐ฅ๏ธ Slurm Cluster Training:
1๏ธโฃ Configure Slurm submission script scripts/slurms/submit_job.sh:
nnodes: Number of nodes required (default 8 nodes)gpus_per_node: GPUs per node (default 8)partition and quotatype: Partition name and QOS type2๏ธโฃ Submit pretraining job:
๐ bash scripts/slurms/submit_job.sh
๐ป Single-node multi-GPU training (non-Slurm):
๐ bash scripts/slurms/pretrain.sh
โ ๏ธ Note: Pretraining requires significant computational resources. Distributed training on Slurm clusters is recommended. Global batch size = 128 ร number of nodes.
LIBERO training fine-tunes on simulation data, supporting single-node multi-GPU training.
๐ Configuration files:
configs/experiments/libero_simulation.yaml - LIBERO training configurationconfigs/datasets/libero_4_tasks.yaml - LIBERO 4-task dataset configuration๐ Run training:
bash train_libero.sh
VLAbench training fine-tunes in the VLAbench simulation environment.
๐ Configuration files:
configs/experiments/vlabench.yaml - VLAbench training configurationconfigs/datasets/vlabench.yaml - VLAbench dataset configuration๐ Run training:
bash train_vlabench.sh
RoboChallenge training uses the train_rc.sh script to fine-tune on specific tasks (e.g., open_the_drawer, put_cup_on_coaster).
๐ Configuration files:
| File | Description |
|---|---|
configs/experiments/rc_open_the_drawer.yaml | ๐๏ธ Open the drawer task config |
configs/experiments/rc_put_cup_on_coaster.yaml | โ Put cup on coaster task config |
configs/datasets/rc_*.yaml | ๐ค Dataset configs (ARX5, etc.) |
๐ Run training:
bash train_rc.sh
๐ก Tip: Modify the
vla_config_pathvariable in the script to switch between different task configurations.
| Model | Description | Checkpoint |
|---|---|---|
| pretrain | Pretrained model on large-scale VLA datasets | Link |
| libero | Fine-tuned on LIBERO simulation tasks | Link |
| libero_exit | LIBERO model with early exit mechanism | Link |
| vlabench | Fine-tuned on VLABench simulation tasks | Link |
| rc_put_cup_on_coaster | Fine-tuned on RoboChallenge put cup task | Link |
| rc_open_the_drawer | Fine-tuned on RoboChallenge open drawer task | Link |
| Dataset | Description | Download |
|---|---|---|
| Droid | DROID dataset for robotic manipulation | Link |
| RoboChallenge | RoboChallenge competition data | Link |
| RoboCOIN | RoboCOIN dataset | Link |
| RoboMIND | RoboMIND benchmark dataset | Link |
| AgiBot | AgiBot dataset | Link |
| LIBERO | LIBERO simulation tasks | Link |
| VlaBench | VLABench simulation environment | Link |
๐ Storage Paths:
model/ directorydata/ directoryโ ๏ธ Important Notes:
- RoboMIND dataset preprocessing: Before using RoboMIND dataset, you need to run the indexing script:
bash scripts/robomind_build_index.sh- LeRobot dataset patch for pretraining: Before pretraining, you must replace the LeRobot dataset file:
Replacecp a1/data/vla/lerobot_datasets_replace.py <CONDA_ENV_PATH>/lib/python3.10/site-packages/lerobot/datasets/lerobot_dataset.py<CONDA_ENV_PATH>with your actual conda environment path (e.g.,/path/to/conda/envs/a1)
๐ Note: Please fill in the actual download links for models and datasets in the table above.
This project is built upon the Molmo project. We thank the Allen Institute for AI for their excellent open-source work.
If you find this work useful for your research, please consider citing:
@misc{zhang2026a1fullytransparentopensource,
title={A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model},
author={Kaidong Zhang and Jian Zhang and Rongtao Xu and Yu Sun and Shuoshuo Xue and Youpeng Wen and Xiaoyu Guo and Minghao Guo and Weijia Liufu and Liu Zihou and Kangyi Ji and Yangsong Zhang and Jiarun Zhu and Jingzhi Liu and Zihang Li and Ruiyi Chen and Meng Cao and Jingming Zhang and Shen Zhao and Xiaojun Chang and Feng Zheng and Ivan Laptev and Xiaodan Liang},
year={2026},
eprint={2604.05672},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2604.05672},
}