cshizhe/robot-PointAct

Official implementation of PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction (RSS'26).

Python

24

5 commits

updated Jul 17, 2026

See the code

README

PointAct

Official implementation of PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction.

PointAct project webpage PointAct paper PointAct models and data

Monolithic 3D-aware VLA teaser
Monolithic 3D-aware VLA
Dual-system 3D-aware VLA teaser
Dual-system 3D-aware VLA
PointAct teaser
PointAct

PointAct is a 3D-aware vision-language-action policy for robot manipulation. It keeps a pretrained vision-language backbone for semantic understanding and adds a dedicated point-action expert so that multi-scale 3D geometry can directly shape action decoding.

Installation

Please follow the main setup guide in INSTALLATION.md. The recommended workflow is:

  1. Create the core pointact environment for training, checkpoint loading, preprocessing, and the policy server.
  2. Create separate simulator environments.
  3. Use the server-client evaluation pipeline since the policy and simulator usually need different environments.

Supported Benchmarks

This repository currently supports three simulators:

Supported Models

This repository includes PointAct and several comparison VLA policies used in our experiments:

ModelDirectoryNotes
PointActpointact/model/vla_pointactSupports both Concerto and Utonia Point Transformer backbones
EO1pointact/model/eo1Monolithic VLA baseline
EO1-Pointpointact/model/eo1EO-1 variant with point-cloud input
QwenGR00Tpointact/model/vla_dualDual-system VLA baseline
QwenGR00T-Pointpointact/model/vla_dualVLA-Dual variant with point-cloud input
Pi0pointact/model/pi0PI0 baseline
Pi0.5pointact/model/pi05PI0.5 baseline

Acknowledgements

This codebase builds on several excellent open-source projects, especially EO-1, GR00T, and LeRobot. We thank the authors and maintainers of these libraries for making their work available to the community.

Citation

If you find PointAct useful in your research, or if you use this code, please cite:

@InProceedings{Chen_2026_PointACT,
    author    = {Chen, Shizhe and Pacaud, Paul and Schmid, Cordelia},
    title     = {PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction},
    booktitle = {Robotics: Science and Systems (RSS)},
    year      = {2026}
}

cshizhe/robot-PointAct

Official implementation of PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction (RSS'26).

Python

24

5 commits

updated Jul 17, 2026

See the code

README

PointAct

Official implementation of PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction.

PointAct project webpage PointAct paper PointAct models and data

Monolithic 3D-aware VLA teaser
Monolithic 3D-aware VLA
Dual-system 3D-aware VLA teaser
Dual-system 3D-aware VLA
PointAct teaser
PointAct

PointAct is a 3D-aware vision-language-action policy for robot manipulation. It keeps a pretrained vision-language backbone for semantic understanding and adds a dedicated point-action expert so that multi-scale 3D geometry can directly shape action decoding.

Installation

Please follow the main setup guide in INSTALLATION.md. The recommended workflow is:

  1. Create the core pointact environment for training, checkpoint loading, preprocessing, and the policy server.
  2. Create separate simulator environments.
  3. Use the server-client evaluation pipeline since the policy and simulator usually need different environments.

Supported Benchmarks

This repository currently supports three simulators:

Supported Models

This repository includes PointAct and several comparison VLA policies used in our experiments:

ModelDirectoryNotes
PointActpointact/model/vla_pointactSupports both Concerto and Utonia Point Transformer backbones
EO1pointact/model/eo1Monolithic VLA baseline
EO1-Pointpointact/model/eo1EO-1 variant with point-cloud input
QwenGR00Tpointact/model/vla_dualDual-system VLA baseline
QwenGR00T-Pointpointact/model/vla_dualVLA-Dual variant with point-cloud input
Pi0pointact/model/pi0PI0 baseline
Pi0.5pointact/model/pi05PI0.5 baseline

Acknowledgements

This codebase builds on several excellent open-source projects, especially EO-1, GR00T, and LeRobot. We thank the authors and maintainers of these libraries for making their work available to the community.

Citation

If you find PointAct useful in your research, or if you use this code, please cite:

@InProceedings{Chen_2026_PointACT,
    author    = {Chen, Shizhe and Pacaud, Paul and Schmid, Cordelia},
    title     = {PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction},
    booktitle = {Robotics: Science and Systems (RSS)},
    year      = {2026}
}

Languages

Python

95.2%

Shell

4.8%