Alibaba-DAMO-Academy/RynnVLA-002

Model

11

stars

15

commits

8

linked in READMEs

Nov 22, 2025

updated

safetensors

README

RynnVLA-002: A Unified Vision-Language-Action and World Model

If our project helps you, please give us a star ⭐ on GitHub to support us. 🙏🙏

arXiv hf_checkpoint License

🌟 Introduction

RynnVLA-002 is an autoregressive action world model that unifies action and image understanding and generation. RynnVLA-002 intergrates Vision-Language-Action (VLA) model (action model) and world model in one single framework. Compared to WorldVLA, RynnVLA-002 adds a continous Action Transformer, wrist camera input and generation, and state input. RynnVLA-002 achieves 97.4% success rate on LIBERO benchmark.


Model Zoo

VLA Model (256 * 256)

World Model (512 * 512)

GoalHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_goal370.022.2577.8419.70
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_goal336.822.1378.1319.43
ObjectHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_object1141.620.3159.5927.30
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_object877.222.1865.0322.60
SpatialHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_spatial405.422.3279.1520.28
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_spatial373.123.8882.4116.33
LongHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_10557.7318.2469.1631.60
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_10427.8619.3672.1927.78

License

All assets and code are under the Apache 2.0 license unless specified otherwise.

Citation

If you find the project helpful for your research, please consider citing our paper:

@article{cen2025WorldVLA,
  title={WorldVLA: Towards Autoregressive Action World Model},
  author={Cen, Jun and Yu, Chaohui and Yuan, Hangjie and Jiang, Yuming and Huang, Siteng and Guo, Jiayan and Li, Xin and Song, Yibing and Luo, Hao and Wang, Fan and Zhao, Deli and Chen, Hao},
  journal={arXiv preprint arXiv:},
  year={2025}
}
💡 Other featured projects from our RynnBot family ✨.

RynnVLA-001: A Vision-Language-Action Model Boosted by Generative Priors
Yuming Jiang, Siteng Huang, Shengke Xue, Yaxi Zhao, Jun Cen, Sicong Leng, Jiayan Guo, Kexiang Wang, Kehan Li, Mingxiu Chen, Fan Wang, Deli Zhao, Xin Li
github github arXiv

RynnEC: Bringing MLLMs into Embodied World
Ronghao Dang*, Yuqian Yuan*, Yunxuan Mao*, Kehan Li*, Jiangpin Liu, Zhikai Wang, Fan Wang, Deli Zhao, Xin Li
github github arXiv

RynnRCP: Open Robotics Context Protocol and RobotMotion
RynnBot Team
github github

Acknowledgment

This project builds upon Lumina-mGPT, Chemeleon, and OpenVLA. We thank these teams for their open-source contributions.

Contributors

jcenaa

15 commits

Alibaba-DAMO-Academy/RynnVLA-002

Model

11

stars

15

commits

8

linked in READMEs

Nov 22, 2025

updated

safetensors

README

RynnVLA-002: A Unified Vision-Language-Action and World Model

If our project helps you, please give us a star ⭐ on GitHub to support us. 🙏🙏

arXiv hf_checkpoint License

🌟 Introduction

RynnVLA-002 is an autoregressive action world model that unifies action and image understanding and generation. RynnVLA-002 intergrates Vision-Language-Action (VLA) model (action model) and world model in one single framework. Compared to WorldVLA, RynnVLA-002 adds a continous Action Transformer, wrist camera input and generation, and state input. RynnVLA-002 achieves 97.4% success rate on LIBERO benchmark.


Model Zoo

VLA Model (256 * 256)

World Model (512 * 512)

GoalHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_goal370.022.2577.8419.70
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_goal336.822.1378.1319.43
ObjectHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_object1141.620.3159.5927.30
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_object877.222.1865.0322.60
SpatialHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_spatial405.422.3279.1520.28
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_spatial373.123.8882.4116.33
LongHF LinkFVD↓PSNR↑SSIM↑LPIPS↓
World ModelAlibaba-DAMO-Academy/RynnVLA-002/World_model_512/libero_10557.7318.2469.1631.60
Action World ModelAlibaba-DAMO-Academy/RynnVLA-002/Action_World_model_512/libero_10427.8619.3672.1927.78

License

All assets and code are under the Apache 2.0 license unless specified otherwise.

Citation

If you find the project helpful for your research, please consider citing our paper:

@article{cen2025WorldVLA,
  title={WorldVLA: Towards Autoregressive Action World Model},
  author={Cen, Jun and Yu, Chaohui and Yuan, Hangjie and Jiang, Yuming and Huang, Siteng and Guo, Jiayan and Li, Xin and Song, Yibing and Luo, Hao and Wang, Fan and Zhao, Deli and Chen, Hao},
  journal={arXiv preprint arXiv:},
  year={2025}
}
💡 Other featured projects from our RynnBot family ✨.

RynnVLA-001: A Vision-Language-Action Model Boosted by Generative Priors
Yuming Jiang, Siteng Huang, Shengke Xue, Yaxi Zhao, Jun Cen, Sicong Leng, Jiayan Guo, Kexiang Wang, Kehan Li, Mingxiu Chen, Fan Wang, Deli Zhao, Xin Li
github github arXiv

RynnEC: Bringing MLLMs into Embodied World
Ronghao Dang*, Yuqian Yuan*, Yunxuan Mao*, Kehan Li*, Jiangpin Liu, Zhikai Wang, Fan Wang, Deli Zhao, Xin Li
github github arXiv

RynnRCP: Open Robotics Context Protocol and RobotMotion
RynnBot Team
github github

Acknowledgment

This project builds upon Lumina-mGPT, Chemeleon, and OpenVLA. We thank these teams for their open-source contributions.

Contributors

jcenaa

15 commits