A curated list of large VLM-based VLA models for robotic manipulation.
452
54 commits
updated Aug 27, 2026
π₯ Large VLM-based Vision-Language-Action (VLA) models have recently emerged as a transformative paradigm for robotic manipulation by tightly coupling perception, language understanding, and action generation. Built upon large Vision-Language Models (VLMs), they enable robots to interpret natural language instructions, perceive complex environments, and perform diverse manipulation tasks with strong generalization.
π We present the first systematic survey on large VLM-based VLA models for robotic manipulation. This repository serves as the companion resource to our survey: "Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey", and includes all the research papers, benchmarks, and resources reviewed in the paper, organized for easy access and reference.
π We will keep updating this repository with newly published works to reflect the latest progress in the field.
If you find this survey helpful for your research or applications, please consider citing it using the following BibTeX entry:
@misc{shao2025largevlmbasedvisionlanguageactionmodels,
title={Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey},
author={Rui Shao and Wei Li and Lingsen Zhang and Renshan Zhang and Zhiyang Liu and Ran Chen and Liqiang Nie},
year={2025},
eprint={2508.13073},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2508.13073},
}
For any questions or suggestions, please feel free to contact us at:
Email: shaorui@hit.edu.cn and liwei2024@stu.hit.edu.cn
A curated list of large VLM-based VLA models for robotic manipulation.
452
54 commits
updated Aug 27, 2026
π₯ Large VLM-based Vision-Language-Action (VLA) models have recently emerged as a transformative paradigm for robotic manipulation by tightly coupling perception, language understanding, and action generation. Built upon large Vision-Language Models (VLMs), they enable robots to interpret natural language instructions, perceive complex environments, and perform diverse manipulation tasks with strong generalization.
π We present the first systematic survey on large VLM-based VLA models for robotic manipulation. This repository serves as the companion resource to our survey: "Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey", and includes all the research papers, benchmarks, and resources reviewed in the paper, organized for easy access and reference.
π We will keep updating this repository with newly published works to reflect the latest progress in the field.
If you find this survey helpful for your research or applications, please consider citing it using the following BibTeX entry:
@misc{shao2025largevlmbasedvisionlanguageactionmodels,
title={Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey},
author={Rui Shao and Wei Li and Lingsen Zhang and Renshan Zhang and Zhiyang Liu and Ran Chen and Liqiang Nie},
year={2025},
eprint={2508.13073},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2508.13073},
}
For any questions or suggestions, please feel free to contact us at:
Email: shaorui@hit.edu.cn and liwei2024@stu.hit.edu.cn