30 repos
Vision Language Action (VLA) models that combine visual perception, language understanding, and robotic control into unified transformer-based architectures. These repositories implement and extend VLA frameworks for robot learning, instruction following, and multi-task manipulation, often leveraging pre-trained transformers and safetensors for efficient model storage and deployment. The cluster includes research implementations, benchmark environments, and real-world robot applications ranging from single-arm manipulation to multi-robot coordination.