Vision Language Action Models for Robotics

30 repos

Vision Language Action (VLA) models that combine visual perception, language understanding, and robotic control into unified transformer-based architectures. These repositories implement and extend VLA frameworks for robot learning, instruction following, and multi-task manipulation, often leveraging pre-trained transformers and safetensors for efficient model storage and deployment. The cluster includes research implementations, benchmark environments, and real-world robot applications ranging from single-arm manipulation to multi-robot coordination.

robotics ·426
safetensors ·360
vla ·349
transformers ·305
custom_code ·297
multimodal ·280
pretraining ·279
image-text-to-text ·279
feature-extraction ·270
openvla ·270