Vision-Language-Action Models for Robotics

2 repos

Vision-language-action (VLA) models that enable robots to understand and execute tasks by combining visual perception with language instructions. This cluster centers on OpenVLA and related fine-tuned variants trained on robotic manipulation datasets like LIBERO, demonstrating how large multimodal models can be adapted for real-world robot control and task learning across different behavioral categories and spatial reasoning requirements.

custom_code ·9
feature-extraction ·9
openvla ·9
robotics ·9
safetensors ·9
transformers ·9