2 repos
Vision-language-action (VLA) models that enable robots to understand and execute tasks by combining visual perception with language instructions. This cluster centers on OpenVLA and related fine-tuned variants trained on robotic manipulation datasets like LIBERO, demonstrating how large multimodal models can be adapted for real-world robot control and task learning across different behavioral categories and spatial reasoning requirements.