10 repos
Vision-language-action (VLA) models that enable robots to understand and act on natural language instructions by integrating visual perception with motor control. This cluster focuses on model architectures and implementations for embodied AI, including optimized variants like SmolVLA and BitVLA designed for efficient deployment on edge devices and robotics platforms like Jetson. Repositories here span model checkpoints, inference frameworks, and robot task implementations that bridge language understanding with robotic manipulation and control.