Robotics and Vision-Language-Action Models

17 repos

Foundation models and learning frameworks for embodied AI systems that combine vision, language, and robotic action. This cluster focuses on training and deploying multimodal models that enable robots to understand visual scenes and natural language instructions to perform physical tasks. The central repositories, particularly LeRobot and related projects, provide datasets, training pipelines, and evaluation frameworks for vision-language-action (VLA) models in robotics, alongside agent architectures and reinforcement learning approaches for autonomous systems.

Python · 17
robots ·159
strands-agents ·159
strands-labs ·159
vision-language-action ·159
world-foundation-models ·159