17 repos
Foundation models and learning frameworks for embodied AI systems that combine vision, language, and robotic action. This cluster focuses on training and deploying multimodal models that enable robots to understand visual scenes and natural language instructions to perform physical tasks. The central repositories, particularly LeRobot and related projects, provide datasets, training pipelines, and evaluation frameworks for vision-language-action (VLA) models in robotics, alongside agent architectures and reinforcement learning approaches for autonomous systems.