23 repos
Vision-language-action (VLA) models and learning frameworks that enable robots to understand and execute physical tasks from visual observations and language instructions. The cluster covers foundation models for robotic manipulation and navigation, reinforcement learning approaches for skill acquisition, and simulation environments for training embodied AI agents. Key themes include behavior cloning from demonstrations, world model learning, and real-to-sim transfer for continuous control in kitchen and household manipulation tasks.