21 repos
Vision-language-action models and robotic foundation models that combine visual perception with language understanding for robotic manipulation and control. These repositories focus on training and deploying models that can understand scenes, follow instructions, and generate robotic actions—bridging computer vision, natural language processing, and robotic control into unified frameworks. The cluster includes implementations of models like InternVLA variants and related approaches for enabling robots to learn from diverse data and generalize across manipulation tasks.