Robotics Vision-Language Models

21 repos

Vision-language-action models and robotic foundation models that combine visual perception with language understanding for robotic manipulation and control. These repositories focus on training and deploying models that can understand scenes, follow instructions, and generate robotic actions—bridging computer vision, natural language processing, and robotic control into unified frameworks. The cluster includes implementations of models like InternVLA variants and related approaches for enabling robots to learn from diverse data and generalize across manipulation tasks.

Python · 7
C++ · 1
robotics ·5,982
vision-language-action-model ·5,646
world-action-model ·3,630
robotic-foundation-model ·3,630
manipulation ·2,198
vision-language-model ·746
florence-2 ·729
cloth-folding ·729
pretrained-models ·729
robotics-dataset ·729