5 repos
Systems and models for teaching robots to manipulate objects and navigate environments through vision-language grounding and learned action policies. The cluster centers on simulated and real-world robotic manipulation tasks, vision-language-action models that ground natural language instructions in continuous control, and datasets/benchmarks like RoboCasa for training embodied AI agents. Python dominates the implementation landscape, reflecting the prevalence of deep learning frameworks and simulation environments for this research area.