10 repos
Fine-tuned vision-language models adapted for robotic perception and control tasks, building on base models like Llama 2 and Vicuna through LoRA adaptation techniques. These repositories represent efforts to ground large language models in visual robotic understanding, enabling models to process images and spatial information relevant to robot manipulation and navigation. The cluster centers on the RoboPoint architecture and its various model variants, using efficient parameter adaptation methods to create task-specific versions without full retraining.