8 repos
Multimodal AI systems that combine vision and language models for autonomous vehicle perception and decision-making. The cluster centers on LMDrive, a framework integrating large language models with visual encoders for end-to-end driving tasks, along with model variants (Vicuna, LLaVA, Llama) and supporting components. Projects here explore how conversational AI and image-text understanding can be applied to autonomous driving scenarios.