Autonomous Driving with Vision-Language Models

8 repos

Multimodal AI systems that combine vision and language models for autonomous vehicle perception and decision-making. The cluster centers on LMDrive, a framework integrating large language models with visual encoders for end-to-end driving tasks, along with model variants (Vicuna, LLaVA, Llama) and supporting components. Projects here explore how conversational AI and image-text understanding can be applied to autonomous driving scenarios.

Jupyter Notebook · 1
autonomous-driving ·0
conversational ·0
en ·0
endpoints_compatible ·0
image-text-to-text ·0
multimodal ·0
qwen3.5 ·0
qwen3_5 ·0
safetensors ·0
trajectory-planning ·0