JohnsonJiang1996/Awesome-VLA4AD

Vision–Language–Action models for Autonomous Driving (VLA4AD) resources, serving as the companion repository to the survey paper “A Survey on Vision–Language–Action Models for Autonomous Driving”.

623

91 commits

updated Nov 20, 2025

See the code

README

Awesome Vision–Language–Action Models for Autonomous Driving 🚗

arXiv GitHub stars GitHub forks Issues Badge License Badge

Welcome to Awesome VLA4AD—a curated, continuously updated collection of research papers and resources on Vision–Language–Action models for Autonomous Driving (VLA4AD). This repository tracks the latest advances in VLA4AD, from explanatory perception modules to end-to-end reasoning and control architectures.

Our latest survey is here. We invite your feedback and discussion!

⭐️ Follow & Star to stay up to date!
🤝 Contributions welcome—if you know of new papers, datasets, or tools, please open an issue or submit a PR.
📬 Questions or suggestions? Reach us at sicong.jiang@mail.mcgill.ca or qka23@mails.tsinghua.edu.cn.


📜 Citation

If this project is useful in your work, we'd appreciate a star 🌟 and a citation of our survey.

@article{jiang2025survey,
  title={A Survey on Vision-Language-Action Models for Autonomous Driving},
  author={Jiang, Sicong and Huang, Zilin and Qian, Kangan and Luo, Ziang and Zhu, Tianze and Zhong, Yang and Tang, Yihong and Kong, Menglin and Wang, Yunlong and Jiao, Siwen and others},
  journal={arXiv preprint arXiv:2506.24044},
  year={2025}
}

📚 Table of Contents


🔥 Motivation & Paradigm Shift

The development of autonomous driving has progressed from modular pipelines to fully integrated systems. This survey summarizes the latest advances into three core paradigms:

  • End-to-End AD: Direct sensor-to-control mapping—efficient but opaque and weak on rare scenarios.

    • Flow: Sensors → Network → Actions
  • VLMs for AD: Adds language reasoning—boosts explainability but doesn’t drive the vehicle.

    • Flow: Sensors → VLM → Answers
  • VLA for AD: Unifies vision, language, and control in one policy—understands instructions, reasons, acts, and explains.

    • Flow: Sensors → Multimodal Encoder → LLM/VLM → Decoder → Actions

Driving Paradigms Comparison
Figure 1. (a) conventional end-to-end AD, (b) vision-language models as explainers, (c) full Vision–Language–Action systems.


🚀 Overview of VLA4AD

A typical VLA4AD model follows an “Input–Process–Output” flow, unifying environment perception, instruction understanding, and vehicle control.

Overview of VLA4AD
Figure 2. Overview of VLA4AD, integrating vision, language, and action modules.

A snapshot of the field’s evolution through four successive stages—from VLM-as-explainer to augmented, reasoning-centric agents:

Progress of VLA Models for AD
Figure 3. Progression of VLA4AD models: (1) VLMs as passive explainers; (2) Modular VLA with intermediate representations; (3) End-to-end VLA mapping sensors directly to actions; (4) Augmented VLA with long-horizon reasoning and tool use.

The following table shows the representative models of VLA4AD and their inside modules:

T
Table 1. Representative VLA4AD Models (2023–2025). Sensor Inputs: Single = single forward-facing camera input; Multi = multi-view camera input; State = vehicle state information & other sensor input. Outputs: LLC= low-level control, Traj.= future trajectory, Multi.= multiple tasks such as perception, prediction or planning.


🏆 Awesome VLA4AD Papers

1️⃣ Pre-VLA: VLM as Explainers

ModelYearKey FeaturesLink
DriveGPT-42023Scene Narration, QAhttps://arxiv.org/abs/2310.01412 / Code
TS-VLM2025Text-guided Attentionhttps://arxiv.org/abs/2505.12670 / Code
DynRsl-VLM2025Adaptive Resolutionhttps://arxiv.org/abs/2503.11265

2️⃣ Modular VLA4AD

ModelYearKey FeaturesLink
RAG-Driver2024Retrieval-Augmentedhttps://arxiv.org/abs/2402.10828 / Code
OpenDriveVLA2025Language-guided Planninghttps://arxiv.org/abs/2503.23463 / Code
DriveMoE2025Expert Routinghttps://arxiv.org/abs/2505.16278 / Code
LangCoop2025V2V Coordinationhttps://arxiv.org/abs/2504.13406 / Code
SafeAuto2025Rule-based Safetyhttps://arxiv.org/abs/2503.00211 / Code
ReCogDrive2025Diffusion + RLhttps://arxiv.org/abs/2506.08052 / Code

3️⃣ End-to-End VLA4AD

ModelYearKey FeaturesLink
ADriver-I2023Diffusion-based World Modelhttps://arxiv.org/abs/2311.13549
EMMA2024Detection + Planninghttps://arxiv.org/abs/2410.23262 / Code
CoVLA-Agent2024Caption + Trajectoryhttps://arxiv.org/abs/2408.10845 / Code
SimLingo2025Action Dreaminghttps://arxiv.org/abs/2503.09594 / Code
DiffVLA2025Sparse-Dense Diffusionhttps://arxiv.org/abs/2505.19381
S4-Driver2025Sparse 3D Representationhttps://arxiv.org/abs/2505.24139

4️⃣ Reasoning-Augmented VLA4AD

ModelYearKey FeaturesLink
ORION2025Memory + Rationaleshttps://arxiv.org/abs/2503.19755 / Code
Impromptu-VLA2025CoT-Aligned Planninghttps://arxiv.org/abs/2505.23757 / Code
FSDrive2025Visual Reasoninghttps://arxiv.org/abs/2505.17685 / Code
AutoVLA2025Drive Tokens + CoThttps://arxiv.org/abs/2506.13757 / Code
Drive-R12025CoT-Aligned Planninghttps://arxiv.org/abs/2506.18234
Alpamayo-R12025Chain of Causation Reasoninghttps://arxiv.org/abs/2511.00088

📊 Datasets & Benchmarks

NameYearModalityTaskURL
BDD100K / BDD-X2018Video + RationalesCaptioning, QAbdd-data.berkeley.edu
nuScenes2020Camera, LiDAR, RadarDetection, QAwww.nuscenes.org
Bench2Drive2024CARLA SimulatorClosed-loop DrivingGithub
Reason2Drive2024Video–QACoT-Chain ConsistencyGithub
Impromptu-VLA Dataset2025Video + QA + TrajCorner-Case TestingGithub
NuInteract2025Multi-view QA3D QAGithub
DriveAction2025In-the-wild QAHigh-level ActionsHuggingFace

⚙️ Installation & Usage

git clone https://github.com/JohnsonJiang1996/Awesome-VLA4AD.git
cd Awesome-VLA4AD
# Browse papers, datasets & code samples in each folder

JohnsonJiang1996/Awesome-VLA4AD

Vision–Language–Action models for Autonomous Driving (VLA4AD) resources, serving as the companion repository to the survey paper “A Survey on Vision–Language–Action Models for Autonomous Driving”.

623

91 commits

updated Nov 20, 2025

See the code

README

Awesome Vision–Language–Action Models for Autonomous Driving 🚗

arXiv GitHub stars GitHub forks Issues Badge License Badge

Welcome to Awesome VLA4AD—a curated, continuously updated collection of research papers and resources on Vision–Language–Action models for Autonomous Driving (VLA4AD). This repository tracks the latest advances in VLA4AD, from explanatory perception modules to end-to-end reasoning and control architectures.

Our latest survey is here. We invite your feedback and discussion!

⭐️ Follow & Star to stay up to date!
🤝 Contributions welcome—if you know of new papers, datasets, or tools, please open an issue or submit a PR.
📬 Questions or suggestions? Reach us at sicong.jiang@mail.mcgill.ca or qka23@mails.tsinghua.edu.cn.


📜 Citation

If this project is useful in your work, we'd appreciate a star 🌟 and a citation of our survey.

@article{jiang2025survey,
  title={A Survey on Vision-Language-Action Models for Autonomous Driving},
  author={Jiang, Sicong and Huang, Zilin and Qian, Kangan and Luo, Ziang and Zhu, Tianze and Zhong, Yang and Tang, Yihong and Kong, Menglin and Wang, Yunlong and Jiao, Siwen and others},
  journal={arXiv preprint arXiv:2506.24044},
  year={2025}
}

📚 Table of Contents


🔥 Motivation & Paradigm Shift

The development of autonomous driving has progressed from modular pipelines to fully integrated systems. This survey summarizes the latest advances into three core paradigms:

  • End-to-End AD: Direct sensor-to-control mapping—efficient but opaque and weak on rare scenarios.

    • Flow: Sensors → Network → Actions
  • VLMs for AD: Adds language reasoning—boosts explainability but doesn’t drive the vehicle.

    • Flow: Sensors → VLM → Answers
  • VLA for AD: Unifies vision, language, and control in one policy—understands instructions, reasons, acts, and explains.

    • Flow: Sensors → Multimodal Encoder → LLM/VLM → Decoder → Actions

Driving Paradigms Comparison
Figure 1. (a) conventional end-to-end AD, (b) vision-language models as explainers, (c) full Vision–Language–Action systems.


🚀 Overview of VLA4AD

A typical VLA4AD model follows an “Input–Process–Output” flow, unifying environment perception, instruction understanding, and vehicle control.

Overview of VLA4AD
Figure 2. Overview of VLA4AD, integrating vision, language, and action modules.

A snapshot of the field’s evolution through four successive stages—from VLM-as-explainer to augmented, reasoning-centric agents:

Progress of VLA Models for AD
Figure 3. Progression of VLA4AD models: (1) VLMs as passive explainers; (2) Modular VLA with intermediate representations; (3) End-to-end VLA mapping sensors directly to actions; (4) Augmented VLA with long-horizon reasoning and tool use.

The following table shows the representative models of VLA4AD and their inside modules:

T
Table 1. Representative VLA4AD Models (2023–2025). Sensor Inputs: Single = single forward-facing camera input; Multi = multi-view camera input; State = vehicle state information & other sensor input. Outputs: LLC= low-level control, Traj.= future trajectory, Multi.= multiple tasks such as perception, prediction or planning.


🏆 Awesome VLA4AD Papers

1️⃣ Pre-VLA: VLM as Explainers

ModelYearKey FeaturesLink
DriveGPT-42023Scene Narration, QAhttps://arxiv.org/abs/2310.01412 / Code
TS-VLM2025Text-guided Attentionhttps://arxiv.org/abs/2505.12670 / Code
DynRsl-VLM2025Adaptive Resolutionhttps://arxiv.org/abs/2503.11265

2️⃣ Modular VLA4AD

ModelYearKey FeaturesLink
RAG-Driver2024Retrieval-Augmentedhttps://arxiv.org/abs/2402.10828 / Code
OpenDriveVLA2025Language-guided Planninghttps://arxiv.org/abs/2503.23463 / Code
DriveMoE2025Expert Routinghttps://arxiv.org/abs/2505.16278 / Code
LangCoop2025V2V Coordinationhttps://arxiv.org/abs/2504.13406 / Code
SafeAuto2025Rule-based Safetyhttps://arxiv.org/abs/2503.00211 / Code
ReCogDrive2025Diffusion + RLhttps://arxiv.org/abs/2506.08052 / Code

3️⃣ End-to-End VLA4AD

ModelYearKey FeaturesLink
ADriver-I2023Diffusion-based World Modelhttps://arxiv.org/abs/2311.13549
EMMA2024Detection + Planninghttps://arxiv.org/abs/2410.23262 / Code
CoVLA-Agent2024Caption + Trajectoryhttps://arxiv.org/abs/2408.10845 / Code
SimLingo2025Action Dreaminghttps://arxiv.org/abs/2503.09594 / Code
DiffVLA2025Sparse-Dense Diffusionhttps://arxiv.org/abs/2505.19381
S4-Driver2025Sparse 3D Representationhttps://arxiv.org/abs/2505.24139

4️⃣ Reasoning-Augmented VLA4AD

ModelYearKey FeaturesLink
ORION2025Memory + Rationaleshttps://arxiv.org/abs/2503.19755 / Code
Impromptu-VLA2025CoT-Aligned Planninghttps://arxiv.org/abs/2505.23757 / Code
FSDrive2025Visual Reasoninghttps://arxiv.org/abs/2505.17685 / Code
AutoVLA2025Drive Tokens + CoThttps://arxiv.org/abs/2506.13757 / Code
Drive-R12025CoT-Aligned Planninghttps://arxiv.org/abs/2506.18234
Alpamayo-R12025Chain of Causation Reasoninghttps://arxiv.org/abs/2511.00088

📊 Datasets & Benchmarks

NameYearModalityTaskURL
BDD100K / BDD-X2018Video + RationalesCaptioning, QAbdd-data.berkeley.edu
nuScenes2020Camera, LiDAR, RadarDetection, QAwww.nuscenes.org
Bench2Drive2024CARLA SimulatorClosed-loop DrivingGithub
Reason2Drive2024Video–QACoT-Chain ConsistencyGithub
Impromptu-VLA Dataset2025Video + QA + TrajCorner-Case TestingGithub
NuInteract2025Multi-view QA3D QAGithub
DriveAction2025In-the-wild QAHigh-level ActionsHuggingFace

⚙️ Installation & Usage

git clone https://github.com/JohnsonJiang1996/Awesome-VLA4AD.git
cd Awesome-VLA4AD
# Browse papers, datasets & code samples in each folder