🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
See the codeAutonomous driving has long relied on modular "Perception-Decision-Action" pipelines, whose hand-crafted interfaces and rule-based components often struggle in complex, dynamic, or long-tailed scenarios. Their cascaded structure also amplifies upstream perception errors, undermining downstream planning and control.
This survey reviews vision-action (VA) models and vision-language-action (VLA) models for autonomous driving. We trace the evolution from early VA approaches to modern VLA frameworks, and organize existing methods into two principal paradigms:
![]() |
|---|
For more details, kindly refer to our :books: Paper, :globe_with_meridians: Project Page, and :hugs: HuggingFace Leaderboard.
If you find this work helpful for your research, please kindly consider citing our paper:
@article{survey_vla4ad,
title = {Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future},
author = {Tianshuai Hu and Xiaolu Liu and Song Wang and Yiyao Zhu and Ao Liang and Lingdong Kong and Guoyang Zhao and Zeying Gong and Jun Cen and Zhiyu Huang and Xiaoshuai Hao and Linfeng Li and Hang Song and Xiangtai Li and Jun Ma and Shaojie Shen and Jianke Zhu and Dacheng Tao and Ziwei Liu and Junwei Liang},
journal = {arXiv preprint arXiv:2512.16760},
year = {2025},
}
@article{survey_3d_4d_world_models,
title = {{3D} and {4D} World Modeling: A Survey},
author = {Lingdong Kong and Wesley Yang and Jianbiao Mei and Youquan Liu and Ao Liang and Dekai Zhu and Dongyue Lu and Wei Yin and Xiaotao Hu and Mingkai Jia and Junyuan Deng and Kaiwen Zhang and Yang Wu and Tianyi Yan and Shenyuan Gao and Song Wang and Linfeng Li and Liang Pan and Yong Liu and Jianke Zhu and Wei Tsang Ooi and Steven C. H. Hoi and Ziwei Liu},
journal = {arXiv preprint arXiv:2509.07996},
year = {2025}
}
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
HTML
100.0%
🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
See the codeAutonomous driving has long relied on modular "Perception-Decision-Action" pipelines, whose hand-crafted interfaces and rule-based components often struggle in complex, dynamic, or long-tailed scenarios. Their cascaded structure also amplifies upstream perception errors, undermining downstream planning and control.
This survey reviews vision-action (VA) models and vision-language-action (VLA) models for autonomous driving. We trace the evolution from early VA approaches to modern VLA frameworks, and organize existing methods into two principal paradigms:
![]() |
|---|
For more details, kindly refer to our :books: Paper, :globe_with_meridians: Project Page, and :hugs: HuggingFace Leaderboard.
If you find this work helpful for your research, please kindly consider citing our paper:
@article{survey_vla4ad,
title = {Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future},
author = {Tianshuai Hu and Xiaolu Liu and Song Wang and Yiyao Zhu and Ao Liang and Lingdong Kong and Guoyang Zhao and Zeying Gong and Jun Cen and Zhiyu Huang and Xiaoshuai Hao and Linfeng Li and Hang Song and Xiangtai Li and Jun Ma and Shaojie Shen and Jianke Zhu and Dacheng Tao and Ziwei Liu and Junwei Liang},
journal = {arXiv preprint arXiv:2512.16760},
year = {2025},
}
@article{survey_3d_4d_world_models,
title = {{3D} and {4D} World Modeling: A Survey},
author = {Lingdong Kong and Wesley Yang and Jianbiao Mei and Youquan Liu and Ao Liang and Dekai Zhu and Dongyue Lu and Wei Yin and Xiaotao Hu and Mingkai Jia and Junyuan Deng and Kaiwen Zhang and Yang Wu and Tianyi Yan and Shenyuan Gao and Song Wang and Linfeng Li and Liang Pan and Yong Liu and Jianke Zhu and Wei Tsang Ooi and Steven C. H. Hoi and Ziwei Liu},
journal = {arXiv preprint arXiv:2509.07996},
year = {2025}
}
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
:timer_clock: In chronological order, from the earliest to the latest.
HTML
100.0%