AutoLab-SAI-SJTU/GE2EAD

Collects papers on autonomous driving E2E learning, VLM/VLA and Hybrid systems, with organized research branches and trends in these fields.

HTML

211

27 commits

updated May 27, 2026

See the code

README

Awesome-GE2EAD

Awesome Logo TechRxiv Project GitHub forks GitHub stars

This is the official repository for "Survey of General End-to-End Autonomous Driving: A Unified Perspective".

This project aims to provide a unified roadmap for the field by:

  • 🗂️ Literature Taxonomy: Classifying methods into Conventional (e.g., UniAD), VLM-centric (e.g., DriveLM), and Hybrid (e.g., Senna) approaches.

  • 💾 Dataset Curation: Collecting both Standard and Vision-Language datasets relevant to end-to-end AD.

  • 📈 Trend Analysis: Outlining main research branches and emerging trends based on our survey.

Citation

If you find this project useful in your research, please consider citing:

@article{yang2025survey,
  title={Survey of General End-to-End Autonomous Driving: A Unified Perspective},
  author={Yang, Yixiang and Han, Chuanrong and Mao, Runhao and others},
  journal={TechRxiv},
  year={2025},
  month={December},
  doi={10.36227/techrxiv.176523315.56439138/v1},
  url={https://doi.org/10.36227/techrxiv.176523315.56439138/v1}
}

📌 Milestones

  • 🚀 2025-12-24: We organize the list of papers in a completely new tabular format.

  • 🚀 2025-12-10: The paper “Survey of General End-to-End Autonomous Driving: A Unified Perspective” was released, and this repository was made publicly available.

Table of Contents

Mindmap, Top Methods

GE2EAD Mindmap Logo

GE2EAD Mindmap

GE2EAD Mindmap Logo

Top Methods

Papers

Conventional End-to-End Methods (VA)

Conventional End-to-End Methods (VA)

2026
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
CLOVER
Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning
2026Self-Distillation · ScoringarXivStars—
RAD-2
Scaling Reinforcement Learning in a Generator-Discriminator Framework
2026RL · Generator-DiscriminatorarXiv——
SparseDriveV2
Scoring is All You Need for End-to-End Autonomous Driving
2026Scoring · Factorized VocabarXivStars—
Latent-WAM
Latent World Action Modeling for End-to-End Autonomous Driving
2026World Model · Latent CompressionarXiv——
FlowAD
Ego-Scene Interactive Modeling for Autonomous Driving
ICLR 2026Flow Matching · Ego-ScenearXivStars—
HDP
Hyper Diffusion Planner: Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
2026Diffusion · Real-VehiclearXivStarsProject
MeanFuser
Fast One-Step Multi-Modal Trajectory Generation via MeanFlow for End-to-End Autonomous Driving
CVPR 2026MeanFlow · One-SteparXivStars—
ResWorld
Temporal Residual World Model for End-to-End Autonomous Driving
ICLR 2026World Model · Temporal ResidualarXivStars—
Drive-JEPA
Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
2026V-JEPA · DistillationarXivStars—
PlannerRFT
Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning
2026Diffusion Planner · Reinforcement Fine-TuningarXiv—Project
DrivoR
Driving on Registers
CVPR 2026Register Tokens · ViT · ScoringarXivStarsProject
AlignDrive
Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
2026Lateral-Longitudinal · Path-ConditionedarXivStarsProject
2025
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
DriveLaW
Unifying Planning and Video Generation in a Latent Driving World
2025World Model · Video GenerationarXivStarsProject
SimScale
Learning to Drive via Real-World Simulation at Scale
2025Scalable Simulation · Neural RenderingarXivStarsProject
FutureX
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
2025World Model · Latent CoTarXiv——
Spatial Retrieval AD
Spatial Retrieval Augmented Autonomous Driving
2025Retrieval · Geo ImagesarXivStarsProject
UniMM-V2X
UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving
2025MoE · Multi-AgentarXivStars—
UniLION
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
2025Linear RNNarXivStars—
DiffusionDriveV2
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
2025Diffusion · RLarXivStars—
LAP
LAP: Fast Latent Diffusion Planner with Fine-Grained Feature Distillation for Autonomous Driving
2025Latent Diffusion · PlanningarXivStars—
GuideFlow
GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
2025Generative · Flow MatchingarXivStars—
DiffRefiner
DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
2025Diffusion · RefinementarXivStars—
ResAD
ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
2025Trajectory ModelingarXivStars—
SeerDrive
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution
NeurIPS 2025World Model · PlanningarXivStars—
BridgeDrive
Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
ICLR 2026Diffusion Bridge · Anchor-to-RefinedarXivStars—
DriveDPO
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
2025DPO · SafetyarXiv——
AnchDrive
AnchDrive: Bootstrapping Diffusion Policies with Hybrid Trajectory Anchors for End-to-End Driving
2025Diffusion · AnchorsarXiv——
AdaThinkDrive
AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
2025RL · CoTarXiv——
VeteranAD
Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving
2025Perception-PlanningarXivStars—
EvaDrive
Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
2025RL · AdversarialarXiv——
ReconDreamer-RL
Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
2025RL · World ModelarXivStars—
GMF-Drive
Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
2025Mamba · FusionarXiv——
DistillDrive
End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning Model
2025DistillationarXivStars—
GEMINUS
Dual-aware Global and Scene-Adaptive Mixture-of-Experts for End-to-End Autonomous Driving
2025MoE · AdaptivearXivStars—
DiVER
Breaking Imitation Bottlenecks: Reinforced Diffusion Powers Diverse Trajectory Generation
2025RL · DiffusionarXiv——
World4Drive
End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
ICCV 2025World ModelarXivStars—
FocalAD
Local Motion Planning for End-to-End Autonomous Driving
2025Motion PlanningarXiv——
GaussianFusion
Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
2025Gaussian Splatting · FusionarXivStars—
CogAD
Cognitive-Hierarchy Guided End-to-End Autonomous Driving
2025Cognitive · HierarchyarXiv——
DiffE2E
Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy
2025Diffusion · HybridarXiv—Project
TransDiffuser
End-to-end Trajectory Generation with Decorrelated Multi-modal Representation for Autonomous Driving
2025Diffusion · MultimodalarXiv——
MomAD
Don’t Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving
CVPR 2025Planning · MomentumarXivStars—
Consistency
Predictive Planner for Autonomous Driving with Consistency Models
2025Consistency · PlanningarXiv——
ARTEMIS
Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving
2025MoE · AutoregressivearXiv——
TTOG
Two Tasks, One Goal: Uniting Motion and Planning for Excellent End To End Autonomous Driving Performance
2025Multi-taskarXiv——
DiffusionDrive
Truncated Diffusion Model for End-to-End Autonomous Driving
CVPR 2025DiffusionarXivStars—
WoTE
End-to-End Driving with Online Trajectory Evaluation via BEV World Model
2025World Model · BEVarXivStars—
DMAD
Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving
2025Multi-taskarXivStars—
Centaur
Robust End-to-End Autonomous Driving with Test-Time Training
2025Test-Time TrainingarXiv——
Drive in Corridors
Enhancing the Safety of End-to-end Autonomous Driving via Corridor Learning and Planning
2025Safety · PlanningarXiv——
BridgeAD
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
CVPR 2025Prediction · PlanningarXivStars—
Hydra-MDP++
Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
2025Distillation · Multi-headarXivStars—
DiffAD
A Unified Diffusion Modeling Approach for Autonomous Driving
2025DiffusionarXiv——
GoalFlow
Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving
CVPR 2025Flow MatchingarXivStars—
HiP-AD
Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single Decoder
ICCV 2025Attention · PlanningarXivStars—
LAW
Enhancing End-to-End Autonomous Driving with Latent World Model
ICLR 2025World ModelarXivStars—
DriveTransformer
Unified Transformer for Scalable End-to-End Autonomous Driving
ICLR 2025TransformerarXivStars—
UncAD
Towards Safe End-to-end Autonomous Driving via Online Map Uncertainty
ICRA 2025Uncertainty · MaparXivStars—
RAD
Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
2025RL · 3DGSarXiv—Project
OAD
Trajectory Offset Learning: A Framework for Enhanced End-to-End Autonomous Driving
2025Trajectory · OffsetResearchGateStars—
2024
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
GaussianAD
Gaussian-Centric End-to-End Autonomous Driving
2024Gaussian Splatting · PerceptionarXivStars—
MA2T
Module-wise Adaptive Adversarial Training for End-to-end Autonomous Driving
2024Adversarial · RobustnessarXiv——
Hint-AD
Holistically Aligned Interpretability in End-to-End Autonomous Driving
2024Interpretability · AlignmentarXivStarsProject
DRAMA
An Efficient End-to-end Motion Planner for Autonomous Driving with Mamba
CVPR 2025Mamba · Motion PlanningarXivStarsProject
PPAD
Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving
ECCV 2024Prediction · PlanningarXivStars—
BEV-Planner
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
CVPR 2024BEV · EvaluationarXivStars—
EfficientFuser
Efficient Fusion and Task Guided Embedding for End-to-end Autonomous Driving
2024Efficient · FusionarXiv——
UAD
End-to-End Autonomous Driving without Costly Modularization and 3D Manual Annotation
2024UnsupervisedarXiv——
Hydra-MDP
End-to-end Multimodal Planning with Multi-target Hydra-Distillation
2024Distillation · MultimodalarXivStars—
DualAD
Disentangling the Dynamic and Static World for End-to-End Driving
CVPR 2025Dual-Stream · DynamicarXivStars—
SparseDrive
End-to-End Autonomous Driving via Sparse Scene Representation
2024Sparse · Scene ReparXivStars—
GAD
GAD-Generative Learning for HD Map-Free Autonomous Driving
2024Generative · Map-FreearXivStars—
SparseAD
Sparse Query-Centric Paradigm for Efficient End-to-End Autonomous Driving
2024Sparse · QueryarXiv——
GenAD
Generative End-to-End Autonomous Driving
ECCV 2024Generative · PredictionarXivStars—
GraphAD
Interaction Scene Graph for End-to-end Autonomous Driving
2024Graph · InteractionarXivStars—
ActiveAD
Planning-Oriented Active Learning for End-to-End Autonomous Driving
2024Active LearningarXiv——
VADv2
End-to-End Vectorized Autonomous Driving via Probabilistic Planning
2024Vectorized · ProbabilisticarXivStars—
2023
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
DriveAdapter
Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving
ICCV 2023Adapter · DecouplingarXivStars—
VAD
Vectorized Scene Representation for Efficient Autonomous Driving
ICCV 2023Vectorized · EfficientarXivStars—
ThinkTwice
Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving
CVPR 2023Decoder · RefinementarXivStars—
ReasonNet
End-to-End Driving with Temporal and Global Reasoning
CVPR 2023Reasoning · TemporalarXivStars—
SuperDriverAI
Towards Design and Implementation for End-to-End Learning-based Autonomous Driving
2023Attention · DNNarXiv——
UniAD
Planning-oriented Autonomous Driving
CVPR 2023Multi-task · UnifiedarXivStars—
E2E Dense
End-to-End Learning of Behavioural Inputs for Autonomous Driving in Dense Traffic
IROS 2023Optimization · Dense TrafficarXiv——
CRCHFL
Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving
IROS 2023Federated LearningarXiv——
PPGeo
Policy pre-training for autonomous driving via self-supervised geometric modeling
ICLR 2023Self-Supervised · GeometricarXivStars—
Before 2023
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
MMFN
Multi-Modal-Fusion-Net for End-to-End Driving
IROS 2022Fusion · Multi-ModalarXivStars—
KEMP
Keyframe-Based Hierarchical End-to-End Deep Model for Long-Term Trajectory Prediction
ICRA 2022Keyframe · HierarchicalarXiv——
TCP
Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline
NeurIPS 2022Trajectory · ControlarXivStars—
ST-P3
End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning
ECCV 2022Spatial-Temporal · InterpretablearXivStars—
MP3
A Unified Model to Map, Perceive, Predict and Plan
CVPR 2021Mapless · PredictionPaper——
Multitask
Multi-task Learning with Attention for End-to-end Autonomous Driving
CVPR 2021Multi-task · AttentionarXivStars—
Transfuser
Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
CVPR 2021Transformer · FusionPaperStars—
NEAT
Neural Attention Fields for End-to-End Autonomous Driving
ICCV 2021Attention Fields · BEVPaperStars—
Fast-LiDARNet
Efficient and Robust LiDAR-Based End-to-End Navigation
ICRA 2021LiDAR · EfficientarXiv——
IVMP
Learning Interpretable End-to-End Vision-Based Motion Planning for Autonomous Driving with Optical Flow Distillation
ICRA 2021Interpretable · Optical FlowarXiv—Project
P3
Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable Semantic Representations
ECCV 2020Semantic · InterpretabilityarXiv——
DARB
Exploring data aggregation in policy learning for vision-based urban autonomous driving
CVPR 2020Data Aggregation · PolicyPaperStars—
Roach
End-to-End Urban Driving by Imitating a Reinforcement Learning Coach
ICCV 2021RL · ImitationarXivStars—
LBC
Learning by cheating
CoRL 2019Knowledge DistillationarXivStars—
CIL
End-to-End driving via conditional imitation learning
CoRL 2018Imitation LearningarXivStars—
Drive in A Day
Learning to drive in a day
2018RLarXivStars—
CNN E2E
End to End Learning for Self-Driving Cars
2016CNN · ImitationarXivStars—
ALVINN
An autonomous land vehicle in a neural network
NeurIPS 1988Neural NetworkPaper——

(back to top)

VLM-Centric End-to-End Methods

VLM-Centric End-to-End Methods

2026
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
ReflectDrive-2
Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving
2026Discrete Diffusion · Self-Edit · RLarXiv——
OneDrive
Unified Multi-Paradigm Driving with Vision-Language-Action Models
2026Single Decoder · Multi-ParadigmarXivStars—
UniDriveVLA
Unifying Understanding, Perception, and Action Planning for Autonomous Driving
2026MoT · Expert DecouplingarXivStarsProject
AutoDrive-P³
Unified Chain of Perception–Prediction–Planning Thought via Reinforcement Fine-Tuning
ICLR 2026P³ CoT · GRPOarXivStars—
DynVLA
Learning World Dynamics for Action Reasoning in Autonomous Driving
2026Dynamics CoT · RFTarXivStars—
LaST-VLA
Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
2026Latent CoT · GRPOarXivStars—
ELF-VLA
Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failures
2026Failure Diagnostics · RLarXiv——
VGGDrive
Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
CVPR 2026VLM · 3D GeometryarXivStars—
HiST-VLA
A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
2026VLA · Spatio-Temporal · Token SparsearXiv——
SparseOccVLA
Bridging Occupancy and Vision-Language Models via Sparse Queries
2026Sparse Occupancy · Unified 4D UnderstandingarXivStarsProject
SGDrive
Scene-to-Goal Hierarchical World Cognition for Autonomous Driving
2026Hierarchical Cognition · Scene-Agent-GoalarXivStars—
2025
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
Counterfactual VLA
Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
2025Self-Reflective · Counterfactual ReasoningarXiv——
ColaVLA
Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning
2025Cognitive Latent Reasoning · Parallel PlanningarXivStarsProject
DrivePI
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
20254D Spatial · OccupancyarXivStars—
WAM-Diff
WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving
2025Masked Diffusion · MoE · Online RLarXivStars—
SpaceDrive
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
2025Spatial EncodingarXivStarsProject
OpenREAD
OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
2025RFT/RL · LLM-as-CriticarXivStars
CoT4AD
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning
2025VLA · CoTarXiv——
MPA
Model-Based Policy Adaptation for Closed-Loop End-to-End Autonomous Driving
NeurIPS 2025Model-Based · SimarXiv—Project
AD-R1
AD-R1: Closed-Loop Reinforcement Learning with Impartial World Models
2025RL · World ModelarXiv——
Alpamayo-R1
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving
2025VLA · ReasoningarXivStars—
DriveVLA-W0
DRIVEVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
2025VLA · World ModelarXivStars—
MTRDrive
MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving
2025VLM · MemoryarXiv——
ReflectDrive
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
2025Diffusion · VLAarXiv——
IRL-VLA
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
2025IRL · VLAarXivStars—
Prune2Drive
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models
2025VLM · PruningarXiv——
FastDriveVLA
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
2025VLA · PruningarXiv——
MCAM
Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding
2025Causal · MultimodalarXivStars—
AutoDrive-R²
Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
2025VLA · ReflectionarXiv——
DriveAgent-R1
Advancing VLM-based Autonomous Driving with Hybrid Thinking and Active Perception
2025VLM · ActivearXiv——
NavigScene
Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving
2025Navigation · PerceptionarXiv——
ADRD
LLM-DRIVEN AUTONOMOUS DRIVING BASED ON RULE-BASED DECISION SYSTEMS
2025LLM · Rule-BasedarXiv——
AutoVLA
A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning
2025VLA · RLarXivStarsProject
Poutine
Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training
2025VLT · RLarXiv——
ReCogDrive
A Reinforced Cognitive Framework for End-to-End Autonomous Driving
2025VLM · DiffusionarXivStarsProject
AD-EE
Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
2025VLM · EfficientarXiv——
FastDrive
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
2025VLM · StructuredarXiv——
HMVLM
Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
2025VLM · Long-TailarXiv——
S4-Driver
Scalable Self-Supervised Driving Multimodal Large Language Model
CVPR 2025Self-Supervised · MLLMPaper——
DiffVLA
Vision-Language Guided Diffusion Planning for Autonomous Driving
2025Diffusion · VLMarXiv——
X-Driver
Explainable Autonomous Driving with Vision-Language Models
2025MLLM · CoTarXiv——
DriveGPT4-V2
Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving
CVPR 2025LLM · Closed-LoopPaper——
DriveMind
A Dual-VLM based Reinforcement Learning Framework for Autonomous Driving
2025Dual-VLM · RLarXiv——
ReasonPlan
Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
2025MLLM · ReasoningarXivStars—
FutureSightDrive
Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
2025CoT · Spatio-TemporalarXivStars—
PADriver
Towards Personalized Autonomous Driving
2025MLLM · PersonalizedarXiv——
LDM
Unlock the Power of Unlabeled Data in Language Driving Model
ICRA 2025Self-Supervised · DistillationarXiv——
DriveMoE
Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
2025MoE · VLAarXivStarsProject
DriveMonkey
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
2025LVLM · InteractivearXivStars—
AgentThink
A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models
2025CoT · ToolsarXiv——
DSDrive
Distilling Large Language Model for Lightweight End-to-End Autonomous Driving
2025Distillation · LightweightarXiv——
LightEMMA
Lightweight End-to-end Multimodal Autonomous Driving
2025Lightweight · MultimodalarXivStars—
THCAD
Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating LLM Guidance with RL
2025LLM · RL · Fast-SlowarXiv——
DriveSOTIF
Advancing Perception SOTIF Through Multimodal Large Language Models
2025SOTIF · MLLMarXiv——
Actor-Reasoner
Interact, Instruct to Improve: A LLM-Driven Parallel Actor-Reasoner Framework
2025LLM · InteractionarXivStars—
MPDrive
Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
CVPR 2025Prompt · SpatialarXiv——
V3LMA
Visual 3D-enhanced Language Model for Autonomous Driving
20253D · LVLMarXiv——
OpenDriveVLA
Towards End-to-end Autonomous Driving with Large Vision Language Action Model
2025VLA · Open-SourcearXivStarsProject
SimLingo
Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
CVPR 2025VLA · Closed-LoopPaperStarsProject
SAFEAUTO
KNOWLEDGE-ENHANCED SAFE AUTONOMOUS DRIVING WITH MULTIMODAL FOUNDATION MODELS
ICLR 2025Safety · MultimodalarXivStars—
NuGrounding
A Multi-View 3D Visual Grounding Framework in Autonomous Driving
2025Grounding · 3DarXiv——
CoT-Drive
Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
2025CoT · ForecastingarXiv——
CoLMDriver
LLM-based Negotiation Benefits Cooperative Autonomous Driving
2025Cooperative · LLMarXivStars—
AlphaDrive
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
2025RL · ReasoningarXivStars—
TrackingMeetsLMM
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
2025Tracking · LMMarXivStars—
BEVDriver
Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving
2025BEV · LLMarXiv——
DynRsl-VLM
Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
2025Dynamic Res · VLMarXiv——
Sce2DriveX
A Generalized MLLM Framework for Scene-to-Drive Learning
2025MLLM · ScenearXiv——
VLM-Assisted-CL
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
2025Continual LearningarXiv——
LeapVAD
A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking
2025Cognitive · Dual-ProcessarXivStarsProject
2024
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
VLM-RL
A Unified Vision Language Model and Reinforcement Learning Framework for Safe Autonomous Driving
2024RL · VLMarXivStarsProject
GPVL
Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving
AAAI 2025Generative · 3D-VLarXivStars—
CALMM-Drive
Confidence-Aware Autonomous Driving with Large Multimodal Model
2024CoT · ConfidencearXiv——
WiseAD
Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
2024VLM · ReasoningarXivStars—
OpenEMMA
Open-Source Multimodal Model for End-to-End Autonomous Driving
WACV 2025Open-Source · MultimodalarXivStars—
FeD
Feedback-Guided Autonomous Driving
CVPR 2024Feedback · LLMPaper—Project
LeapAD
Continuously learning, adapting, and improving: A dual-process approach to autonomous driving
NeurIPS 2024Dual-Process · ContinualarXivStarsProject
DriveMM
All-in-One Large Multimodal Model for Autonomous Driving
2024Multimodal · GeneralizationarXivStarsProject
Exp-Planning
Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving
ECCV 2024Explainability · PlanningarXiv——
LaVida Drive
Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
2024VQA · InteractionarXiv——
EMMA
End-to-End Multimodal Model for Autonomous Driving
2024End-to-End · MultimodalarXiv——
DriVLMe
Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experiences
IROS 2024Embodied · SocialarXivStarsProject
OccLLaMA
An Occupancy-Language-Action Generative World Model for Autonomous Driving
2024World Model · OccupancyarXiv——
MiniDrive
More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens
2024Efficient · MoEarXiv——
RDA-Driver
Making Large Language Models Better Planners with Reasoning-Decision Alignment
ECCV 2024Reasoning · AlignmentarXiv——
EC-Drive
Edge-Cloud Collaborative Motion Planning for Autonomous Driving with Large Language Models
ICCT 2024Edge-Cloud · CollaborativearXiv—Project
V2X-VLM
End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models
2024V2X · CooperativearXivStarsProject
Cube-LLM
Language-Image Models with 3D Understanding
20243D · Language-ImagearXiv—Project
VLM-MPC
Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC)
2024MPC · ControlarXiv——
SimpleLLM4AD
An End-to-End Vision-Language Model with Graph Visual Question Answering
IEIT SystemsGraph VQA · PipelinearXiv——
AsyncDriver
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
ECCV 2024Asynchronous · Closed-LooparXivStars—
AD-H
AUTONOMOUS DRIVING WITH HIERARCHICAL AGENTS
ICLR 2025Hierarchical · AgentsPaper——
CarLLaVA
Vision language models for camera-only closed-loop driving
2024Camera-only · Closed-LooparXiv—Project
PlanAgent
A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning
2024Agent · Closed-LooparXiv——
Atlas
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
20243D-Tokenized · LLMarXiv——
TRR Agent
Interpretable Decision-Making for Autonomous Vehicles with Retrieval-Augmented Reasoning via LLM
2024RAG · Rule-BasedarXiv——
OmniDrive
A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
CVPR 2025Counterfactual · 3DarXivStars—
Co-driver
VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding
2024Assistant · Human-likearXiv——
AgentsCoDriver
Large Language Model Empowered Collaborative Driving with Lifelong Learning
2024Collaborative · LifelongarXiv——
EM-VLM4AD
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering
CVPR 2024Efficient · VQAarXivStars—
LeGo-Drive
Language-enhanced Goal-oriented Closed-Loop End-to-End Autonomous Driving
IROS 2024Goal-oriented · Closed-LooparXivStarsProject
Hybrid Reasoning
Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving
ICCMA 2024Reasoning · MatharXiv——
VLAAD
Vision and Language Assistant for Autonomous Driving
WACV 2024Assistant · ExplainabilityPaper——
ELM
Embodied Understanding of Driving Scenarios
ECCV 2024Embodied · Scene UnderstandingarXiv——
RAG-Driver
Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning
RSS 2024RAG · In-ContextarXivStarsProject
BEV-TSR
Text-Scene Retrieval in BEV Space for Autonomous Driving
AAAI 2025Retrieval · BEVarXiv——
LLaDA
Driving Everywhere with Large Language Model Policy Adaptation
CVPR 2024Adaptation · Traffic RulesarXivStarsProject
2023
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
LingoQA
Visual Question Answering for Autonomous Driving
ECCV 2024VQA · LLMarXivStars—
LaMPilot
An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
CVPR 2024Benchmark · LLMarXivStars—
LLM-ASSIST
Enhancing Closed-Loop Planning with Language-Based Reasoning
2023Planning · ReasoningarXiv—Project
DriveLM
Driving with Graph Visual Question Answering
ECCV 2024Graph VQA · ReasoningarXivStars—
DriveMLM
Aligning Multi-Modal Large Language Models with Behavioral Planning States
2023MLLM · PlanningarXivStars—
LiDAR-LLM
Exploring the Potential of Large Language Models for 3D LiDAR Understanding
2023LiDAR · LLMarXiv—Project
Talk2BEV
Language-enhanced Bird's-eye View Maps for Autonomous Driving
2023BEV · LVLMarXivStarsProject
Talk2Drive
Personalized Autonomous Driving with Large Language Models: Field Experiments
2023Personalized · LLMarXiv—Project
LMDrive
Closed-Loop End-to-End Driving with Large Language Models
CVPR 2024Closed-Loop · LLMarXivStars—
Reason2Drive
Towards Interpretable and Chain-based Reasoning for Autonomous Driving
ECCV 2024Reasoning · InterpretabilityarXivStars—
CAVG
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving
2023Grounding · GPT-4arXivStars—
Dolphins
Multimodal Language Model for Driving
ECCV 2024Multimodal · VLMarXivStarsProject
Agent-Driver
A Language Agent for Autonomous Driving
COLM 2024Agent · MemoryarXivStarsProject
LLM-Safety
Empowering Autonomous Driving with Large Language Models: A Safety Perspective
ICLR 2024Safety · MPCarXivStars—
Co-Pilot
ChatGPT as Your Vehicle Co-Pilot: An Initial Attempt
2023Co-Pilot · LLMPaper——
RRR
Receive, Reason, and React: Drive as You Say with Large Language Models
ITSM 2024Tools · LLMarXiv——
LanguageMPC
Large Language Models as Decision Makers for Autonomous Driving
2023MPC · CoTarXiv——
Driving with LLMs
Fusing Object-Level Vector Modality for Explainable Autonomous Driving
2023Object-Level · ExplainablearXivStars—
DriveGPT4
Interpretable End-to-end Autonomous Driving via Large Language Model
RALInterpretable · LLMPaper—Project
GPT-Driver
Learning to Drive with GPT
NeurIPS 2023Planner · GPTarXivStarsProject
DiLu
A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
ICLR 2024Knowledge-Driven · ReflectionarXivStarsProject
Drive as You Speak
Enabling Human-Like Interaction with Large Language Models in Autonomous Vehicles
2023Interaction · LLMarXiv——
HiLM-D
Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
IJCVHigh-Res · MLLMarXiv——
SurrealDriver
Designing LLM-powered Generative Driver Agent Framework based on Human Data
2023Generative · AgentarXiv——
Drive Like a Human
Rethinking Autonomous Driving with Large Language Models
2023Reasoning · ReflectionarXivStars—
ADAPT
Action-aware Driving Caption Transformer
ICRA 2023Captioning · TransformerarXivStars—

(back to top)

Hybrid End-to-End Methods

Hybrid End-to-End Methods

2026
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
DriveDreamer-Policy
A Geometry-Grounded World–Action Model for Unified Generation and Planning
2026WAM · Geometry-GroundedarXiv—Project
Uni-World VLA
Interleaved World Modeling and Planning for Autonomous Driving
2026Interleaved · World ModelarXivStars—
HybridDriveVLA
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
2026Dual System · Fast-SlowarXiv——
DriveWorld-VLA
Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving
2026World Model · VLA · LatentarXivStars—
LatentVLA
Efficient Vision-Language Models for Autonomous Driving via Latent Action Prediction
2026Latent Action · Knowledge DistillationarXiv——
2025
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
MindDrive
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
2025World Model · VLM EvaluatorarXivStarsProject
AdaDrive
AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
ICCV 2025Slow-Fast · LLMarXivStars—
ReAL-AD
Towards Human-Like Reasoning in End-to-End Autonomous Driving
2025Reasoning · VLMarXiv—Project
VLAD
A VLM-Augmented Autonomous Driving Framework with Hierarchical Planning and Interpretable Decision Process
ITSC 2025VLM · HierarchicalarXiv——
LeAD
The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving
2025LLM · E2EarXiv——
NetRoller
Interfacing General and Specialized Models for End-to-End Autonomous Driving
2025Adapter · VLMarXivStars—
SOLVE
Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
CVPR 2025VLM · FusionarXiv——
VERDI
VLM-Embedded Reasoning for Autonomous Driving
2025VLM · ReasoningarXiv——
ALN-P3
Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
2025Alignment · LanguagearXiv——
VLM-E2E
Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
2025VLM · AttentionarXiv——
DIMA
Distilling Multi-modal Large Language Models for Autonomous Driving
CVPR 2025Distillation · MLLMarXiv——
2024
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
VLM-AD
End-to-End Autonomous Driving through Vision-Language Model Supervision
2024Supervision · VLMarXiv——
FASIONAD
FAst and Slow FusION Thinking Systems for Human-Like Autonomous Driving
2024Fast-Slow · FusionarXiv——
Senna
Bridging Large Vision-Language Models and End-to-End Autonomous Driving
2024VLM · RobustnessarXivStars—
Hint-AD
Holistically Aligned Interpretability in End-to-End Autonomous Driving
CoRL 2024Interpretability · AlignmentarXivStarsProject
DriveVLM
The Convergence of Autonomous Driving and Large Vision-Language Models
CoRL 2024Hybrid · VLMarXiv—Project
DME-Driver
Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
AAAI 2025Logic · PerceptionarXiv——
VLP
Vision Language Planning for Autonomous Driving
CVPR 2024Planning · ReasoningarXiv——

(back to top)

Dataset

Normal Dataset
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code
NAVSIM
Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
NeurIPS 2024Closed-Loop · Planning BenchmarkarXivStars
Bench2Drive
Towards Multi-Ability Benchmarking of Closed-Loop End-to-End Autonomous Driving
NeurIPS 2024CARLA · Closed-Loop · Multi-AbilityarXivStars
nuScenes
A Multimodal Dataset for Autonomous Driving
CVPR 2020Multimodal · LiDAR · RadararXivDataset
Waymo
Waymo Open Dataset: Scalability in Perception
CVPR 2020Perception · LiDARPaperDataset
HUGSIM
A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving
2024Gaussian Splatting · Closed-Loop SimarXivStars
ONCE
One Million Scenes for Autonomous Driving
NeurIPS 2021Unsupervised · 3D DetectionarXivDataset
Lyft
One Thousand and One Hours: Self-driving Motion Prediction Dataset
2020Motion PredictionarXivDataset
BDD100K
A Diverse Driving Dataset for Heterogeneous Multitask Learning
CVPR 2020Multitask · VideoarXivStars
Argoverse
3D Tracking and Forecasting with Rich Maps
CVPR 2019Tracking · Forecasting · MapsarXivDataset
ApolloScape
The ApolloScape Open Dataset for Autonomous Driving
CVPR 2018Segmentation · LiDARarXivStars
CARLA
An Open Urban Driving Simulator
CoRL 2017Simulator · Urban DrivingarXivDataset
Mapillary Vistas
Semantic Understanding of Street Scenes
ICCV 2017Semantic SegmentationPaperDataset
KITTI
The KITTI Vision Benchmark Suite
CVPR 20123D Detection · TrackingPaperDataset
Vision Language Dataset

Vision Language Dataset

2025
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
nuScenesR²-6K
Incentivizing Reasoning and Self-Reflection Capacity for VLA Model
2025CoT · ReasoningarXiv——
Bench2ADVLM
A Closed-Loop Benchmark for Vision-language Models
2025Benchmark · Closed-LooparXiv——
Drive-R1
Bridging Reasoning and Planning in VLMs with RL
2025RL · ReasoningarXiv——
STSBench
A Spatio-temporal Scenario Benchmark for MLLMs
2025Spatio-Temporal · 3DarXivDatasetProject
DriveAction
A Benchmark for Exploring Human-like Driving Decisions in VLA Models
2025Action-Driven · VLAarXivDatasetProject
S4-Driver
WOMD-Planning-ADE Benchmark: Scalable Self-Supervised Driving MLLM
CVPR 2025Self-Supervised · PlanningarXiv——
ImpromptuVLA
Open Weights and Open Data for Driving Vision-Language-Action Models
2025Open Data · VLAarXivDatasetProject
NuInteract
Extending Large Vision-Language Model for Diverse Interactive Tasks
2025Interaction · VLMarXivStarsProject
VLADBench
Fine-Grained Evaluation of Large Vision-Language Models
2025Evaluation · ReasoningarXivDatasetProject
DriveLMM-o1
A Step-by-Step Reasoning Dataset and Large Multimodal Model
2025Reasoning · MLLMarXivDatasetProject
SimLingo
Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
CVPR 2025Alignment · Closed-LooparXivDatasetProject
Robusto-1
Comparing Humans and VLMs on real out-of-distribution AD VQA
2025OOD · VQAarXivDataset—
DrivingVQA
RIV-CoT: Retrieval-Based Interleaved Visual Chain-of-Thought
2025VQA · CoTarXivDatasetProject
DriveBench
Are VLMs Ready for Autonomous Driving? An Empirical Study
ICCV 2025Reliability · EvaluationarXivDatasetProject
CoVLA
Comprehensive Vision-Language-Action Dataset
WACV 2025VLA · VideoarXivDatasetProject
WOMD-Reasoning
A Large-Scale Dataset for Interaction Reasoning in Driving
ICML 2025Interaction · ReasoningarXivDatasetProject
OmniDrive
LLM-Agent for Autonomous Driving with 3D Perception
CVPR 20253D Perception · AgentarXivStarsProject
CODA-LM
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
WACV 2025Corner Cases · EvaluationarXivDatasetProject
HiLM-D
(DRAMA-ROLISP) Enhancing MLLMs with Multi-Scale High-Resolution Details
IJCV 2025Risk · High-ResarXivStarsProject
nuPrompt
Language Prompt for Autonomous Driving
AAAI 2025Prompt · 3DarXivStarsProject
2024
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
SURDS
Benchmarking Spatial Understanding and Reasoning in Driving Scenarios
2024Spatial · ReasoningarXivDatasetProject
ContextVLM
Zero-Shot and Few-Shot Context Understanding
ITSC 2024Context · Few-ShotarXivDatasetProject
DriveCoT
Integrating Chain-of-Thought Reasoning with End-to-End Driving
2024CoT · ReasoningarXivDatasetProject
DriveVLM
SUP-AD Dataset: The Convergence of Autonomous Driving and VLMs
CoRL 2024Scene Understanding · PlanningarXiv—Project
NuInstruct
Holistic Autonomous Driving Understanding by BEV Injected Multi-Modal Large Models
CVPR 2024Instruction · BEVarXivDatasetProject
DriveLM
Driving with Graph Visual Question Answering
ECCV 2024Graph VQA · GrapharXivDatasetProject
LingoQA
Visual Question Answering for Autonomous Driving
ECCV 2024VQA · FreeformarXivStarsProject
LMDrive
Closed-Loop End-to-End Driving with Large Language Models
2024Closed-Loop · LanguagearXivDatasetProject
NuScenes-MQA
Integrated Evaluation of Captions and QA using Markup Annotations
WACV 2024Captioning · QAarXivDatasetProject
Talk2BEV
Language-enhanced Bird’s-eye View Maps
ICRA 2024BEV · MapsarXivDatasetProject
DriveGPT4
Interpretable End-to-end Autonomous Driving via LLM
RA-L 2024Interpretable · InstructionarXivDatasetProject
Rank2Tell
A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning
WACV 2024Ranking · ReasoningarXivDataset—
NuScenes-QA
A Multi-Modal Visual Question Answering Benchmark
AAAI 2024VQA · BenchmarkarXivDatasetProject
MAPLM
A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene
CVPR 2024Map · TrafficPaperDatasetProject
2023
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
DriveMLM
Aligning Multi-Modal Large Language Models with Behavioral Planning States
2023Planning · ExplanationarXivStars—
Reason2Drive
Towards Interpretable and Chain-based Reasoning for Autonomous Driving
2023Reasoning · Chain-basedarXivDatasetProject
Refer-KITTI
Referring Multi-Object Tracking
CVPR 2023Tracking · ReferringarXivStarsProject
DRAMA
Joint Risk Localization and Captioning in Driving
WACV 2023Risk · CaptioningarXivDatasetProject
Before 2023
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
SUTD-TrafficQA
A Question Answering Benchmark and an Efficient Network for Video Reasoning
CVPR 2021Video QA · ReasoningarXivDatasetProject
BDD-OIA
Explainable Object-induced Action Decision for Autonomous Vehicles
CVPR 2020Explainable · DecisionarXivDatasetProject
HAD
Grounding Human-to-Vehicle Advice for Self-driving Vehicles
CVPR 2019Advice · GroundingarXivDataset—
Talk2Car
Taking Control of Your Self-Driving Car
EMNLP 2019Commands · ReferralarXivStarsProject
BDD-X
Textual Explanations for Self-Driving Vehicles
ECCV 2018Explanation · CaptioningarXivStarsProject

(back to top)

License

The GE2EAD resources is released under the Apache 2.0 license.

(back to top)

Significant stargazers

Zichen Wen

67 followers · starred Dec 2025

AutoLab-SAI-SJTU/GE2EAD

Collects papers on autonomous driving E2E learning, VLM/VLA and Hybrid systems, with organized research branches and trends in these fields.

HTML

211

27 commits

updated May 27, 2026

See the code

README

Awesome-GE2EAD

Awesome Logo TechRxiv Project GitHub forks GitHub stars

This is the official repository for "Survey of General End-to-End Autonomous Driving: A Unified Perspective".

This project aims to provide a unified roadmap for the field by:

  • 🗂️ Literature Taxonomy: Classifying methods into Conventional (e.g., UniAD), VLM-centric (e.g., DriveLM), and Hybrid (e.g., Senna) approaches.

  • 💾 Dataset Curation: Collecting both Standard and Vision-Language datasets relevant to end-to-end AD.

  • 📈 Trend Analysis: Outlining main research branches and emerging trends based on our survey.

Citation

If you find this project useful in your research, please consider citing:

@article{yang2025survey,
  title={Survey of General End-to-End Autonomous Driving: A Unified Perspective},
  author={Yang, Yixiang and Han, Chuanrong and Mao, Runhao and others},
  journal={TechRxiv},
  year={2025},
  month={December},
  doi={10.36227/techrxiv.176523315.56439138/v1},
  url={https://doi.org/10.36227/techrxiv.176523315.56439138/v1}
}

📌 Milestones

  • 🚀 2025-12-24: We organize the list of papers in a completely new tabular format.

  • 🚀 2025-12-10: The paper “Survey of General End-to-End Autonomous Driving: A Unified Perspective” was released, and this repository was made publicly available.

Table of Contents

Mindmap, Top Methods

GE2EAD Mindmap Logo

GE2EAD Mindmap

GE2EAD Mindmap Logo

Top Methods

Papers

Conventional End-to-End Methods (VA)

Conventional End-to-End Methods (VA)

2026
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
CLOVER
Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning
2026Self-Distillation · ScoringarXivStars—
RAD-2
Scaling Reinforcement Learning in a Generator-Discriminator Framework
2026RL · Generator-DiscriminatorarXiv——
SparseDriveV2
Scoring is All You Need for End-to-End Autonomous Driving
2026Scoring · Factorized VocabarXivStars—
Latent-WAM
Latent World Action Modeling for End-to-End Autonomous Driving
2026World Model · Latent CompressionarXiv——
FlowAD
Ego-Scene Interactive Modeling for Autonomous Driving
ICLR 2026Flow Matching · Ego-ScenearXivStars—
HDP
Hyper Diffusion Planner: Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
2026Diffusion · Real-VehiclearXivStarsProject
MeanFuser
Fast One-Step Multi-Modal Trajectory Generation via MeanFlow for End-to-End Autonomous Driving
CVPR 2026MeanFlow · One-SteparXivStars—
ResWorld
Temporal Residual World Model for End-to-End Autonomous Driving
ICLR 2026World Model · Temporal ResidualarXivStars—
Drive-JEPA
Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
2026V-JEPA · DistillationarXivStars—
PlannerRFT
Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning
2026Diffusion Planner · Reinforcement Fine-TuningarXiv—Project
DrivoR
Driving on Registers
CVPR 2026Register Tokens · ViT · ScoringarXivStarsProject
AlignDrive
Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
2026Lateral-Longitudinal · Path-ConditionedarXivStarsProject
2025
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
DriveLaW
Unifying Planning and Video Generation in a Latent Driving World
2025World Model · Video GenerationarXivStarsProject
SimScale
Learning to Drive via Real-World Simulation at Scale
2025Scalable Simulation · Neural RenderingarXivStarsProject
FutureX
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
2025World Model · Latent CoTarXiv——
Spatial Retrieval AD
Spatial Retrieval Augmented Autonomous Driving
2025Retrieval · Geo ImagesarXivStarsProject
UniMM-V2X
UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving
2025MoE · Multi-AgentarXivStars—
UniLION
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
2025Linear RNNarXivStars—
DiffusionDriveV2
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
2025Diffusion · RLarXivStars—
LAP
LAP: Fast Latent Diffusion Planner with Fine-Grained Feature Distillation for Autonomous Driving
2025Latent Diffusion · PlanningarXivStars—
GuideFlow
GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
2025Generative · Flow MatchingarXivStars—
DiffRefiner
DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
2025Diffusion · RefinementarXivStars—
ResAD
ResAD: Normalized Residual Trajectory Modeling for End-to-End Autonomous Driving
2025Trajectory ModelingarXivStars—
SeerDrive
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution
NeurIPS 2025World Model · PlanningarXivStars—
BridgeDrive
Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
ICLR 2026Diffusion Bridge · Anchor-to-RefinedarXivStars—
DriveDPO
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
2025DPO · SafetyarXiv——
AnchDrive
AnchDrive: Bootstrapping Diffusion Policies with Hybrid Trajectory Anchors for End-to-End Driving
2025Diffusion · AnchorsarXiv——
AdaThinkDrive
AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
2025RL · CoTarXiv——
VeteranAD
Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving
2025Perception-PlanningarXivStars—
EvaDrive
Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
2025RL · AdversarialarXiv——
ReconDreamer-RL
Enhancing Reinforcement Learning via Diffusion-based Scene Reconstruction
2025RL · World ModelarXivStars—
GMF-Drive
Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
2025Mamba · FusionarXiv——
DistillDrive
End-to-End Multi-Mode Autonomous Driving Distillation by Isomorphic Hetero-Source Planning Model
2025DistillationarXivStars—
GEMINUS
Dual-aware Global and Scene-Adaptive Mixture-of-Experts for End-to-End Autonomous Driving
2025MoE · AdaptivearXivStars—
DiVER
Breaking Imitation Bottlenecks: Reinforced Diffusion Powers Diverse Trajectory Generation
2025RL · DiffusionarXiv——
World4Drive
End-to-End Autonomous Driving via Intention-aware Physical Latent World Model
ICCV 2025World ModelarXivStars—
FocalAD
Local Motion Planning for End-to-End Autonomous Driving
2025Motion PlanningarXiv——
GaussianFusion
Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
2025Gaussian Splatting · FusionarXivStars—
CogAD
Cognitive-Hierarchy Guided End-to-End Autonomous Driving
2025Cognitive · HierarchyarXiv——
DiffE2E
Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy
2025Diffusion · HybridarXiv—Project
TransDiffuser
End-to-end Trajectory Generation with Decorrelated Multi-modal Representation for Autonomous Driving
2025Diffusion · MultimodalarXiv——
MomAD
Don’t Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving
CVPR 2025Planning · MomentumarXivStars—
Consistency
Predictive Planner for Autonomous Driving with Consistency Models
2025Consistency · PlanningarXiv——
ARTEMIS
Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving
2025MoE · AutoregressivearXiv——
TTOG
Two Tasks, One Goal: Uniting Motion and Planning for Excellent End To End Autonomous Driving Performance
2025Multi-taskarXiv——
DiffusionDrive
Truncated Diffusion Model for End-to-End Autonomous Driving
CVPR 2025DiffusionarXivStars—
WoTE
End-to-End Driving with Online Trajectory Evaluation via BEV World Model
2025World Model · BEVarXivStars—
DMAD
Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving
2025Multi-taskarXivStars—
Centaur
Robust End-to-End Autonomous Driving with Test-Time Training
2025Test-Time TrainingarXiv——
Drive in Corridors
Enhancing the Safety of End-to-end Autonomous Driving via Corridor Learning and Planning
2025Safety · PlanningarXiv——
BridgeAD
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
CVPR 2025Prediction · PlanningarXivStars—
Hydra-MDP++
Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
2025Distillation · Multi-headarXivStars—
DiffAD
A Unified Diffusion Modeling Approach for Autonomous Driving
2025DiffusionarXiv——
GoalFlow
Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving
CVPR 2025Flow MatchingarXivStars—
HiP-AD
Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single Decoder
ICCV 2025Attention · PlanningarXivStars—
LAW
Enhancing End-to-End Autonomous Driving with Latent World Model
ICLR 2025World ModelarXivStars—
DriveTransformer
Unified Transformer for Scalable End-to-End Autonomous Driving
ICLR 2025TransformerarXivStars—
UncAD
Towards Safe End-to-end Autonomous Driving via Online Map Uncertainty
ICRA 2025Uncertainty · MaparXivStars—
RAD
Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
2025RL · 3DGSarXiv—Project
OAD
Trajectory Offset Learning: A Framework for Enhanced End-to-End Autonomous Driving
2025Trajectory · OffsetResearchGateStars—
2024
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
GaussianAD
Gaussian-Centric End-to-End Autonomous Driving
2024Gaussian Splatting · PerceptionarXivStars—
MA2T
Module-wise Adaptive Adversarial Training for End-to-end Autonomous Driving
2024Adversarial · RobustnessarXiv——
Hint-AD
Holistically Aligned Interpretability in End-to-End Autonomous Driving
2024Interpretability · AlignmentarXivStarsProject
DRAMA
An Efficient End-to-end Motion Planner for Autonomous Driving with Mamba
CVPR 2025Mamba · Motion PlanningarXivStarsProject
PPAD
Iterative Interactions of Prediction and Planning for End-to-end Autonomous Driving
ECCV 2024Prediction · PlanningarXivStars—
BEV-Planner
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
CVPR 2024BEV · EvaluationarXivStars—
EfficientFuser
Efficient Fusion and Task Guided Embedding for End-to-end Autonomous Driving
2024Efficient · FusionarXiv——
UAD
End-to-End Autonomous Driving without Costly Modularization and 3D Manual Annotation
2024UnsupervisedarXiv——
Hydra-MDP
End-to-end Multimodal Planning with Multi-target Hydra-Distillation
2024Distillation · MultimodalarXivStars—
DualAD
Disentangling the Dynamic and Static World for End-to-End Driving
CVPR 2025Dual-Stream · DynamicarXivStars—
SparseDrive
End-to-End Autonomous Driving via Sparse Scene Representation
2024Sparse · Scene ReparXivStars—
GAD
GAD-Generative Learning for HD Map-Free Autonomous Driving
2024Generative · Map-FreearXivStars—
SparseAD
Sparse Query-Centric Paradigm for Efficient End-to-End Autonomous Driving
2024Sparse · QueryarXiv——
GenAD
Generative End-to-End Autonomous Driving
ECCV 2024Generative · PredictionarXivStars—
GraphAD
Interaction Scene Graph for End-to-end Autonomous Driving
2024Graph · InteractionarXivStars—
ActiveAD
Planning-Oriented Active Learning for End-to-End Autonomous Driving
2024Active LearningarXiv——
VADv2
End-to-End Vectorized Autonomous Driving via Probabilistic Planning
2024Vectorized · ProbabilisticarXivStars—
2023
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
DriveAdapter
Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous Driving
ICCV 2023Adapter · DecouplingarXivStars—
VAD
Vectorized Scene Representation for Efficient Autonomous Driving
ICCV 2023Vectorized · EfficientarXivStars—
ThinkTwice
Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous Driving
CVPR 2023Decoder · RefinementarXivStars—
ReasonNet
End-to-End Driving with Temporal and Global Reasoning
CVPR 2023Reasoning · TemporalarXivStars—
SuperDriverAI
Towards Design and Implementation for End-to-End Learning-based Autonomous Driving
2023Attention · DNNarXiv——
UniAD
Planning-oriented Autonomous Driving
CVPR 2023Multi-task · UnifiedarXivStars—
E2E Dense
End-to-End Learning of Behavioural Inputs for Autonomous Driving in Dense Traffic
IROS 2023Optimization · Dense TrafficarXiv——
CRCHFL
Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving
IROS 2023Federated LearningarXiv——
PPGeo
Policy pre-training for autonomous driving via self-supervised geometric modeling
ICLR 2023Self-Supervised · GeometricarXivStars—
Before 2023
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
MMFN
Multi-Modal-Fusion-Net for End-to-End Driving
IROS 2022Fusion · Multi-ModalarXivStars—
KEMP
Keyframe-Based Hierarchical End-to-End Deep Model for Long-Term Trajectory Prediction
ICRA 2022Keyframe · HierarchicalarXiv——
TCP
Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong Baseline
NeurIPS 2022Trajectory · ControlarXivStars—
ST-P3
End-to-end Vision-based Autonomous Driving via Spatial-Temporal Feature Learning
ECCV 2022Spatial-Temporal · InterpretablearXivStars—
MP3
A Unified Model to Map, Perceive, Predict and Plan
CVPR 2021Mapless · PredictionPaper——
Multitask
Multi-task Learning with Attention for End-to-end Autonomous Driving
CVPR 2021Multi-task · AttentionarXivStars—
Transfuser
Multi-Modal Fusion Transformer for End-to-End Autonomous Driving
CVPR 2021Transformer · FusionPaperStars—
NEAT
Neural Attention Fields for End-to-End Autonomous Driving
ICCV 2021Attention Fields · BEVPaperStars—
Fast-LiDARNet
Efficient and Robust LiDAR-Based End-to-End Navigation
ICRA 2021LiDAR · EfficientarXiv——
IVMP
Learning Interpretable End-to-End Vision-Based Motion Planning for Autonomous Driving with Optical Flow Distillation
ICRA 2021Interpretable · Optical FlowarXiv—Project
P3
Perceive, Predict, and Plan: Safe Motion Planning Through Interpretable Semantic Representations
ECCV 2020Semantic · InterpretabilityarXiv——
DARB
Exploring data aggregation in policy learning for vision-based urban autonomous driving
CVPR 2020Data Aggregation · PolicyPaperStars—
Roach
End-to-End Urban Driving by Imitating a Reinforcement Learning Coach
ICCV 2021RL · ImitationarXivStars—
LBC
Learning by cheating
CoRL 2019Knowledge DistillationarXivStars—
CIL
End-to-End driving via conditional imitation learning
CoRL 2018Imitation LearningarXivStars—
Drive in A Day
Learning to drive in a day
2018RLarXivStars—
CNN E2E
End to End Learning for Self-Driving Cars
2016CNN · ImitationarXivStars—
ALVINN
An autonomous land vehicle in a neural network
NeurIPS 1988Neural NetworkPaper——

(back to top)

VLM-Centric End-to-End Methods

VLM-Centric End-to-End Methods

2026
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
ReflectDrive-2
Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving
2026Discrete Diffusion · Self-Edit · RLarXiv——
OneDrive
Unified Multi-Paradigm Driving with Vision-Language-Action Models
2026Single Decoder · Multi-ParadigmarXivStars—
UniDriveVLA
Unifying Understanding, Perception, and Action Planning for Autonomous Driving
2026MoT · Expert DecouplingarXivStarsProject
AutoDrive-P³
Unified Chain of Perception–Prediction–Planning Thought via Reinforcement Fine-Tuning
ICLR 2026P³ CoT · GRPOarXivStars—
DynVLA
Learning World Dynamics for Action Reasoning in Autonomous Driving
2026Dynamics CoT · RFTarXivStars—
LaST-VLA
Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
2026Latent CoT · GRPOarXivStars—
ELF-VLA
Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failures
2026Failure Diagnostics · RLarXiv——
VGGDrive
Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
CVPR 2026VLM · 3D GeometryarXivStars—
HiST-VLA
A Hierarchical Spatio-Temporal Vision-Language-Action Model for End-to-End Autonomous Driving
2026VLA · Spatio-Temporal · Token SparsearXiv——
SparseOccVLA
Bridging Occupancy and Vision-Language Models via Sparse Queries
2026Sparse Occupancy · Unified 4D UnderstandingarXivStarsProject
SGDrive
Scene-to-Goal Hierarchical World Cognition for Autonomous Driving
2026Hierarchical Cognition · Scene-Agent-GoalarXivStars—
2025
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
Counterfactual VLA
Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
2025Self-Reflective · Counterfactual ReasoningarXiv——
ColaVLA
Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning
2025Cognitive Latent Reasoning · Parallel PlanningarXivStarsProject
DrivePI
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
20254D Spatial · OccupancyarXivStars—
WAM-Diff
WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving
2025Masked Diffusion · MoE · Online RLarXivStars—
SpaceDrive
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
2025Spatial EncodingarXivStarsProject
OpenREAD
OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
2025RFT/RL · LLM-as-CriticarXivStars
CoT4AD
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning
2025VLA · CoTarXiv——
MPA
Model-Based Policy Adaptation for Closed-Loop End-to-End Autonomous Driving
NeurIPS 2025Model-Based · SimarXiv—Project
AD-R1
AD-R1: Closed-Loop Reinforcement Learning with Impartial World Models
2025RL · World ModelarXiv——
Alpamayo-R1
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving
2025VLA · ReasoningarXivStars—
DriveVLA-W0
DRIVEVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
2025VLA · World ModelarXivStars—
MTRDrive
MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving
2025VLM · MemoryarXiv——
ReflectDrive
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
2025Diffusion · VLAarXiv——
IRL-VLA
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
2025IRL · VLAarXivStars—
Prune2Drive
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models
2025VLM · PruningarXiv——
FastDriveVLA
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
2025VLA · PruningarXiv——
MCAM
Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding
2025Causal · MultimodalarXivStars—
AutoDrive-R²
Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
2025VLA · ReflectionarXiv——
DriveAgent-R1
Advancing VLM-based Autonomous Driving with Hybrid Thinking and Active Perception
2025VLM · ActivearXiv——
NavigScene
Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving
2025Navigation · PerceptionarXiv——
ADRD
LLM-DRIVEN AUTONOMOUS DRIVING BASED ON RULE-BASED DECISION SYSTEMS
2025LLM · Rule-BasedarXiv——
AutoVLA
A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning
2025VLA · RLarXivStarsProject
Poutine
Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training
2025VLT · RLarXiv——
ReCogDrive
A Reinforced Cognitive Framework for End-to-End Autonomous Driving
2025VLM · DiffusionarXivStarsProject
AD-EE
Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
2025VLM · EfficientarXiv——
FastDrive
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
2025VLM · StructuredarXiv——
HMVLM
Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
2025VLM · Long-TailarXiv——
S4-Driver
Scalable Self-Supervised Driving Multimodal Large Language Model
CVPR 2025Self-Supervised · MLLMPaper——
DiffVLA
Vision-Language Guided Diffusion Planning for Autonomous Driving
2025Diffusion · VLMarXiv——
X-Driver
Explainable Autonomous Driving with Vision-Language Models
2025MLLM · CoTarXiv——
DriveGPT4-V2
Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving
CVPR 2025LLM · Closed-LoopPaper——
DriveMind
A Dual-VLM based Reinforcement Learning Framework for Autonomous Driving
2025Dual-VLM · RLarXiv——
ReasonPlan
Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
2025MLLM · ReasoningarXivStars—
FutureSightDrive
Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
2025CoT · Spatio-TemporalarXivStars—
PADriver
Towards Personalized Autonomous Driving
2025MLLM · PersonalizedarXiv——
LDM
Unlock the Power of Unlabeled Data in Language Driving Model
ICRA 2025Self-Supervised · DistillationarXiv——
DriveMoE
Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
2025MoE · VLAarXivStarsProject
DriveMonkey
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
2025LVLM · InteractivearXivStars—
AgentThink
A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models
2025CoT · ToolsarXiv——
DSDrive
Distilling Large Language Model for Lightweight End-to-End Autonomous Driving
2025Distillation · LightweightarXiv——
LightEMMA
Lightweight End-to-end Multimodal Autonomous Driving
2025Lightweight · MultimodalarXivStars—
THCAD
Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating LLM Guidance with RL
2025LLM · RL · Fast-SlowarXiv——
DriveSOTIF
Advancing Perception SOTIF Through Multimodal Large Language Models
2025SOTIF · MLLMarXiv——
Actor-Reasoner
Interact, Instruct to Improve: A LLM-Driven Parallel Actor-Reasoner Framework
2025LLM · InteractionarXivStars—
MPDrive
Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
CVPR 2025Prompt · SpatialarXiv——
V3LMA
Visual 3D-enhanced Language Model for Autonomous Driving
20253D · LVLMarXiv——
OpenDriveVLA
Towards End-to-end Autonomous Driving with Large Vision Language Action Model
2025VLA · Open-SourcearXivStarsProject
SimLingo
Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
CVPR 2025VLA · Closed-LoopPaperStarsProject
SAFEAUTO
KNOWLEDGE-ENHANCED SAFE AUTONOMOUS DRIVING WITH MULTIMODAL FOUNDATION MODELS
ICLR 2025Safety · MultimodalarXivStars—
NuGrounding
A Multi-View 3D Visual Grounding Framework in Autonomous Driving
2025Grounding · 3DarXiv——
CoT-Drive
Efficient Motion Forecasting for Autonomous Driving with LLMs and Chain-of-Thought Prompting
2025CoT · ForecastingarXiv——
CoLMDriver
LLM-based Negotiation Benefits Cooperative Autonomous Driving
2025Cooperative · LLMarXivStars—
AlphaDrive
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
2025RL · ReasoningarXivStars—
TrackingMeetsLMM
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
2025Tracking · LMMarXivStars—
BEVDriver
Leveraging BEV Maps in LLMs for Robust Closed-Loop Driving
2025BEV · LLMarXiv——
DynRsl-VLM
Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models
2025Dynamic Res · VLMarXiv——
Sce2DriveX
A Generalized MLLM Framework for Scene-to-Drive Learning
2025MLLM · ScenearXiv——
VLM-Assisted-CL
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
2025Continual LearningarXiv——
LeapVAD
A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking
2025Cognitive · Dual-ProcessarXivStarsProject
2024
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
VLM-RL
A Unified Vision Language Model and Reinforcement Learning Framework for Safe Autonomous Driving
2024RL · VLMarXivStarsProject
GPVL
Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving
AAAI 2025Generative · 3D-VLarXivStars—
CALMM-Drive
Confidence-Aware Autonomous Driving with Large Multimodal Model
2024CoT · ConfidencearXiv——
WiseAD
Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
2024VLM · ReasoningarXivStars—
OpenEMMA
Open-Source Multimodal Model for End-to-End Autonomous Driving
WACV 2025Open-Source · MultimodalarXivStars—
FeD
Feedback-Guided Autonomous Driving
CVPR 2024Feedback · LLMPaper—Project
LeapAD
Continuously learning, adapting, and improving: A dual-process approach to autonomous driving
NeurIPS 2024Dual-Process · ContinualarXivStarsProject
DriveMM
All-in-One Large Multimodal Model for Autonomous Driving
2024Multimodal · GeneralizationarXivStarsProject
Exp-Planning
Explanation for Trajectory Planning using Multi-modal Large Language Model for Autonomous Driving
ECCV 2024Explainability · PlanningarXiv——
LaVida Drive
Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement
2024VQA · InteractionarXiv——
EMMA
End-to-End Multimodal Model for Autonomous Driving
2024End-to-End · MultimodalarXiv——
DriVLMe
Enhancing LLM-based Autonomous Driving Agents with Embodied and Social Experiences
IROS 2024Embodied · SocialarXivStarsProject
OccLLaMA
An Occupancy-Language-Action Generative World Model for Autonomous Driving
2024World Model · OccupancyarXiv——
MiniDrive
More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens
2024Efficient · MoEarXiv——
RDA-Driver
Making Large Language Models Better Planners with Reasoning-Decision Alignment
ECCV 2024Reasoning · AlignmentarXiv——
EC-Drive
Edge-Cloud Collaborative Motion Planning for Autonomous Driving with Large Language Models
ICCT 2024Edge-Cloud · CollaborativearXiv—Project
V2X-VLM
End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models
2024V2X · CooperativearXivStarsProject
Cube-LLM
Language-Image Models with 3D Understanding
20243D · Language-ImagearXiv—Project
VLM-MPC
Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC)
2024MPC · ControlarXiv——
SimpleLLM4AD
An End-to-End Vision-Language Model with Graph Visual Question Answering
IEIT SystemsGraph VQA · PipelinearXiv——
AsyncDriver
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
ECCV 2024Asynchronous · Closed-LooparXivStars—
AD-H
AUTONOMOUS DRIVING WITH HIERARCHICAL AGENTS
ICLR 2025Hierarchical · AgentsPaper——
CarLLaVA
Vision language models for camera-only closed-loop driving
2024Camera-only · Closed-LooparXiv—Project
PlanAgent
A Multi-modal Large Language Agent for Closed-loop Vehicle Motion Planning
2024Agent · Closed-LooparXiv——
Atlas
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
20243D-Tokenized · LLMarXiv——
TRR Agent
Interpretable Decision-Making for Autonomous Vehicles with Retrieval-Augmented Reasoning via LLM
2024RAG · Rule-BasedarXiv——
OmniDrive
A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
CVPR 2025Counterfactual · 3DarXivStars—
Co-driver
VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding
2024Assistant · Human-likearXiv——
AgentsCoDriver
Large Language Model Empowered Collaborative Driving with Lifelong Learning
2024Collaborative · LifelongarXiv——
EM-VLM4AD
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering
CVPR 2024Efficient · VQAarXivStars—
LeGo-Drive
Language-enhanced Goal-oriented Closed-Loop End-to-End Autonomous Driving
IROS 2024Goal-oriented · Closed-LooparXivStarsProject
Hybrid Reasoning
Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving
ICCMA 2024Reasoning · MatharXiv——
VLAAD
Vision and Language Assistant for Autonomous Driving
WACV 2024Assistant · ExplainabilityPaper——
ELM
Embodied Understanding of Driving Scenarios
ECCV 2024Embodied · Scene UnderstandingarXiv——
RAG-Driver
Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning
RSS 2024RAG · In-ContextarXivStarsProject
BEV-TSR
Text-Scene Retrieval in BEV Space for Autonomous Driving
AAAI 2025Retrieval · BEVarXiv——
LLaDA
Driving Everywhere with Large Language Model Policy Adaptation
CVPR 2024Adaptation · Traffic RulesarXivStarsProject
2023
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
LingoQA
Visual Question Answering for Autonomous Driving
ECCV 2024VQA · LLMarXivStars—
LaMPilot
An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
CVPR 2024Benchmark · LLMarXivStars—
LLM-ASSIST
Enhancing Closed-Loop Planning with Language-Based Reasoning
2023Planning · ReasoningarXiv—Project
DriveLM
Driving with Graph Visual Question Answering
ECCV 2024Graph VQA · ReasoningarXivStars—
DriveMLM
Aligning Multi-Modal Large Language Models with Behavioral Planning States
2023MLLM · PlanningarXivStars—
LiDAR-LLM
Exploring the Potential of Large Language Models for 3D LiDAR Understanding
2023LiDAR · LLMarXiv—Project
Talk2BEV
Language-enhanced Bird's-eye View Maps for Autonomous Driving
2023BEV · LVLMarXivStarsProject
Talk2Drive
Personalized Autonomous Driving with Large Language Models: Field Experiments
2023Personalized · LLMarXiv—Project
LMDrive
Closed-Loop End-to-End Driving with Large Language Models
CVPR 2024Closed-Loop · LLMarXivStars—
Reason2Drive
Towards Interpretable and Chain-based Reasoning for Autonomous Driving
ECCV 2024Reasoning · InterpretabilityarXivStars—
CAVG
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving
2023Grounding · GPT-4arXivStars—
Dolphins
Multimodal Language Model for Driving
ECCV 2024Multimodal · VLMarXivStarsProject
Agent-Driver
A Language Agent for Autonomous Driving
COLM 2024Agent · MemoryarXivStarsProject
LLM-Safety
Empowering Autonomous Driving with Large Language Models: A Safety Perspective
ICLR 2024Safety · MPCarXivStars—
Co-Pilot
ChatGPT as Your Vehicle Co-Pilot: An Initial Attempt
2023Co-Pilot · LLMPaper——
RRR
Receive, Reason, and React: Drive as You Say with Large Language Models
ITSM 2024Tools · LLMarXiv——
LanguageMPC
Large Language Models as Decision Makers for Autonomous Driving
2023MPC · CoTarXiv——
Driving with LLMs
Fusing Object-Level Vector Modality for Explainable Autonomous Driving
2023Object-Level · ExplainablearXivStars—
DriveGPT4
Interpretable End-to-end Autonomous Driving via Large Language Model
RALInterpretable · LLMPaper—Project
GPT-Driver
Learning to Drive with GPT
NeurIPS 2023Planner · GPTarXivStarsProject
DiLu
A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
ICLR 2024Knowledge-Driven · ReflectionarXivStarsProject
Drive as You Speak
Enabling Human-Like Interaction with Large Language Models in Autonomous Vehicles
2023Interaction · LLMarXiv——
HiLM-D
Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
IJCVHigh-Res · MLLMarXiv——
SurrealDriver
Designing LLM-powered Generative Driver Agent Framework based on Human Data
2023Generative · AgentarXiv——
Drive Like a Human
Rethinking Autonomous Driving with Large Language Models
2023Reasoning · ReflectionarXivStars—
ADAPT
Action-aware Driving Caption Transformer
ICRA 2023Captioning · TransformerarXivStars—

(back to top)

Hybrid End-to-End Methods

Hybrid End-to-End Methods

2026
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
DriveDreamer-Policy
A Geometry-Grounded World–Action Model for Unified Generation and Planning
2026WAM · Geometry-GroundedarXiv—Project
Uni-World VLA
Interleaved World Modeling and Planning for Autonomous Driving
2026Interleaved · World ModelarXivStars—
HybridDriveVLA
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
2026Dual System · Fast-SlowarXiv——
DriveWorld-VLA
Unified Latent-Space World Modeling with Vision–Language–Action for Autonomous Driving
2026World Model · VLA · LatentarXivStars—
LatentVLA
Efficient Vision-Language Models for Autonomous Driving via Latent Action Prediction
2026Latent Action · Knowledge DistillationarXiv——
2025
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
MindDrive
MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving
2025World Model · VLM EvaluatorarXivStarsProject
AdaDrive
AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving
ICCV 2025Slow-Fast · LLMarXivStars—
ReAL-AD
Towards Human-Like Reasoning in End-to-End Autonomous Driving
2025Reasoning · VLMarXiv—Project
VLAD
A VLM-Augmented Autonomous Driving Framework with Hierarchical Planning and Interpretable Decision Process
ITSC 2025VLM · HierarchicalarXiv——
LeAD
The LLM Enhanced Planning System Converged with End-to-end Autonomous Driving
2025LLM · E2EarXiv——
NetRoller
Interfacing General and Specialized Models for End-to-End Autonomous Driving
2025Adapter · VLMarXivStars—
SOLVE
Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
CVPR 2025VLM · FusionarXiv——
VERDI
VLM-Embedded Reasoning for Autonomous Driving
2025VLM · ReasoningarXiv——
ALN-P3
Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
2025Alignment · LanguagearXiv——
VLM-E2E
Enhancing End-to-End Autonomous Driving with Multimodal Driver Attention Fusion
2025VLM · AttentionarXiv——
DIMA
Distilling Multi-modal Large Language Models for Autonomous Driving
CVPR 2025Distillation · MLLMarXiv——
2024
🧠 Method🗓️ Year / Venue🏷️ Tags📄 Paper💻 GitHub🌐 Project
VLM-AD
End-to-End Autonomous Driving through Vision-Language Model Supervision
2024Supervision · VLMarXiv——
FASIONAD
FAst and Slow FusION Thinking Systems for Human-Like Autonomous Driving
2024Fast-Slow · FusionarXiv——
Senna
Bridging Large Vision-Language Models and End-to-End Autonomous Driving
2024VLM · RobustnessarXivStars—
Hint-AD
Holistically Aligned Interpretability in End-to-End Autonomous Driving
CoRL 2024Interpretability · AlignmentarXivStarsProject
DriveVLM
The Convergence of Autonomous Driving and Large Vision-Language Models
CoRL 2024Hybrid · VLMarXiv—Project
DME-Driver
Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
AAAI 2025Logic · PerceptionarXiv——
VLP
Vision Language Planning for Autonomous Driving
CVPR 2024Planning · ReasoningarXiv——

(back to top)

Dataset

Normal Dataset
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code
NAVSIM
Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking
NeurIPS 2024Closed-Loop · Planning BenchmarkarXivStars
Bench2Drive
Towards Multi-Ability Benchmarking of Closed-Loop End-to-End Autonomous Driving
NeurIPS 2024CARLA · Closed-Loop · Multi-AbilityarXivStars
nuScenes
A Multimodal Dataset for Autonomous Driving
CVPR 2020Multimodal · LiDAR · RadararXivDataset
Waymo
Waymo Open Dataset: Scalability in Perception
CVPR 2020Perception · LiDARPaperDataset
HUGSIM
A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous Driving
2024Gaussian Splatting · Closed-Loop SimarXivStars
ONCE
One Million Scenes for Autonomous Driving
NeurIPS 2021Unsupervised · 3D DetectionarXivDataset
Lyft
One Thousand and One Hours: Self-driving Motion Prediction Dataset
2020Motion PredictionarXivDataset
BDD100K
A Diverse Driving Dataset for Heterogeneous Multitask Learning
CVPR 2020Multitask · VideoarXivStars
Argoverse
3D Tracking and Forecasting with Rich Maps
CVPR 2019Tracking · Forecasting · MapsarXivDataset
ApolloScape
The ApolloScape Open Dataset for Autonomous Driving
CVPR 2018Segmentation · LiDARarXivStars
CARLA
An Open Urban Driving Simulator
CoRL 2017Simulator · Urban DrivingarXivDataset
Mapillary Vistas
Semantic Understanding of Street Scenes
ICCV 2017Semantic SegmentationPaperDataset
KITTI
The KITTI Vision Benchmark Suite
CVPR 20123D Detection · TrackingPaperDataset
Vision Language Dataset

Vision Language Dataset

2025
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
nuScenesR²-6K
Incentivizing Reasoning and Self-Reflection Capacity for VLA Model
2025CoT · ReasoningarXiv——
Bench2ADVLM
A Closed-Loop Benchmark for Vision-language Models
2025Benchmark · Closed-LooparXiv——
Drive-R1
Bridging Reasoning and Planning in VLMs with RL
2025RL · ReasoningarXiv——
STSBench
A Spatio-temporal Scenario Benchmark for MLLMs
2025Spatio-Temporal · 3DarXivDatasetProject
DriveAction
A Benchmark for Exploring Human-like Driving Decisions in VLA Models
2025Action-Driven · VLAarXivDatasetProject
S4-Driver
WOMD-Planning-ADE Benchmark: Scalable Self-Supervised Driving MLLM
CVPR 2025Self-Supervised · PlanningarXiv——
ImpromptuVLA
Open Weights and Open Data for Driving Vision-Language-Action Models
2025Open Data · VLAarXivDatasetProject
NuInteract
Extending Large Vision-Language Model for Diverse Interactive Tasks
2025Interaction · VLMarXivStarsProject
VLADBench
Fine-Grained Evaluation of Large Vision-Language Models
2025Evaluation · ReasoningarXivDatasetProject
DriveLMM-o1
A Step-by-Step Reasoning Dataset and Large Multimodal Model
2025Reasoning · MLLMarXivDatasetProject
SimLingo
Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
CVPR 2025Alignment · Closed-LooparXivDatasetProject
Robusto-1
Comparing Humans and VLMs on real out-of-distribution AD VQA
2025OOD · VQAarXivDataset—
DrivingVQA
RIV-CoT: Retrieval-Based Interleaved Visual Chain-of-Thought
2025VQA · CoTarXivDatasetProject
DriveBench
Are VLMs Ready for Autonomous Driving? An Empirical Study
ICCV 2025Reliability · EvaluationarXivDatasetProject
CoVLA
Comprehensive Vision-Language-Action Dataset
WACV 2025VLA · VideoarXivDatasetProject
WOMD-Reasoning
A Large-Scale Dataset for Interaction Reasoning in Driving
ICML 2025Interaction · ReasoningarXivDatasetProject
OmniDrive
LLM-Agent for Autonomous Driving with 3D Perception
CVPR 20253D Perception · AgentarXivStarsProject
CODA-LM
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
WACV 2025Corner Cases · EvaluationarXivDatasetProject
HiLM-D
(DRAMA-ROLISP) Enhancing MLLMs with Multi-Scale High-Resolution Details
IJCV 2025Risk · High-ResarXivStarsProject
nuPrompt
Language Prompt for Autonomous Driving
AAAI 2025Prompt · 3DarXivStarsProject
2024
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
SURDS
Benchmarking Spatial Understanding and Reasoning in Driving Scenarios
2024Spatial · ReasoningarXivDatasetProject
ContextVLM
Zero-Shot and Few-Shot Context Understanding
ITSC 2024Context · Few-ShotarXivDatasetProject
DriveCoT
Integrating Chain-of-Thought Reasoning with End-to-End Driving
2024CoT · ReasoningarXivDatasetProject
DriveVLM
SUP-AD Dataset: The Convergence of Autonomous Driving and VLMs
CoRL 2024Scene Understanding · PlanningarXiv—Project
NuInstruct
Holistic Autonomous Driving Understanding by BEV Injected Multi-Modal Large Models
CVPR 2024Instruction · BEVarXivDatasetProject
DriveLM
Driving with Graph Visual Question Answering
ECCV 2024Graph VQA · GrapharXivDatasetProject
LingoQA
Visual Question Answering for Autonomous Driving
ECCV 2024VQA · FreeformarXivStarsProject
LMDrive
Closed-Loop End-to-End Driving with Large Language Models
2024Closed-Loop · LanguagearXivDatasetProject
NuScenes-MQA
Integrated Evaluation of Captions and QA using Markup Annotations
WACV 2024Captioning · QAarXivDatasetProject
Talk2BEV
Language-enhanced Bird’s-eye View Maps
ICRA 2024BEV · MapsarXivDatasetProject
DriveGPT4
Interpretable End-to-end Autonomous Driving via LLM
RA-L 2024Interpretable · InstructionarXivDatasetProject
Rank2Tell
A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning
WACV 2024Ranking · ReasoningarXivDataset—
NuScenes-QA
A Multi-Modal Visual Question Answering Benchmark
AAAI 2024VQA · BenchmarkarXivDatasetProject
MAPLM
A Real-World Large-Scale Vision-Language Dataset for Map and Traffic Scene
CVPR 2024Map · TrafficPaperDatasetProject
2023
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
DriveMLM
Aligning Multi-Modal Large Language Models with Behavioral Planning States
2023Planning · ExplanationarXivStars—
Reason2Drive
Towards Interpretable and Chain-based Reasoning for Autonomous Driving
2023Reasoning · Chain-basedarXivDatasetProject
Refer-KITTI
Referring Multi-Object Tracking
CVPR 2023Tracking · ReferringarXivStarsProject
DRAMA
Joint Risk Localization and Captioning in Driving
WACV 2023Risk · CaptioningarXivDatasetProject
Before 2023
📦 Dataset🗓️ Year / Venue🏷️ Tags📄 Paper💾 Dataset / Code🌐 Project
SUTD-TrafficQA
A Question Answering Benchmark and an Efficient Network for Video Reasoning
CVPR 2021Video QA · ReasoningarXivDatasetProject
BDD-OIA
Explainable Object-induced Action Decision for Autonomous Vehicles
CVPR 2020Explainable · DecisionarXivDatasetProject
HAD
Grounding Human-to-Vehicle Advice for Self-driving Vehicles
CVPR 2019Advice · GroundingarXivDataset—
Talk2Car
Taking Control of Your Self-Driving Car
EMNLP 2019Commands · ReferralarXivStarsProject
BDD-X
Textual Explanations for Self-Driving Vehicles
ECCV 2018Explanation · CaptioningarXivStarsProject

(back to top)

License

The GE2EAD resources is released under the Apache 2.0 license.

(back to top)

Significant stargazers

Zichen Wen

67 followers · starred Dec 2025

Languages

HTML

51.9%

CSS

44.8%

JavaScript

3.2%