A curated collection of papers on E2E-AD, aimed at researchers, engineers, and enthusiasts in the field of autonomous driving systems. This repository provides a comprehensive selection of papers, focusing primarily on the training methods and ecosystems that drive the development of intelligent autonomous vehicles.
101
88 commits
updated Jun 3, 2026
Welcome to this curated collection of papers on End-to-End Autonomous Driving (E2E-AD), aimed at researchers, engineers, and enthusiasts in the field of autonomous driving systems. This repository provides a comprehensive selection of papers, focusing primarily on the training methods and ecosystems that drive the development of intelligent autonomous vehicles.
In particular, we focus on the Data-Strategy-Platform framework for E2E-AD systems, offering insights into:
Each paper in this repository has been selected for its relevance and contribution to the field, and we hope it serves as a valuable resource for anyone working in or learning about autonomous driving technology.
| 🎉 2026.05.20 | Our survey was accepted by IEEE Transactions on Intelligent Transportation Systems (TITS). |
| 🔄 2026.05.22 | Repository updated with the latest E2E-AD training ecosystem papers and resources. |
We warmly welcome pull requests and suggestions for adding new papers, benchmarks, datasets, and useful resources.
If you find this repository helpful, please consider citing our survey, starring this repository ⭐, and sharing it with the community.
| Title | Abstract | Year | Project |
|---|---|---|---|
| Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training | DetailsProposes AutoScale, a closed-loop data engine that optimizes real-synthetic driving data mixtures with scene representations, cluster reweighting, retrieval, training, and evaluation feedback. | arXiv 2026 | |
| 4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving | DetailsIntroduces a multi-modal 4D FMCW LiDAR dataset with point-wise velocity, multi-LiDAR and camera streams, 3D boxes, tracks, and benchmarks for detection, BEV flow, forecasting, and planning. | arXiv 2026 | |
| XWOD: A Real-World Benchmark for Object Detection under Extreme Weather Conditions | DetailsBuilds an extreme-weather object-detection benchmark with real traffic images across rain, snow, fog, flooding, tornado, wildfire, and other adverse conditions. | arXiv 2026 | |
| ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving | DetailsProvides a first-person driving VLM benchmark built on ScenePilot-4K, evaluating scene understanding, spatial perception, motion planning, safety reasoning, and regional generalization. | arXiv 2026 | Dataset |
| VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting | DetailsUses feed-forward 3D Gaussian Splatting to synthesize viewpoint-robust driving observations, improving end-to-end policy robustness under camera pose and viewpoint changes. | NeurIPS 2025 | Project |
| WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios | DetailsAdapts Waymo Open Dataset into an end-to-end driving benchmark emphasizing challenging long-tail scenes for planning-oriented policy evaluation. | arXiv 2025 | Project |
| SimScale: Learning to Drive via Real-World Simulation at Scale | DetailsScales real-world simulation for closed-loop end-to-end driving by converting large-scale logs into interactive training environments for policy learning and evaluation. | CVPR 2026 Oral | Project / Code |
| SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data | DetailsStudies synthetic-data augmentation for real-world end-to-end driving models, targeting better robustness and generalization under data scarcity and long-tail scenarios. | ICCV 2025 | |
| CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving | Details | WACV 2025 | Project |
| Argoverse 2: Next generation datasets for self-driving perception and forecasting | Details | arXiv 2023 | Project |
| WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving | Details | ICML 2025 | Code / Project |
| nuscenes: A multimodal dataset for autonomous driving | Details | CVPR 2020 | Project |
| One million scenes for autonomous driving: Once dataset | Details | arXiv 2021 | Project |
| Scalability in perception for autonomous driving: Waymo open dataset | DetailsIntroduces the Waymo Open Dataset for scalable autonomous-driving perception research, including synchronized LiDAR, camera, labels, and benchmarks for detection and tracking. | CVPR 2020 | Project |
| Zenseact Open Dataset: A Large-Scale and Diverse Multimodal Dataset for Autonomous Driving | DetailsIntroduces a large-scale multimodal autonomous-driving dataset with diverse European driving scenes and annotations for perception, prediction, and planning research. | ICCV 2023 | Project |
| Scaling out-of-distribution detection for real-world settings | Details | PMLR 2022 | Code |
| SHIFT: a synthetic driving dataset for continuous multi-task domain adaptation | DetailsIntroduces a synthetic driving dataset for continuous domain adaptation across weather, time, and scene changes, supporting multiple perception tasks. | CVPR 2022 | Project |
| V2x-vit: Vehicle-to-everything cooperative perception with vision transformer | Details | ECCV 2022 | Code |
| Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving | DetailsProvides a V2X benchmark for accident and motion prediction, focusing on safety-critical cooperative driving scenarios. | AAAI 2024 | |
| Bdd100k: A diverse driving dataset for heterogeneous multitask learning | Details | CVPR 2020 | Project |
| The apolloscape dataset for autonomous driving | Details | CVPR 2018 | Project |
| Tumtraf v2x cooperative perception dataset | Details | CVPR 2024 | Project |
| Tumtraf intersection dataset: All you need for urban 3d camera-lidar roadside perception | DetailsPresents an urban roadside camera-LiDAR dataset for 3D perception at intersections, supporting infrastructure-side cooperative perception research. | ITSC 2023 | Code / Project |
| V2v4real: A real-world large-scale dataset for vehicle-to-vehicle cooperative perception | DetailsIntroduces a real-world V2V cooperative perception dataset with synchronized multi-vehicle sensing for studying collaboration under realistic communication and viewpoint constraints. | CVPR 2023 | Project |
| Rope3d: The roadside perception dataset for autonomous driving and monocular 3d object detection task | Details | CVPR 2022 | Project |
| Cooperative perception for 3D object detection in driving scenarios using infrastructure sensors | Details | TITS 2020 | |
| Lumpi: The leibniz university multi-perspective intersection dataset | Details | IV 2022 | Project |
| An automated driving systems data acquisition and analytics platform | Details | TRC 2023 | |
| S-nerf++: Autonomous driving simulation via neural reconstruction and generation | Details | TPAMI 2025 | Code |
| ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration | DetailsBuilds driving world models for scene reconstruction through online restoration, improving reconstruction quality and temporal consistency for simulation and data generation. | CVPR 2025 | |
| Scene reconstruction techniques for autonomous driving: a review of 3D Gaussian splatting | Details | AIR | |
| Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving | Details | CVPR 2024 | Code / Project |
| Scenecontrol: Diffusion for controllable traffic scene generation | Details | ICRA 2024 | |
| Diffscene: Diffusion-based safety-critical scenario generation for autonomous vehicles | Details | AAAI 2025 | |
| Simulation-based reinforcement learning for real-world autonomous driving | Details | ICRA 2020 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| CLOVER: Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning | DetailsRanks candidate plans with closed-loop value estimation, improving planning selection by considering downstream interactive outcomes rather than only open-loop trajectory error. | arXiv 2026 | Code |
| Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives | DetailsImproves imitation-based driving by mining hard negative behaviors and training the policy to avoid unsafe actions in challenging scenarios. | arXiv 2026 | |
| Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling | DetailsModels ego-centric scene causality jointly with driving policy learning, aiming to separate causal factors from spurious correlations for safer end-to-end planning. | arXiv 2026 | |
| Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction | DetailsAnalyzes how temporal sampling frequency interacts with model capacity in trajectory prediction, offering guidance for data construction and policy training. | arXiv 2026 | |
| Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution | DetailsJointly models future scene evolution and ego trajectory planning, using bidirectional interactions between prediction and planning to improve driving decisions. | NeurIPS 2025 | Code |
| Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving | DetailsCouples perception and planning in a single end-to-end framework so planning supervision can shape perception features toward driving-relevant scene understanding. | AAAI 2026 | |
| Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning | DetailsUses historical prediction and future planning jointly, bridging temporal context and planning targets for stronger end-to-end driving performance. | CVPR 2025 | Code |
| Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving | DetailsIntroduces momentum-aware planning to reduce unstable steering and improve trajectory smoothness while preserving planning accuracy. | CVPR 2025 | Code |
| Planning-oriented autonomous driving | Details | CVPR 2023 | Code |
| Transfuser: Imitation with transformer-based sensor fusion for autonomous driving | Details | TPAMI 2022 | Code |
| Safety-enhanced autonomous driving using interpretable sensor fusion transformer | Details | CoRL 2022 | Code |
| Reasonnet: End-to-end driving with temporal and global reasoning | Details | CVPR 2023 | Code |
| Ppad: Iterative interactions of prediction and planning for end-to-end autonomous driving | Details | ECCV 2024 | Code |
| Multi-modal fusion transformer for end-to-end autonomous driving | Details | CVPR 2021 | Code |
| Think twice before driving: Towards scalable decoders for end-to-end autonomous driving | Details | CVPR 2023 | Code |
| Learning from all vehicles | Details | CVPR 2022 | Code |
| Neat: Neural attention fields for end-to-end autonomous driving | Details | ICCV 2021 | Code |
| Learning to steer by mimicking features from heterogeneous auxiliary networks | Details | AAAI 2019 | Code |
| Driving on Registers | DetailsInvestigates register-like internal representations for end-to-end driving, improving how models store and use scene context for planning. | arXiv 2026 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| DriveSafer: End-to-End Autonomous Driving with Safety Guidance | DetailsAdds explicit safety guidance to end-to-end driving, steering trajectory generation toward safer behavior in complex traffic scenes. | arXiv 2026 | |
| MISTY: High-Throughput Motion Planning via Mixer-based Single-step Drifting | DetailsUses mixer-based single-step trajectory generation to accelerate motion planning while preserving multi-modal planning quality. | arXiv 2026 | |
| FeaXDrive: Feasibility-aware Trajectory-Centric Diffusion Planning for End-to-End Autonomous Driving | DetailsIntroduces feasibility-aware diffusion planning centered on trajectory generation, improving the physical and driving-rule validity of predicted plans. | arXiv 2026 | |
| RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework | DetailsScales reinforcement learning with a generator-discriminator framework, targeting stronger policy optimization for autonomous driving. | arXiv 2026 | |
| HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving | DetailsCombines hierarchical diffusion planning with metric-decoupled reinforcement learning to improve safety, comfort, and task progress separately. | arXiv 2026 | |
| Temporally Decoupled Diffusion Planning for Autonomous Driving | DetailsDecouples temporal components in diffusion-based planning, enabling more flexible long-horizon trajectory generation for autonomous driving. | arXiv 2026 | |
| DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving | DetailsRefines coarse trajectory proposals through diffusion with semantic interaction modeling, improving fine-grained planning in end-to-end driving. | AAAI 2026 | |
| Driving with Advice: Large Model as Motion Advisor for Joint Planning | DetailsUses a large model as a motion advisor to guide joint planning, injecting high-level semantic advice into trajectory generation. | AAAI 2026 | |
| DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy | DetailsCombines action diffusion with supervised policy learning, balancing generative trajectory diversity with stable end-to-end driving behavior. | NeurIPS 2025 | Project |
| WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving | DetailsGenerates future waypoints through parallel coarse-to-fine discrete flow matching and further optimizes closed-loop behavior with simulator-guided rewards. | CVPR 2026 | Code |
| Dichotomous Diffusion Policy Optimization | DetailsProposes DIPOLE, a stable RL method for diffusion policies that decomposes policy improvement into reward-maximizing and reward-minimizing branches, enabling controllable inference and VLA driving experiments on NAVSIM. | ICLR 2026 | Project / Code |
| Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving | Details | CVPR 2025 | Code |
| Diffvla: Vision-language guided diffusion planning for autonomous driving | Details | arXiv 2025 | |
| DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving | Details | arXiv 2025 | |
| A Knowledge-Driven Diffusion Policy for End-to-End Autonomous Driving Based on Expert Routing | Details | arXiv 2025 | Code / Project |
| Diffusion-based planning for autonomous driving with flexible guidance | Details | arXiv 2025 | |
| Diffusion-ES: Gradient-free planning with diffusion for autonomous and instruction-guided driving | Details | CVPR 2024 | |
| Uncertainty-Based Alternative Diffusion Policy for Safe Autonomous Driving | Details | TITS 2025 | |
| Recogdrive: A reinforced cognitive framework for end-to-end autonomous driving | Details | arXiv 2025 | |
| FlowDrive: Energy Flow Field for End-to-End Autonomous Driving | Details | arXiv 2025 | Project |
| Diffusion-based planning for autonomous driving with flexible guidance | Details | arXiv 2025 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs | DetailsEvaluates driving VLA robustness under sensor perturbations such as fog, showing how perception degradation can expose fragile reasoning and planning behavior. | arXiv 2026 | |
| SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving | DetailsAligns VLA driving behavior with safety preferences by emphasizing negative examples and risk-aware supervision during training. | arXiv 2026 | |
| CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving | DetailsOptimizes prompts in latent space with contrastive objectives, improving end-to-end driving policy adaptation without heavy model retraining. | arXiv 2026 | |
| MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving | DetailsIntroduces a unified streaming VLA architecture for autonomous driving, integrating perception, language understanding, and action generation in an online setting. | arXiv 2026 | |
| OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation | DetailsPerforms one-step latent reasoning for planning while producing vision-language explanations, reducing multi-step reasoning overhead in VLA driving. | arXiv 2026 | Project |
| OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models | DetailsUnifies multiple driving paradigms in a VLA framework, connecting perception, reasoning, and action generation across different supervision modes. | arXiv 2026 | Code |
| SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model | DetailsImproves VLA driving through efficient action bridging and negative-recovery samples, helping policies learn to recover from unsafe or suboptimal actions. | arXiv 2026 | |
| UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving | DetailsUnifies scene understanding, perception, and action planning in a VLA architecture for end-to-end autonomous driving. | arXiv 2026 | Code |
| ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving | DetailsEvaluates and improves instruction-conditioned driving robustness with counterfactual language and scene perturbations. | arXiv 2026 | |
| Vega: Learning to Drive with Natural Language Instructions | DetailsTrains driving policies conditioned on natural language instructions, connecting high-level command following with trajectory-level planning. | arXiv 2026 | |
| NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning | DetailsShows that a data-efficient VLA policy can drive effectively without explicit chain-of-thought reasoning, reducing inference cost for deployment. | CVPR 2026 | Project |
| VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving | DetailsAdds cross-view geometric grounding to VLMs, improving spatial understanding and planning reliability in multi-view autonomous driving. | CVPR 2026 | Project / Code |
| FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning | DetailsProposes ReconPruner, a plug-and-play MAE-style visual token pruner that preserves foreground driving information; introduces nuScenes-FG with 241K image-mask pairs and improves nuScenes open-loop planning across pruning ratios. | AAAI 2026 | |
| VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events | DetailsAdapts pretrained VLMs to safety-critical dashcam events with metadata captions, LLM descriptions, VQA pairs, and CoT supervision, improving collision and near-collision detection with interpretable reasoning traces. | arXiv 2026 | |
| Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning | DetailsAligns VLM reasoning with end-to-end policy learning to reduce decision/planning inconsistency and improve reliability in complex driving scenes. | arXiv 2026 | |
| SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving | DetailsProposes a scene-adaptive MoE VLA architecture that routes computation by scene context for stronger robustness and efficiency in end-to-end driving. | arXiv 2026 | |
| LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving | DetailsIntroduces latent spatio-temporal reasoning for driving VLA models to improve long-horizon planning quality and robustness under complex scene dynamics. | arXiv 2026 | |
| Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models | DetailsAnalyzes narrow-policy collapse in driving VLAs and proposes exploration-centric training to improve robustness in long-tail scenarios. | arXiv 2026 | |
| Modular Autonomy with Conversational Interaction: An LLM-driven Framework for Decision Making in Autonomous Driving | DetailsConnects an LLM-based conversational interface to modular autonomy software, translating passenger commands into validated driving-system actions. | IV 2026 | |
| SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving | DetailsStructures VLM driving cognition into scene, agent, and goal levels, producing compact representations for trajectory planning. | arXiv 2026 | |
| LatentVLA: Efficient Vision-Language Models for Autonomous Driving via Latent Action Prediction | DetailsUses self-supervised latent action prediction and distillation to build efficient VLA driving policies with reduced language-annotation dependence. | arXiv 2026 | |
| A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving | DetailsIntroduces OFF-EMMA, an off-road VLA driving model that uses visual prompts and self-consistent reasoning to improve trajectory planning on rough terrain. | arXiv 2026 | |
| FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder | Details | WACV 2026 Workshop | |
| Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning | DetailsAdds counterfactual self-reflection to VLA driving, allowing the model to revise planned actions before trajectory generation in challenging scenes. | arXiv 2025 | |
| KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System | DetailsCombines driving knowledge retrieval with a value model to guide interpretable, value-aligned trajectory assessment and planning. | arXiv 2025 | |
| FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving | Details | arXiv 2025 | Code |
| ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving | Details | arXiv 2025 | Project / Code |
| DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving | Details | arXiv 2025 | Project |
| DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning | Details | arXiv 2025 | |
| OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model | DetailsDevelops a large vision-language-action model for end-to-end autonomous driving, connecting multi-view perception, language reasoning, and trajectory output. | AAAI 2026 | Code |
| Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models | DetailsReleases open weights and open data for driving VLA models, supporting reproducible research on language-conditioned autonomous driving. | arXiv 2025 | Code / Project |
| Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating Large Language Model Guidance with Reinforcement Learning | DetailsCombines fast low-level control with slower LLM-guided reasoning and reinforcement learning to improve human-centric driving decisions. | ITSC 2025 | Project |
| DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving | DetailsDevelops a vision-language reasoning model for cross-task autonomous driving, connecting perception, reasoning, and planning tasks. | arXiv 2025 | Project |
| DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving | DetailsUses vision-language guidance to condition diffusion-based trajectory planning, combining semantic reasoning with generative action prediction. | arXiv 2025 | |
| AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning | DetailsCombines adaptive reasoning with reinforcement fine-tuning in a VLA driving model to improve planning under complex scene context. | NeurIPS 2025 | Code / Project |
| Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving | DetailsExtends large vision-language models to interactive autonomous-driving tasks such as question answering, decision support, and scene-grounded reasoning. | arXiv 2025 | |
| X-Driver: Explainable Autonomous Driving with Vision-Language Models | DetailsUses vision-language models to produce explainable driving decisions, connecting visual evidence with action-level reasoning. | arXiv 2025 | |
| AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning | DetailsCombines VLM reasoning with reinforcement learning to improve autonomous-driving decisions under complex scene context. | arXiv 2025 | Code |
| Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning | DetailsBuilds a generalized MLLM framework for translating scene understanding into driving decisions and trajectories. | arXiv 2025 | |
| VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving | DetailsGuides model predictive control with VLM-derived scene understanding and driving intent for autonomous driving. | ICML 2025 | |
| VLM-E2E: Enhancing End-to-End Autonomous Driving with Multi-modal Driver Attention Fusion | DetailsFuses driver attention with multi-modal VLM features to enhance end-to-end autonomous-driving prediction and planning. | arXiv 2025 | |
| VLM-Assisted Continual learning for Visual Question Answering in Self-Driving | DetailsUses VLM assistance for continual learning in self-driving visual question answering, reducing forgetting across driving domains. | arXiv 2025 | |
| WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model | Details | arXiv 2024 | Code / Project |
| CALMM-Drive: Confidence-Aware Autonomous Driving with Large Multimodal Model | DetailsUses confidence-aware multimodal reasoning to improve reliability in end-to-end autonomous-driving decisions. | arXiv 2024 | |
| OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving | Details | WACV 2025 | Code |
| VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision | Details | arXiv 2024 | |
| TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning | DetailsIntroduces text-guided SoftSort pooling to improve multi-view driving reasoning in VLMs. | arXiv 2025 | |
| LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving | DetailsBuilds a lightweight multimodal end-to-end driving model for efficient planning and reasoning. | arXiv 2025 | Code |
| DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model | Details | RAL 2024 | Project |
| ADAPT: Action-aware Driving Caption Transformer | Details | ICRA 2023 | Code |
| Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving | Details | NeurIPS 2024 | Code |
| DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models | Details | arXiv 2024 | Project |
| LingoQA: Visual Question Answering for Autonomous Driving | Details | ECCV 2024 | Code |
| Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts | Details | NeurIPS 2024 | |
| ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation | DetailsGenerates driving actions from vision-language instructions in a holistic end-to-end autonomous-driving framework. | arXiv 2025 | Code |
| Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving | DetailsUses 3D vision-language pre-training to support generative trajectory planning for end-to-end autonomous driving. | arXiv 2025 | |
| FutureSightDrive: Visualizing Trajectory Planning with Spatio-Temporal CoT for Autonomous Driving | DetailsUses spatio-temporal chain-of-thought visualization to make trajectory planning more interpretable and reasoning-aware. | arXiv 2025 | Code |
| Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning | DetailsUses reinforcement learning to bridge VLM reasoning and trajectory planning for autonomous driving. | arXiv 2025 | |
| ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving | DetailsIntroduces a reinforced cognitive framework that aligns perception, reasoning, and planning in end-to-end driving. | arXiv 2025 | Project / Code |
| ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving | DetailsUses cognitive latent reasoning for hierarchical parallel trajectory planning in driving VLA models. | arXiv 2025 | Project |
| Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving? | DetailsDiscusses embodied intelligent driving with large multimodal models, combining semantic understanding with policy optimization for continuous decision learning. | arXiv 2026 |
🔗Refer to Link
| Title | Abstract | Year | Project |
|---|---|---|---|
| HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models | DetailsUses trajectory-guided world modeling to connect heterogeneous sensor inputs with end-to-end planning for autonomous driving. | arXiv 2026 | |
| Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving | DetailsIntegrates scene reconstruction and future generation in a joint driving world model, supporting both representation learning and planning-oriented simulation. | arXiv 2026 | |
| EponaV2: Driving World Model with Comprehensive Future Reasoning | DetailsExtends driving world modeling with comprehensive future reasoning, improving long-horizon scene prediction and planning awareness. | arXiv 2026 | |
| The DAWN of World-Action Interactive Models | DetailsStudies world-action interaction models that jointly learn how actions affect future scene evolution, bridging prediction and control. | arXiv 2026 | |
| DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning | DetailsUnifies video generation and driving planning with a geometry-grounded world-action model for action-conditioned future simulation. | arXiv 2026 | |
| Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving | DetailsTrains a latent world-action model that predicts action-conditioned future dynamics for end-to-end autonomous driving. | arXiv 2026 | |
| Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation | DetailsUnifies visual scene generation and motion planning representations so generated futures can directly support driving decisions. | arXiv 2026 | |
| Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving | DetailsCombines world-model prediction with risk-aware MPC to improve generalization and safety in end-to-end driving. | arXiv 2026 | |
| ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving | DetailsUses temporal residual world modeling to improve future scene prediction and planning-relevant representation learning. | ICLR 2026 | Code |
| Learning Vision-Language-Action World Models for Autonomous Driving | DetailsBuilds a VLA world model that learns action-conditioned future prediction and planning-relevant representations from driving data. | CVPR 2026 Findings | Project |
| LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving | DetailsConnects multimodal scene understanding with generative world modeling, using future generation to support end-to-end driving. | arXiv 2026 | |
| ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving | DetailsCombines dense world modeling with exploration-oriented training to improve VLA driving performance in diverse scenarios. | arXiv 2026 | |
| Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving | DetailsInterleaves world modeling and planning in a unified VLA framework so imagined futures and actions can refine each other. | arXiv 2026 | |
| OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models | DetailsUses long-horizon occupancy world models to simulate multi-kilometer driving scenes for scalable evaluation and training. | arXiv 2026 | |
| AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models | DetailsScales multi-agent traffic simulation with self-supervised world models, enabling realistic interaction modeling for driving policy training. | arXiv 2026 | |
| DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving | DetailsUses dual latent world models for Gaussian-centric pre-training, aligning perception, reconstruction, and future prediction in autonomous driving. | arXiv 2026 | |
| Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception | DetailsBuilds roadside infrastructure-centric world models that combine long temporal context with broad spatial coverage for cooperative perception. | arXiv 2026 | |
| DriveVA: Video Action Models are Zero-Shot Drivers | DetailsExplores whether video action models can act as zero-shot drivers by mapping visual context directly to driving actions. | arXiv 2026 | |
| Latent Chain-of-Thought World Modeling for End-to-End Driving | DetailsIntroduces latent chain-of-thought reasoning inside driving world models to improve future prediction and downstream planning. | arXiv 2025 | |
| GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation | DetailsUses 4D occupancy guidance for physics-aware driving video generation, improving spatial and temporal consistency in world-model rollouts. | arXiv 2025 | |
| X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving | DetailsBuilds a controllable ego-centric multi-camera world model to scale closed-loop evaluation and training for end-to-end driving policies. | arXiv 2026 | |
| DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving | DetailsIntroduces a flow-based dynamic world model to better capture multi-modal future scene evolution for robust planning. | arXiv 2026 | Code |
| Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges | DetailsProvides a structured taxonomy and evaluation framework for latent driving world models, highlighting open challenges for reliable deployment. | arXiv 2026 | |
| Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving | DetailsIntroduces kinematics-aware latent world modeling for improved data efficiency and physically consistent planning in autonomous driving. | arXiv 2026 | |
| MAD: Motion Appearance Decoupling for efficient Driving World Models | DetailsDecouples structured motion learning from appearance synthesis, adapting video diffusion models into controllable driving world models more efficiently. | arXiv 2026 | Project |
| WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving | DetailsAligns latent world-model representation learning with planning through hierarchical decomposition and reinforcement fine-tuning for safer end-to-end driving. | AAAI 2026 | |
| Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space | DetailsUnifies ego and surrounding-vehicle trajectory modeling in video latent space, improving interaction-aware driving world prediction. | AAAI 2026 | |
| InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement | DetailsUses latent ensemble disagreement as intrinsic motivation for reward-free world-model pretraining, enabling reusable exploration policies for driving. | arXiv 2025 | |
| DriveLaW:Unifying Planning and Video Generation in a Latent Driving World | DetailsUnifies video generation and motion planning by sharing latent world representations between future prediction and trajectory generation. | arXiv 2025 | |
| 3D-VLA: A 3D Vision-Language-Action Generative World Model | Details | ICML 2024 | Code |
| CarDreamer: Open-source learning platform for world-model-based autonomous driving | DetailsProvides an open-source learning platform for autonomous-driving research with world-model-based simulation and policy training. | IOTJ 2025 | Code |
| VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving | Details | arXiv 2025 | Project |
| Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks | Details | arXiv 2025 | |
| Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs | Details | arXiv 2025 | |
| From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction | DetailsPredicts collaborative future states and actions with a policy world model, linking multi-agent forecasting to planning. | NeurIPS 2025 | Code |
| Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks | Details | arXiv 2025 | Project |
| SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries | Details | arXiv 2025 | Code |
| OmniNWM: Omniscient Driving Navigation World Models | Details | arXiv 2025 | Project |
| DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving | Details | arXiv 2025 | Code |
| IRL-VLA: Training a Vision-Language-Action Policy via Reward World Model | DetailsTrains a VLA policy with a reward world model, using learned future feedback to guide action optimization. | arXiv 2025 | Project / Code |
| TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving | Details | arXiv 2025 | Project |
| World4Drive: End-to-end autonomous driving via intention-aware physical latent world model | DetailsUses an intention-aware physical latent world model to connect future dynamics prediction with end-to-end trajectory planning. | ICCV 2025 | Code |
| World model-based end-to-end scene generation for accident anticipation in autonomous driving | Details | arXiv 2025 | |
| DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation | Details | AAAI 2025 | Project |
| GAIA-1: A Generative World Model for Autonomous Driving | Details | arXiv 2023 | |
| Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving | Details | CVPR 2024 | Code / Project. |
| TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction | Details | ICRA 2023 | Code |
| MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction | DetailsLearns a generalizable driving world model through video mask reconstruction, improving future generation and representation transfer. | CVPR 2025 | Project / Code |
| DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation | DetailsUses world models as 4D data machines to generate temporally consistent driving scene representations for downstream perception tasks. | CVPR 2025 | Project |
| X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability | DetailsGenerates large-scale driving scenes with high fidelity and flexible controls, supporting simulation, data synthesis, and policy evaluation. | NeurIPS 2025 | Project |
| Epona: Autoregressive Diffusion World Model for Autonomous Driving | DetailsBuilds an autoregressive diffusion world model for autonomous driving, generating future scene rollouts conditioned on prior context. | ICCV 2025 | Project / Code |
| SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model | DetailsUses a generative world model for city-scale traffic simulation, enabling controllable multi-agent scenario synthesis. | CVPR 2025 | |
| ReSim: Reliable World Simulation for Autonomous Driving | DetailsProvides a reliable world simulation framework for autonomous driving that emphasizes realistic closed-loop behavior and policy evaluation. | NeurIPS 2025 | Project / Code |
| End-to-end driving with online trajectory evaluation via BEV world model | DetailsUses a BEV world model to evaluate trajectories online, improving end-to-end driving decisions through future-scene assessment. | ICCV 2025 | Code |
| Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) | DetailsAligns world models with reinforcement learning in CARLA v2, training end-to-end policies from raw observations through imagined and closed-loop feedback. | NeurIPS 2025 | |
| Semi-supervised vision-centric 3d occupancy world model for autonomous driving | Details | ICLR 2025 | |
| VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving | Details | arXiv 2025 | Project |
| AdaWM: Adaptive World-Model-Based Planning for Autonomous Driving | DetailsUses adaptive world-model-based planning to improve future prediction and trajectory selection in autonomous driving. | ICLR 2025 | |
| Genad: Generative end-to-end autonomous driving | Details | ECCV 2024 | Code |
| COME: Adding Scene-Centric Forecasting Control to Occupancy World Model | DetailsAdds scene-centric forecasting control to occupancy world models, improving controllable future prediction for autonomous driving. | NeurIPS 2025 | Code |
| UniWorld: Autonomous Driving Pre-training via World Models | Details | arXiv 2023 | |
| Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models | Details | ICRA 2025 | |
| UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving | DetailsUnifies VLM-based scene understanding, trajectory planning, and trajectory-conditioned future image generation within one driving world model. | arXiv 2026 | Project |
These repositories offer broader collections of resources that may overlap with or complement the focus of this list.
If you find this repository helpful, a citation to our paper would be greatly appreciated:
@ARTICLE{xu2026survey,
author={Xu, Chengkai and Cui, Yiming and Liu, Jiaqi and Guo, Yicheng and Qin, Cheng and Zhang, Geyuan and Dong, Xinwei and Fang, Shiyu and Hang, Peng and Sun, Jian},
journal={IEEE Transactions on Intelligent Transportation Systems},
title={A Survey on End-to-End Autonomous Driving Training From the Perspectives of Data, Strategy, and Platform},
year={2026},
volume={},
number={},
pages={1-20},
keywords={Modeling;Training;Optimization;Autonomous driving;Safety;Surveys;Vehicles;Testing;Learning (artificial intelligence);Reinforcement learning;Autonomous vehicle;end-to-end;artificial intelligence;intelligent transportation system},
doi={10.1109/TITS.2026.3695999}
}
A curated collection of papers on E2E-AD, aimed at researchers, engineers, and enthusiasts in the field of autonomous driving systems. This repository provides a comprehensive selection of papers, focusing primarily on the training methods and ecosystems that drive the development of intelligent autonomous vehicles.
101
88 commits
updated Jun 3, 2026
Welcome to this curated collection of papers on End-to-End Autonomous Driving (E2E-AD), aimed at researchers, engineers, and enthusiasts in the field of autonomous driving systems. This repository provides a comprehensive selection of papers, focusing primarily on the training methods and ecosystems that drive the development of intelligent autonomous vehicles.
In particular, we focus on the Data-Strategy-Platform framework for E2E-AD systems, offering insights into:
Each paper in this repository has been selected for its relevance and contribution to the field, and we hope it serves as a valuable resource for anyone working in or learning about autonomous driving technology.
| 🎉 2026.05.20 | Our survey was accepted by IEEE Transactions on Intelligent Transportation Systems (TITS). |
| 🔄 2026.05.22 | Repository updated with the latest E2E-AD training ecosystem papers and resources. |
We warmly welcome pull requests and suggestions for adding new papers, benchmarks, datasets, and useful resources.
If you find this repository helpful, please consider citing our survey, starring this repository ⭐, and sharing it with the community.
| Title | Abstract | Year | Project |
|---|---|---|---|
| Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training | DetailsProposes AutoScale, a closed-loop data engine that optimizes real-synthetic driving data mixtures with scene representations, cluster reweighting, retrieval, training, and evaluation feedback. | arXiv 2026 | |
| 4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving | DetailsIntroduces a multi-modal 4D FMCW LiDAR dataset with point-wise velocity, multi-LiDAR and camera streams, 3D boxes, tracks, and benchmarks for detection, BEV flow, forecasting, and planning. | arXiv 2026 | |
| XWOD: A Real-World Benchmark for Object Detection under Extreme Weather Conditions | DetailsBuilds an extreme-weather object-detection benchmark with real traffic images across rain, snow, fog, flooding, tornado, wildfire, and other adverse conditions. | arXiv 2026 | |
| ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving | DetailsProvides a first-person driving VLM benchmark built on ScenePilot-4K, evaluating scene understanding, spatial perception, motion planning, safety reasoning, and regional generalization. | arXiv 2026 | Dataset |
| VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting | DetailsUses feed-forward 3D Gaussian Splatting to synthesize viewpoint-robust driving observations, improving end-to-end policy robustness under camera pose and viewpoint changes. | NeurIPS 2025 | Project |
| WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios | DetailsAdapts Waymo Open Dataset into an end-to-end driving benchmark emphasizing challenging long-tail scenes for planning-oriented policy evaluation. | arXiv 2025 | Project |
| SimScale: Learning to Drive via Real-World Simulation at Scale | DetailsScales real-world simulation for closed-loop end-to-end driving by converting large-scale logs into interactive training environments for policy learning and evaluation. | CVPR 2026 Oral | Project / Code |
| SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data | DetailsStudies synthetic-data augmentation for real-world end-to-end driving models, targeting better robustness and generalization under data scarcity and long-tail scenarios. | ICCV 2025 | |
| CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving | Details | WACV 2025 | Project |
| Argoverse 2: Next generation datasets for self-driving perception and forecasting | Details | arXiv 2023 | Project |
| WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving | Details | ICML 2025 | Code / Project |
| nuscenes: A multimodal dataset for autonomous driving | Details | CVPR 2020 | Project |
| One million scenes for autonomous driving: Once dataset | Details | arXiv 2021 | Project |
| Scalability in perception for autonomous driving: Waymo open dataset | DetailsIntroduces the Waymo Open Dataset for scalable autonomous-driving perception research, including synchronized LiDAR, camera, labels, and benchmarks for detection and tracking. | CVPR 2020 | Project |
| Zenseact Open Dataset: A Large-Scale and Diverse Multimodal Dataset for Autonomous Driving | DetailsIntroduces a large-scale multimodal autonomous-driving dataset with diverse European driving scenes and annotations for perception, prediction, and planning research. | ICCV 2023 | Project |
| Scaling out-of-distribution detection for real-world settings | Details | PMLR 2022 | Code |
| SHIFT: a synthetic driving dataset for continuous multi-task domain adaptation | DetailsIntroduces a synthetic driving dataset for continuous domain adaptation across weather, time, and scene changes, supporting multiple perception tasks. | CVPR 2022 | Project |
| V2x-vit: Vehicle-to-everything cooperative perception with vision transformer | Details | ECCV 2022 | Code |
| Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving | DetailsProvides a V2X benchmark for accident and motion prediction, focusing on safety-critical cooperative driving scenarios. | AAAI 2024 | |
| Bdd100k: A diverse driving dataset for heterogeneous multitask learning | Details | CVPR 2020 | Project |
| The apolloscape dataset for autonomous driving | Details | CVPR 2018 | Project |
| Tumtraf v2x cooperative perception dataset | Details | CVPR 2024 | Project |
| Tumtraf intersection dataset: All you need for urban 3d camera-lidar roadside perception | DetailsPresents an urban roadside camera-LiDAR dataset for 3D perception at intersections, supporting infrastructure-side cooperative perception research. | ITSC 2023 | Code / Project |
| V2v4real: A real-world large-scale dataset for vehicle-to-vehicle cooperative perception | DetailsIntroduces a real-world V2V cooperative perception dataset with synchronized multi-vehicle sensing for studying collaboration under realistic communication and viewpoint constraints. | CVPR 2023 | Project |
| Rope3d: The roadside perception dataset for autonomous driving and monocular 3d object detection task | Details | CVPR 2022 | Project |
| Cooperative perception for 3D object detection in driving scenarios using infrastructure sensors | Details | TITS 2020 | |
| Lumpi: The leibniz university multi-perspective intersection dataset | Details | IV 2022 | Project |
| An automated driving systems data acquisition and analytics platform | Details | TRC 2023 | |
| S-nerf++: Autonomous driving simulation via neural reconstruction and generation | Details | TPAMI 2025 | Code |
| ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration | DetailsBuilds driving world models for scene reconstruction through online restoration, improving reconstruction quality and temporal consistency for simulation and data generation. | CVPR 2025 | |
| Scene reconstruction techniques for autonomous driving: a review of 3D Gaussian splatting | Details | AIR | |
| Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving | Details | CVPR 2024 | Code / Project |
| Scenecontrol: Diffusion for controllable traffic scene generation | Details | ICRA 2024 | |
| Diffscene: Diffusion-based safety-critical scenario generation for autonomous vehicles | Details | AAAI 2025 | |
| Simulation-based reinforcement learning for real-world autonomous driving | Details | ICRA 2020 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| CLOVER: Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning | DetailsRanks candidate plans with closed-loop value estimation, improving planning selection by considering downstream interactive outcomes rather than only open-loop trajectory error. | arXiv 2026 | Code |
| Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives | DetailsImproves imitation-based driving by mining hard negative behaviors and training the policy to avoid unsafe actions in challenging scenarios. | arXiv 2026 | |
| Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling | DetailsModels ego-centric scene causality jointly with driving policy learning, aiming to separate causal factors from spurious correlations for safer end-to-end planning. | arXiv 2026 | |
| Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction | DetailsAnalyzes how temporal sampling frequency interacts with model capacity in trajectory prediction, offering guidance for data construction and policy training. | arXiv 2026 | |
| Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution | DetailsJointly models future scene evolution and ego trajectory planning, using bidirectional interactions between prediction and planning to improve driving decisions. | NeurIPS 2025 | Code |
| Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving | DetailsCouples perception and planning in a single end-to-end framework so planning supervision can shape perception features toward driving-relevant scene understanding. | AAAI 2026 | |
| Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning | DetailsUses historical prediction and future planning jointly, bridging temporal context and planning targets for stronger end-to-end driving performance. | CVPR 2025 | Code |
| Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving | DetailsIntroduces momentum-aware planning to reduce unstable steering and improve trajectory smoothness while preserving planning accuracy. | CVPR 2025 | Code |
| Planning-oriented autonomous driving | Details | CVPR 2023 | Code |
| Transfuser: Imitation with transformer-based sensor fusion for autonomous driving | Details | TPAMI 2022 | Code |
| Safety-enhanced autonomous driving using interpretable sensor fusion transformer | Details | CoRL 2022 | Code |
| Reasonnet: End-to-end driving with temporal and global reasoning | Details | CVPR 2023 | Code |
| Ppad: Iterative interactions of prediction and planning for end-to-end autonomous driving | Details | ECCV 2024 | Code |
| Multi-modal fusion transformer for end-to-end autonomous driving | Details | CVPR 2021 | Code |
| Think twice before driving: Towards scalable decoders for end-to-end autonomous driving | Details | CVPR 2023 | Code |
| Learning from all vehicles | Details | CVPR 2022 | Code |
| Neat: Neural attention fields for end-to-end autonomous driving | Details | ICCV 2021 | Code |
| Learning to steer by mimicking features from heterogeneous auxiliary networks | Details | AAAI 2019 | Code |
| Driving on Registers | DetailsInvestigates register-like internal representations for end-to-end driving, improving how models store and use scene context for planning. | arXiv 2026 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| DriveSafer: End-to-End Autonomous Driving with Safety Guidance | DetailsAdds explicit safety guidance to end-to-end driving, steering trajectory generation toward safer behavior in complex traffic scenes. | arXiv 2026 | |
| MISTY: High-Throughput Motion Planning via Mixer-based Single-step Drifting | DetailsUses mixer-based single-step trajectory generation to accelerate motion planning while preserving multi-modal planning quality. | arXiv 2026 | |
| FeaXDrive: Feasibility-aware Trajectory-Centric Diffusion Planning for End-to-End Autonomous Driving | DetailsIntroduces feasibility-aware diffusion planning centered on trajectory generation, improving the physical and driving-rule validity of predicted plans. | arXiv 2026 | |
| RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework | DetailsScales reinforcement learning with a generator-discriminator framework, targeting stronger policy optimization for autonomous driving. | arXiv 2026 | |
| HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving | DetailsCombines hierarchical diffusion planning with metric-decoupled reinforcement learning to improve safety, comfort, and task progress separately. | arXiv 2026 | |
| Temporally Decoupled Diffusion Planning for Autonomous Driving | DetailsDecouples temporal components in diffusion-based planning, enabling more flexible long-horizon trajectory generation for autonomous driving. | arXiv 2026 | |
| DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving | DetailsRefines coarse trajectory proposals through diffusion with semantic interaction modeling, improving fine-grained planning in end-to-end driving. | AAAI 2026 | |
| Driving with Advice: Large Model as Motion Advisor for Joint Planning | DetailsUses a large model as a motion advisor to guide joint planning, injecting high-level semantic advice into trajectory generation. | AAAI 2026 | |
| DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy | DetailsCombines action diffusion with supervised policy learning, balancing generative trajectory diversity with stable end-to-end driving behavior. | NeurIPS 2025 | Project |
| WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving | DetailsGenerates future waypoints through parallel coarse-to-fine discrete flow matching and further optimizes closed-loop behavior with simulator-guided rewards. | CVPR 2026 | Code |
| Dichotomous Diffusion Policy Optimization | DetailsProposes DIPOLE, a stable RL method for diffusion policies that decomposes policy improvement into reward-maximizing and reward-minimizing branches, enabling controllable inference and VLA driving experiments on NAVSIM. | ICLR 2026 | Project / Code |
| Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving | Details | CVPR 2025 | Code |
| Diffvla: Vision-language guided diffusion planning for autonomous driving | Details | arXiv 2025 | |
| DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving | Details | arXiv 2025 | |
| A Knowledge-Driven Diffusion Policy for End-to-End Autonomous Driving Based on Expert Routing | Details | arXiv 2025 | Code / Project |
| Diffusion-based planning for autonomous driving with flexible guidance | Details | arXiv 2025 | |
| Diffusion-ES: Gradient-free planning with diffusion for autonomous and instruction-guided driving | Details | CVPR 2024 | |
| Uncertainty-Based Alternative Diffusion Policy for Safe Autonomous Driving | Details | TITS 2025 | |
| Recogdrive: A reinforced cognitive framework for end-to-end autonomous driving | Details | arXiv 2025 | |
| FlowDrive: Energy Flow Field for End-to-End Autonomous Driving | Details | arXiv 2025 | Project |
| Diffusion-based planning for autonomous driving with flexible guidance | Details | arXiv 2025 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs | DetailsEvaluates driving VLA robustness under sensor perturbations such as fog, showing how perception degradation can expose fragile reasoning and planning behavior. | arXiv 2026 | |
| SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving | DetailsAligns VLA driving behavior with safety preferences by emphasizing negative examples and risk-aware supervision during training. | arXiv 2026 | |
| CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving | DetailsOptimizes prompts in latent space with contrastive objectives, improving end-to-end driving policy adaptation without heavy model retraining. | arXiv 2026 | |
| MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving | DetailsIntroduces a unified streaming VLA architecture for autonomous driving, integrating perception, language understanding, and action generation in an online setting. | arXiv 2026 | |
| OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation | DetailsPerforms one-step latent reasoning for planning while producing vision-language explanations, reducing multi-step reasoning overhead in VLA driving. | arXiv 2026 | Project |
| OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models | DetailsUnifies multiple driving paradigms in a VLA framework, connecting perception, reasoning, and action generation across different supervision modes. | arXiv 2026 | Code |
| SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model | DetailsImproves VLA driving through efficient action bridging and negative-recovery samples, helping policies learn to recover from unsafe or suboptimal actions. | arXiv 2026 | |
| UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving | DetailsUnifies scene understanding, perception, and action planning in a VLA architecture for end-to-end autonomous driving. | arXiv 2026 | Code |
| ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving | DetailsEvaluates and improves instruction-conditioned driving robustness with counterfactual language and scene perturbations. | arXiv 2026 | |
| Vega: Learning to Drive with Natural Language Instructions | DetailsTrains driving policies conditioned on natural language instructions, connecting high-level command following with trajectory-level planning. | arXiv 2026 | |
| NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning | DetailsShows that a data-efficient VLA policy can drive effectively without explicit chain-of-thought reasoning, reducing inference cost for deployment. | CVPR 2026 | Project |
| VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving | DetailsAdds cross-view geometric grounding to VLMs, improving spatial understanding and planning reliability in multi-view autonomous driving. | CVPR 2026 | Project / Code |
| FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning | DetailsProposes ReconPruner, a plug-and-play MAE-style visual token pruner that preserves foreground driving information; introduces nuScenes-FG with 241K image-mask pairs and improves nuScenes open-loop planning across pruning ratios. | AAAI 2026 | |
| VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events | DetailsAdapts pretrained VLMs to safety-critical dashcam events with metadata captions, LLM descriptions, VQA pairs, and CoT supervision, improving collision and near-collision detection with interpretable reasoning traces. | arXiv 2026 | |
| Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning | DetailsAligns VLM reasoning with end-to-end policy learning to reduce decision/planning inconsistency and improve reliability in complex driving scenes. | arXiv 2026 | |
| SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving | DetailsProposes a scene-adaptive MoE VLA architecture that routes computation by scene context for stronger robustness and efficiency in end-to-end driving. | arXiv 2026 | |
| LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving | DetailsIntroduces latent spatio-temporal reasoning for driving VLA models to improve long-horizon planning quality and robustness under complex scene dynamics. | arXiv 2026 | |
| Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models | DetailsAnalyzes narrow-policy collapse in driving VLAs and proposes exploration-centric training to improve robustness in long-tail scenarios. | arXiv 2026 | |
| Modular Autonomy with Conversational Interaction: An LLM-driven Framework for Decision Making in Autonomous Driving | DetailsConnects an LLM-based conversational interface to modular autonomy software, translating passenger commands into validated driving-system actions. | IV 2026 | |
| SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving | DetailsStructures VLM driving cognition into scene, agent, and goal levels, producing compact representations for trajectory planning. | arXiv 2026 | |
| LatentVLA: Efficient Vision-Language Models for Autonomous Driving via Latent Action Prediction | DetailsUses self-supervised latent action prediction and distillation to build efficient VLA driving policies with reduced language-annotation dependence. | arXiv 2026 | |
| A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving | DetailsIntroduces OFF-EMMA, an off-road VLA driving model that uses visual prompts and self-consistent reasoning to improve trajectory planning on rough terrain. | arXiv 2026 | |
| FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder | Details | WACV 2026 Workshop | |
| Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning | DetailsAdds counterfactual self-reflection to VLA driving, allowing the model to revise planned actions before trajectory generation in challenging scenes. | arXiv 2025 | |
| KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System | DetailsCombines driving knowledge retrieval with a value model to guide interpretable, value-aligned trajectory assessment and planning. | arXiv 2025 | |
| FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving | Details | arXiv 2025 | Code |
| ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving | Details | arXiv 2025 | Project / Code |
| DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving | Details | arXiv 2025 | Project |
| DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning | Details | arXiv 2025 | |
| OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model | DetailsDevelops a large vision-language-action model for end-to-end autonomous driving, connecting multi-view perception, language reasoning, and trajectory output. | AAAI 2026 | Code |
| Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models | DetailsReleases open weights and open data for driving VLA models, supporting reproducible research on language-conditioned autonomous driving. | arXiv 2025 | Code / Project |
| Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating Large Language Model Guidance with Reinforcement Learning | DetailsCombines fast low-level control with slower LLM-guided reasoning and reinforcement learning to improve human-centric driving decisions. | ITSC 2025 | Project |
| DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving | DetailsDevelops a vision-language reasoning model for cross-task autonomous driving, connecting perception, reasoning, and planning tasks. | arXiv 2025 | Project |
| DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving | DetailsUses vision-language guidance to condition diffusion-based trajectory planning, combining semantic reasoning with generative action prediction. | arXiv 2025 | |
| AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning | DetailsCombines adaptive reasoning with reinforcement fine-tuning in a VLA driving model to improve planning under complex scene context. | NeurIPS 2025 | Code / Project |
| Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving | DetailsExtends large vision-language models to interactive autonomous-driving tasks such as question answering, decision support, and scene-grounded reasoning. | arXiv 2025 | |
| X-Driver: Explainable Autonomous Driving with Vision-Language Models | DetailsUses vision-language models to produce explainable driving decisions, connecting visual evidence with action-level reasoning. | arXiv 2025 | |
| AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning | DetailsCombines VLM reasoning with reinforcement learning to improve autonomous-driving decisions under complex scene context. | arXiv 2025 | Code |
| Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning | DetailsBuilds a generalized MLLM framework for translating scene understanding into driving decisions and trajectories. | arXiv 2025 | |
| VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving | DetailsGuides model predictive control with VLM-derived scene understanding and driving intent for autonomous driving. | ICML 2025 | |
| VLM-E2E: Enhancing End-to-End Autonomous Driving with Multi-modal Driver Attention Fusion | DetailsFuses driver attention with multi-modal VLM features to enhance end-to-end autonomous-driving prediction and planning. | arXiv 2025 | |
| VLM-Assisted Continual learning for Visual Question Answering in Self-Driving | DetailsUses VLM assistance for continual learning in self-driving visual question answering, reducing forgetting across driving domains. | arXiv 2025 | |
| WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model | Details | arXiv 2024 | Code / Project |
| CALMM-Drive: Confidence-Aware Autonomous Driving with Large Multimodal Model | DetailsUses confidence-aware multimodal reasoning to improve reliability in end-to-end autonomous-driving decisions. | arXiv 2024 | |
| OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving | Details | WACV 2025 | Code |
| VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision | Details | arXiv 2024 | |
| TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning | DetailsIntroduces text-guided SoftSort pooling to improve multi-view driving reasoning in VLMs. | arXiv 2025 | |
| LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving | DetailsBuilds a lightweight multimodal end-to-end driving model for efficient planning and reasoning. | arXiv 2025 | Code |
| DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model | Details | RAL 2024 | Project |
| ADAPT: Action-aware Driving Caption Transformer | Details | ICRA 2023 | Code |
| Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving | Details | NeurIPS 2024 | Code |
| DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models | Details | arXiv 2024 | Project |
| LingoQA: Visual Question Answering for Autonomous Driving | Details | ECCV 2024 | Code |
| Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts | Details | NeurIPS 2024 | |
| ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation | DetailsGenerates driving actions from vision-language instructions in a holistic end-to-end autonomous-driving framework. | arXiv 2025 | Code |
| Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving | DetailsUses 3D vision-language pre-training to support generative trajectory planning for end-to-end autonomous driving. | arXiv 2025 | |
| FutureSightDrive: Visualizing Trajectory Planning with Spatio-Temporal CoT for Autonomous Driving | DetailsUses spatio-temporal chain-of-thought visualization to make trajectory planning more interpretable and reasoning-aware. | arXiv 2025 | Code |
| Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning | DetailsUses reinforcement learning to bridge VLM reasoning and trajectory planning for autonomous driving. | arXiv 2025 | |
| ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving | DetailsIntroduces a reinforced cognitive framework that aligns perception, reasoning, and planning in end-to-end driving. | arXiv 2025 | Project / Code |
| ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving | DetailsUses cognitive latent reasoning for hierarchical parallel trajectory planning in driving VLA models. | arXiv 2025 | Project |
| Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving? | DetailsDiscusses embodied intelligent driving with large multimodal models, combining semantic understanding with policy optimization for continuous decision learning. | arXiv 2026 |
🔗Refer to Link
| Title | Abstract | Year | Project |
|---|---|---|---|
| HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models | DetailsUses trajectory-guided world modeling to connect heterogeneous sensor inputs with end-to-end planning for autonomous driving. | arXiv 2026 | |
| Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving | DetailsIntegrates scene reconstruction and future generation in a joint driving world model, supporting both representation learning and planning-oriented simulation. | arXiv 2026 | |
| EponaV2: Driving World Model with Comprehensive Future Reasoning | DetailsExtends driving world modeling with comprehensive future reasoning, improving long-horizon scene prediction and planning awareness. | arXiv 2026 | |
| The DAWN of World-Action Interactive Models | DetailsStudies world-action interaction models that jointly learn how actions affect future scene evolution, bridging prediction and control. | arXiv 2026 | |
| DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning | DetailsUnifies video generation and driving planning with a geometry-grounded world-action model for action-conditioned future simulation. | arXiv 2026 | |
| Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving | DetailsTrains a latent world-action model that predicts action-conditioned future dynamics for end-to-end autonomous driving. | arXiv 2026 | |
| Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation | DetailsUnifies visual scene generation and motion planning representations so generated futures can directly support driving decisions. | arXiv 2026 | |
| Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving | DetailsCombines world-model prediction with risk-aware MPC to improve generalization and safety in end-to-end driving. | arXiv 2026 | |
| ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving | DetailsUses temporal residual world modeling to improve future scene prediction and planning-relevant representation learning. | ICLR 2026 | Code |
| Learning Vision-Language-Action World Models for Autonomous Driving | DetailsBuilds a VLA world model that learns action-conditioned future prediction and planning-relevant representations from driving data. | CVPR 2026 Findings | Project |
| LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving | DetailsConnects multimodal scene understanding with generative world modeling, using future generation to support end-to-end driving. | arXiv 2026 | |
| ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving | DetailsCombines dense world modeling with exploration-oriented training to improve VLA driving performance in diverse scenarios. | arXiv 2026 | |
| Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving | DetailsInterleaves world modeling and planning in a unified VLA framework so imagined futures and actions can refine each other. | arXiv 2026 | |
| OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models | DetailsUses long-horizon occupancy world models to simulate multi-kilometer driving scenes for scalable evaluation and training. | arXiv 2026 | |
| AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models | DetailsScales multi-agent traffic simulation with self-supervised world models, enabling realistic interaction modeling for driving policy training. | arXiv 2026 | |
| DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving | DetailsUses dual latent world models for Gaussian-centric pre-training, aligning perception, reconstruction, and future prediction in autonomous driving. | arXiv 2026 | |
| Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception | DetailsBuilds roadside infrastructure-centric world models that combine long temporal context with broad spatial coverage for cooperative perception. | arXiv 2026 | |
| DriveVA: Video Action Models are Zero-Shot Drivers | DetailsExplores whether video action models can act as zero-shot drivers by mapping visual context directly to driving actions. | arXiv 2026 | |
| Latent Chain-of-Thought World Modeling for End-to-End Driving | DetailsIntroduces latent chain-of-thought reasoning inside driving world models to improve future prediction and downstream planning. | arXiv 2025 | |
| GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation | DetailsUses 4D occupancy guidance for physics-aware driving video generation, improving spatial and temporal consistency in world-model rollouts. | arXiv 2025 | |
| X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving | DetailsBuilds a controllable ego-centric multi-camera world model to scale closed-loop evaluation and training for end-to-end driving policies. | arXiv 2026 | |
| DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving | DetailsIntroduces a flow-based dynamic world model to better capture multi-modal future scene evolution for robust planning. | arXiv 2026 | Code |
| Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges | DetailsProvides a structured taxonomy and evaluation framework for latent driving world models, highlighting open challenges for reliable deployment. | arXiv 2026 | |
| Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving | DetailsIntroduces kinematics-aware latent world modeling for improved data efficiency and physically consistent planning in autonomous driving. | arXiv 2026 | |
| MAD: Motion Appearance Decoupling for efficient Driving World Models | DetailsDecouples structured motion learning from appearance synthesis, adapting video diffusion models into controllable driving world models more efficiently. | arXiv 2026 | Project |
| WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving | DetailsAligns latent world-model representation learning with planning through hierarchical decomposition and reinforcement fine-tuning for safer end-to-end driving. | AAAI 2026 | |
| Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space | DetailsUnifies ego and surrounding-vehicle trajectory modeling in video latent space, improving interaction-aware driving world prediction. | AAAI 2026 | |
| InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement | DetailsUses latent ensemble disagreement as intrinsic motivation for reward-free world-model pretraining, enabling reusable exploration policies for driving. | arXiv 2025 | |
| DriveLaW:Unifying Planning and Video Generation in a Latent Driving World | DetailsUnifies video generation and motion planning by sharing latent world representations between future prediction and trajectory generation. | arXiv 2025 | |
| 3D-VLA: A 3D Vision-Language-Action Generative World Model | Details | ICML 2024 | Code |
| CarDreamer: Open-source learning platform for world-model-based autonomous driving | DetailsProvides an open-source learning platform for autonomous-driving research with world-model-based simulation and policy training. | IOTJ 2025 | Code |
| VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving | Details | arXiv 2025 | Project |
| Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks | Details | arXiv 2025 | |
| Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs | Details | arXiv 2025 | |
| From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction | DetailsPredicts collaborative future states and actions with a policy world model, linking multi-agent forecasting to planning. | NeurIPS 2025 | Code |
| Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks | Details | arXiv 2025 | Project |
| SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries | Details | arXiv 2025 | Code |
| OmniNWM: Omniscient Driving Navigation World Models | Details | arXiv 2025 | Project |
| DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving | Details | arXiv 2025 | Code |
| IRL-VLA: Training a Vision-Language-Action Policy via Reward World Model | DetailsTrains a VLA policy with a reward world model, using learned future feedback to guide action optimization. | arXiv 2025 | Project / Code |
| TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving | Details | arXiv 2025 | Project |
| World4Drive: End-to-end autonomous driving via intention-aware physical latent world model | DetailsUses an intention-aware physical latent world model to connect future dynamics prediction with end-to-end trajectory planning. | ICCV 2025 | Code |
| World model-based end-to-end scene generation for accident anticipation in autonomous driving | Details | arXiv 2025 | |
| DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation | Details | AAAI 2025 | Project |
| GAIA-1: A Generative World Model for Autonomous Driving | Details | arXiv 2023 | |
| Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving | Details | CVPR 2024 | Code / Project. |
| TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction | Details | ICRA 2023 | Code |
| MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction | DetailsLearns a generalizable driving world model through video mask reconstruction, improving future generation and representation transfer. | CVPR 2025 | Project / Code |
| DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation | DetailsUses world models as 4D data machines to generate temporally consistent driving scene representations for downstream perception tasks. | CVPR 2025 | Project |
| X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability | DetailsGenerates large-scale driving scenes with high fidelity and flexible controls, supporting simulation, data synthesis, and policy evaluation. | NeurIPS 2025 | Project |
| Epona: Autoregressive Diffusion World Model for Autonomous Driving | DetailsBuilds an autoregressive diffusion world model for autonomous driving, generating future scene rollouts conditioned on prior context. | ICCV 2025 | Project / Code |
| SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model | DetailsUses a generative world model for city-scale traffic simulation, enabling controllable multi-agent scenario synthesis. | CVPR 2025 | |
| ReSim: Reliable World Simulation for Autonomous Driving | DetailsProvides a reliable world simulation framework for autonomous driving that emphasizes realistic closed-loop behavior and policy evaluation. | NeurIPS 2025 | Project / Code |
| End-to-end driving with online trajectory evaluation via BEV world model | DetailsUses a BEV world model to evaluate trajectories online, improving end-to-end driving decisions through future-scene assessment. | ICCV 2025 | Code |
| Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) | DetailsAligns world models with reinforcement learning in CARLA v2, training end-to-end policies from raw observations through imagined and closed-loop feedback. | NeurIPS 2025 | |
| Semi-supervised vision-centric 3d occupancy world model for autonomous driving | Details | ICLR 2025 | |
| VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving | Details | arXiv 2025 | Project |
| AdaWM: Adaptive World-Model-Based Planning for Autonomous Driving | DetailsUses adaptive world-model-based planning to improve future prediction and trajectory selection in autonomous driving. | ICLR 2025 | |
| Genad: Generative end-to-end autonomous driving | Details | ECCV 2024 | Code |
| COME: Adding Scene-Centric Forecasting Control to Occupancy World Model | DetailsAdds scene-centric forecasting control to occupancy world models, improving controllable future prediction for autonomous driving. | NeurIPS 2025 | Code |
| UniWorld: Autonomous Driving Pre-training via World Models | Details | arXiv 2023 | |
| Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models | Details | ICRA 2025 | |
| UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving | DetailsUnifies VLM-based scene understanding, trajectory planning, and trajectory-conditioned future image generation within one driving world model. | arXiv 2026 | Project |
These repositories offer broader collections of resources that may overlap with or complement the focus of this list.
If you find this repository helpful, a citation to our paper would be greatly appreciated:
@ARTICLE{xu2026survey,
author={Xu, Chengkai and Cui, Yiming and Liu, Jiaqi and Guo, Yicheng and Qin, Cheng and Zhang, Geyuan and Dong, Xinwei and Fang, Shiyu and Hang, Peng and Sun, Jian},
journal={IEEE Transactions on Intelligent Transportation Systems},
title={A Survey on End-to-End Autonomous Driving Training From the Perspectives of Data, Strategy, and Platform},
year={2026},
volume={},
number={},
pages={1-20},
keywords={Modeling;Training;Optimization;Autonomous driving;Safety;Surveys;Vehicles;Testing;Learning (artificial intelligence);Reinforcement learning;Autonomous vehicle;end-to-end;artificial intelligence;intelligent transportation system},
doi={10.1109/TITS.2026.3695999}
}