A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform
Organizations: College of Transportation, Tongji University, Shanghai 201804, China
Abstract
Autonomous driving is a cornerstone technology for the future of intelligent transportation, where end-to-end learning has emerged as a transformative paradigm that directly maps multimodal sensory inputs to driving actions through unified differentiable models. While offering advantages, the effectiveness of end-to-end autonomous driving (E2E-AD) is ultimately determined by the quality of its training ecosystem. This paper provides a comprehensive review of training methods and ecosystem for E2E-AD. We introduce a Data-Strategy-Platform taxonomy that conceptualizes training as an interdependent system. The data layer defines what can be learned, the strategy layer governs how learning aligns with driving objectives, and the platform layer supports scalability and continuous evolution. Within this framework, we survey recent advances across data-centric pipelines, learning paradigms, and training infrastructures, and analyze their interplay in shaping model performance, robustness, and deployability. Finally, we reflect on current limitations and articulate a forward-looking vision that emphasizes a shift from data quantity to data value, from isolated optimization to foundation-driven generalization, and from static training to integrated training-testing loops, aiming toward robust, scalable, and trustworthy autonomous driving systems. We maintain a continuously updated repository tracking cutting-edge literature and works at \href{https://github.com/Jiaaqiliu/Awesome-Training-Ecosystem-for-E2E-AD}{Our Project Page}.
Figures & tables
| Survey | Data | Str. | Plat. | Main focus |
| Zhou et al. [ 15 ] | ✗ | ✓ | ✗ | MLLMs |
| Cui et al. [ 9 ] | ✗ | ✓ | ✗ | MLLMs |
| Chen et al. [ 7 ] | ✓ | ✓ | ✗ | E2E Architecture |
| Zhao et al. [ 17 ] | ✗ | ✓ | ✗ | Deep Learning |
| Wu et al. [ 18 ] | ✗ | ✓ | ✓ | Reinforcement Learning |
| Jiang et al. [ 16 ] | ✗ | ✓ | ✗ | MLLMs |
| Dataset | Year | Scale | Camera | LiDAR | RaDAR | QA | Type | Data Source |
| (A) Emphasizing Data Scale | ||||||||
| DDD17 [ 69 ] | 2017 | 12 hours | ✓ | ✗ | ✗ | ✗ | Real | Europe |
| BDD [ 70 ] | 2017 | 10000+ km | ✓ | ✓ | ✗ | ✗ | Real | China |
| DDD20 [ 71 ] | 2020 | 51 hours | ✓ | ✗ | ✗ | ✗ | Real | USA |
| A2D2 [ 72 ] | 2020 | 41,277 clips | ✓ | ✓ | ✗ | ✗ | Real | Germany |
| Nuscenes [ 51 ] | 2021 | 40K clips | ✓ | ✓ | ✓ | ✓* | Real | USA & Singapore |
| Paradigm | Multi modal | AR error | Traj- level | RT ready | Long- horizon | Low intera. | Interp. | Best role | Bottleneck |
|---|---|---|---|---|---|---|---|---|---|
| IL | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | Fast policy prior; Supervised initialization | Closed-loop drift; Long-tail blind spots |
| RL | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | Objective-aligned refinement; Closed-loop optimization | Reward and exploration risk; Sim-to-real gap |
| DP | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | Multimodal trajectory proposer; Timing consistency | Sampling latency; Hard-constraint guarantees |
| MLLM | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | Semantic reasoning; Intent inference | Hallucination and control gap; Real-time cost |
| WM | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✗ | Imagination-based planning; Closed-loop generation | Rollout drift; Model bias and hallucination |
| Method | Year | Paradigm | Training Hardware * |
| InterFuser [ 241 ] | CoRL 2022 | IL | 8 V100 |
| UniAD [ 4 ] | CVPR 2023 | IL | 16 A100 |
| DiffusionDrive [ 162 ] | CVPR 2025 | DP | 8 RTX4090 |
| AutoVLA [ 195 ] | NeurIPS 2025 | MLLM | 8 L40 |
| ORION [ 183 ] | ICCV 2025 | MLLM | 32 A800 |
| FSDrive [ 196 ] | NeurIPS 2025 | MLLM | 32 A6000 |
| Company | Learning Mechanism | Testing Approach | Key Focus | Advantages | Challenges |
| Waymo | Continuous retraining with real-world validation | Simulation-first using Carcraft platform | High-fidelity, rule-based scenario generation for safety validation | High reproducibility and interpretability; rigorous safety validation | High scalability costs; limited adaptability to real-world conditions |
| Tesla | Continuous fleet learning with online retraining | Real-world testing through fleet data collection | Event-triggered data collection; large-scale fleet feedback; automatic data annotation | Scalable exposure to real-world conditions; rapid iteration of training | Weak reproducibility; regulatory challenges regarding data privacy and safety |
| Cruise | Incremental city-level training updates with controlled testing zones | Hybrid real-world and simulation testing | City-specific training and safety validation | Robust version control; effective safety monitoring within urban environments | Limited scalability across regions; expensive infrastructure for testing and validation |
| Baidu Apollo | Periodic retraining with continuous deployment in open-road pilots | Hybrid real-world and simulation testing | Government-integrated data governance and safety compliance | Strong regulatory compliance; consistent model reproducibility | Slower feedback adaptation; challenges in handling diverse driving environments |
| XPeng | Federated learning and asynchronous updates for cross-regional adaptation | Cloud-edge collaborative training and testing | Cross-region data federation for model adaptation and deployment | Efficient cross-domain updates; real-time adaptability to regional differences | High infrastructure costs; limited transparency in testing methodologies |
| Huawei ADS | Continuous learning with gray release and rollback mechanisms | Cloud-edge collaborative training and testing | Stable asynchronous updates and safety-focused rollback strategies | Reliable model adaptation with real-time updates; controllable risk management | Limited openness and transparency; high computational cost for training |