SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving
Authors: Chengjun Zhang, Yuhao Zhang, Jie Yang, Mohamad Sawan
Organizations: Zhejiang Key Laboratory of 3D Micro/Nano Fabrication and Characterization, Westlake Institute for Optoelectronics, Fuyang, Hangzhou, China · Integrated-On-Chips Brain-Computer Interfaces Zhejiang Engineering Research Center, Hangzhou, Zhejiang, China · CenBRAIN Neurotech, School of Engineering, Westlake University, Hangzhou, Zhejiang, China
End-to-end autonomous driving demands trajectory planners that are both highly accurate and cheap enough for edge deployment. State-of-the-art artificial neural network (ANN) planners meet the accuracy requirement at the cost of heavy dense computation, while spiking neural networks (SNNs)---though promising orders-of-magnitude energy savings through sparse, event-driven arithmetic---still lag far behind in planning accuracy. We present \textbf{SDPAD}, a fully spike-driven end-to-end planning pipeline that closes this gap. SDPAD converts a pre-trained ANN perception stack into integer-spike form via quantized ANN2SNN conversion, lifts multi-view images into the bird's-eye-view (BEV) space with a spike-driven-max (SDM) depth distribution (Spike-3D-Lift), and plans through the Spike-QFormer, a spiking query transformer in which ego, agent, and map queries distilled from the BEV scene are fused by learnable waypoint queries via cross-attention, followed by deformable spike-cross-attention refinement. Every operation is gated by integer spikes and inference is a single feed-forward pass without temporal simulation loops. On the nuScenes open-loop benchmark, SDPAD achieves an average L2 error of 0.40,m and a collision rate of 0.12%, on par with strong ANN planners while consuming 69.9,mJ---less than 2% of recent ANN baselines. In closed-loop evaluation on the NAVSIM navtest split, SDPAD reaches 86.3 PDMS, surpassing the previous SNN planner SAD by 4.3 points and matching mainstream ANN planners at a fraction of their energy. To our knowledge, SDPAD is the first fully spike-driven planner evaluated in end-to-end autonomous driving, demonstrating that SNNs can rival dense ANNs in complex driving tasks.
Figures & tables
Figure 1: The comparison of different end-to-end paradigms. (a) Dense ANN pipelines compute in floating point throughout. (b) Existing SNN planners still rely on continuous membrane potentials or softmax depth scores for BEV lifting. (c) SDPAD is fully spike-driven end to end: the proposed Spike-3D-Lift turn BEV sampling into spike BEV representation.
Figure 2: Overall architecture of the proposed Pipline. Multi - view images are encoded into a Spiking BEV representation via E - Spikeformer and Spike - 3D - Lift, which simultaneously supports semantic segmentation and 3D object detection. The planning stage then refines waypoints with a 3 - mode reference curve and Deformable SCA, and finally outputs the ego - vehicle trajectory.
Figure 3: The illustration of the Spike - 3D-Lift for Spike-BEV generation process. Image features are weighted by the accumulated spiking depth scores to form 3D frustum features (membrane potentials), followed by camera - based interpolation and spiking neural conversion to Spiking BEV Representation.
Method
SNN
Img. Backbone
Auxiliary Task
L2 (m) ↓
Collision Rate (%) ↓
Energy (mJ) ↓
1s
2s
3s
Avg.
1s
2s
3s
Avg.
ST-P3 Hu et al. (2022)
✗
EffNet-b4
Det& Map
1.33
2.11
2.90
2.11
0.23
0.62
1.27
0.71
3520.40
UniAD Hu et al. (2023)
✗
ResNet-101
Det&Track&Map&Motion&Occ
0.45
0.70
1.04
0.73
0.62
0.58
0.63
0.61
-
VAD Jiang et al. (2023)
✗
ResNet-50
Det&Map&Motion
0.41
0.70
1.05
0.72
0.07
0.17
0.41
0.22
-
SparseDrive Sun et al. (2025)
✗
ResNet-50
Det&Track&Map&Motion
0.29
0.58
0.96
0.61
0.01
0.05
0.18
0.08
-
DiffusionDrive Liao et al. (2025)
✗
ResNet-50
Det&Track&Map&Motion
0.27
0.54
0.90
0.57
0.03
0.05
0.16
0.08
-
Table 1: Comparison on nuScenes dataset with open-loop metrics. metric calculated in the same way as VAD
Method
SNN
Img. Backbone
Auxiliary Task
NAC ↑
DAC ↑
TTC ↑
Comf. ↑
EP ↑
PDMS ↑
Energy (mJ) ↓
Transfuser Chitta et al. (2022)
✗
ResNet-34
Det&Map
97.7
92.8
92.8
100
79.2
84.0
114.46
UniAD Hu et al. (2023)
✗
ResNet-34
Det&Map
97.8
91.9
92.9
100
78.8
83.4
-
VADv2 Chen et al. (2024b)
✗
ResNet-34
Det&Map
97.2
89.1
91.6
100
76.0
80.9
-
Hydra-MDP-W-EP Li et al. (2024)
✗
ResNet-34
Det&Map
98.3
96.0
94.6
100
78.7
86.5
-
LAW Li et al. (2025)
✗
ResNet-34
None
96.4
95.4
88.7
99.9
81.7
84.6
-
World4Drive Zheng et al. (2025)
✗
ResNet-34
None
97.4
94.3
92.8
100
79.9
85.1
-
Table 2: Comparison of state-of-the-art methods on the NAVSIM navtest split.
ID
Ego Query
Agent Query
Map Query
Deformable -SCA
Param. (M) ↓
L2 (m) ↓
Collision Rate (%) ↓
Energy (mJ) ↓
1s
2s
3s
1s
2s
3s
1
✗
✗
✓
✗
34.91
0.28
0.57
0.98
0.04
0.19
0.44
69.27
2
✓
✗
✓
✗
38.11
0.27
0.56
0.97
0.02
0.16
0.51
68.71
3
✗
✓
✓
✗
38.38
0.20
0.39
0.70
0.02
0.12
0.35
69.58
4
✓
✓
✗
✗
35.21
0.21
0.38
0.67
0.02
0.11
0.34
70.02
5
✓
✓
✓
✗
38.39
0.20
0.38
0.66
0.02
0.12
0.30
69.64
Table 3: Ablation for design choices.
Table 7
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Module
Specification
Image backbone
Spike-driven transformer (large), 19.0M
embed dims [64,128,256,360] , stride-16 output
Spike-3D-Lift
Radial–Cartesian sampling Zhang et al. [2025]
SDM depth neuron, Vmax=8 , BEV channels 80
BEV encoder
channels [128,256,360,256,256] , output 256
agent-view features 10×10
Appendix
Table B.6: Per-stage architecture specification of SDPAD.
Method
Backbone
Image Size
Frames
mAP ↑
NDS ↑
mATE ↓
mASE ↓
mAOE ↓
mAVE ↓
mAAE ↓
BEVDet (arXiv. 2021)
ResNet50
256 × 704
1
0.298
0.379
0.725
0.279
0.589
0.860
0.245
PETRv2 (ICCV. 2023)
ResNet50
256 × 704
2
0.349
0.456
0.700
0.275
0.580
0.437
0.187
BEVDepth (AAAI. 2023)
ResNet50
256 × 704
2
0.351
0.475
0.639
0.267
0.479
0.428
0.198
BEVStereo (AAAI. 2023)
ResNet50
256 × 704
2
0.372
0.500
0.598
0.270
0.438
0.367
0.190
SA-BEV (ICCV. 2023)
ResNet50
256 × 704
2
0.387
0.512
0.613
0.266
0.352
0.382
0.199
BEFormerv2 (CVPR. 2023)
ResNet50
-
-
0.423
0.529
0.618
0.273
0.413
0.333
0.188
Appendix
Table C.7: Comparison with previous state-of-the-art multi-view 3D detectors on the nuScenes val set.
Figure D.4: Predicted trajectories of SDPAD versus ground truth in diverse nuScenes scenarios, together with the BEV segmentation and agent detection outputs. We also show the intermediate Spiking BEV features.
Figure D.5: Predicted trajectories of SDPAD versus ground truth in diverse NAVSIM.
This letter presents PrismAD, a decoupled end-to-end autonomous driving framework based on a Semantic Mixture-of-Planners. Existing planners usually aggregate heterogeneous scene tokens into a coupled representation space, forcing a single planning branch to jointly model agent interaction, road geometry, and driving intention. Such coupling may weaken factor-specific reasoning and obscure the contribution of different planning cues. To address this limitation, PrismAD partitions scene tokens into interaction, geometry, and intent groups, and assigns them to independent planning experts with the same architecture but separate parameters. Each expert learns a specialized motion-planning representation, while a semantics-aware router adaptively aggregates expert predictions with separate routing weights for motion prediction and ego planning. Sparse top-K activation with noisy gating is further introduced to improve routing robustness and reduce unnecessary expert computation. Extensive experiments on the nuScenes open-loop dataset and NeuroNCAP closed-loop benchmark demonstrate that PrismAD exhibits competitive performance. Our code will be released soon.
Kang Ding, Zhigui Lin, Hongsong Wang +5
School of Vehicle and Mobility, Tsinghua University, Beijing 100084, China. · State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China. · School of Cyberspace Security, Southeast University, Nanjing 210096, China. +3
Autonomous driving perception demands accurate and efficient processing of three-dimensional sensor data under strict power constraints. Traditional convolutional neural networks achieve strong detection accuracy but are computationally intensive, limiting their suitability for deployment on resource-constrained neuromorphic platforms. Spiking neural networks offer a compelling alternative through event-driven sparse computation, yet their application to complex real-world perception tasks such as three-dimensional object detection remains limited. In this work, we propose an end-to-end spiking encoder-decoder network for object detection in bird's eye view representations of LiDAR point clouds, trained using surrogate gradient backpropagation. We train two variants: a membrane potential variant that reads continuous neuron state at the output stage for maximum accuracy, achieving 92.05/87.04/86.51 AP at IoU=0.5 (Easy/Moderate/Hard), and, a fully binary spiking variant that operates exclusively on spike trains at every layer for direct neuromorphic deployment. We evaluate four input spike encoding strategies and demonstrate that allowing the network to learn spike representations directly from data outperforms hand-crafted Poisson, latency, and z-axis encoding schemes on the KITTI benchmark, where sequential frames are unavailable and the BEV input is presented repeatedly across timesteps as a proxy for temporal streaming. A block-wise energy analysis demonstrates a 3.33× reduction in synaptic operation energy over an equivalent CNN under conservative loop-based operation. Together, these results demonstrate the viability of spiking neural networks for accurate and energy-efficient neuromorphic perception in autonomous driving.
Sambit Mohapatra, Senthil Yogamani, Heinrich Gotzig +1
Valeo, Germany · TU Ilmenau, Germany · Valeo, Ireland
End-to-end autonomous driving increasingly leverages self-supervised video pretraining to learn transferable planning representations. However, pretraining video world models for scene understanding has so far brought only limited improvements. This limitation is compounded by the inherent ambiguity of driving: each scene typically provides only a single human trajectory, making it difficult to learn multimodal behaviors. In this work, we propose Drive-JEPA, a framework that integrates Video Joint-Embedding Predictive Architecture (V-JEPA) with multimodal trajectory distillation for end-to-end driving. First, we adapt V-JEPA for end-to-end driving, pretraining a ViT encoder on large-scale driving videos to produce predictive representations aligned with trajectory planning. Second, we introduce a proposal-centric planner that distills diverse simulator-generated trajectories alongside human trajectories, with a momentum-aware selection mechanism to promote stable and safe behavior. When evaluated on NAVSIM, the V-JEPA representation combined with a simple transformer-based decoder outperforms prior methods by 3 PDMS in the perception-free setting. The complete Drive-JEPA framework achieves 93.3 PDMS on v1 and 87.8 EPDMS on v2, setting a new state-of-the-art.
Linhan Wang, Zichong Yang, Chen Bai +6
Virginia Tech · Purdue University · XPENG Motors +1