Online vectorized HD map construction is essential for scaling safe autonomous driving and requires accurate, real-time inference. Prior methods typically rely on dense bird's-eye-view (BEV) grids as the intermediate representation. We propose \textit{MapLightning}, which replaces the dense BEV grid with a compact set of 1D learnable map tokens. To construct map tokens from image features, we choose self-attention over vanilla cross-attention because it enables joint interactions and contextual aggregation among image and map tokens. Our transformer-based mapper concatenates map and image tokens, applies full self-attention, discards the image tokens, and retains the updated map tokens for decoding. This design offers three advantages. First, our representation is efficient, using fewer tokens, consuming less memory, and running faster. Second, the lightweight design allows the map decoder to use full rather than deformable cross-attention for better global context. Third, unlike BEV-based methods, our network does not use camera projection parameters, making it robust to camera-extrinsic perturbations. MapLightning uses up to 16.7× fewer intermediate tokens than dense BEV-based methods and achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse2. Its lightweight variant surpasses MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse2, while delivering 1.73× faster inference (40+ FPS) with 53% less memory. We further show improvements on uncertainty-aware map construction and downstream trajectory prediction. Code and models will be released.
Figures & tables
Figure 1: Aggregating multi-view image tokens into a scene representation. (a) BEV-based methods ( Liao et al., 2022 ; Liao et al., 2025 ) geometrically project image features onto a dense grid of fixed BEV cells, few of which contain map elements. (b) Vanilla cross-attention compresses image tokens into compact map tokens, but image tokens serve only as fixed keys and values, with no image-to-image or map-to-map interaction. (c) Self-attention (ours, Section 3.3 ) attends over the concatenated map and image tokens, refining all tokens jointly. Gray and violet edges denote image–map and intra-group attention, respectively. In the masks, rows are queries, columns are keys/values, shaded cells mark allowed attention, and M / I denote map/image tokens. See Table 1 for comparison.
Figure 2: Overview of MapLightning . A shared image backbone encodes each of the C camera views over T timesteps into image tokens, which are flattened and augmented with learned positional, camera, and time embeddings. Instead of projecting these tokens into a dense BEV grid, we concatenate them with K learnable map tokens and jointly refine both with a self-attention mapper (Figure 1 (c)), without camera parameters or depth. The image tokens are then discarded, and the compact map tokens serve as the scene representation, which task-specific decoders attend to with full cross-attention. Beyond online vectorized HD map construction, this representation supports 3D object detection and downstream trajectory prediction from the predicted maps.
Aggregation
#Tokens (k) ↓
FPS ↑
nuSc. mAP ↑
Argo2 mAP ↑
ResNet-50 backbone
BEV Projection
20.0
14.1
26.7
53.1
Ours (cross-attn)
0.3
19.1
25.5
52.6
Ours (self-attn)
0.3
19.2
30.1
58.3
MobileNetV3-Large backbone
BEV Projection
20.0
23.3
20.5
43.7
Table 1: Comparison of aggregation mechanisms under a common decoder framework in the single-frame setting ( T=1 ). Self-attention achieves the best accuracy with 66.7 × fewer intermediate tokens than dense BEV projection ( Liao et al., 2025 ) , and matches cross-attention in token count and FPS, demonstrating effective token compression and interaction. Token counts (0.3k) are for nuScenes; Argoverse 2 uses 0.35k tokens for 7 cameras (same per-camera budget).
Figure 3: Image-to-image attention in the self-attention mapper. We select one image token as the query (magenta cross) and visualize its attention scores over the other image tokens, which serve as keys (red: high, blue: low), at the first and last ( L=4 ) mapper layers. From left to right: (1) a query on a lane divider attends along that divider and to the parallel divider across the lane; (2) a query on a lane divider attends along the divider’s full extent toward the horizon; (3) a query on a pedestrian crossing attends across the crossing’s full width; and (4) a query on a road boundary attends along the curb. From Layer 1 to Layer 4, responses on these structures become stronger and more complete, while diffuse background responses (e.g., the sky in column 4) fade, illustrating progressive refinement of the image tokens. Such image-to-image interaction is absent in cross-attention, where image tokens serve only as fixed keys and values. Best viewed zoomed in.
Figure 4: Qualitative comparisons of online vectorized HD map construction for Dense BEV Projection ( Liao et al., 2025 ) , ours (cross-attention), and ours (self-attention). Self-attention produces higher-quality and more consistent vectorized map reconstructions, with fewer spurious elements, more complete map geometries, and better-preserved intersection geometry. The first three rows are from nuScenes, while the last row is from Argoverse 2. Orange, blue, and green denote lane dividers, pedestrian crossings, and road boundaries, respectively. Best viewed zoomed in.
Figure 5: Qualitative comparison of uncertainty-aware online vectorized HD map construction for Dense BEV Projection ( Liao et al., 2025 ) , ours (cross-attention), and ours (self-attention) on nuScenes. Uncertainty is visualized at each predicted vector point as a semi-transparent ellipse whose size reflects the predicted Laplace variance. Self-attention produces more localized predictions with substantially less diffuse uncertainty, providing cleaner uncertainty-aware maps for downstream trajectory prediction. Orange, blue, and green denote lane dividers, pedestrian crossings, and road boundaries, respectively, while the red box marks the ego vehicle.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Qualitative comparisons of online vectorized HD map construction for Dense BEV Projection ( Liao et al., 2025 ) , ours (cross-attention), and ours (self-attention). Self-attention produces higher-quality and more consistent vectorized map reconstructions, with fewer spurious elements, more complete map geometries, and better-preserved intersection geometry. The top two examples are from nuScenes, while the bottom two are from Argoverse 2. Orange, blue, and green denote lane dividers, pedestrian crossings, and road boundaries, respectively. Best viewed zoomed in.
Method
NDS ↑
mAP ↑
mATE ↓
mASE ↓
mAOE ↓
mAVE ↓
mAAE ↓
FPS ↑
BEVDepth ( 2023b )
0.475
0.351
0.639
0.267
0.479
0.428
0.198
15.7
Sparse4Dv2 ( 2023 )
0.539
0.439
0.598
0.270
0.475
0.282
0.179
20.3
StreamPETR ( 2023 )
0.540
0.432
0.581
0.272
0.413
0.295
0.195
26.7
SparseBEV ( 2023a )
0.545
0.432
0.606
0.274
0.387
0.251
0.186
23.5
BEVFormer v2 ( 2023 )
0.529
0.423
0.618
0.273
0.413
0.333
0.181
8.1
VideoBEV ( 2024 )
0.535
0.422
0.564
0.276
0.440
0.286
0.198
–
Appendix
Table 8: Quantitative comparisons of 3D object detection on the nuScenes original split with a ResNet-50 backbone. Ours with self-attention surpasses both the cross-attention variant and state-of-the-art methods in accuracy, and achieves the highest inference speed (FPS) among methods that report it. Following common 3D object detection protocol, NDS and mAP are reported as fractions in [0,1] , whereas improvements in the text are in percentage points. All listed results are without perspective-view pre-training on nuImages. All methods use 256 × 704 inputs, except BEVFormer v2, which uses 640 × 1600 inputs and future frames as in its original paper. FPS of BEVFormer v2 and ours are measured on an NVIDIA RTX 3090; other FPS values are as reported in the original papers on RTX 3090. “–” indicates FPS not reported for this setting.
Figure 7: Visualization of train–val geographic overlap across nuScenes regions. The Near-Extrapolation split ( Lilja et al., 2024 ) , used in our online vectorized HD map construction experiments as the default setting, shows almost no train–val overlap. In contrast, the original split exhibits substantial train–val geographic overlap (79.4%), and the StreamMapNet split retains residual overlap (2.1%), indicating data leakage that undermines their reliability for evaluating online vectorized HD map construction. Red dashed boxes highlight regions of train–val geographic overlap within 5 m. Best viewed zoomed in.
Figure 8: Robustness to camera extrinsic perturbations. We independently add zero-mean Gaussian noise to the rotation (θx,θy,θz) and translation (Δx,Δy,Δz) components of all camera extrinsics. Our model does not use camera extrinsics and therefore shows no degradation, whereas Dense BEV Projection ( Liao et al., 2025 ) explicitly relies on them and suffers substantial mAP drops under increasing perturbations.
Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local vectorized maps from onboard sensor observations. Existing methods commonly adopt bird's-eye-view (BEV) features as the intermediate scene representation, encoding the surrounding space with fixed-resolution dense grids. However, map elements are spatially sparse yet require fine-grained geometric localization, making uniformly allocated BEV representations redundant and less effective for vectorized map prediction. In this work, we propose GaussianMap, an online HD map construction framework that learns an adaptive Gaussian representation of the surrounding scene. This representation consists of a set of Gaussian primitives on the BEV plane, each encoding a flexible local region with geometric properties and a feature vector, allowing the model to allocate representational capacity to map-relevant regions. To generate such a representation from sensor observations, we introduce a feed-forward Gaussian encoder that progressively refines these primitives through Gaussian interaction modeling and multi-sensor feature aggregation. The refined Gaussian representation is then splatted into a BEV feature map and decoded into vectorized map predictions. Extensive experiments on nuScenes and Argoverse 2 datasets demonstrate that GaussianMap achieves state-of-the-art performance in both camera-only and camera-LiDAR fusion settings. Our code will be made publicly available.
Hongyu Lyu, Julie Stephany Berrio Perez, Mao Shan +1
The University of Sydney, Australian Centre for Robotics, Sydney, Australia. · The University of Queensland, Australia.
Accurate High-Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade-off: vectorization-based approaches preserve topology but struggle with geometric fidelity, while rasterization-based approaches enable precise geometric supervision but produce unstructured outputs. To bridge this gap, we propose GSMap, a novel framework that unifies both paradigms via a learnable 2D Gaussian representation. Each map element is modeled as an ordered sequence of 2D Gaussians, whose centers correspond to the vertices of the vectorized polyline/polygon. This formulation enables simultaneous optimization through: (1) Differentiable rasterization that enforces pixel-level geometric constraints, and (2) Topology-aware vectorization that maintains structural regularity. Experiments on both nuScenes and Argoverse2 demonstrate that our Gaussian-based representation effectively unifies geometric and topological learning, achieving significant performance improvements and demonstrating strong compatibility with existing HD mapping architectures. Code will be available at https://github.com/peakpang/GSMap
Zhenxuan Zeng, Lingxuan Wang, Sheng Yang +4
School of Computer Science, Northwestern Polytechnical University, China · Unmanned Vehicle Dept, Cainiao Inc., Alibaba Group, China
Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while memory-based methods extend observation time by aggregating ego-trajectory frames, they lack the spatial diversity needed to reveal occluded regions. Incorporating views from surrounding vehicles offers complementary perspectives, yet naive fusion introduces three key challenges: computational cost from large candidate pools, redundancy from near-collinear viewpoints, and noise from pose errors and occlusion artifacts. We present OptiMVMap, which reformulates multi-vehicle mapping as a select-then-fuse problem to address these challenges systematically. An Optimal Vehicle Selection (OVS) module strategically identifies a compact subset of helpers that maximally reduce ego-centric uncertainty in occluded regions, addressing computation and redundancy challenges. Cross-Vehicle Attention (CVA) and Semantic-aware Noise Filter (SNF) then perform pose-tolerant alignment and artifact suppression before BEV-level fusion, addressing the noise challenge. This targeted pipeline yields more complete and topologically faithful maps with substantially fewer views than indiscriminate aggregation. On nuScenes and Argoverse2, OptiMVMap improves MapTRv2 by +10.5 mAP and +9.3 mAP, respectively, and surpasses memory-augmented baselines MVMap and HRMapNet by +6.2 mAP and +3.8 mAP on nuScenes. These results demonstrate that uncertainty-guided selection of helper vehicles is essential for efficient and accurate multi-vehicle vectorized mapping. The code is released at https://github.com/DanZeDong/OptiMVMap.
Zedong Dan, Zijie Wang, Wei Zhang +6
Sun Yat-sen University · Zhongguancun Academy · Shenzhen Loop Area Institute +2