Refine Connections, Close the Gap: A Reliable Enhancement Framework for Driving Scene Topology
Organizations: Bosch Research North America · Bosch Center for Artificial Intelligence (BCAI) · Massachusetts Institute of Technology · Bosch XC China
Abstract
In autonomous driving, understanding scene topology - the connectivity between lanes and traffic elements - is critical for safe path planning and motion control. While current methods excel at detecting individual map elements, their connectivity reasoning often falls short of its theoretical potential, leaving a significant performance gap relative to the theoretical upper-bound achievable given the underlying detections. Furthermore, the decision-ready topology graphs passed to downstream tasks often remain unreliable. Current approaches typically derive connectivity by thresholding continuous topology scores; however, these scores often fail to reflect the true logical likelihood of connectivity, resulting in false positives or missing connections. Existing benchmarks further overlook this issue by primarily evaluating continuous metrics, rather than assessing the discrete connectivity required for decision-making. To bridge these gaps, we propose TopoEnhance, a novel topology enhancement framework designed to unlock the latent potential of existing methods and improve the reliability of decision-ready topology. We formulate topology enhancement as a denoising-based reconstruction process, where the model learns to recover structural consistency from stochastically corrupted ground-truth graphs. This formulation enables the model to resolve logical inconsistencies and rectify unreliable connections, producing robust discrete topology graphs that closely approach theoretical maximum performance. Extensive experiments across different baselines show that TopoEnhance consistently improves both continuous topology metrics (TOP score), and discrete connectivity measured by our adapted Topology Jaccard Similarity (TJS) metric. As a flexible, source-agnostic framework, TopoEnhance delivers substantial gains across diverse state-of-the-art baselines without requiring retraining.
Figures & tables
| Dataset | Method | Before TopoEnhance | After TopoEnhance | Upper-Bound | |||
| TOP | TOP | TOP | TOP | TOP | TOP | ||
| Argoverse2 (Subset_A) | TopoNet ( Li et al., 2026 ) | 10.9 | 23.8 | 21.8 | 25.8 | 24.4 | 28.8 |
| TopoMLP ( Wu et al., 2024 ) | 21.6 | 26.9 | 24.4 | 28.7 | 25.0 | 31.4 | |
| Topo2D ( Li et al., 2024 ) | 22.3 | 26.2 | 24.3 | 27.3 | 25.0 | 29.8 | |
| TopoLogic ( Fu et al., 2024 ) | 23.9 | 25.4 | 24.5 | 27.2 | 26.0 | 30.6 | |
| SMART ( Ye et al., 2025 ) | 37.0 | 33.0 | 40.1 | 35.4 | 40.9 | 38.9 | |
| Dataset | Method | Before TopoEnhance | After TopoEnhance | Upper-bound | |||
| TJS ll | TJS lt | TJS ll | TJS lt | TJS ll | TJS lt | ||
| Argoverse2 (Subset_A) | TopoNet ( Li et al., 2026 ) | 16.3 | 31.0 | 32.9 | 55.8 | 44.3 | 63.7 |
| TopoMLP ( Wu et al., 2024 ) | 18.5 | 31.5 | 59.7 | 59.8 | 69.7 | 72.8 | |
| Topo2D ( Li et al., 2024 ) | 18.9 | 34.5 | 55.0 | 51.5 | 68.7 | 70.8 | |
| TopoLogic ( Fu et al., 2024 ) | 21.0 | 32.5 | 43.9 | 53.3 | 50.9 | 71.4 | |
| SMART ( Ye et al., 2025 ) | 38.1 | 36.6 | 87.7 | 82.8 | 92.9 | 90.1 | |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | TOP ll | TOP lt | TJS ll | TJS lt |
| TopoNet | 3.34 | 14.40 | 11.84 | 20.49 |
| + TopoEnhance | 10.23 | 16.37 | 12.87 | 35.49 |
| Upper-Bound | 11.47 | 18.42 | 16.90 | 37.37 |
| Dataset | Method | DET l | DET t | TOP ll | TOP lt | TJS ll | TJS lt |
| Argoverse2 (Subset_A) | TopoPoint, reported | 31.40 | 55.30 | 28.70 | 30.00 | – | – |
| TopoPoint, our reproduction | 28.82 | 54.20 | 28.10 | 30.36 | 22.31 | 35.94 | |
| + TopoEnhance | 28.82 | 54.20 | 28.21 | 32.28 | 25.03 | 55.13 | |
| Upper-Bound | 28.82 | 54.20 | 29.65 | 36.58 | 31.66 | 56.67 | |
| nuScenes (Subset_B) | TopoPoint, reported | 31.20 | 60.20 | 28.30 | 27.10 | – | – |
| TopoPoint, our reproduction | 23.67 | 60.66 | 24.73 | 22.42 | 14.09 | 35.13 |
| Dataset | Method | Detection | Before TopoEnhance | After TopoEnhance | Upper-Bound | |||||||
| DET | DET | TOP | TOP | OLS | TOP | TOP | OLS | TOP | TOP | OLS | ||
| Argoverse2 (Subset_A) | TopoNet ( Li et al., 2026 ) | 28.6 | 48.6 | 10.9 | 23.8 | 39.8 | 21.8 | 25.8 | 43.7 | 24.4 | 28.8 | 45.1 |
| TopoMLP ( Wu et al., 2024 ) | 28.5 | 49.5 | 21.6 | 26.9 | 44.1 | 24.4 | 28.7 | 45.2 | 25.0 | 31.4 | 46.0 | |
| Topo2D ( Li et al., 2024 ) | 29.1 | 50.6 | 22.3 | 26.2 | 44.5 | 24.3 | 27.3 | 45.3 | 25.0 | 29.8 | 46.1 | |
| TopoLogic ( Fu et al., 2024 ) | 29.9 | 47.2 | 23.9 | 25.4 | 44.1 | 24.5 | 27.2 | 44.7 | 26.0 | 30.6 | 45.9 | |
| SMART ( Ye et al., 2025 ) | 46.6 | 47.7 | 37.0 | 33.0 | 53.1 | 40.1 | 35.4 | 54.3 | 40.9 | 38.9 | 55.2 | |
| Backbone Model | Dim | ||
| DINOv2-ViT-S | 384 | 20.3 | 24.64 |
| DINOv2-ViT-B | 768 | 21.8 | 25.06 |
| DINOv2-ViT-L | 1024 | 21.84 | 25.8 |
| DINOv2-ViT-G | 1536 | 21.84 | 25.75 |
| DINOv3-ViT-S | 384 | 21.56 | 24.48 |
| DINOv3-ViT-B | 768 | 21.83 | 24.35 |
| Loss TOP ll TOP lt BCE 12.1 22.0 BCE + Focal 22.2 22.0 Normalized BCE (Ours) 21.8 25.8 |
| Strategy TOP ll TOP lt TopoEnhance 21.8 25.8 TopoEnhance without Jitter 21.7 25.4 |
| Model | Noise | TOP ll | TOP lt | ||
| Before TopoEnhance | After TopoEnhance | Before TopoEnhance | After TopoEnhance | ||
| TopoNet | 10.90 | 21.80 | 23.80 | 25.80 | |
| 4.62 | 9.99 | 15.81 | 17.11 | ||
| 2.78 | 6.14 | 11.31 | 12.35 | ||
| 1.99 | 4.39 | 8.48 | 9.39 | ||
| 1.30 | 2.77 | 4.94 | 5.64 | ||
| Model | Before TopoEnhance | After TopoEnhance |
| TopoNet | 24.48% | 24.21% |
| TopoMLP | 12.62% | 7.77% |
| TopoLogic | 10.68% | 8.27% |
| SMART | 10.52% | 6.38% |
| Model | Before TopoEnhance | After TopoEnhance |
| TopoNet | 0.166 | 0.118 |
| TopoMLP | 0.111 | 0.080 |
| TopoLogic | 0.285 | 0.228 |
| SMART | 0.104 | 0.075 |
| Model | Relation | Error Type | Total | Corrected | Broken |
| TopoNet | Lane–Lane | Missed ego successor | 243 | 60 | 0 |
| Missed intersection turn | 5,378 | 1,502 | 0 | ||
| Lane–Traffic | Missed red light | 153 | 103 | 0 | |
| Missed green light | 398 | 311 | 0 | ||
| Missed yellow light | 54 | 23 | 0 | ||
| Missed turn signal | 49 | 36 | 0 |
| Model | Before TopoEnhance | After TopoEnhance |
| TopoNet | 39.19% | 68.92% |
| TopoMLP | 94.81% | 100% |
| TopoLogic | 92.41% | 93.67% |
| SMART | 91.36% | 98.77% |
| Component | Latency (ms) | Percentage |
| Feature Extraction | 0.21 | 3.2% |
| GNN Message Passing | 5.21 | 80.9% |
| Score Prediction | 0.64 | 9.9% |
| Total | 6.43 | 100% |
| Component | Memory Overhead |
| Feature Extraction | +0.1 MB |
| GNN Message Passing | +0.0 MB |
| Score Prediction | +0.2 MB |
| Total | +0.3 MB |