Organizations: Shenzhen Research Institute of Big Data, Shenzhen, China · Center for Cloud Computing, Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences, Shenzhen, China · School of Electrical Engineering and Telecommunications, The University of New South Wales, Sydney, NSW 2052, Australia · State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), University of Macau, Macau, China · Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong
Large vision-language models (VLMs) provide powerful open-world perception and reasoning for autonomous driving, but their high computational cost and inference latency make continuous cloud-side use impractical. This motivates fast--slow collaboration, where efficient onboard modules handle real-time perception and control while cloud models provide high-level reasoning only when needed. The key challenge is deciding when cloud reasoning should influence time-critical driving decisions. Existing methods often rely on perception uncertainty, heuristic triggers, or resource-driven policies, without assessing whether resolving an uncertainty will improve planning. We propose \textbf{SIGMA}, a simulation-in-the-loop framework for task-oriented fast--slow collaboration. SIGMA embeds the planner into uncertainty assessment and evaluates how plausible scene realizations under semantic and geometric uncertainty affect feasible trajectories and planning cost. Based on these outcomes, it estimates the expected reduction in planning cost from resolving uncertainty. We further introduce expected planning gain (EPG), a decision-level metric for cloud invocation, cloud-guidance integration, and request prioritization under deadline and resource constraints. Experiments in CARLA show that SIGMA reduces unnecessary cloud interactions while improving planning, efficiency, and navigation success in static and dynamic obstacle scenarios. Compared with fixed-period collaboration, SIGMA reduces unnecessary cloud interactions by 50%, improves navigation success by more than 6%, and cuts finish time by up to 26.2% in dynamic scenarios.
Figures & tables
Fig. 1: Overview of the proposed planning-gain-guided fast–slow collaborative driving framework.
Work / Ref.
Uncertainty Modeling
Capability / constraints
Planning-level decision awareness
Semantic uncertain
Geometric uncertain
Plan impact
Real time
Open-world capacity
Resource aware
Planner integration
Collaboration criterion
Planning-gain modeling
Xiao et al. (TMC’24) [ 13 ]
–
–
–
High
Low
✓
–
Resource
–
Ye et al. (TMC’24) [ 14 ]
✓
×
×
High
Low
✓
–
Accuracy/resource
–
Fang et al. (TMC’24) [ 15 ]
–
–
–
High
Low
✓
–
Perception priority
–
Lin et al. (TMC’25) [ 16 ]
✓
×
×
Medium
Low
✓
–
Security
–
Liang et al. (TMC’26) [ 17 ]
–
–
–
High
Medium
✓
–
Blind spot
–
TABLE I: Feature-level comparison with representative related works.
Fig. 2: Edge-side uncertainty-aware perception and onboard planning.
G(cth)=recall of OOD detections∑j=1NvalI(y^j∈/Ytrain)∑j=1NvalI(y^j∈/Ytrain)⋅I(cj<cth)+recall of in-distribution detections∑j=1NvalI(y^j∈Ytrain)∑j=1NvalI(y^j∈Ytrain)⋅I(cj>cth).
(7)
Table 4
Fig. 3: Cloud-side slow-reasoning for planning guidance.
Fig. 4: Cloud-assisted forward simulation for decision evaluation.
Fig. 5: Cloud-based LVM perception pipeline with SAM-assisted.
Fig. 6: Cloud-based large model decision-making pipeline.
Fig. 7: Comparison of local and cloud perception results.
Split
# Samples
Known
Unknown
Total
Local_Train
9,647
38
0
38
Local_Val
2,060
21
5
26
Local_Test
3,400
19
6
25
Cloud_Train
15,107
49
0
49
TABLE II: Dataset splits for known- and unknown-object evaluation.
Fig. 8: ODCT threshold selection on the validation and test sets.
Fig. 9: Qualitative examples of 3D object detection with predicted geometric uncertainty.
Fig. 10: Predicted geometric uncertainty versus localization error on KITTI.
Fig. 11: Relationship with predicted uncertainty, and localization error.
Metric
Easy
Moderate
Hard
3D AP@0.70
89.48
79.20
78.66
BEV AP@0.70
90.33
88.15
87.71
BBox AP@0.70
98.61
89.60
89.24
3D AP_R40@0.70
92.83
83.57
81.14
BEV AP_R40@0.70
96.09
89.65
87.24
BBox AP_R40@0.70
99.27
95.33
92.95
TABLE III: KITTI detection accuracy of the uncertainty-aware detector.
Fig. 12: Static-obstacle scenario configuration for Experiment 2.
Fig. 13: Static-scenario experimental results: trajectories, control profiles, and collaboration states.
Fig. 14: Trajectory and control profiles of PCS-2.
Method
FTime
TLen
AvgLD
SVar
MLat
Unit
(s)
(m)
(m)
(m/s)
(m)
LOS
25.94
124.54
0.53
1.09
1.52
PCS
24.09
122.40
0.59
1.13
1.89
PCS-2
23.53
123.84
0.57
1.39
1.69
SIGMA
22.35
121.85
0.60
1.30
1.76
TABLE IV: Quantitative comparison in Experiment 2.
Fig. 15: Simulation-in-the-loop MPC rollouts under different plausible scene samples.
U
Method
FTime
TLen
AvgLD
SVar
MLat
(Unit)
(s)
(m)
(m)
( m/s )
(m)
0
LOS
42.47
114.27
0.52
2.34
1.59
0
PCS
42.31
114.63
0.53
2.64
1.41
0
SIGMA
43.00
113.56
0.51
2.41
1.62
1
LOS
48.05
119.95
0.49
2.57
1.54
1
PCS
44.02
114.44
0.53
2.68
1.43
TABLE V: Quantitative comparison of Experiment 3.
Fig. 16: Comparison of success rate and finish time in Experiment 3.
Fig. 17: Dynamic-obstacle scenario configuration for Experiment 4.
Fig. 18: Trajectory and control profiles in Experiment 4.
Method
FTime(s)
TLen(m)
AvgLD(m)
SVar(m/s)
MLat(m)
PCS
23.40
71.47
0.47
2.99
1.68
SIGMA
17.28
72.90
0.62
2.29
1.44
TABLE VI: Quantitative comparison in Experiment 4.
Institute for AI Industry Research (AIR), Tsinghua University · Department of Automation, University of Science and Technology of China · Beijing University of Aeronautics and Astronautics +1
Department of Information Management, Peking University, Beijing 100871, China · School of Intelligence Science and Technology, Peking University, Beijing 100871, China · State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing 100080, China +3