Organizations: Shenzhen Research Institute of Big Data, Shenzhen, China · Center for Cloud Computing, Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences, Shenzhen, China · School of Electrical Engineering and Telecommunications, The University of New South Wales, Sydney, NSW 2052, Australia · State Key Laboratory of Internet of Things for Smart City (SKL-IOTSC), University of Macau, Macau, China · Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong
Large vision-language models (VLMs) provide powerful open-world perception and reasoning for autonomous driving, but their high computational cost and inference latency make continuous cloud-side use impractical. This motivates fast--slow collaboration, where efficient onboard modules handle real-time perception and control while cloud models provide high-level reasoning only when needed. The key challenge is deciding when cloud reasoning should influence time-critical driving decisions. Existing methods often rely on perception uncertainty, heuristic triggers, or resource-driven policies, without assessing whether resolving an uncertainty will improve planning. We propose \textbf{SIGMA}, a simulation-in-the-loop framework for task-oriented fast--slow collaboration. SIGMA embeds the planner into uncertainty assessment and evaluates how plausible scene realizations under semantic and geometric uncertainty affect feasible trajectories and planning cost. Based on these outcomes, it estimates the expected reduction in planning cost from resolving uncertainty. We further introduce expected planning gain (EPG), a decision-level metric for cloud invocation, cloud-guidance integration, and request prioritization under deadline and resource constraints. Experiments in CARLA show that SIGMA reduces unnecessary cloud interactions while improving planning, efficiency, and navigation success in static and dynamic obstacle scenarios. Compared with fixed-period collaboration, SIGMA reduces unnecessary cloud interactions by 50%, improves navigation success by more than 6%, and cuts finish time by up to 26.2% in dynamic scenarios.
Figures & tables
Fig. 1: Overview of the proposed planning-gain-guided fast–slow collaborative driving framework.
Work / Ref.
Uncertainty Modeling
Capability / constraints
Planning-level decision awareness
Semantic uncertain
Geometric uncertain
Plan impact
Real time
Open-world capacity
Resource aware
Planner integration
Collaboration criterion
Planning-gain modeling
Xiao et al. (TMC’24) [ 13 ]
–
–
–
High
Low
✓
–
Resource
–
Ye et al. (TMC’24) [ 14 ]
✓
×
×
High
Low
✓
–
Accuracy/resource
–
Fang et al. (TMC’24) [ 15 ]
–
–
–
High
Low
✓
–
Perception priority
–
Lin et al. (TMC’25) [ 16 ]
✓
×
×
Medium
Low
✓
–
Security
–
Liang et al. (TMC’26) [ 17 ]
–
–
–
High
Medium
✓
–
Blind spot
–
TABLE I: Feature-level comparison with representative related works.
Fig. 2: Edge-side uncertainty-aware perception and onboard planning.
G(cth)=recall of OOD detections∑j=1NvalI(y^j∈/Ytrain)∑j=1NvalI(y^j∈/Ytrain)⋅I(cj<cth)+recall of in-distribution detections∑j=1NvalI(y^j∈Ytrain)∑j=1NvalI(y^j∈Ytrain)⋅I(cj>cth).
(7)
Table 4
Fig. 3: Cloud-side slow-reasoning for planning guidance.
Fig. 4: Cloud-assisted forward simulation for decision evaluation.
Fig. 5: Cloud-based LVM perception pipeline with SAM-assisted.
Fig. 6: Cloud-based large model decision-making pipeline.
Fig. 7: Comparison of local and cloud perception results.
Split
# Samples
Known
Unknown
Total
Local_Train
9,647
38
0
38
Local_Val
2,060
21
5
26
Local_Test
3,400
19
6
25
Cloud_Train
15,107
49
0
49
TABLE II: Dataset splits for known- and unknown-object evaluation.
Fig. 8: ODCT threshold selection on the validation and test sets.
Fig. 9: Qualitative examples of 3D object detection with predicted geometric uncertainty.
Fig. 10: Predicted geometric uncertainty versus localization error on KITTI.
Fig. 11: Relationship with predicted uncertainty, and localization error.
Metric
Easy
Moderate
Hard
3D AP@0.70
89.48
79.20
78.66
BEV AP@0.70
90.33
88.15
87.71
BBox AP@0.70
98.61
89.60
89.24
3D AP_R40@0.70
92.83
83.57
81.14
BEV AP_R40@0.70
96.09
89.65
87.24
BBox AP_R40@0.70
99.27
95.33
92.95
TABLE III: KITTI detection accuracy of the uncertainty-aware detector.
Fig. 12: Static-obstacle scenario configuration for Experiment 2.
Fig. 13: Static-scenario experimental results: trajectories, control profiles, and collaboration states.
Fig. 14: Trajectory and control profiles of PCS-2.
Method
FTime
TLen
AvgLD
SVar
MLat
Unit
(s)
(m)
(m)
(m/s)
(m)
LOS
25.94
124.54
0.53
1.09
1.52
PCS
24.09
122.40
0.59
1.13
1.89
PCS-2
23.53
123.84
0.57
1.39
1.69
SIGMA
22.35
121.85
0.60
1.30
1.76
TABLE IV: Quantitative comparison in Experiment 2.
Fig. 15: Simulation-in-the-loop MPC rollouts under different plausible scene samples.
U
Method
FTime
TLen
AvgLD
SVar
MLat
(Unit)
(s)
(m)
(m)
( m/s )
(m)
0
LOS
42.47
114.27
0.52
2.34
1.59
0
PCS
42.31
114.63
0.53
2.64
1.41
0
SIGMA
43.00
113.56
0.51
2.41
1.62
1
LOS
48.05
119.95
0.49
2.57
1.54
1
PCS
44.02
114.44
0.53
2.68
1.43
TABLE V: Quantitative comparison of Experiment 3.
Fig. 16: Comparison of success rate and finish time in Experiment 3.
Fig. 17: Dynamic-obstacle scenario configuration for Experiment 4.
Fig. 18: Trajectory and control profiles in Experiment 4.
Method
FTime(s)
TLen(m)
AvgLD(m)
SVar(m/s)
MLat(m)
PCS
23.40
71.47
0.47
2.99
1.68
SIGMA
17.28
72.90
0.62
2.29
1.44
TABLE VI: Quantitative comparison in Experiment 4.
Large language models (LLMs) can improve autonomous driving planning but are costly to query online, and existing fast-slow planners often rely on hand-designed triggering rules that either over-call the slow system or call it at the wrong times. We formulate slow-system invocation as a resource-aware sequential decision problem and propose the Adaptive Slow-System Control Gate (ASSCG), which makes frame-level Query/Cache/Drop decisions to refresh, reuse, or suppress slow guidance. ASSCG uses an RWKV backbone for efficient long-horizon gating and is trained with supervised fine-tuning followed by GRPO-style compute-aware reinforcement fine-tuning. We apply ASSCG to two different fast-slow architectures: (i) AsyncDriver on nuPlan Hard20 closed-loop evaluation, where ASSCG improves score to 67.28 (+2.28) while reducing average end-to-end inference latency by 60%; and (ii) a RecogDrive-based dual system that we build by replacing its original VLM-2B module with a lightweight ViT-based fast planner and adding an LLM slow planner, evaluated on NAVSIM, where ASSCG achieves 91.4 PDMS (+0.6) and increases average speed by 25%. The project page, including video visualizations and additional results, is available at https://williamxuanyu.github.io/asscg/.
Sining Ang, Yuan Chen, Liu Haiyan +5
Institute for AI Industry Research (AIR), Tsinghua University · Department of Automation, University of Science and Technology of China · Beijing University of Aeronautics and Astronautics +1
Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate simulation ticks and replaying the previous command in between, so half of all control outputs ignore the newest observations. We present a fast-slow architecture that removes this compromise. A frozen 7B vision-language backbone acts as the slow system, digesting navigation instructions and visual history at low frequency while exposing its per-layer key-value cache as a standing representation of the scene. A lightweight action expert acts as the fast system, attending to this cache and to the current camera frame at every simulation tick to regress waypoints in a single forward pass. Since the cache lags behind the world at deployment, we train the expert under randomized staleness, aligning training with asynchronous execution. On LangAuto-Short routes in CARLA, our system produces fresh control at every 50 ms simulation tick and lifts route completion from 37.0 to 94.0 over the frame-skipping baseline. A frame-skip ablation with the same expert separates the two factors at work: the expert raises the driving score on its own, while per-tick freshness raises completion from 82.1 to 94.0 and cuts red-light violations by a third. Trained on a single town, the expert transfers zero-shot to two unseen towns, holding 84-94% route completion where the baseline reaches 31-41%. It reduces open-loop waypoint error by nearly a factor of four compared to the backbone's own action head, at a per-tick model cost of 32 ms that is independent of history length on a single consumer GPU.
Yun Li, Jiachen Gong, Simon Thompson +7
The University of Tokyo, Tokyo, Japan · TIER IV, Inc., Tokyo, Japan · Huawei
Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic traffic. Existing methods either perform online language reasoning without explicit dynamics verification or use world models mainly in offline pipelines, leaving a gap between semantic intent and physical feasibility at decision time. We propose Reason--Imagine--Act (RIA), a closed-loop framework that couples an LLM reasoner with an action-conditioned world model for online safety verification. At each step, the LLM proposes an action template and candidate sub-actions, the world model performs short-horizon rollouts, and a safety scorer selects the safest executable action with feedback to the next reasoning step. Under a unified CARLA point-goal protocol (1000 episodes), RIA achieves 80.05% route completion, 51.10% arrival rate, and 0.20% collision rate. Under the same closed-loop interface, RIA consistently outperforms training-free baselines, including CARLA TM and MADA, on core closed-loop metrics. For reproducibility, code is available at https://github.com/pku-smart-city/source_code/tree/main/RIA.
Zhengqi Sun, Yiwen Sun, Boxuan Liu +3
Department of Information Management, Peking University, Beijing 100871, China · School of Intelligence Science and Technology, Peking University, Beijing 100871, China · State Key Laboratory of General Artificial Intelligence, BIGAI, Beijing 100080, China +3