TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks
Authors: Yiming Gao, Shaocheng Luo
Organizations: Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI, USA. · Department of Electrical and Computer Engineering, Duke University, Durham, NC, USA.
Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's native point clouds or internal processing. Local perception provides metric vehicle states, while the MLLM combines these measurements with roadside imagery to infer relational traffic context and construct a semantic scene graph. Based on this representation, TACTIC selects and configures two complementary primitives: \emph{push-away}, which shifts the perceived range of a lead vehicle, and \emph{phantom-obstacle braking}, which triggers emergency braking through obstacle injection. Measured traffic states and empirically calibrated constraints ground the generated tactics in physically feasible operating regions. To accommodate MLLM latency, TACTIC overlaps reasoning and execution asynchronously while high-rate local perception detects scene changes and triggers replanning. Across 280 randomized CARLA trials, the full policy achieves a 100% collision rate, compared with 35% for a fixed rule, 60% for random selection, and 75% for a restricted LLM using mode selection with default parameters. Joint physical-and-image input achieves 100% success, versus 65% with physical measurements alone and 75% with imagery alone, while asynchronous Δ refresh reduces scene-mutation response from 7.4 s to 2.0 s. These results show that scene-dependent tactical planning can expose context-sensitive LiDAR failure modes that fixed attack policies may miss.
Figures & tables
Fig. 1 : TACTIC overview. The attacker-operated roadside perception stack observes local traffic with co-located LiDAR and camera and provides sensing state S to the multimodal LLM agent. The agent constructs a semantic scene graph, selects a feasible tactic P , and updates the policy when significant scene changes occur. Cached topology and asynchronous reasoning reduce repeated scene generation and replanning latency.
Fig. 2 : Multimodal scene-graph generation in TACTIC . Roadside imagery is encoded into visual tokens, while tracked vehicle states are serialized as text tokens. The MLLM jointly conditions on both modalities to infer semantic traffic relations and decodes them as a structured scene graph. When the topology remains unchanged, the steady-state manager updates only metric edge attributes without full multimodal regeneration.
Fig. 3 : Example phantom-obstacle braking attack. At T1 the system monitors a nominal scene; execution begins at T2=8.0 s and collision occurs at T3=10.4 s. Windows A–C show the driving scene, target perception, and scene graph.
Fig. 4 : Example scene change and replanning. Execution is reconsidered at T2=3.9 s and T4=13.3 s as traffic evolves, ultimately causing a collision between T and lead vehicle A1 .
Fig. 5 : Measured push-away feasibility bounds. (a) Displacement range and policy operating points. (b) Ramp rates bounded by completion and the ρmax=3.0 m/s consistency limit.
Δd /m
Succ.
d<7
R<1
t90 /s
Δd /m
Succ.
d<7
R<1
t90 /s
5
0/20
0/20
0/20
5.83
15
20/20
20/20
18/20
4.87
10
5/20
20/20
9/20
4.30
20
20/20
20/20
6/20
6.20
TABLE I : Locked-displacement response of the push-away primitive (20 trials per level; ramp 3.0 m/s).
Fig. 6 : Dose–response of the two primitives (20 trials per level; 95% Wilson CI). (a) push-away success versus relay displacement. (b) phantom-braking success versus emission duration at a fixed 13–15 m gap.
Decision source
Attack decision
Succ. [95% CI]
t90 /s
G
η
LLM calls
Duty
Dose
Fisher p
LLM policy
LLM-selected mode and (Δd,ρ,o,w,τ)
20/20 [84,100]
3.6
0.61
0.26
2.1
0.39
3.9
(ref)
LLM restricted
LLM-selected mode, default parameters
15/20 [53,89]
3.6
0.35
0.08
2.6
0.65
9.7
4.7×10−2
Fixed rule
fixed push-away, per-trial random dose
7/20 [18,57]
4.3
0.33
0.07
–
0.67
4.7
1.3×10−5
Random
committed draw over mode and parameters
12/20 [39,78]
4.2
0.39
0.15
–
0.62
4.0
3.3×10−3
TABLE II : Decision-source comparison under the hybrid implementation (20 trials per group).
Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and potentially adversarial sensor anomalies. This gap is especially critical for LiDAR, where external actors can physically manipulate the sensing process to induce black-box perception failures without accessing the model. Existing LiDAR benchmarks provide little visibility into this failure mode. Prior adversarial LiDAR studies have largely centered on attack hardware, geometric and algorithmic defenses, and early-generation detectors, leaving the robustness of modern perception systems unexplored. To address this evaluation gap, we introduce ATLAS (Adversarial Temporal LiDAR Attack Suite), the first large-scale, physically grounded evaluation benchmark for LiDAR perception models under black-box sensor attacks, simulating the two primary attack modes -- point injection and point removal -- across real driving sequences. Evaluating a broad cross-section of current state-of-the-art LiDAR perception models, ATLAS reveals a surprising robustness asymmetry: models with stronger performance on standard benchmarks tend to better withstand removal attacks, yet are actually more vulnerable to injection attacks than weaker models. We trace this vulnerability to standard object database sampling augmentations, revealing how current training practices can induce architecture-agnostic robustness failures, and study initial directions for mitigating both attack modes. We release the ATLAS generation code to support extensible, reproducible evaluations as attack capabilities evolve, helping make black-box sensor robustness an explicit consideration in future LiDAR perception development.
The structural vulnerabilities of point cloud-based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primarily on isolated 3D object models, while recent LiDAR spoofing attacks target richer and more realistic driving scenes but focus mainly on physical realizability rather than understanding detector behavior or attack efficiency. In this work, we investigate how LiDAR-based detectors rely on spatial evidence in complex scenes and whether these reliance patterns can be exploited to induce failures more efficiently. To this end, we propose an explainability-guided adversarial analysis methodology. We introduce the Saliency-LiDAR (SALL) method, which aggregates Integrated Gradient attributions across scenes to produce universal saliency maps for LiDAR-based 3D object detectors. Guided by these maps, we design the Explainability-aware Frustum Attack (EFA), which selectively perturbs only the most influential frustums rather than uniformly attacking entire object regions. Experiments on KITTI and nuScenes, across detectors such as PointPillars and SECOND, show that EFA reduces detection recall by more than 15 percentage points while requiring 25-50% fewer perturbed frustums than the state-of-the-art non-saliency-aware baseline. These findings reveal that modern 3D detectors concentrate discriminative evidence in a small subset of spatial regions, exposing a structural robustness vulnerability in current LiDAR perception systems. Our code is released at https://github.com/SecMindLab/Saliency_LiDAR.
LiDAR semantic segmentation is a key perception task in autonomous driving, where false predictions can affect downstream planning and safety-critical decision-making. Although adversarial attacks, and specifically adversarial examples, have been widely studied for image classification and 3D point cloud segmentation, unrestricted adversarial examples remain largely unexplored in the space of 2D range images, which are projections of 3D point clouds. The proposed method is, to the best of our knowledge, the first diffusion-based unrestricted adversarial attack against 2D range-image segmentation, using adversarial guidance from a segmentation loss. By applying guidance directly during sampling, the method produces unrestricted adversarial examples that remain close to the learned LiDAR data manifold while inducing structured segmentation errors. Experiments on the SemanticKITTI dataset using RangeNet++ and CENet segmentation networks demonstrate that the attack provides adjustable degradation across guidance strengths and transfers across segmentation architectures. Compared with norm-bounded FGSM and SegPGD baselines, the proposed attack offers a distinct effectiveness-realism trade-off, achieving controllable white-box and transfer degradation while maintaining competitive distributional and visual realism.
School of Electrical and Computer Engineering, National Technical University of Athens, Greece · Industrial Systems Institute, Athena Research Center, Patras Science Park, Greece