Backdoored vision-language-action (VLA) policies can preserve benign task performance while producing malicious actions when a trigger appears. Detecting such activation is difficult because malicious behavior can comprise individually plausible actions, while unfamiliar tasks introduce legitimate changes in observations and behavior. We introduce TMT, a runtime backdoor detector based on Token Manifold and latent Transition modeling. Trained on benign rollouts, its two branches assess input-token structure and prediction errors in adjacent-layer latent dynamics. A suspicious rollout identified by the token manifold branch, once confirmed through latent deviations, guides transition selection for subsequent monitoring. We further explore policy purification through self-distillation: a frozen copy of the backdoored policy provides benign-input actions to supervise a student on paired benign and triggered observations, without requiring a separate clean reference policy. For evaluation, we adapt traditional backdoor detectors and repurpose anomaly and failure detection methods as VLA backdoor detectors. In a post-hoc comparison with ten baselines, TMT achieves state-of-the-art backdoor detection performance on unseen tasks across three VLA backdoor attacks. Our project page is available at https://zzr42.github.io/tmt/.
Figures & tables
Figure 1: Overview of TMT for runtime backdoor detection. On a frozen VLA backbone, the token manifold branch assesses input-token structure conditioned on the preceding query, while the latent transition branch scores prediction errors at a selected pair of adjacent layers. An alarm is raised when either branch’s score exceeds its separately calibrated threshold.
Figure 2: Visualizations of input tokens from backdoored OpenVLA-OFT ( Kim et al., 2025 ) under BadVLA, GoBA, and DropVLA. Blue circles denote clean queries, and red crosses denote backdoor-active queries.
BadVLA
GoBA
DropVLA
Method
AUC ↑
TDR ↑
FRR ↓
Query ↓
AUC ↑
TDR ↑
FRR ↓
Query ↓
AUC ↑
TDR ↑
FRR ↓
Query ↓
Backdoor Detection
STRIP
0.770
22.57
0.00
13.44
0.940
4.00
0.00
14.00
0.749
0.80
2.00
30.50
TeCo
0.251
0.57
7.71
48.00
0.662
5.71
0.86
2.65
0.673
6.00
1.60
24.27
DeDe
0.535
0.57
3.43
30.00
0.631
2.00
6.86
36.14
0.367
0.00
0.00
–
DUP-MS
0.001
0.00
0.29
–
0.841
4.86
0.00
13.71
0.570
1.60
2.00
39.00
Table 1: Post-hoc comparison on tasks unseen during detector fitting, using backdoored OpenVLA-OFT. We report rollout-level AUC, TDR, FRR, and mean detection latency in policy queries (Query), measured from trigger onset to the first alarm at or after onset. Latency is averaged over malicious rollouts with such an alarm; – indicates that none occur. Bold and underlining indicate the highest and second-highest distinct values, respectively, for AUC and TDR.
Figure 3: Benign and GoBA-triggered real-world behavior with detection-score traces. Dashed lines denote the calibrated thresholds of the token manifold and latent transition branches. The red box highlights threshold exceedances that trigger an alarm. Frames are drawn from separate demonstration recordings.
Table 2: Real-world GoBA detection on unseen tasks.
Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China. · Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China. · City University of Hong Kong, Hong Kong SAR, China.
Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences · Chongqing School, University of Chinese Academy of Sciences · Faculty of Information Technology and Electrical Engineering, University of Oulu, Finland