cs.ROOct 6, 2026

Adapting Vision-Language-Action Models to Unknown Visual Disruptions During Execution

Authors: Ahin Lee, Jinwoo Seo, Youngsoo Jang, Taesik Gong

Organizations: UNIST, Ulsan, Republic of Korea

Abstract

Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy's own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FailPatch: Failure Residual Patching for Vision-Language-Action Models

    Sep 28, 2026Peng Yu, Jiacheng Wang, Ziheng Zhang +6Robot Policy LearningRobot Policy Adaptation

  2. Semi-Supervised Vision-Language-Action Model

    Jun 19, 2026Hongyang He, Jiuming Liu, Victor SanchezVLM AdaptationVision-Language-Action Models

  3. ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models

    Jun 24, 2026Kejing Wang, Toan Nguyen, Minh Hoang Nguyen +2Robot Policy AdaptationVLM Adaptation