Visual disruptions can arise while a robot is executing a task, leaving a vision-language-action (VLA) policy to respond without knowing the disruption type or timing. We introduce Self-supervised Adaptation from Leftover Trajectories (SALT), which uses the leftover trajectory, the unexecuted part of the previous action chunk, as self-supervision for test-time adaptation. Because consecutive chunks overlap in time, the leftover provides a temporally aligned target for the current prediction over the same future control interval. At the onset of a visual shift, the leftover can retain a plan formed before the corruption, so updating the policy toward it anchors the adaptation across the shift (Transition Anchoring). SALT keeps the adapted policy and regenerates the current chunk, whose leftover becomes the target at the next replan, carrying the correction forward along the execution trajectory (Sequential Correction Propagation). Supervision comes entirely from the policy's own predictions, requiring no disruption annotations, expert actions, or target-domain demonstrations, and a lightweight adaptation gate calibrated only on nominal trajectories decides when updates begin. On LIBERO-10, SALT increases average success across five persistent visual corruptions from 43.9% to 53.2% with SmolVLA and from 58.7% to 66.0% with GR00T N1.7, while largely preserving nominal performance. On a real robot, it raises task progress averaged over digital and physical disruptions from 0.49 to 0.61.
Figures & tables
Figure 1: Visual disruption during an ongoing task. A VLA begins executing under nominal observations, but an unexpected visual corruption can arise during execution and alter its subsequent action predictions. Left: an example of task failure under the visual disruption. Right: deviation between each new action chunk and the aligned leftover of the previous chunk. It stays small under nominal observations (gray) but jumps at the first corrupted replan (red). See Appendix H for details.
Figure 2: Temporal overlap between consecutive chunks. The leftover trajectory At[E:H] and the next chunk’s prefix At+1[0:L] cover the same future control interval. Indices mark action positions in At+1 , so position j of At+1 and position E+j of At refer to the same future control step.
Figure 3: SALT adapts to a visual disruption during execution. Illustrated with a chunk of H=6 actions, of which E=3 are executed and L=3 remain as the leftover. (a) Transition Anchoring : at the shift boundary, the leftover of the chunk planned from the last nominal observation serves as the supervision target. SALT updates the policy under the shifted observation toward this leftover, pulling it toward the plan formed before the shift. (b) Sequential Correction Propagation : at later replans, the target is the leftover of the chunk regenerated by the adapted policy at the previous replan. Because the adapted policy is carried forward, each update refines the correction along the execution trajectory.
Figure 4: Leftover quality at transition. Median 42-step ℓ2 distance to the clean-input prediction across 1,000 pairs: 0.643 for the leftover and 1.294 for the corrupted-input prediction.
Backbone
Method
Clean
Motion
Gaussian
Zoom
Glass
Fog
Avg.
SmolVLA
Frozen
71.5
63.0
25.0
29.0
56.5
46.0
43.9
ActMAD ( Mirza et al., 2023 )
74.5
60.0
29.0
31.0
54.0
48.5
44.5
BID ( Liu et al., 2025 )
69.5
63.0
27.0
32.5
57.0
50.5
46.0
Self-training
71.0
61.0
23.5
26.5
55.5
48.0
42.9
SALT
70.5
68.5
43.0
42.0
58.5
54.0
53.2
GR00T N1.7
Frozen
81.5
76.5
46.0
53.0
40.0
78.0
58.7
Table 1: Success rates (%) under five persistent visual corruptions on LIBERO-10 . Avg. averages the five corrupted conditions. Clean is reported separately. Bold marks the best corrupted-condition result within each backbone.
Figure 5: Representative real-world visual disruptions. Nominal and disrupted observations are shown for Two Cats in Cups. Gaussian blur and fog are applied digitally to the camera stream; lights off and camera cover are physical disturbances introduced during execution.
Figure 6: Real-world task-progress score under digital and physical disruptions. Each bar averages 20 trials per method. A score of 1.0 means the task was completed.
Figure 7: Leftover target properties and shift timing on SmolVLA. (a) Temporal alignment. Shifting the leftover away from its aligned position at offset 0 reduces success to or below Frozen. (b) Target identity. Another episode’s plan and a time-shuffled previous chunk remain at or below Frozen, while the aligned leftover approaches the clean-input prediction target. (c) Shift timing. Success when the shift boundary is fixed at four points of the clean rollout. Shaded bands show 95% confidence intervals. Exact values for (c) are reported in Table 11 .
Figure 8: One update narrows the clean-plan gap (SmolVLA, Gaussian blur; thin: tasks, thick: pooled).
Figure 9: Roles and propagation of leftover-supervised updates on SmolVLA. (a) Ablations of the two roles: success over the 1,000 episodes of Table 1 ; without anchoring skips only the update at the shift boundary. (b) Sequential correction: normalized ℓ2 -distance reduction for policies updated through replan t−3,…,t (40 replayed clean trajectories, five corruptions; 95% bootstrap intervals). All policies are evaluated on the same ot .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Corruption
Parameters
Level 3
Level 5
Level 7
Motion blur
radius, σ (px)
10, 4
15, 6
20, 10
Gaussian blur
σ (px)
3
5
7
Zoom blur
max. zoom (step)
1.20 (0.02)
1.30 (0.03)
1.40 (0.01)
Glass blur
σ , displacement, iterations
0.9, 2, 3
1.1, 3, 2
1.5, 4, 2
Fog
strength, decay
1.5, 2.5
2.5, 2.0
3.5, 1.6
Appendix
Table 2: Corruption parameters by severity level , taken from the LIBERO-Plus tables. Zoom blur averages the image over zoom factors from 1.0 to the listed maximum in the listed step.
Figure 10: Visual corruption examples by severity. The five LIBERO-Plus corruption families at severity levels 3, 5, and 7.
Setting
SmolVLA
GR00T N1.7
Checkpoint
SmolVLA fine-tuned on LIBERO
GR00T N1.7 LIBERO-10
Chunk H / executed prefix E / overlap L
50/8/42
16/8/8
Adapted projections
keys of the 8 cross-attention layers of the action expert
keys of the 16 cross-attention blocks of the action-head DiT
Trainable scales
2,560
24,576
Bound η
0.15
0.15
Optimizer
Adam, lr 0.06, 5 steps
Adam, lr 0.06, 5 steps
Appendix
Table 3: Model-specific implementation.
Figure 11: SmolVLA success rate (%) on LIBERO-10 across corruption severity. Frozen and SALT are evaluated on 200 episodes per corruption and severity. Mot. and Gaus. abbreviate motion and Gaussian blur; Avg. excludes clean episodes. Numbers above Avg. are SALT gains in percentage points.
Method
Motion
Gaussian
Zoom
Glass
Fog
Avg.
Frozen
70.0
43.0
39.0
65.5
52.5
54.0
SALT
70.5
57.0
54.0
64.5
56.0
60.4
Appendix
Table 4: Agent-view-only corruption (SmolVLA, LIBERO-10, success %). The wrist image stays clean. 200 episodes per cell.
Backbone
Suite
Method
Motion
Gaussian
Zoom
Glass
Fog
Avg.
SmolVLA
Spatial
Frozen
74.5
56.5
48.0
75.0
74.5
65.7
SALT
75.5
70.5
68.0
75.5
72.0
72.3
Object
Frozen
93.0
75.5
67.5
92.5
92.5
84.2
SALT
91.5
88.0
86.5
91.0
89.5
89.3
Goal
Frozen
80.5
67.5
57.5
79.5
80.5
73.1
SALT
83.0
67.5
72.0
81.0
79.5
76.6
Appendix
Table 5: Results across LIBERO suites. Success rate (%) under persistent severity-5 corruption, 1,000 episodes per suite and method.
Figure 12: Real-world tabletop tasks. From top to bottom: Two Cats in Cups places the orange and gray cats in the left and right cups, respectively; Stack places the middle block on the right block, then the left block in a cup; and Drawer places a block in the drawer and closes it. Each row shows successive stages from the starting scene to task completion.
Two Cats in Cups
Drawer
Stack
Method
Gaussian
Fog
Lights off
Camera cover
Gaussian
Fog
Lights off
Camera cover
Gaussian
Fog
Lights off
Camera cover
Frozen
2
0
1
1
0
3
9
5
0
0
0
0
ActMAD
3
0
0
3
0
4
8
4
0
1
0
0
BID
1
1
2
2
0
8
8
9
0
0
0
1
Self-training
1
0
1
2
0
3
7
4
0
0
0
0
SALT
7
0
3
5
0
9
14
8
0
0
0
1
Appendix
Table 6: Real-world success counts. Successful trials out of 20 per task, condition, and method.
Backbone
Adaptation starts
Clean
Motion
Gaussian
Zoom
Glass
Fog
Avg.
SmolVLA
Never (Frozen)
71.5
63.0
25.0
29.0
56.5
46.0
43.9
At activation gate
70.5
68.5
43.0
42.0
58.5
54.0
53.2
At every replan
69.5
67.5
42.0
45.5
61.0
56.5
54.5
Appendix
Table 7: When SALT starts adapting (SmolVLA, LIBERO-10, success %). Every row except Never runs the same leftover-supervised update; only its start differs. The activation-gate row is the configuration of Table 1 . Avg. excludes clean episodes.
SmolVLA
GR00T N1.7
Method
Nominal
Corrupted
Nominal
Corrupted
Frozen
1.00
1.00
1.00
1.00
BID
33.7
33.7
24.3
24.3
ActMAD
2.7
2.7
3.1
3.1
Self-training
1.4
2.8
1.5
3.2
SALT (gate)
1.5
2.9
1.4
3.4
Appendix
Table 8: Average per-replan latency across methods (multiples of the frozen policy; update frequencies from the evaluation of Table 1 ).
Model
Frozen
Quiet
Optimization
Regeneration
Update replan
SmolVLA
134.8
143.9
269.3
143.4
559.3
GR00T N1.7
139.8
142.8
414.7
154.4
723.7
Appendix
Table 9: Per-replan latency breakdown. Median milliseconds on one H200, measured in the rollout loop. Optimization uses five gradient steps and eight noise–time draws; Quiet is a replan without an update, gate included.
Figure 13: Supervision length on SmolVLA. Success when only the first n of the 42 aligned leftover actions are supervised (the 1,000 episodes of Table 1 over five persistent corruptions). Supervising 16 to 42 positions performs similarly; shorter targets are less effective.
Figure 14: One update narrows the clean-plan gap on GR00T N1.7. Gaussian blur: distance to the clean-input prediction at the same shift-boundary state and flow-noise seed before and after one update ( n=200 , eight aligned positions). Thin lines are task medians; the thick line is the pooled median.
Figure 15: Action-consistency gap around the shift boundary. Median (line) and interquartile range (band) of the normalized RMSE between the previous chunk’s leftover and the current chunk’s aligned prefix, over 150 shift boundaries per backbone. Gray marks clean-to-clean transitions and red marks transitions at or after t∗ . The SmolVLA panel is the right panel of Figure 1 .
Backbone
Clean alarm
Early
Motion
Gaussian
Zoom
Glass
Fog
Overall
Mean delay
SmolVLA
26.5
9.0
91.0
91.0
91.0
91.0
90.5
90.9
0.13
GR00T N1.7
12.5
8.0
31.5
92.0
90.5
92.0
31.0
67.4
0.74
Appendix
Table 10: Gate timing on LIBERO-10 (%). Overall divides activations at or after t∗ by all 1,000 corrupted episodes; delay is the mean number of replans between t∗ and activation.
Shift boundary
Frozen
SALT
Gain
Active replans
10%
33.5 [30.2, 36.9]
50.0 [45.7, 54.3]
+16.5 [13.1, 19.9]
44.3
25%
40.0 [36.0, 44.1]
54.1 [49.2, 59.0]
+14.1 [10.8, 17.4]
37.5
50%
45.1 [40.3, 49.9]
57.0 [51.9, 62.3]
+11.9 [8.4, 15.4]
28.4
75%
53.9 [48.2, 59.5]
62.5 [56.6, 68.5]
+8.6 [5.6, 11.7]
17.3
Appendix
Table 11: Success by shift-boundary position (SmolVLA, %, 1,000 episodes per row, 95% confidence intervals in brackets). Active replans is the mean number of replans per episode at which SALT updates.
Schedule
Method
Motion
Gaussian
Zoom
Glass
Fog
Avg.
Temporary
Frozen
63.5
48.5
51.5
64.5
58.0
57.2
SALT
68.0
59.5
62.0
69.0
66.0
64.9
Gradual
Frozen
62.5
13.5
21.0
50.5
40.5
37.6
SALT
64.5
28.0
35.5
59.0
43.0
46.0
Appendix
Table 12: Temporary and gradual corruption by family (SmolVLA, success %, 200 episodes per corruption and method). SALT updates at every replan once a previous chunk is available. Avg. averages the five corrupted conditions.
Method
Motion
Gaussian
Zoom
Glass
Fog
Avg.
Frozen
58.0
3.5
11.0
51.0
39.0
32.5
SALT
56.5
8.0
14.0
51.0
35.0
32.9
Appendix
Table 13: Corruption from the first replan (SmolVLA, success %, 200 episodes per family). SALT updates at every replan once a previous chunk is available. Avg. averages the five corrupted conditions.
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose FailPatch, a failure-driven residual patching framework that decouples action supervision from execution-reliability supervision. Successful demonstrations ground how the policy should act, while deployment trajectories indicate when its behavior becomes unreliable. We further observe that action hidden representations exhibit clear linear separability between reliable and failure-associated states while directly conditioning action generation. Building on these insights, FailPatch introduces a Null-gated Residual Expert Bank into the action hidden space of a frozen VLA policy. A unified Preserve--Redirect--Trust objective retains the original policy in reliable states, selects residual experts in failure-associated states and redirects representations from failure regions toward success-associated regions under bounded intervention. With only 0.52% trainable parameters, FailPatch improves success rates by 11.0 percentage points on four long-horizon RoboTwin tasks under clean evaluation, 9.5 percentage points under clean-to-random generalization, and 16.7 percentage points over the baseline across three real-world tasks. Project and code: https://github.com/yupeng-2003/FailPatch.
Peng Yu, Jiacheng Wang, Ziheng Zhang +6
Xi’an Jiaotong University · Dexmal · Nanjing University
Vision-Language-Action (VLA) models enable robots to predict actions directly from visual observations and language instructions, but adapting them to new environments still depends on costly action-labeled demonstrations. To reduce this dependence, we study semi-supervised VLA adaptation under limited supervision signals, where only a small portion of trajectories contain robot actions and the remaining trajectories provide action-unlabeled vision-language observations. Unlike standard semi-supervised learning, the missing supervision is an embodied action signal that must be visually grounded, language-consistent, physically feasible, and temporally stable. To address this problem, we propose SemiVLA, a self-distilled teacher-student framework that learns from reliable pseudo-actions on unlabeled trajectories. SemiVLA introduces a VLA-specific reliability controller to assess vision-language alignment, action feasibility, and temporal transition consistency, and further updates the teacher through a Bottleneck-Projected Alignment Update to avoid noisy feedback contamination. With OpenVLA as the backbone, SemiVLA consistently improves multiple PEFT strategies across LIBERO and CALVIN. Under 10% labeled trajectories, SemiVLA with Selective LoRA achieves 89.0% average success on LIBERO, outperforming supervised LoRA by 8.0 points without extra inference cost.
Effective online adaptation of vision-language-action (VLA) models remains challenging, as sparse rewards provide weak supervision for high-dimensional autoregressive action policies. Although self-distillation can in principle provide denser training signals, we find that text-based privileged teachers conditioned on demonstrations, retrieved experiences, or high-level plans are ineffective for VLA adaptation, exposing a modality gap between symbolic guidance and low-level robot actions. We propose ROAD-VLA, an advantage-guided self-distillation framework that constructs a proximal teacher directly in action space by perturbing action-token logits with calibrated advantage estimates. This converts sparse rewards into dense token-level supervision while keeping the teacher close to the current policy. We further derive a policy-improvement lower bound under calibrated advantages and accurate teacher matching. Across seven robotic manipulation environments with in-distribution and out-of-distribution shifts, ROADVLA outperforms PPO in nearly all settings, demonstrating robust online VLA adaptation.
Kejing Wang, Toan Nguyen, Minh Hoang Nguyen +2
Applied Artificial Intelligence Initiative Deakin University Geelong, Australia · Air Force Research Laboratory USA