Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the action commanded by a VLA and the motion executed by the robot. To stress-test VLA robustness across execution conditions that are impractical to cover with physical robots alone, we introduce RoboStress, a controlled simulation benchmark. It combines established joint-level models of friction, backlash, compliance, and gravity-compensation error into seven deployment scenarios whose execution errors depend on the robot's state and motion history. On RoboStress, self-compensating VLA achieves higher average task success than both the base policies and methods that build in robustness during training. On two physical robot arms with different usage histories, it raises the average task success rate by more than 30 percentage points on each arm, and the gains extend to objects not seen in the task demonstrations.
Figures & tables
Figure 1: Overview of our study. (a) Self-compensating VLA updates the policy online using command-execution residuals and generates subsequent commands that pre-compensate for execution errors. (b) The RoboStress benchmark combines four established joint-level models into controlled deployment scenarios in simulation to stress-test VLA robustness.
Figure 2: Overview of self-compensating VLA. A frozen VLA with trainable LoRA adapters predicts an action chunk, which the robot executes. The residual ηobs between the commanded and executed motion forms a pseudo-target attarget , which the compensation objective Lcomp uses to update only the LoRA adapters online during deployment, leaving the base policy frozen.
Figure 3: Overview of the joint-level noise injection. The operational space controller ( Khatib, 1987 ) converts end-effector commands into joint torques. Torque-level noise (friction, gravity-compensation error) is added to joint torques, and position-level noise (backlash, compliance) to joint positions before physics integration. The next observation is rendered from the updated state.
Scenario
Physical cause
Fric.
Grav.
Back.
Comp.
Severity over episode
Joints
Heavy Payload
grasped load
✓
✓
While grasping
all
Thermal Drift-Stribeck
temperature change
✓
Linearly increasing
all
Thermal Drift-Backlash
temperature change
✓
Linearly increasing
all
Aged Transmission
gear wear
✓
✓
Fixed
all
Aged Joint-Uniform
gear + bearing wear
✓
✓
✓
Fixed
all
Aged Joint-Shoulder
gear + bearing wear
✓
✓
✓
Fixed
J1–J3
Table 1: Deployment scenarios in RoboStress. Checked noise components have increased severity, while unmarked components retain their nominal settings.
π0.5
π0
Scenario
Base
DR
RobustVLA
Ours
Base
DR
RobustVLA
Ours
Clean
97.0
97.1
96.6
95.8
91.1
90.5
92.0
92.6
Heavy Payload
44.1
42.1
38.9
56.8
42.3
32.0
30.9
42.8
Thermal Drift-Stribeck
54.1
54.6
55.9
60.5
46.1
46.3
46.0
50.6
Thermal Drift-Backlash
83.4
84.5
83.1
84.8
75.5
77.3
77.3
78.8
Aged Transmission
33.2
39.7
40.2
41.8
26.8
22.9
23.8
28.9
Table 2: Success rates (%) on RoboStress deployment scenarios, averaged over LIBERO tasks. Averages exclude Clean. Best results within each backbone are shown in bold for each row.
Figure 6
Robot A
Robot B
π0.5
π0
π0.5
π0
Task
Base
RobustVLA
Ours
Base
RobustVLA
Ours
Base
RobustVLA
Ours
Base
RobustVLA
Ours
1
48
52
84
44
52
76
44
56
80
36
56
60
2
36
60
76
40
28
80
36
48
72
32
52
84
3
48
48
76
44
68
72
56
52
84
52
44
84
4
40
48
68
32
40
60
36
44
60
32
44
56
Table 6: Real-world success rates (%) on two Piper arms. Tasks 1–3 involve picking up the gray bowl from next to the plate, on the plastic cabinet, and on the gift box, respectively, and placing it on the plate. Task 4 involves opening the cabinet’s top drawer. Best results are shown in bold.
Figure 5: Rollouts of self-compensating VLA (Ours) and the base policy (Base) on Tasks 2 and 4, for Robot A and Robot B. Green and red borders mark success and failure.
Figure 6: Generalization to unseen objects, evaluated on Robot A. Left: Our rollouts with the gray bowl (ID) and object variants (OOD). Right: Success rates on all OOD variants.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Fc,j,0
Fs,j,0
vs,j
σv,j,0
Bj,0
Mjeff
Dj,0
Kj,0
βj,0
Joint
(Nm)
(Nm)
(rad/s)
(Nm ⋅ s/rad)
(rad)
(kg ⋅ m 2 )
(Nm ⋅ s/rad)
(Nm/rad)
J1
2.175
3.480
0.15
0.348
0.00087
0.70
90.73
6000
0.05
J2
2.175
3.480
0.15
0.348
0.00087
0.70
90.73
6000
0.10
J3
2.175
3.480
0.15
0.348
0.00087
0.70
90.73
6000
0.06
J4
2.175
3.480
0.15
0.348
0.00087
0.70
90.73
6000
0.10
J5
0.480
0.720
0.10
0.048
0.00130
0.20
39.60
4000
0.08
Appendix
Table A1: Nominal noise parameters per joint j , defined in § 4.1 .
Component
Applied to
Full setting
Stribeck friction ( wf )
Fc,j , Fs,j , σv,j ( vs,j fixed)
×2
Backlash ( wb )
Bj
×15
Compliance ( wc )
Kj=Kj,0/wjc , Dj=Dj,0/wjc ( Mjeff fixed)
K÷5 , D÷5
Gravity ( wg )
βj=wjgβj,0
×4
Appendix
Table A2: Full setting of each severity factor of Eq. ( 7 ). wjf , wjb , and wjc model wear, and wjg models payload mismatch.
all joints: the above Stribeck and backlash parameters plus K=1200/800 Nm/rad, D=40.58/17.71 Nm ⋅ s/rad.
Aged Joint-Shoulder
Stribeck + Backlash + Compliance
J1–J3 at the full setting; J4–J7 nominal.
Appendix
Table A4: Deployment scenarios. Values are reported as J1–J4 / J5–J7 unless stated otherwise. “Full setting” applies the factors of Table A2 , and “nominal” denotes the values of Table A1 .
Scenario
Residual (%)
Simulation (RoboStress)
Clean
23.4
Heavy Payload
48.6
Thermal Drift-Stribeck
47.5
Thermal Drift-Backlash
46.1
Aged Transmission
38.7
Appendix
Table C5: Residual magnitude across scenarios, computed as ∥ηobs∥2/∥apolicy∥2 on π0 rollouts and expressed as a percentage. The simulation rows use the end-effector residual in normalized action units and the real-robot rows use the joint-space residual of § B.6 in radians.
Scenario
w/o anchor
w/ anchor
Heavy Payload
38.2
42.8
Thermal Drift-Stribeck
50.4
50.6
Thermal Drift-Backlash
77.9
78.8
Aged Transmission
25.2
28.9
Aged Joint-Uniform
22.8
24.0
Aged Joint-Shoulder
35.5
39.2
Appendix
Table D6: Effect of the anchor loss on π0 . Success rates (%) per deployment scenario, averaged over the four LIBERO suites.
Figure D1: Success rate (%) of the base policy and self-compensating VLA with π0.5 across five severity levels (L1 to L5) for each deployment scenario (rows) and LIBERO suite (columns) in RoboStress, averaged over the same three seeds as Table 2 .
Wear factors
Heavy Payload
Thermal Drift ramp end
Level
Stribeck wf
Backlash wb
Compliance K÷
K÷
β×
Stribeck
Backlash
L1 (mild)
1.5
5
2
2
2
1.5 → 1.5
5 → 5
L2
1.625
7.5
2.75
2.75
2.5
1.5 → 1.625
5 → 7.5
L3
1.75
10
3.5
3.5
3
1.5 → 1.75
5 → 10
L4
1.875
12.5
4.25
4.25
3.5
1.5 → 1.875
5 → 12.5
L5 (full)
2
15
5
5
4
1.5 → 2
5 → 15
Appendix
Table D7: Severity levels. Each column moves linearly from L1 to L5, and L5 is the setting of Table 2 . Compliance divides the stiffness by the listed factor and the damping by its square root ( Kj=Kj,0/wjc , Dj=Dj,0/wjc ), which keeps ζ=0.70 at every level. Wear factors apply to the worn joints of each scenario, that is, all joints for Aged Transmission and Aged Joint-Uniform, J1–J3 for Shoulder, and J4 for Elbow. Thermal Drift entries give the ramp as onset → end multiplier.
While Vision-Language-Action (VLA) models demonstrate impressive capabilities in robotic manipulation, their memoryless nature renders them brittle to test-time environment shifts, particularly hardware shifts caused by wear or imperfect calibration. Enabling these models to self-adapt during deployment without requiring continuous on-site recalibration remains a critical bottleneck for real-world scalability. In this work, we introduce Self-Adaptive VLA, a novel post-training recipe that enables the policy to iteratively adapt to deployment-time hardware shifts leveraging its own rollouts as context. To do so, we first collect policy rollouts under deliberately injected hardware shifts. We then transform the base policy's training data into shift-conditioned expert demonstrations by pre-compensating the expert actions for these known shifts. Next, we introduce a lightweight, plug-in context encoder that compresses the context, including visual observation, proprioception, and actions in the shifted environment, into a latent context token. This token modulates the policy through adaptive layer normalization (AdaLN). Furthermore, we find that context tokens can be ensembled, allowing the policy to iteratively self-correct and mitigate failures step by step. Extensive experiments across four precision-critical bi-manual and dexterous manipulation tasks show that Self-Adaptive VLA recovers over 80% of the base policy's performance under hardware shifts, such as actuation bias and joint encoder offsets. Moreover, Self-Adaptive VLA enables more robust deployment to new workstations compared to the base policy. Our approach provides a pathway for robust large-scale real-world robot deployments and easier maintenance. See videos at https://icefoxzhx.github.io/self-adaptive-vla.
Vision-Language-Action (VLA) models are increasingly used as generalist robot policies, yet their evaluation still relies largely on static benchmarks that randomly sample task scenes. In high-dimensional embodied spaces, failures are sparse and clustered, so static benchmarking can underestimate robustness risks. We reframe VLA evaluation as an active failure-discovery problem and propose a failure-aware test-generation approach that combines diversity-driven exploration with surrogate models learned from observed executions. The method steers testing toward high-risk yet diverse scene regions. Across four state-of-the-art VLA models, it uncovers substantially more failures (up to +29.7 % over selected baselines) while revealing more diverse failure modes. This mean that, for instance, in the case of GR00T-N1.6, success rate dropped from 64.4% to 34.7%. More broadly, our findings call for a shift in VLA evaluation: from passive measurement on fixed task suites to adaptive, failure-seeking test generation that exposes the structure of model weaknesses before deployment.
Arusa Kanwal, Pablo Valle, Shaukat Ali +1
Mondragon University Mondragon, Spain · Simula Research Laboratory Oslo, Norway
Vision-Language-Action (VLA) models follow a data-driven paradigm and are constrained by the coverage of training data, making them prone to failure on edge-case configurations after deployment. To mitigate such risks, it is essential to expose high-quality failure modes and convert the resulting failures into supervisory data for model enhancement. Existing studies largely stop at failure detection and lack a mechanism for leveraging discovered failures for model repair. We propose VLAMotor, the first analysis framework for VLA enhancement, which integrates distance-aware model testing for failure exposure and agent-based data synthesis for model finetunning. First, VLAMotor estimates input uncertainty based on the distance to training samples, and combines uncertainty ranking with redundancy elimination to build compact test sets that expose diverse failures. Then, VLAMotor abstracts failure trajectories into structured semantic representations, and plans parameterized repair-skill sequences, which are then realized as executable trajectories through inverse kinematics and motion execution. The resulting successful trajectories are automatically labeled and used to fine-tune the original VLA model, yielding an enhanced VLA model. Evaluation on four representative robotic manipulation tasks shows that 92.33% of the in-simulation test cases generated by VLAMotor trigger VLA failures, and VLAMotor improves test coverage over the state-of-the-art tool by 18.93%. By fine-tuning VLA models with synthetic data derived from failed test cases, VLAMotor further enhances the overall success rate of VLA models by 49.25%. When deployed on real hardware, the simulation-enhanced models improve the success rate over the original VLA models by 57.50%, demonstrating an effective and low-cost direction for VLA enhancement.
Zeqin Liao, Peifan Ren, Zixu Gao +6
School of computing and data science, Nanyang Technological University · School of Software Engineering, Sun Yat-sen University, China and GuangDong Engineering Technology Research Center of Blockchain, China · Northwest A&F University