Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks
Authors: Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia
Organizations: School of Informatics, University of Edinburgh, UK · Department of Electronics, Informatics and Bioengineering, Politecnico di Milano, Italy
Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.
Figures & tables
Fig. 1 : An illustrative example of two successful trajectories, collected using the π0.5 model on the “pick up the black bowl on the ramekin and place it on the plate” task from the LIBERO-Spatial suite. Under the LIBERO-Plus camera angle perturbation, the goal condition is met, yet the bowl is dropped and re-picked during the trajectory. TSR alone does not capture this behavioural impact; instead, it is reflected in the mean Cartesian jerk and cumulative gripper movement trajectory metrics.
Work
TSR
Behavioural metrics
Robustness perturbations
LIBERO [ 19 ]
✓
✗
✗
LIBERO-Plus [ 7 ]
✓
✗
✓
BEHAVIOR [ 27 ]
✓
✓
✗
RoboEval [ 12 ]
✓
✓
✗ *
RLBench [ 23 ]
✓
✗
✗
CALVIN [ 24 ]
✓
✗
✗
TABLE I : Embodied AI evaluation approaches.
Median task-level change (1.d.p) (%) (#sig)
Success
Model
Pert.
Dur.
C. Path
J. Path
C. Jerk
J. Jerk
Grip.
P95 C. Jerk
TSR
n
LIBERO-Spatial
π0.5 Original TSR 98.8% Mean Pert.* TSR 90.3%
Back.
+10.3% (9)
+5.9%(5)
+8.7%(5)
-1.6%(0)
-0.8%(1)
+26.4% (7)
+3.8%(0)
97.1%
501/516
Cam.
+22.5% (10)
+6.1% (7)
+11.1% (6)
+24.1% (7)
+18.8% (7)
+37.8% (9)
+48.7% (7)
69.7%
524/752
Lan.
+6.7% (7)
+1.9%(4)
+2.3%(3)
-3.1%(1)
-1.3%(3)
+7.9%(3)
-3.8%(0)
90.8%
708/780
Light
+5.4% (7)
+1.9% (6)
+5.6% (7)
-3.1%(1)
-1.9%(2)
0.0%(2)
-3.9%(0)
98.6%
576/584
Obj.
+2.4%(5)
-1.3%(2)
-3.6%(1)
-1.0%(1)
-1.2%(1)
-0.6%(0)
-1.6%(1)
97.9%
377/385
TABLE II : Cells report the median task-level percentage change in the mean successful trajectory metric from baseline (original LIBERO) to the perturbation condition, with the number of tasks (out of 10) in the suite exhibiting a significant increase shown in parentheses. Cells are shaded when a significant increase is observed in the majority of tasks ( ≥6 ), with shading intensity indicating the size of the median percentage increase ( 0 – 5% , 5 – 15% , 15 – 30% and >30% ). Decreases and effects significant in fewer than 6 tasks are left unshaded. Dur. = Duration ( s ), C. Path = Cartesian path length ( m ), J. Path = joint path length ( rad ), C. Jerk = mean Cartesian jerk ( m/s3 ), J. Jerk = mean joint jerk ( rad/s3 ), Grip. = Total gripper movement ( m ), P95 C. Jerk = P95 Cartesian jerk ( m/s3 ). Models are jointly fine-tuned across all LIBERO suites.
Dur. (%)
Cart. jerk (%)
Grip. (%)
Pert.
π0.5
VLN
OFT
π0.5
VLN
OFT
π0.5
VLN
OFT
LIBERO-Spatial
Back.
+50
+27
-33
+27
-18
-9
+80
+34
-21
Cam.
+142
+58
+317
+160
+22
+122
+335
+38
+452
Lan.
0
+50
-29
+14
-1
-19
-1
-2
-13
Light
+6
0
-50
+12
-5
-28
+34
-25
-33
Obj.
-25
0
-15
+3
+6
+19
+4
+4
-1
TABLE III : Median percentage change in median absolute deviation (MAD) relative to baseline for successful trajectories. Positive values indicate increased variability; negative values indicate reduced variability. Cell colour indicates the magnitude and direction of change using symmetric thresholds: 25 – 49% , 50 – 99% , and ≥100% , with darker shading indicating larger changes. Values <25% are unshaded. VLN = VLANeXt.