Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks
Authors: Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia
Organizations: School of Informatics, University of Edinburgh, UK · Department of Electronics, Informatics and Bioengineering, Politecnico di Milano, Italy
Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.
Figures & tables
Fig. 1 : An illustrative example of two successful trajectories, collected using the π0.5 model on the “pick up the black bowl on the ramekin and place it on the plate” task from the LIBERO-Spatial suite. Under the LIBERO-Plus camera angle perturbation, the goal condition is met, yet the bowl is dropped and re-picked during the trajectory. TSR alone does not capture this behavioural impact; instead, it is reflected in the mean Cartesian jerk and cumulative gripper movement trajectory metrics.
Work
TSR
Behavioural metrics
Robustness perturbations
LIBERO [ 19 ]
✓
✗
✗
LIBERO-Plus [ 7 ]
✓
✗
✓
BEHAVIOR [ 27 ]
✓
✓
✗
RoboEval [ 12 ]
✓
✓
✗ *
RLBench [ 23 ]
✓
✗
✗
CALVIN [ 24 ]
✓
✗
✗
TABLE I : Embodied AI evaluation approaches.
Median task-level change (1.d.p) (%) (#sig)
Success
Model
Pert.
Dur.
C. Path
J. Path
C. Jerk
J. Jerk
Grip.
P95 C. Jerk
TSR
n
LIBERO-Spatial
π0.5 Original TSR 98.8% Mean Pert.* TSR 90.3%
Back.
+10.3% (9)
+5.9%(5)
+8.7%(5)
-1.6%(0)
-0.8%(1)
+26.4% (7)
+3.8%(0)
97.1%
501/516
Cam.
+22.5% (10)
+6.1% (7)
+11.1% (6)
+24.1% (7)
+18.8% (7)
+37.8% (9)
+48.7% (7)
69.7%
524/752
Lan.
+6.7% (7)
+1.9%(4)
+2.3%(3)
-3.1%(1)
-1.3%(3)
+7.9%(3)
-3.8%(0)
90.8%
708/780
Light
+5.4% (7)
+1.9% (6)
+5.6% (7)
-3.1%(1)
-1.9%(2)
0.0%(2)
-3.9%(0)
98.6%
576/584
Obj.
+2.4%(5)
-1.3%(2)
-3.6%(1)
-1.0%(1)
-1.2%(1)
-0.6%(0)
-1.6%(1)
97.9%
377/385
TABLE II : Cells report the median task-level percentage change in the mean successful trajectory metric from baseline (original LIBERO) to the perturbation condition, with the number of tasks (out of 10) in the suite exhibiting a significant increase shown in parentheses. Cells are shaded when a significant increase is observed in the majority of tasks ( ≥6 ), with shading intensity indicating the size of the median percentage increase ( 0 – 5% , 5 – 15% , 15 – 30% and >30% ). Decreases and effects significant in fewer than 6 tasks are left unshaded. Dur. = Duration ( s ), C. Path = Cartesian path length ( m ), J. Path = joint path length ( rad ), C. Jerk = mean Cartesian jerk ( m/s3 ), J. Jerk = mean joint jerk ( rad/s3 ), Grip. = Total gripper movement ( m ), P95 C. Jerk = P95 Cartesian jerk ( m/s3 ). Models are jointly fine-tuned across all LIBERO suites.
Dur. (%)
Cart. jerk (%)
Grip. (%)
Pert.
π0.5
VLN
OFT
π0.5
VLN
OFT
π0.5
VLN
OFT
LIBERO-Spatial
Back.
+50
+27
-33
+27
-18
-9
+80
+34
-21
Cam.
+142
+58
+317
+160
+22
+122
+335
+38
+452
Lan.
0
+50
-29
+14
-1
-19
-1
-2
-13
Light
+6
0
-50
+12
-5
-28
+34
-25
-33
Obj.
-25
0
-15
+3
+6
+19
+4
+4
-1
TABLE III : Median percentage change in median absolute deviation (MAD) relative to baseline for successful trajectories. Positive values indicate increased variability; negative values indicate reduced variability. Cell colour indicates the magnitude and direction of change using symmetric thresholds: 25 – 49% , 50 – 99% , and ≥100% , with darker shading indicating larger changes. Values <25% are unshaded. VLN = VLANeXt.
Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety violations along the trajectory: a policy may reach the goal while applying excessive contact, disturbing bystander objects, destabilizing a held object, or entering robot self-contact. We present SafeVLA-Bench, a post-hoc safety-evaluation framework for existing simulator-based VLA benchmarks that reveals violations missed by success-only evaluation. It encodes task-aware safety requirements as Signal Temporal Logic (STL) invariants with quantitative robustness semantics. Alongside native success, it reports the safety rate and the success-but-unsafe rate (SBU) used in prior safety evaluations, and introduces the Violation Severity Index (VSI), a bounded worst-violation depth score. We instantiate SafeVLA-Bench on LIBERO and RoboCasa-365, evaluating twenty-seven policy-benchmark entries across tabletop and kitchen manipulation tasks. High task success does not imply safe execution: the fifteen tabletop policies above 90% mean success still have 18-28% unsafe-episode rates, and 38-56% of successful RoboCasa-365 rollouts violate at least one active safety clause. A post-training case study further shows that SafeVLA-Bench can be used to improve policy safety. Project page: https://safevla.org
Jialiang Fan, Weizhe Xu, Zijun Wang +3
University of Notre Dame · University of Pennsylvania
Vision-language-action models (VLAs) have been extensively used in robotics applications, achieving great success in various manipulation problems. More recently, VLAs have been used in long-horizon tasks and evaluated on benchmarks, such as BEHAVIOR1K (B1K), for solving complex household chores. The common metric for measuring progress in such benchmarks is success rate or partial score based on satisfaction of progress-agnostic criteria, meaning only the final states of the objects are considered, regardless of the events that lead to such states. In this paper, we argue that using such evaluation protocols say little about safety aspects of operation and can potentially exaggerate reported performance, undermining core challenges for future real-world deployment. To this end, we conduct a thorough analysis of state-of-the-art models on the B1K Challenge and evaluate policies in terms of robustness via reproducibility and consistency of performance, safety aspects of policies operations, task awareness, and key elements leading to the incompletion of tasks. We then propose evaluation protocols to capture safety violations to better measure the true performance of the policies in more complex and interactive scenarios. At the end, we discuss the limitations of the existing VLAs and motivate future research.
Amir Rasouli, Yangzheng Wu, Zhiyuan Li +4
Noahs Ark Laboratory, Huawei Technologies Canada. · while at Huawei Canada.
Vision-Language-Action (VLA) models are increasingly used as generalist robot policies, yet their evaluation still relies largely on static benchmarks that randomly sample task scenes. In high-dimensional embodied spaces, failures are sparse and clustered, so static benchmarking can underestimate robustness risks. We reframe VLA evaluation as an active failure-discovery problem and propose a failure-aware test-generation approach that combines diversity-driven exploration with surrogate models learned from observed executions. The method steers testing toward high-risk yet diverse scene regions. Across four state-of-the-art VLA models, it uncovers substantially more failures (up to +29.7 % over selected baselines) while revealing more diverse failure modes. This mean that, for instance, in the case of GR00T-N1.6, success rate dropped from 64.4% to 34.7%. More broadly, our findings call for a shift in VLA evaluation: from passive measurement on fixed task suites to adaptive, failure-seeking test generation that exposes the structure of model weaknesses before deployment.
Arusa Kanwal, Pablo Valle, Shaukat Ali +1
Mondragon University Mondragon, Spain · Simula Research Laboratory Oslo, Norway