cs.ROOct 1, 2026

Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks

Authors: Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia

Organizations: School of Informatics, University of Edinburgh, UK · Department of Electronics, Informatics and Bioengineering, Politecnico di Milano, Italy

Abstract

Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.

Figures & tables

Explore similar work

CardsList
  1. SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models

    May 30, 2026Jialiang Fan, Weizhe Xu, Zijun Wang +3

  2. How VLAs (Really) Work In Open-World Environments

    Apr 23, 2026Amir Rasouli, Yangzheng Wu, Zhiyuan Li +4Evaluation ProtocolsLong-Horizon Task Planning

  3. FATE-VLA:Failue-aware test generation for vision-language-action models

    Jun 1, 2026Arusa Kanwal, Pablo Valle, Shaukat Ali +1Vision-Language-Action FrameworkRobot Policies