cs.ROSep 24, 2026

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

Authors: Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro, Subramanian Ramamoorthy, Matteo Matteucci, Alessandro Suglia

Organizations: Department of Electronics, Informatics, and Bioengineering, Politecnico di Milano, Milan, Italy · School of Informatics, University of Edinburgh, Edinburgh, UK

Abstract

Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their very high number of parameters, their computational requirements are often prohibitive for robotics control. To mitigate these inefficiencies, existing methods predominantly skip VLM backbone layers with early exits or reduce denoising steps, while leaving action expert depth untouched. We propose a framework that exposes backbone depth VV, action expert depth AA, and denoising steps DD as three jointly configurable compute axes in a VLA. Starting from a pretrained VLA, we attach lightweight Exit Transformers (ET) at intermediate depths in both the backbone and the action expert, trained to distil the last layer of the policy into each exit. Furthermore, we introduce a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deeper than the backbone. Finally, we show that the optimal compute budget is task-dependent, with different tasks benefiting from different axes and depths. Notably, our method does not require training the original policy from scratch, and for each exit, it increases the number of parameters by only 2.1%2.1\% for SmolVLA and 4.1%4.1\% for π0.5π_{0.5}. We validate our approach across two flow-matching VLAs (SmolVLA, π0.5π_{0.5}) and two benchmarks (LIBERO, Meta-World), revealing complementary effects: VV and AA respectively reduce FLOPs and latency, while DD improves both. Our joint configurations (V,A,D)(V,A,D) reduce latency by 79.2%79.2\% and computation (FLOPs) by 31.8%31.8\%, while improving mean success rate by 5.6%5.6\%.

Figures & tables

Explore similar work

CardsList
  1. Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

    Jun 18, 2026Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18Diffusion-Based Vision-Language-ActionsScalable Robot Learning

  2. ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

    May 28, 2026Ye Li, Huanan Liu, Kangye Ji +7Diffusion-Based Vision-Language-ActionsAction Generation

  3. AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

    Date pendingSunghwan Han, Youngtae Han, Youngmin YiFlow-Matching Vision-Language-ActionAction Generation