Organizations: Department of Electronics, Informatics, and Bioengineering, Politecnico di Milano, Milan, Italy · School of Informatics, University of Edinburgh, Edinburgh, UK
Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their very high number of parameters, their computational requirements are often prohibitive for robotics control. To mitigate these inefficiencies, existing methods predominantly skip VLM backbone layers with early exits or reduce denoising steps, while leaving action expert depth untouched. We propose a framework that exposes backbone depth V, action expert depth A, and denoising steps D as three jointly configurable compute axes in a VLA. Starting from a pretrained VLA, we attach lightweight Exit Transformers (ET) at intermediate depths in both the backbone and the action expert, trained to distil the last layer of the policy into each exit. Furthermore, we introduce a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deeper than the backbone. Finally, we show that the optimal compute budget is task-dependent, with different tasks benefiting from different axes and depths. Notably, our method does not require training the original policy from scratch, and for each exit, it increases the number of parameters by only 2.1% for SmolVLA and 4.1% for π0.5. We validate our approach across two flow-matching VLAs (SmolVLA, π0.5) and two benchmarks (LIBERO, Meta-World), revealing complementary effects: V and A respectively reduce FLOPs and latency, while D improves both. Our joint configurations (V,A,D) reduce latency by 79.2% and computation (FLOPs) by 31.8%, while improving mean success rate by 5.6%.
Figures & tables
Fig. 1: (V,A,D) compute axes on a flow-matching VLA. Left: VLM depth V , the action expert depth A and the denoising steps D . Right: multi-objective Pareto curves for two exemplary tasks on Meta-World MT50 [ 4 ] , stick-push and box-close , across SmolVLA and π0.5 . Our joint configuration outperforms the base policy in latency, FLOPs and success rate.
Checkpoint
N
Exits
LR
Warm.
Batch
Steps
HuggingFaceVLA/smolvla_libero
32
{8,16,24}
10−4
1000
8
31,184
lerobot/smolvla_metaworld
16
{4,8,12}
10−4
1000
8
25,600
lerobot/pi05_libero_finetuned_v044
18
{6,9,12,14,16}
10−4
500
8
15,592
tiantianx/pi05_metaworld
18
{6,9,12,14,16}
10−4
1000
8
25,600
TABLE I: Training hyperparameters of the ET for each checkpoint.
Model / Benchmark
Version
(V,A,D)
SR (%)
Latency (s)
GFLOPs
SmolVLA / LIBERO
Base
(32,32,10)
66.6
0.798
556.6
Joint
(16,24,3)
68.2
0.159
406.3
SmolVLA / Meta-World
Base
(16,16,10)
60.7
0.397
318.0
Joint
(8,8,3)
74.9
0.081
173.0
π0.5 / LIBERO
Base
(18,18,10)
87.6
0.537
5000.2
Joint
(14,14,2)
86.7
0.095
3541.0
TABLE II: Base and joint configurations. Joint configurations (V,A,D) outperform the full compute baseline on average in success rate, latency, and FLOPs.
SmolVLA (32 layers)
π0.5 (18 layers)
V
8
16
24
6
9
12
14
16
VLM GFLOPs w/o
86.3
86.3
86.3
3939
3939
3939
3939
3939
VLM GFLOPs w/
21.6
43.1
64.7
1094
1970
2845
3283
3648
Speedup ( × )
4.0
2.0
1.3
3.6
2.0
1.4
1.2
1.1
VLA GFLOPs w/o
559
559
559
5213
5213
5213
5213
5213
VLA GFLOPs w/
495
516
538
2368
3244
4119
4557
4894
TABLE III: KV cache synthesis ablation. Reported gains of the KV cache synthesis mechanism over the standard approach on full LIBERO (300 episodes per suite for SmolVLA, 400 for π0.5 ), for each VLM early exit V .
Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose prohibitive computational burdens during downstream fine-tuning and real-time inference. In this work, we reveal a highly non-trivial architectural characteristic of these continuous control foundation policies (e.g., pi_0, GR00T-N1.5): despite being trained on diverse physical trajectories, they exhibit severe layer-wise representational redundancy. To exploit this, we introduce a structural compression pipeline that is entirely training-free, bypassing the need of existing methods to load full-scale models to learn optimized token reductions or dynamic layer selectors. Instead, using only a single forward pass via Centered Kernel Alignment to identify redundant layer features, we remove twin layers to permanently compress the model depth by up to 50% across both the VLM backbone and the continuous control policy head. Downstream fine-tuning of this streamlined architecture yields a dual acceleration benefit: a 40-50% reduction in training time and up to 30% faster real-time inference, while matching or exceeding full-scale base model performance. We comprehensively validate our method across three simulation benchmarks (LIBERO, RoboCasa, SimplerEnv) and 10 diverse real-world manipulation tasks across 4 unique robotic embodiments. These results prove that advanced VLAs require significantly fewer layers than previously assumed, offering a highly compute-efficient paradigm for scalable robot learning.
Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18
Center for AI Research, VinUniversity · VinRobotics · University of Arkansas +10
Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control frequency hinder real-time robotic manipulation, especially when large vision-language backbones and iterative action heads run at every control step. Existing VLA acceleration methods often optimize individual components or rely on fixed acceleration rules, treating different control steps with largely fixed computation and overlooking the non-uniform reasoning demands of sequential embodied control. Inspired by human motor control, where cognitive and feedback resources concentrate on goal-sensitive stages, we argue that VLA models should learn when to invest full computation and when to reuse prior computation. We propose ElegantVLA, a plug-in phase-adaptive inference framework that accelerates VLA models through intra-model dynamic compute scheduling. ElegantVLA introduces a lightweight scheduler that observes temporal representation similarity, robot-motion cues, and episode progress to jointly allocate computation across the vision encoder, LLM, and action head. For perception-language reasoning, the scheduler selects a five-level Vision-LLM compute mode, from full recomputation to multi-step temporal reuse, based on visual-language representation stability. For action generation, it selects a three-level denoising mode, reusing intermediate denoising states during stable motion while preserving full refinement for goal-sensitive stages. By coordinating these decisions, ElegantVLA offers a general acceleration framework for modern VLA pipelines with explicit action-generation modules, without modifying or retraining the base model. Experiments on GR00T and CogACT achieve up to 2.55x and 3.77x speedup, and on six real-world GR00T tasks ElegantVLA cuts computation by 2.18x while raising control frequency from 13.8 Hz to 26.3 Hz.
Ye Li, Huanan Liu, Kangye Ji +7
Tsinghua University · University of Illinois at Urbana-Champaign
Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes. While various acceleration techniques have been proposed, they often rely on fine-tuning or access to training datasets, which are frequently unavailable due to privacy and proprietary concerns. Moreover, although flow-matching-based VLAs have emerged as efficient alternatives to standard diffusion models, current acceleration efforts largely target VLM inference costs, failing to address the iterative ODE solving process inherent in flow matching inference. To address these limitations, we propose AdaVLA, an online, training-free adaptive framework for fast yet accurate flow-matching-based Vision-Language-Action models. We introduce a novel metric derived from the flow matching trajectory curvature to quantify action generation confidence during inference. This metric enables the dynamic reduction of inference steps and the adaptive adjustment of MLP pruning ratios through an efficiently computed importance evaluation, requiring no access to training data. Experimental results on the LIBERO benchmark using a Jetson AGX Orin device demonstrate that our method achieves 1.87× and 2.24× speedups for π0.5 and X-VLA, respectively, with negligible degradation in success rates. Furthermore, we validate the robustness of our approach on real-world robotic tasks using SmolVLA.
Sunghwan Han, Youngtae Han, Youngmin Yi
Dept. of Artificial Intelligence, Sogang University, Seoul, 04107, Republic of Korea