Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spatial coverage radius, defined as the largest spatial blind spot induced by the retained token set after pruning. Our analysis reveals a strong correlation between the spatial structure of retained tokens and task success, suggesting that reliable VLA pruning requires preserving not only task-relevant tokens but also the spatial scaffold of the scene. Motivated by this diagnosis, we propose GeoScaffold, a training-free visual token pruning method that partitions each image into spatial regions, allocates inter-region token budgets using task-relevance weights, and selects intra-region scaffold tokens via farthest point sampling to reduce the local coverage radius. On pi 0.5 and LIBERO, GeoScaffold retains only 20% of visual tokens while preserving a 93.2% average success rate, and achieves a 1.78 times prefill speedup over the unpruned baseline.
Figures & tables
Figure 1: (a) Structured stride pruning can outperform random and attention-based pruning at moderate rates, yet fails abruptly at a nearby pruning rate. (b) Different pruning policies induce distinct spatial layouts of retained tokens, exposing large uncovered regions under aggressive pruning. (c) The coverage radius, measuring the largest spatial blind spot, shows a strong positive correlation with pruning-induced error across pruning policies and rates.
Figure 2: Overview of GeoScaffold. Three steps implement the two-level decomposition: Step 1: Region Scoring , aggregate token attention into region weights; Step 2: Budget Allocation , assign budgets with a one-token spatial floor per region; Step 3: Scaffold Selection , use center-seeded farthest point sampling to minimize local coverage radius.
Method
Spatial
Object
Goal
Long
Avg.
prefill (ms) ↓
total (ms) ↓
keep 256/256 tokens per image, pruning rates = 0%
π0.5 baseline
98.6
98.6
98.6
92.4
97.1
27.68
61.03
keep 51/256 tokens per image, pruning rates = 80%
FastV [ 6 ]
75.4
86.6
74.0
64.2
75.0
13.51
45.52
SparseVLM [ 26 ]
39.4
80.6
66.2
52.8
59.8
13.53
45.51
GeoScaffold (ours)
97.6
96.6
90.6
88.2
93.2
15.58 ( × 1.78)
47.56 ( × 1.28)
Table 1: Comparison of different VLA acceleration methods on the LIBERO benchmark.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Spatial
Object
Goal
Long
Avg.
keep 512/512 tokens per image, pruning rates = 0%
OpenVLA-OFT baseline
95.8
98.7
96.3
90.7
95.4
keep 128/512 tokens per image, pruning rates = 75%
VLA-Cache [ 21 ]
85.6
85.3
84.1
80.3
83.8
FastV [ 6 ]
92.6
89.8
96.2
83.2
90.5
GeoScaffold (ours)
90.4
86.0
96.8
90.0
90.8
Appendix
Table 3: Performance on LIBERO at OpenVLA-OFT [ 9 ] across two aggressive pruning rates. Bold = column best within each block.
Method
image_encode
prefill
denoise
total
prefill ↑
total ↑
Baseline (no pruning)
11.28
27.68
22.08
61.03
1.00×
1.00×
GeoScaffold R=0.80
11.19
15.58
20.80
47.56
1.78×
1.28×
GeoScaffold R=0.875
11.20
14.22
20.68
46.10
1.95×
1.32×
GeoScaffold R=0.90
11.20
14.11
20.63
45.92
1.96×
1.33×
Appendix
Table 4: Per-stage single-step latency (ms) of π0.5 on RTX 4090 under different pruning rates.
Asset
License or Terms
Usage in This Work
π0.5 / OpenPI
Apache License 2.0
Primary VLA model for inference-time pruning evaluation; no redistribution of model weights.
OpenVLA-OFT
MIT License
Cross-architecture evaluation of the proposed pruning principle; no redistribution of model weights.
LIBERO
MIT License
Simulation benchmark and evaluation task suites.
FastV
License not specified
Algorithmic baseline for visual token pruning; no FastV code is redistributed.
SparseVLM
Apache License 2.0
Visual-token pruning baseline.
VLA-Cache
Apache License 2.0
Training-free VLA acceleration baseline.
Appendix
Table 5: Existing assets used in this work and their licenses or terms of use.
Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an adaptive subtask division strategy to detect abrupt attention shifts, thereby improving forecasting accuracy and pruning reliability. Extensive experiments in simulation and real-world settings demonstrate that our method achieves up to 1.89x speedup with a minimal degradation in success rate of less than 1.7%, while outperforming state-of-the-art methods by up to 1.9%.
Shilin Ma, Chubin Zhang, Changyuan Wang +6
Tsinghua Shenzhen International Graduate School, Tsinghua University · GigaAI
Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning methods suffer from a fundamental trade-off: aggressive compression using pruning inevitably discards critical geometric details like contact points, leading to severe performance degradation. This forces a compromise, limiting the achievable compression rate and thus the potential speedup. We argue that breaking this trade-off requires rethinking compression as a geometry-aware, continuous token resampling in the vision encoder. To this end, we propose the Differentiable Grid Sampler (GridS), a plug-and-play module that performs task-aware, continuous resampling of visual tokens in VLA. By adaptively predicting a minimal set of salient coordinates and extracting features via differentiable interpolation, GridS preserves essential spatial information while achieving drastic compression (with fewer than 10% original visual tokens). Experiments on both LIBERO benchmark and a real robotic platform demonstrate that validating the lowest feasible visual token count reported to date, GridS achieves a 76% reduction in FLOPs with no degradation in the success rate. The code is available at https://github.com/Fediory/Grid-Sampler.
Yixu Feng, Zinan Zhao, Yanxiang Ma +4
The University of Sydney · City University of Hong Kong · StellarEdge Robotics
Vision-Language Navigation (VLN) enables robots to follow natural-language instructions in visually grounded environments, serving as a key capability for embodied robotic systems. Recent Vision-Language-Action (VLA) models have demonstrated strong navigation performance, but their high computational cost introduces latency that limits real-time deployment. We propose a training-free spatio-temporal vision token pruning framework tailored to VLA-based VLN. We apply spatial token selection to the current view, alongside spatio-temporal compression for historical memories, enabling efficient long-horizon inference while reducing redundant computation. Leveraging attention-based token importance and query-guided spatio-temporal filtering, the proposed approach preserves navigation-relevant information without retraining or modifying pretrained models, allowing plug-and-play integration into existing VLA systems. Through experiments on standard VLN benchmarks, we confirm that our method significantly outperforms existing pruning strategies. It successfully preserves superior navigation accuracy under extreme pruning scenarios, all while maintaining the highly competitive inference efficiency. Real-world deployment on a Unitree Go2 quadruped robot further validates reliable and low-latency instruction-following navigation under practical robotic constraints. We hope this work helps bridge the gap between large-scale multimodal modeling and efficient, real-time embodied deployment in robotic navigation systems. Project Page: https://wqtwjt1996.github.io/publications/2026-vln.html
Qitong Wang, Yijun Liang, Ming Li +2
Department of Computer and Information Sciences at the University of Delaware, Newark, DE 19711, USA · University of Maryland’s Department of Computer Science, College Park, MD 20742, USA · Mohamed bin Zayed University of Artificial Intelligence, Masdar City, Abu Dhabi 7909, UAE