Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.
Figures & tables
Fig. 1 : LIBERO-Long success rates vs. computational budget. VLA-ACL achieves a better performance-efficiency trade-off than existing frozen-VLA pruning and caching methods.
Fig. 2 : Overview of VLA-ACL. Modules with a snowflake symbol are kept frozen, modules with a flame symbol receive gradient updates. a) During training, two forward passes of the frozen LLM are performed. On the left, the full-vision-context teacher action is generated. On the right, vision embeddings enter our differentiable soft pruning gate and produce the soft-pruned action. b) During inference, the hard pruning gate only lets through K patches per view with the highest patch scores, obtained by the trained patch scorer. The pruned tokens never enter the VLA language backbone.
Fig. 3 : The soft pruning gate consists of the trainable patch scorer, which outputs a logit for each vision patch. The logits pass through a soft Top- K function, which maps logits to [0,1] scores summing to K . These scores regulate the information content per patch inside the gating module, where patches with low scores are perturbed using embeddings b of a pitch-black image.
Method
Success Rate ↑
TFLOPs
Latency (ms)
Spatial
Object
Goal
Long
Average
OpenVLA-OFT
97.8
97.6
97.6
94.2
96.8
4.013
63.49
+ VLA-Cache
98.3
97.5
98.3
95.4
97.4
3.097
66.79
+ VLA-ADP
99.4
98.0
96.4
91.2
96.3
3.139
58.26
+ SAFE-Pruner
98.0
98.0
96.2
93.4
96.4
1.946
56.61
+ SpecPrune-VLA
97.4
95.8
97.7
93.4
96.1
1.726
60.36
TABLE I : Results of VLA-ACL on the LIBERO suite compared with OpenVLA-OFT and other frozen-VLA pruning and caching methods.
Fig. 4 : Visualization of pruning/caching policies on LIBERO-Spatial. Grayed patches are pruned (VLA-ACL, VLA-ADP) or used as previously cached representations (VLA-Cache). VLA-ADP and VLA-Cache both rely on full-context forwards during inference, while VLA-ACL sustains a much more aggressive pruning schedule at every control step, enabling lower latency while retaining similar performance across all LIBERO suites.
Method
Success Rate ↑
TFLOPs
Latency (ms)
Task 1
Task 2
Task 3
Task 4
Average
OpenVLA-OFT
75
95
85
40
73.75
6.257
91.5
+ VLA-ACL
72.5
95
80
42.5
72.5
2.953
62.1
TABLE II : Results of VLA-ACL at K=64 on real-world tasks compared with OpenVLA-OFT.
Fig. 6 : ℓ1 distance between VLA-ACL ( K=64 ), teacher, and ground-truth (GT) actions over a LIBERO-Long training episode.
Teacher
GT
K
Success Rate ↑
✓
✗
32
95.2
✗
✓
95.1
✓
✓
95.4
✓
✗
64
96.4
✗
✓
95.9
✓
✓
97.0
TABLE III : LIBERO success rates of different training objectives.
Method
K
Pruning layer
Success Rate ↑ / Action Loss ↓
Spatial
Object
Goal
Long
Average
Attention
32
pre
61.8 / 0.127
55.0 / 0.131
59.0 / 0.096
3.2 / 0.122
44.8 / 0.119
Attention + RoPE
pre
96.6 / 0.061
95.0 / 0.065
96.6 / 0.058
69.6 / 0.064
89.5 / 0.062
VLA-Pruner
3
88.1 / 0.064
87.6 / 0.100
84.9 / 0.070
68.8 / 0.052
82.4 / 0.071
VLA-ACL
pre
98.2 / 0.033
98.2 / 0.043
96.6 / 0.047
88.7 / 0.037
95.4 / 0.040
Attention
64
pre
77.4 / 0.110
59.8 / 0.111
76.6 / 0.079
33.2 / 0.096
61.8 / 0.099
TABLE IV : Comparison of selection methods at equal budget. Action loss denotes Lacl with λ=0.5 and μ=0.5 . K denotes per-view visual token budget.
Fig. 7 : ℓ1 between soft-pruned actions of different gating modes vs. hard-pruned actions at different K . Smaller is better.
Method
Success Rate ↑
TFLOPs
Latency (ms)
Spatial
Object
Goal
Long
Average
π0.5
98.8
98.2
98.0
92.4
96.9
2.115
25.9
+ VLA-Cache
94.8
96.0
96.4
88.2
93.9
1.632
-
+ VLA-Pruner
95.8
97.2
95.6
88.0
94.2
1.583
-
+ SAFE-Pruner
97.0
98.8
96.0
90.2
95.5
1.482
-
+ VLA-ACL (K=64)
97.4
98.0
96.6
92.3
96.1
0.701
11.6
TABLE V : Results of VLA-ACL on the LIBERO suite compared with π0.5 and other frozen-VLA pruning and caching methods. Success rates and TFLOPs of the baselines are taken from the results reported by SAFE-Pruner [ 8 ] ; their latency cannot be measured as the code is not public.
Fig. 8 : Our real-world platform: a bimanual AgileX PiPER-X with a main static camera and left and right wrist cameras (RealSense D435).
Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an adaptive subtask division strategy to detect abrupt attention shifts, thereby improving forecasting accuracy and pruning reliability. Extensive experiments in simulation and real-world settings demonstrate that our method achieves up to 1.89x speedup with a minimal degradation in success rate of less than 1.7%, while outperforming state-of-the-art methods by up to 1.9%.
Shilin Ma, Chubin Zhang, Changyuan Wang +6
Tsinghua Shenzhen International Graduate School, Tsinghua University · GigaAI
Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.
Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spatial coverage radius, defined as the largest spatial blind spot induced by the retained token set after pruning. Our analysis reveals a strong correlation between the spatial structure of retained tokens and task success, suggesting that reliable VLA pruning requires preserving not only task-relevant tokens but also the spatial scaffold of the scene. Motivated by this diagnosis, we propose GeoScaffold, a training-free visual token pruning method that partitions each image into spatial regions, allocates inter-region token budgets using task-relevance weights, and selects intra-region scaffold tokens via farthest point sampling to reduce the local coverage radius. On pi 0.5 and LIBERO, GeoScaffold retains only 20% of visual tokens while preserving a 93.2% average success rate, and achieves a 1.78 times prefill speedup over the unpruned baseline.
Jiayu Chen, Shuyong Gao, Jingkai Jia +6
Fudan University · The Hong Kong Polytechnic University