Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.
Figures & tables
Fig. 1 : LIBERO-Long success rates vs. computational budget. VLA-ACL achieves a better performance-efficiency trade-off than existing frozen-VLA pruning and caching methods.
Fig. 2 : Overview of VLA-ACL. Modules with a snowflake symbol are kept frozen, modules with a flame symbol receive gradient updates. a) During training, two forward passes of the frozen LLM are performed. On the left, the full-vision-context teacher action is generated. On the right, vision embeddings enter our differentiable soft pruning gate and produce the soft-pruned action. b) During inference, the hard pruning gate only lets through K patches per view with the highest patch scores, obtained by the trained patch scorer. The pruned tokens never enter the VLA language backbone.
Fig. 3 : The soft pruning gate consists of the trainable patch scorer, which outputs a logit for each vision patch. The logits pass through a soft Top- K function, which maps logits to [0,1] scores summing to K . These scores regulate the information content per patch inside the gating module, where patches with low scores are perturbed using embeddings b of a pitch-black image.
Method
Success Rate ↑
TFLOPs
Latency (ms)
Spatial
Object
Goal
Long
Average
OpenVLA-OFT
97.8
97.6
97.6
94.2
96.8
4.013
63.49
+ VLA-Cache
98.3
97.5
98.3
95.4
97.4
3.097
66.79
+ VLA-ADP
99.4
98.0
96.4
91.2
96.3
3.139
58.26
+ SAFE-Pruner
98.0
98.0
96.2
93.4
96.4
1.946
56.61
+ SpecPrune-VLA
97.4
95.8
97.7
93.4
96.1
1.726
60.36
TABLE I : Results of VLA-ACL on the LIBERO suite compared with OpenVLA-OFT and other frozen-VLA pruning and caching methods.
Fig. 4 : Visualization of pruning/caching policies on LIBERO-Spatial. Grayed patches are pruned (VLA-ACL, VLA-ADP) or used as previously cached representations (VLA-Cache). VLA-ADP and VLA-Cache both rely on full-context forwards during inference, while VLA-ACL sustains a much more aggressive pruning schedule at every control step, enabling lower latency while retaining similar performance across all LIBERO suites.
Method
Success Rate ↑
TFLOPs
Latency (ms)
Task 1
Task 2
Task 3
Task 4
Average
OpenVLA-OFT
75
95
85
40
73.75
6.257
91.5
+ VLA-ACL
72.5
95
80
42.5
72.5
2.953
62.1
TABLE II : Results of VLA-ACL at K=64 on real-world tasks compared with OpenVLA-OFT.
Fig. 6 : ℓ1 distance between VLA-ACL ( K=64 ), teacher, and ground-truth (GT) actions over a LIBERO-Long training episode.
Teacher
GT
K
Success Rate ↑
✓
✗
32
95.2
✗
✓
95.1
✓
✓
95.4
✓
✗
64
96.4
✗
✓
95.9
✓
✓
97.0
TABLE III : LIBERO success rates of different training objectives.
Method
K
Pruning layer
Success Rate ↑ / Action Loss ↓
Spatial
Object
Goal
Long
Average
Attention
32
pre
61.8 / 0.127
55.0 / 0.131
59.0 / 0.096
3.2 / 0.122
44.8 / 0.119
Attention + RoPE
pre
96.6 / 0.061
95.0 / 0.065
96.6 / 0.058
69.6 / 0.064
89.5 / 0.062
VLA-Pruner
3
88.1 / 0.064
87.6 / 0.100
84.9 / 0.070
68.8 / 0.052
82.4 / 0.071
VLA-ACL
pre
98.2 / 0.033
98.2 / 0.043
96.6 / 0.047
88.7 / 0.037
95.4 / 0.040
Attention
64
pre
77.4 / 0.110
59.8 / 0.111
76.6 / 0.079
33.2 / 0.096
61.8 / 0.099
TABLE IV : Comparison of selection methods at equal budget. Action loss denotes Lacl with λ=0.5 and μ=0.5 . K denotes per-view visual token budget.
Fig. 7 : ℓ1 between soft-pruned actions of different gating modes vs. hard-pruned actions at different K . Smaller is better.
Method
Success Rate ↑
TFLOPs
Latency (ms)
Spatial
Object
Goal
Long
Average
π0.5
98.8
98.2
98.0
92.4
96.9
2.115
25.9
+ VLA-Cache
94.8
96.0
96.4
88.2
93.9
1.632
-
+ VLA-Pruner
95.8
97.2
95.6
88.0
94.2
1.583
-
+ SAFE-Pruner
97.0
98.8
96.0
90.2
95.5
1.482
-
+ VLA-ACL (K=64)
97.4
98.0
96.6
92.3
96.1
0.701
11.6
TABLE V : Results of VLA-ACL on the LIBERO suite compared with π0.5 and other frozen-VLA pruning and caching methods. Success rates and TFLOPs of the baselines are taken from the results reported by SAFE-Pruner [ 8 ] ; their latency cannot be measured as the code is not public.
Fig. 8 : Our real-world platform: a bimanual AgileX PiPER-X with a main static camera and left and right wrist cameras (RealSense D435).