Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.
Figures & tables
Figure 1 : Layer-wise visualization of attention in VLA.
Figure 2 : Preliminary observations on top-k attention concentration and token pruning strategy comparison.
Figure 3 : Overview of the framework.
Figure 4 : Task-level variation of layer-wise action-to-visual attention mass. The x-axis denotes the layer index, the y-axis denotes the calibration task index, and the z-axis denotes the vision attention mass.
Method
Tokens
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
Speedup ( ↑ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
OpenVLA [ 5 ]
OpenVLA [ 5 ]
256
85.3
85.9
76.1
52.3
74.9
178.46
1.000 ×
100.0
+ FastV [ 13 ]
128
81.7
80.5
73.1
47.9
70.8
139.88
1.276 ×
59.2
+ DivPrune [ 17 ]
128
81.1
78.8
71.7
47.6
69.8
136.10
1.311 ×
62.3
+ SP-VLA [ 21 ]
128
83.5
82.6
74.1
48.7
72.2
127.07
1.404 ×
68.6
Table 1: Comparison under reported/default token settings on LIBERO [ 19 ] . Red bold marks the best result among pruned methods within each backbone.
Method
Tokens
Success Rate (%, ↑ )
PickCoke
MoveNear
Open/Close
Avg.
OpenVLA [ 5 ]
OpenVLA [ 5 ]
256
54.7
57.8
48.1
53.5
+ FastV [ 13 ]
32
19.6
20.1
16.5
18.7
+ SP-VLA [ 21 ]
32
35.4
36.1
35.4
35.6
+ VLA-ADP [ 22 ]
32
39.4
42.5
38.2
40.0
Table 2: Comparison of SAPrune within the OpenVLA [ 5 ] model in the SIMPLER [ 20 ] environment. Red bold marks the best result among pruned methods.
Method
Spatial SR
Object SR
Goal SR
Long SR
Avg. SR
Visual tokens / image
Retention
E2E latency (ms)
π0.5 [ 35 ]
97.4
98.4
97.6
93.0
96.6
256
100%
91.47
+ FastV [ 13 ]
95.8
95.6
94.1
89.3
93.7
128
50%
70.84
+ VLA-Pruner [ 9 ]
95.8
96.2
97.1
91.8
95.2
64
25%
62.73
+ SAPrune
97.2
98.0
96.9
91.5
95.9
32
12.5%
59.68
Table 3: Comparison on π0.5 under reported/default token settings on LIBERO. Red bold marks the best result among pruned methods.
Method
Success Rate ↑
Latency (ms) ↓
FLOPs Ratio (%) ↓
Cube Grasping
Pick&Place
Select-Pick&Place
Multi-Cube Stacking
Average
π0 [ 6 ]
75.0
63.0
54.0
30.0
55.5
78.16
100.0
+ FastV [ 13 ]
69.0
59.0
47.0
23.0
49.5
58.92
60.6
+ DivPrune [ 17 ]
70.0
60.0
47.0
24.0
50.3
61.38
64.8
+ SP-VLA [ 21 ]
72.0
61.0
52.0
25.0
52.5
58.40
63.4
+ VLA-ADP [ 22 ]
74.0
64.0
54.0
26.0
54.5
59.30
47.3
Table 4 : Real-world manipulation results on AgileX Piper with the π0 [ 6 ] backbone. Each task is evaluated over 100 trials. Bold marks the best result among pruned methods.
Method
Pruning location
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
OpenVLA [ 5 ]
–
85.3
85.9
76.1
52.3
74.9
178.46
100.0
Uniform-A
[1, 11, 23]
78.2
77.1
67.5
40.2
65.8
104.40
34.7
Uniform-B
[5, 19, 28]
79.1
78.4
69.6
42.5
67.4
117.85
42.7
Single-layer
[2]
68.2
67.1
58.5
26.7
55.1
95.64
25.5
Random three layers
–
71.5
72.3
64.1
32.9
60.2
119.19
41.1
Table 5 : Ablation on pruning-location strategy under OpenVLA [ 5 ] on LIBERO [ 19 ] . Red bold marks the best result among pruning variants.
Figure 6 : Tasks on LIBERO Benchmark, SIMPLER and Real World Environment.
Method
Tokens
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
Speedup ( ↑ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
OpenVLA [ 5 ]
OpenVLA
256
85.3
85.9
76.1
52.3
74.9
178.46
1.000 ×
100.0
+ FastV [ 13 ]
32
55.9
52.0
51.2
18.1
44.3
102.84
1.735 ×
29.9
+ DivPrune [ 17 ]
32
50.1
49.8
47.8
14.9
40.7
103.76
1.720 ×
33.0
+ SP-VLA [ 21 ]
32
68.0
69.5
63.4
30.0
57.7
101.32
1.761 ×
32.1
Appendix
Table 6: Controlled comparison under a unified 32-token budget on LIBERO [ 19 ] . All applicable pruning methods are evaluated with the same final retained-token budget of 32 visual tokens. Red bold marks the best result among pruned methods within each backbone.
Component
Shallow 256 → 128
Middle 128 → 64
Deep 64 → 32
Total
Candidate ranking / Top- M
0.08
0.05
0.04
0.17
Reliability computation ( Pl , Al , λl )
0.49
0.32
0.17
0.98
Functional signature + greedy selection
1.90
0.82
0.37
3.09
Token indexing / gather
0.25
0.16
0.11
0.52
Algorithm 2 total
2.72
1.35
0.69
4.76
Appendix
Table 7: Wall-clock latency breakdown of Algorithm 2 on OpenVLA. Latency is reported in milliseconds.
Method
Preprocess (ms)
VLA inference (ms)
Action decoding (ms)
Communication (ms)
E2E latency (ms)
E2E frequency (Hz)
π0 [ 6 ]
6.08
78.16
2.46
1.36
88.68
11.28
+ FastV [ 13 ]
6.05
58.92
2.44
1.35
69.37
14.42
+ SP-VLA [ 21 ]
6.04
58.40
2.47
1.37
68.89
14.52
+ VLA-ADP [ 22 ]
6.07
59.30
2.48
1.36
69.83
14.32
+ SAPrune
6.10
51.45
2.46
1.35
61.97
16.14
Appendix
Table 9: Real-robot policy-pipeline latency and E2E frequency on π0 . E2E frequency is computed as fE2E=1000/TE2E . E2E latency also includes 0.61–0.62 ms of orchestration and synchronization overhead. Red bold marks the best result among pruned methods.
Calibration
Layers
Long SR (%)
LIBERO (original)
[2,16,26]
50.8
SIMPLER ( N=30 )
[2,15,26]
49.3
Real-world ( N=60 )
[2,17,26]
51.1
Real-world ( N=30 )
[2,17,27]
47.6
Appendix
Table 10: Calibration transfer on OpenVLA. Calibration is performed on different data sources and evaluated on LIBERO-Long without LIBERO-specific recalibration.
Method
Tokens
Camera
Robot
Language
Light
Background
Noise
Layout
Overall
π0.5 [ 35 ]
256
75.7
77.1
85.9
96.4
94.9
90.1
85.2
85.7
+ SAPrune
32
74.5
76.1
84.8
95.6
93.8
88.6
83.9
84.5
Appendix
Table 11: Success rates (%) under LIBERO-Plus distribution shifts. Overall is computed over all 10,030 task variants, accounting for different category sizes. Tokens denotes visual tokens per image.
Figure 7 : Performance under different final prune ratios on OpenVLA. Success rates on the four LIBERO suites are reported as the final prune ratio increases. Our method consistently preserves stronger performance, especially under aggressive pruning regimes.
Figure 8 : Performance under different final prune ratios on OpenVLA-OFT. Success rates on the four LIBERO suites are reported as the final prune ratio increases. Our method remains the most robust among the compared pruning methods across a wide range of pruning levels.
Figure 9 : Performance under different final prune ratios on π0 . Success rates on the four LIBERO suites are reported as the final prune ratio increases. The results show that our method better maintains policy performance as pruning becomes more aggressive.
Variant
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
Full SAPrune
84.9
85.1
75.3
50.8
74.0
110.96
39.5
w/o Attention core
75.9
75.1
65.7
39.0
63.9
111.47
40.2
w/o Functional diversity
78.1
77.6
68.9
43.2
67.0
107.93
39.1
Appendix
Table 12 : Ablation study of token-retention paths on LIBERO [ 19 ] . All variants use the same stage-aware pruning layers and the same 32-token budget; only the retention rule is changed. Red bold marks the best success rate among pruning variants.
Figure 10 : Sensitivity and allocation analysis of SAPrune. (a) The average LIBERO success rate increases with the candidate-pool expansion coefficient γ and becomes stable when γ is moderately large. (b) The layer-wise allocation coefficient λ remains small in shallow layers and increases in middle and deep layers.
Figure 11 : Visualization of SAPrune token pruning results across LIBERO tasks .
Figure 12 : Real-world rollout comparison between the full π0 baseline and SAPrune. We compare baseline and SAPrune rollouts on four real-robot manipulation tasks, including cube grasping, pick-and-place, multi-cube stacking, and select-pick-and-place. Green check marks indicate successful task completion.
Real-time inference of vision-language-action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across execution steps. Based on this observation, we design a forward-looking strategy to forecast the token saliency in deep layers, which prevents the premature removal of critical tokens and leads to more stable acceleration. We further introduce an adaptive subtask division strategy to detect abrupt attention shifts, thereby improving forecasting accuracy and pruning reliability. Extensive experiments in simulation and real-world settings demonstrate that our method achieves up to 1.89x speedup with a minimal degradation in success rate of less than 1.7%, while outperforming state-of-the-art methods by up to 1.9%.
Shilin Ma, Chubin Zhang, Changyuan Wang +6
Tsinghua Shenzhen International Graduate School, Tsinghua University · GigaAI
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.
Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spatial coverage radius, defined as the largest spatial blind spot induced by the retained token set after pruning. Our analysis reveals a strong correlation between the spatial structure of retained tokens and task success, suggesting that reliable VLA pruning requires preserving not only task-relevant tokens but also the spatial scaffold of the scene. Motivated by this diagnosis, we propose GeoScaffold, a training-free visual token pruning method that partitions each image into spatial regions, allocates inter-region token budgets using task-relevance weights, and selects intra-region scaffold tokens via farthest point sampling to reduce the local coverage radius. On pi 0.5 and LIBERO, GeoScaffold retains only 20% of visual tokens while preserving a 93.2% average success rate, and achieves a 1.78 times prefill speedup over the unpruned baseline.
Jiayu Chen, Shuyong Gao, Jingkai Jia +6
Fudan University · The Hong Kong Polytechnic University