Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.
Figures & tables
Figure 1 : Layer-wise visualization of attention in VLA.
Figure 2 : Preliminary observations on top-k attention concentration and token pruning strategy comparison.
Figure 3 : Overview of the framework.
Figure 4 : Task-level variation of layer-wise action-to-visual attention mass. The x-axis denotes the layer index, the y-axis denotes the calibration task index, and the z-axis denotes the vision attention mass.
Method
Tokens
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
Speedup ( ↑ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
OpenVLA [ 5 ]
OpenVLA [ 5 ]
256
85.3
85.9
76.1
52.3
74.9
178.46
1.000 ×
100.0
+ FastV [ 13 ]
128
81.7
80.5
73.1
47.9
70.8
139.88
1.276 ×
59.2
+ DivPrune [ 17 ]
128
81.1
78.8
71.7
47.6
69.8
136.10
1.311 ×
62.3
+ SP-VLA [ 21 ]
128
83.5
82.6
74.1
48.7
72.2
127.07
1.404 ×
68.6
Table 1: Comparison under reported/default token settings on LIBERO [ 19 ] . Red bold marks the best result among pruned methods within each backbone.
Method
Tokens
Success Rate (%, ↑ )
PickCoke
MoveNear
Open/Close
Avg.
OpenVLA [ 5 ]
OpenVLA [ 5 ]
256
54.7
57.8
48.1
53.5
+ FastV [ 13 ]
32
19.6
20.1
16.5
18.7
+ SP-VLA [ 21 ]
32
35.4
36.1
35.4
35.6
+ VLA-ADP [ 22 ]
32
39.4
42.5
38.2
40.0
Table 2: Comparison of SAPrune within the OpenVLA [ 5 ] model in the SIMPLER [ 20 ] environment. Red bold marks the best result among pruned methods.
Method
Spatial SR
Object SR
Goal SR
Long SR
Avg. SR
Visual tokens / image
Retention
E2E latency (ms)
π0.5 [ 35 ]
97.4
98.4
97.6
93.0
96.6
256
100%
91.47
+ FastV [ 13 ]
95.8
95.6
94.1
89.3
93.7
128
50%
70.84
+ VLA-Pruner [ 9 ]
95.8
96.2
97.1
91.8
95.2
64
25%
62.73
+ SAPrune
97.2
98.0
96.9
91.5
95.9
32
12.5%
59.68
Table 3: Comparison on π0.5 under reported/default token settings on LIBERO. Red bold marks the best result among pruned methods.
Method
Success Rate ↑
Latency (ms) ↓
FLOPs Ratio (%) ↓
Cube Grasping
Pick&Place
Select-Pick&Place
Multi-Cube Stacking
Average
π0 [ 6 ]
75.0
63.0
54.0
30.0
55.5
78.16
100.0
+ FastV [ 13 ]
69.0
59.0
47.0
23.0
49.5
58.92
60.6
+ DivPrune [ 17 ]
70.0
60.0
47.0
24.0
50.3
61.38
64.8
+ SP-VLA [ 21 ]
72.0
61.0
52.0
25.0
52.5
58.40
63.4
+ VLA-ADP [ 22 ]
74.0
64.0
54.0
26.0
54.5
59.30
47.3
Table 4 : Real-world manipulation results on AgileX Piper with the π0 [ 6 ] backbone. Each task is evaluated over 100 trials. Bold marks the best result among pruned methods.
Method
Pruning location
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
OpenVLA [ 5 ]
–
85.3
85.9
76.1
52.3
74.9
178.46
100.0
Uniform-A
[1, 11, 23]
78.2
77.1
67.5
40.2
65.8
104.40
34.7
Uniform-B
[5, 19, 28]
79.1
78.4
69.6
42.5
67.4
117.85
42.7
Single-layer
[2]
68.2
67.1
58.5
26.7
55.1
95.64
25.5
Random three layers
–
71.5
72.3
64.1
32.9
60.2
119.19
41.1
Table 5 : Ablation on pruning-location strategy under OpenVLA [ 5 ] on LIBERO [ 19 ] . Red bold marks the best result among pruning variants.
Figure 6 : Tasks on LIBERO Benchmark, SIMPLER and Real World Environment.
Method
Tokens
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
Speedup ( ↑ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
OpenVLA [ 5 ]
OpenVLA
256
85.3
85.9
76.1
52.3
74.9
178.46
1.000 ×
100.0
+ FastV [ 13 ]
32
55.9
52.0
51.2
18.1
44.3
102.84
1.735 ×
29.9
+ DivPrune [ 17 ]
32
50.1
49.8
47.8
14.9
40.7
103.76
1.720 ×
33.0
+ SP-VLA [ 21 ]
32
68.0
69.5
63.4
30.0
57.7
101.32
1.761 ×
32.1
Appendix
Table 6: Controlled comparison under a unified 32-token budget on LIBERO [ 19 ] . All applicable pruning methods are evaluated with the same final retained-token budget of 32 visual tokens. Red bold marks the best result among pruned methods within each backbone.
Component
Shallow 256 → 128
Middle 128 → 64
Deep 64 → 32
Total
Candidate ranking / Top- M
0.08
0.05
0.04
0.17
Reliability computation ( Pl , Al , λl )
0.49
0.32
0.17
0.98
Functional signature + greedy selection
1.90
0.82
0.37
3.09
Token indexing / gather
0.25
0.16
0.11
0.52
Algorithm 2 total
2.72
1.35
0.69
4.76
Appendix
Table 7: Wall-clock latency breakdown of Algorithm 2 on OpenVLA. Latency is reported in milliseconds.
Method
Preprocess (ms)
VLA inference (ms)
Action decoding (ms)
Communication (ms)
E2E latency (ms)
E2E frequency (Hz)
π0 [ 6 ]
6.08
78.16
2.46
1.36
88.68
11.28
+ FastV [ 13 ]
6.05
58.92
2.44
1.35
69.37
14.42
+ SP-VLA [ 21 ]
6.04
58.40
2.47
1.37
68.89
14.52
+ VLA-ADP [ 22 ]
6.07
59.30
2.48
1.36
69.83
14.32
+ SAPrune
6.10
51.45
2.46
1.35
61.97
16.14
Appendix
Table 9: Real-robot policy-pipeline latency and E2E frequency on π0 . E2E frequency is computed as fE2E=1000/TE2E . E2E latency also includes 0.61–0.62 ms of orchestration and synchronization overhead. Red bold marks the best result among pruned methods.
Calibration
Layers
Long SR (%)
LIBERO (original)
[2,16,26]
50.8
SIMPLER ( N=30 )
[2,15,26]
49.3
Real-world ( N=60 )
[2,17,26]
51.1
Real-world ( N=30 )
[2,17,27]
47.6
Appendix
Table 10: Calibration transfer on OpenVLA. Calibration is performed on different data sources and evaluated on LIBERO-Long without LIBERO-specific recalibration.
Method
Tokens
Camera
Robot
Language
Light
Background
Noise
Layout
Overall
π0.5 [ 35 ]
256
75.7
77.1
85.9
96.4
94.9
90.1
85.2
85.7
+ SAPrune
32
74.5
76.1
84.8
95.6
93.8
88.6
83.9
84.5
Appendix
Table 11: Success rates (%) under LIBERO-Plus distribution shifts. Overall is computed over all 10,030 task variants, accounting for different category sizes. Tokens denotes visual tokens per image.
Figure 7 : Performance under different final prune ratios on OpenVLA. Success rates on the four LIBERO suites are reported as the final prune ratio increases. Our method consistently preserves stronger performance, especially under aggressive pruning regimes.
Figure 8 : Performance under different final prune ratios on OpenVLA-OFT. Success rates on the four LIBERO suites are reported as the final prune ratio increases. Our method remains the most robust among the compared pruning methods across a wide range of pruning levels.
Figure 9 : Performance under different final prune ratios on π0 . Success rates on the four LIBERO suites are reported as the final prune ratio increases. The results show that our method better maintains policy performance as pruning becomes more aggressive.
Variant
LIBERO Success Rate (%, ↑ )
Latency (ms, ↓ )
FLOPs Ratio (%, ↓ )
Spatial
Object
Goal
Long
Avg.
Full SAPrune
84.9
85.1
75.3
50.8
74.0
110.96
39.5
w/o Attention core
75.9
75.1
65.7
39.0
63.9
111.47
40.2
w/o Functional diversity
78.1
77.6
68.9
43.2
67.0
107.93
39.1
Appendix
Table 12 : Ablation study of token-retention paths on LIBERO [ 19 ] . All variants use the same stage-aware pruning layers and the same 32-token budget; only the retention rule is changed. Red bold marks the best success rate among pruning variants.
Figure 10 : Sensitivity and allocation analysis of SAPrune. (a) The average LIBERO success rate increases with the candidate-pool expansion coefficient γ and becomes stable when γ is moderately large. (b) The layer-wise allocation coefficient λ remains small in shallow layers and increases in middle and deep layers.
Figure 11 : Visualization of SAPrune token pruning results across LIBERO tasks .
Figure 12 : Real-world rollout comparison between the full π0 baseline and SAPrune. We compare baseline and SAPrune rollouts on four real-robot manipulation tasks, including cube grasping, pick-and-place, multi-cube stacking, and select-pick-and-place. Green check marks indicate successful task completion.