Beyond Token Importance: Preserving Spatial Scaffolds for Efficient Vision-Language-Action Inference
Organizations: Fudan University · The Hong Kong Polytechnic University
Abstract
Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spatial coverage radius, defined as the largest spatial blind spot induced by the retained token set after pruning. Our analysis reveals a strong correlation between the spatial structure of retained tokens and task success, suggesting that reliable VLA pruning requires preserving not only task-relevant tokens but also the spatial scaffold of the scene. Motivated by this diagnosis, we propose GeoScaffold, a training-free visual token pruning method that partitions each image into spatial regions, allocates inter-region token budgets using task-relevance weights, and selects intra-region scaffold tokens via farthest point sampling to reduce the local coverage radius. On pi 0.5 and LIBERO, GeoScaffold retains only 20% of visual tokens while preserving a 93.2% average success rate, and achieves a 1.78 times prefill speedup over the unpruned baseline.
Figures & tables
| Method | Spatial | Object | Goal | Long | Avg. | prefill (ms) | total (ms) |
|---|---|---|---|---|---|---|---|
| keep tokens per image, pruning rates = 0% | |||||||
| baseline | 98.6 | 98.6 | 98.6 | 92.4 | 27.68 | 61.03 | |
| keep tokens per image, pruning rates = 80% | |||||||
| FastV [ 6 ] | 75.4 | 86.6 | 74.0 | 64.2 | 75.0 | 13.51 | 45.52 |
| SparseVLM [ 26 ] | 39.4 | 80.6 | 66.2 | 52.8 | 59.8 | 13.53 | 45.51 |
| GeoScaffold (ours) | 97.6 | 96.6 | 90.6 | 88.2 | 93.2 | 15.58 ( 1.78) | 47.56 ( 1.28) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Spatial | Object | Goal | Long | Avg. |
|---|---|---|---|---|---|
| keep tokens per image, pruning rates = 0% | |||||
| OpenVLA-OFT baseline | 95.8 | 98.7 | 96.3 | 90.7 | 95.4 |
| keep tokens per image, pruning rates = 75% | |||||
| VLA-Cache [ 21 ] | 85.6 | 85.3 | 84.1 | 80.3 | 83.8 |
| FastV [ 6 ] | 92.6 | 89.8 | 96.2 | 83.2 | 90.5 |
| GeoScaffold (ours) | 90.4 | 86.0 | 96.8 | 90.0 | 90.8 |
| Method | image_encode | prefill | denoise | total | prefill | total |
|---|---|---|---|---|---|---|
| Baseline (no pruning) | 11.28 | 27.68 | 22.08 | 61.03 | ||
| GeoScaffold | 11.19 | 15.58 | 20.80 | 47.56 | ||
| GeoScaffold | 11.20 | 14.22 | 20.68 | 46.10 | ||
| GeoScaffold | 11.20 | 14.11 | 20.63 | 45.92 |
| Asset | License or Terms | Usage in This Work |
|---|---|---|
| / OpenPI | Apache License 2.0 | Primary VLA model for inference-time pruning evaluation; no redistribution of model weights. |
| OpenVLA-OFT | MIT License | Cross-architecture evaluation of the proposed pruning principle; no redistribution of model weights. |
| LIBERO | MIT License | Simulation benchmark and evaluation task suites. |
| FastV | License not specified | Algorithmic baseline for visual token pruning; no FastV code is redistributed. |
| SparseVLM | Apache License 2.0 | Visual-token pruning baseline. |
| VLA-Cache | Apache License 2.0 | Training-free VLA acceleration baseline. |