Organizations: College of Engineering, Southern University of Science and Technology (SUSTech), Shenzhen 518055, China · Master of Science in Data Science and Machine Learning, Department of Mathematics, Faculty of Science, National University of Singapore, Singapore
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.
Figures & tables
Fig. 1: Two bottlenecks of existing VLA token caching. (Left) Token-level spatial misalignment: indiscriminate attention-head aggregation introduces task-irrelevant responses and degrades visual grounding. (Right) Layer-level depth mismatch: reusing representations at unstable layers can propagate noisy cross-modal features, whereas delaying reuse after representations stabilize incurs redundant computation.
Fig. 2: Principle of KV cache reuse in VLA inference. At timestep T , visual inputs are decoupled into static (blue) and dynamic (yellow) tokens. By reusing historical caches for invariant instructions and static tokens, the LLM allocates computation exclusively to new dynamic tokens, significantly accelerating autoregressive decoding.
Fig. 3: Overview of the proposed framework.
Method
Spatial (%)
Object (%)
Goal (%)
Long (%)
Avg Success Rate (%) ↑
FLOPs (T) ↓
Control Freq. (Hz) ↑
Latency (ms) ↓
OpenVLA-OFT
100% Tokens
Vanilla
99.00
98.00
97.20
94.40
97.15 (100.0%)
4.0139 (1.00 × )
10.43 (1.00 × )
95.87 (1.00 × )
Retain 50% Tokens
FastV
95.40
98.60
96.80
93.40
96.05 (98.9%)
2.4134 ( 1.66 × )
11.84 ( 1.13 × )
84.48 ( 1.13 × )
SparseVLM
96.20
88.40
96.20
75.80
89.15 (91.8%)
2.6809 ( 1.50 × )
11.71 ( 1.12 × )
85.41 ( 1.12 × )
TABLE I: Comparison of latency and performance between OpenVLA-OFT and BitVLA evaluated on the LIBERO Benchmark. The task success rate retention ratios (%) and speedup multipliers ( × ) are calculated relative to the Vanilla baseline of each respective model at 100% tokens.
Method
Spatial (%)
Object (%)
Goal (%)
Long (%)
Avg Success Rate (%) ↑
FLOPs (T) ↓
Control Freq. (Hz) ↑
Latency (ms) ↓
π0.5
100% Tokens
Vanilla
98.00
97.20
97.00
93.00
96.30 (100.0%)
2.1153 (1.00 × )
8.14 (1.00 × )
122.85 (1.00 × )
Retain 50% Tokens
FastV
57.00
67.20
60.80
59.40
61.10 (63.4%)
1.6061 ( 1.32 × )
8.90 ( 1.09 × )
112.34 ( 1.09 × )
SparseVLM
78.20
73.00
86.40
43.60
70.30 (73.0%)
1.5429 ( 1.37 × )
8.82 ( 1.08 × )
113.38 ( 1.08 × )
TABLE II: Comparison of latency and performance between π0.5 and VLA-Adapter evaluated on the LIBERO Benchmark.
Method
Task 1 (%)
Task 2 (%)
Task 3 (%)
Task 4 (%)
Task 5 (%)
Avg Success Rate (%) ↑
FLOPs (T) ↓
Control Freq. (Hz) ↑
Latency (ms) ↓
VLA-Adapter
100% Tokens
Vanilla
98.50
94.60
89.20
84.90
78.90
89.22 (100.0%)
0.2597 (1.00 × )
21.45 (1.00 × )
46.62 (1.00 × )
Retain 50% Tokens
VLA-Cache
95.60
91.00
84.30
76.80
69.40
83.42 (93.5%)
0.1831 ( 1.42 × )
22.72 ( 1.06 × )
44.01 ( 1.06 × )
TVCache (Ours)
97.10
92.50
87.90
81.10
73.10
86.34 (96.8%)
0.1751 ( 1.48 × )
22.73 ( 1.06 × )
43.99 ( 1.06 × )
TABLE III: Comparison of performance for VLA-Adapter evaluated on the CALVIN Benchmark.
Task Name
Description
SetKettle
Pick up the kettle and place it on the plate.
PackPear
Pick up the pear and place it inside the basket.
HangCup
Hang the cup on the wooden mug tree.
PotDuck
Put the duck into the pot and cover it with the lid.
TABLE IV: Descriptions of the evaluation tasks.
TVCache Ret.
Patch sim.
Head filtering
Token selection
Reuse-layer selection
Bookkeeping
Total (ms)
50%
0.742
0.746
0.322
1.012
0.066
2.888
25%
0.651
0.747
0.316
1.052
0.073
2.839
12.5%
0.782
0.708
0.304
0.968
0.075
2.837
TABLE V: Runtime overhead breakdown of TVCache under different token retention ratios.
Fig. 4: Real-world robot evaluation across four manipulation tasks. The figure shows the physical task scenarios and the corresponding success rates of TVCache and VLA-Cache under different token retention ratios.
Method
Spatial
Object
Goal
Long
Avg. ↑
OpenVLA-OFT
100% Tokens
Vanilla
99.00
98.00
97.20
94.40
97.15
Retain 50% Tokens
VLA-Cache
98.20
98.20
97.40
93.40
96.80
+ Head
98.60 ( ↑0.40 )
98.40 ( ↑0.20 )
97.80 ( ↑0.40 )
93.80 ( ↑0.40 )
97.15 ( ↑0.35 )
TABLE VI: Ablation study on OpenVLA-OFT and BitVLA.
Fig. 5: Qualitative comparison of cross-modal attention. TVCache filters degraded heads, reducing background noise and spatial misalignment for more task-relevant visual grounding.
Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-value(KV) representations for visual tokens that barely change across neighboring frames. Recent work such as VLA-Cache reduces that cost by reusing KV states for visually static patches, but its policy relies only on observation-space heuristics and does not account for the model's own uncertainty. We propose Gated VLA-Cache, a lightweight, training-free extension that augments visual-similarity caching with neural introspection. The method monitors the logit margin between the top two predicted action tokens, a zero-cost confidence signal available during decoding. When the margin drops below a threshold, the cache is invalidated and a full recompute is triggered. Evaluated on four LIBERO benchmark suites with both OpenVLA and OpenVLA-OFT, Gated VLA-Cache improves reliability when blind caching hurts. On LIBERO-Goal and LIBERO-Long, it recovers over 100% of the lost accuracy while retaining 80% of the compute savings.
Zhijie Wu, Kento Kawaharazuka, Kei Okada
Graduate School of Information Science and Technology, The University of Tokyo, 7-3-1 Hongo, Bunkyo-ku, Tokyo, 113-8656, Japan
Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based VLA models have shown remarkable success due to their capability to generate precise and smooth action sequences and capture multimodal distributions. However, the iterative denoising process in the action head acts as a major computational bottleneck, posing a critical challenge for real-time deployment. To address this challenge, we propose ActionCache, a plug-and-play external cache that opportunistically reuses past intermediate actions to warm-start generations from the vicinity of target actions, thereby drastically reducing the inference latency. Specifically, ActionCache stores the intermediate actions with compact multimodal keys, which enables retrieval from similar past contexts across different episodes or even different tasks. Experimental results in simulation and real-world environments demonstrate that ActionCache maintains high task success rates in a low-latency regime, achieving inference acceleration of up to 11.75× and 34.43× for representative flow-based VLA models, π0.5 and GR00T-N1.6, respectively.
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.