Organizations: College of Engineering, Southern University of Science and Technology (SUSTech), Shenzhen 518055, China · Master of Science in Data Science and Machine Learning, Department of Mathematics, Faculty of Science, National University of Singapore, Singapore
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.
Figures & tables
Fig. 1: Two bottlenecks of existing VLA token caching. (Left) Token-level spatial misalignment: indiscriminate attention-head aggregation introduces task-irrelevant responses and degrades visual grounding. (Right) Layer-level depth mismatch: reusing representations at unstable layers can propagate noisy cross-modal features, whereas delaying reuse after representations stabilize incurs redundant computation.
Fig. 2: Principle of KV cache reuse in VLA inference. At timestep T , visual inputs are decoupled into static (blue) and dynamic (yellow) tokens. By reusing historical caches for invariant instructions and static tokens, the LLM allocates computation exclusively to new dynamic tokens, significantly accelerating autoregressive decoding.
Fig. 3: Overview of the proposed framework.
Method
Spatial (%)
Object (%)
Goal (%)
Long (%)
Avg Success Rate (%) ↑
FLOPs (T) ↓
Control Freq. (Hz) ↑
Latency (ms) ↓
OpenVLA-OFT
100% Tokens
Vanilla
99.00
98.00
97.20
94.40
97.15 (100.0%)
4.0139 (1.00 × )
10.43 (1.00 × )
95.87 (1.00 × )
Retain 50% Tokens
FastV
95.40
98.60
96.80
93.40
96.05 (98.9%)
2.4134 ( 1.66 × )
11.84 ( 1.13 × )
84.48 ( 1.13 × )
SparseVLM
96.20
88.40
96.20
75.80
89.15 (91.8%)
2.6809 ( 1.50 × )
11.71 ( 1.12 × )
85.41 ( 1.12 × )
TABLE I: Comparison of latency and performance between OpenVLA-OFT and BitVLA evaluated on the LIBERO Benchmark. The task success rate retention ratios (%) and speedup multipliers ( × ) are calculated relative to the Vanilla baseline of each respective model at 100% tokens.
Method
Spatial (%)
Object (%)
Goal (%)
Long (%)
Avg Success Rate (%) ↑
FLOPs (T) ↓
Control Freq. (Hz) ↑
Latency (ms) ↓
π0.5
100% Tokens
Vanilla
98.00
97.20
97.00
93.00
96.30 (100.0%)
2.1153 (1.00 × )
8.14 (1.00 × )
122.85 (1.00 × )
Retain 50% Tokens
FastV
57.00
67.20
60.80
59.40
61.10 (63.4%)
1.6061 ( 1.32 × )
8.90 ( 1.09 × )
112.34 ( 1.09 × )
SparseVLM
78.20
73.00
86.40
43.60
70.30 (73.0%)
1.5429 ( 1.37 × )
8.82 ( 1.08 × )
113.38 ( 1.08 × )
TABLE II: Comparison of latency and performance between π0.5 and VLA-Adapter evaluated on the LIBERO Benchmark.
Method
Task 1 (%)
Task 2 (%)
Task 3 (%)
Task 4 (%)
Task 5 (%)
Avg Success Rate (%) ↑
FLOPs (T) ↓
Control Freq. (Hz) ↑
Latency (ms) ↓
VLA-Adapter
100% Tokens
Vanilla
98.50
94.60
89.20
84.90
78.90
89.22 (100.0%)
0.2597 (1.00 × )
21.45 (1.00 × )
46.62 (1.00 × )
Retain 50% Tokens
VLA-Cache
95.60
91.00
84.30
76.80
69.40
83.42 (93.5%)
0.1831 ( 1.42 × )
22.72 ( 1.06 × )
44.01 ( 1.06 × )
TVCache (Ours)
97.10
92.50
87.90
81.10
73.10
86.34 (96.8%)
0.1751 ( 1.48 × )
22.73 ( 1.06 × )
43.99 ( 1.06 × )
TABLE III: Comparison of performance for VLA-Adapter evaluated on the CALVIN Benchmark.
Task Name
Description
SetKettle
Pick up the kettle and place it on the plate.
PackPear
Pick up the pear and place it inside the basket.
HangCup
Hang the cup on the wooden mug tree.
PotDuck
Put the duck into the pot and cover it with the lid.
TABLE IV: Descriptions of the evaluation tasks.
TVCache Ret.
Patch sim.
Head filtering
Token selection
Reuse-layer selection
Bookkeeping
Total (ms)
50%
0.742
0.746
0.322
1.012
0.066
2.888
25%
0.651
0.747
0.316
1.052
0.073
2.839
12.5%
0.782
0.708
0.304
0.968
0.075
2.837
TABLE V: Runtime overhead breakdown of TVCache under different token retention ratios.
Fig. 4: Real-world robot evaluation across four manipulation tasks. The figure shows the physical task scenarios and the corresponding success rates of TVCache and VLA-Cache under different token retention ratios.
Method
Spatial
Object
Goal
Long
Avg. ↑
OpenVLA-OFT
100% Tokens
Vanilla
99.00
98.00
97.20
94.40
97.15
Retain 50% Tokens
VLA-Cache
98.20
98.20
97.40
93.40
96.80
+ Head
98.60 ( ↑0.40 )
98.40 ( ↑0.20 )
97.80 ( ↑0.40 )
93.80 ( ↑0.40 )
97.15 ( ↑0.35 )
TABLE VI: Ablation study on OpenVLA-OFT and BitVLA.
Fig. 5: Qualitative comparison of cross-modal attention. TVCache filters degraded heads, reducing background noise and spatial misalignment for more task-relevant visual grounding.