Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.
Figures & tables
Figure 1: Conceptual comparison with related visual relevance mechanisms in VLA policies. Implicit grounding methods guide visual representations through internal objectives or attention-side signals; explicit grounding methods localize task-relevant visual regions; token pruning methods remove weakly relevant tokens. In contrast, DIVA preserves the full visual token sequence and softly attenuates visual representations in both input space and latent space.
Figure 2: Overview of DIVA on the task “turn on the stove and put the moka pot on it.” DIVA first obtains a compact task representation through Grounded Text Pooling and estimates patch-wise relevance anchors from task-conditioned visual tokens. The same anchors are then used for dual-space visual attenuation: explicit input-space attenuation reweights projected visual tokens before backbone input, while implicit latent-space attenuation persistently weakens low-relevance visual hidden states across backbone layers.
Method
Spatial
Object
Goal
Long
Avg.
Diffusion Policy
78.3
92.5
68.3
50.5
72.4
TraceVLA
84.6
85.2
75.1
54.1
74.8
Octo
78.9
85.7
84.6
51.1
75.1
OpenVLA
84.7
88.4
79.2
53.7
76.5
DiTA
84.2
96.3
85.4
63.8
82.4
CoT-VLA
87.5
91.6
87.6
69.0
83.9
Table 1: LIBERO success rates (%) across four task suites. OpenVLA-OFT † is our evaluation of the official checkpoint.
Method
Cam.
Robot
Lang.
Light
Backg.
Noise
Layout
Total
OpenVLA [ 11 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
NORA [ 8 ]
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
UniVLA [ 27 ]
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
RIPT-VLA [ 25 ]
55.2
31.2
77.6
88.4
91.6
73.5
74.2
68.4
OpenVLA-OFT [ 10 ]
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
DIVA (Ours)
62.4
33.7
80.6
92.6
96.1
79.1
76.3
72.6
Table 2: Zero-shot LIBERO-Plus success scores (%) across perturbation factors. Total is task-weighted over all evaluation instances.
Figure 3: Real-world evaluation on the UF850 platform. Successful executions out of 25 physical trials for OpenVLA-OFT and DIVA under the standard setup, unseen objects, and an unseen background.
Figure 4: Dual-space attenuation dynamics on the LIBERO task “turn on the stove and put the moka pot on it.” For each timestep, the left image in each pair shows explicit relevance anchors αi , and the right image shows the corresponding latent attenuation map.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Value
Grounded Text Pooling
Visual temperature τv
0.10
Text temperature τt
0.05
Relevance estimator
Hidden width
512
Attention heads
8
Self-attention blocks
1
Budget objective
Target relevance ρ
0.40
Appendix
Table 5: DIVA hyperparameters used in the reported experiments.
Perturbation factor
Variants
Camera viewpoints
1,599
Robot initial states
1,550
Language instructions
1,537
Light conditions
1,142
Background textures
1,076
Sensor noise
1,601
Appendix
Table 6: Composition of the LIBERO-Plus evaluation set.
Figure 5: DIVA dynamics on two LIBERO long-horizon tasks across eight policy-query steps. Each task uses rows for RGB, relevance, and latent attenuation.
Figure 6: DIVA dynamics on two additional LIBERO tasks across eight policy-query steps. Relevance shifts from candidate objects and targets toward the manipulated object and active interaction region as execution progresses.
Figure 7: Real-world trajectories under the standard setup for instructed-cube placement, instructed-cup stacking, and sequential button pressing. Each task is shown across all available saved queries, up to eight policy-query steps, with RGB, explicit, and latent attenuation rows.
Figure 8: Real-world cup-stacking trajectories under unseen object distractors and an unseen background. Each condition is shown across eight policy-query steps with RGB, explicit, and latent attenuation rows.
Department of Mechanical Engineering University College London London, United Kingdom · Department of Engineering Science University of Oxford Oxford, United Kingdom