cs.AIOct 6, 2026

DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies

Authors: Kaixi Feng, Guoheng Sun, Ziyao Wang, Yexiao He, Zheyu Shen, Ang Li

Organizations: University of Maryland, College Park

Abstract

Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

    Jun 10, 2026Kechun Xu, Zhenjie Zhu, Anzhe Chen +2Visuomotor Policy LearningVision-Language-Action Models

  2. Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation

    May 15, 2026Jin Shi, Brady Zhang, Yishun LuVLM DistillationEfficient VLA Model Inference

  3. Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead

    Sep 29, 2026Junghyun Kim, Ngseo Kim, ChungWoo Lee +7Domain GeneralizationVision-Language-Action Models