cs.ROSep 29, 2026

LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks

Authors: Zeming Wei, Jianheng Ye, Xinshuai Song, Sirui Chen, Yang Liu, Liang Lin

Organizations: Sun Yat-sen University · X-Era AI Lab · Pengcheng Laboratory

Abstract

Vision-language-action (VLA) models struggle to reuse recurring interactions in unseen tasks. Our diagnostic study reveals that reliable task completion does not imply consistent execution of constituent atomic actions across task contexts. We present LexiconVLA, a retrievable atomic-action lexicon for cross-task reuse. Global and detail codebooks capture shared interaction structure and fine-grained execution variation, respectively, preserving both reusable patterns and execution details. Visual-Atomic Action Alignment couples trajectory reconstruction from visual state changes with visual outcome prediction from action codes, grounding the lexicon in motion and its effects. We learn these codebooks with trajectory reconstruction and visual alignment on our AtomAction Dataset of 57,803 segments from 69 tasks. A planner and scene-aware adapter translate new goals into code-conditioned subtasks for a shared policy, without skill-specific experts or deployment-time parameter updates. Across five policy backbones on 26 RLBench tasks, LexiconVLA largely maintains performance on 18 seen tasks while improving success on 8 tasks held out from policy training. With BridgeVLA, unseen-task success rises from 16.67% to 34.17% (+17.50 percentage points), and overall success reaches 71.08%, the highest among methods with reported results. Real-robot experiments demonstrate stepwise execution and failure recovery.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

    May 12, 2026Xiaosong Jia, Bowen Yang, Zuhao Ge +17Recent Vision-Language ModelsGuidance

  2. PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models

    May 19, 2026Peizheng Guo, Jingyao Wang, Changwen Zheng +1Generalizable Vision-Language-Action PoliciesVision-Language-Action Framework

  3. ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models

    Sep 7, 2026Songhua Yang, Ziyu Liu, Xuetao Li +3In-ContextMultimodal Fusion