cs.CVOct 1, 2026

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

Authors: Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, +3 more

Organizations: Harbin Institute of Technology, Shenzhen · Institute for Artificial Intelligence, Great Bay University · Shenzhen Loop Area Institute · Macao Polytechnic University · Shenzhen Technology University · Dongguan Key Laboratory for Intelligence and Information Technology

Abstract

Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LARA: Latent Action Representation Alignment for Vision-Language-Action Models

    Jun 5, 2026Mengya Liu, Baoxiong Jia, Jiangyong Huang +2Latent Action ModelsLarge-Scale Robot Demonstration Datasets

  2. VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

    Jul 7, 2025Juyi Lin, Amir Taherin, Arash Akbari +11Diffusion-Based Vision-Language-ActionsRobotic Manipulation

  3. GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

    May 12, 2026Xiaosong Jia, Bowen Yang, Zuhao Ge +17Recent Vision-Language ModelsGuidance