cs.LGSep 28, 2026

Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning

Authors: Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng, Yuqi Huang, Li Shen, Ya Zhang, Dacheng Tao

Organizations: Shanghai Jiao Tong University · Shenzhen Campus of Sun Yat-sen University · Nanyang Technological University

Abstract

Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action grounding and hurts generalization and fine-tuning efficiency. In this work, we present Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. While prior research has predominantly emphasized bi-modal vision--language alignment, we systematically formalize and study tri-modal alignment in VLA models, and provide both ablations and analysis to isolate its role in improving adaptation and robustness. To further accelerate deployment, we adopt a flow-matching objective, enabling substantially fewer inference steps than diffusion-based policies while maintaining accuracy. Theoretically, we establish a quantitative connection between the tri-modal alignment gap and the optimization tightness of flow matching; empirically, experiments on the extensive benchmark show that AGFT achieves superior success rates and lower inference latency compared to SOTA baselines, underscoring tri-modal alignment as a key ingredient for scaling robust VLA manipulation.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

    Jul 2, 2026Guoyang Xia, Fengfa Li, Hongjin Ji +4Vision-Language-Action FrameworkIterative Co-Training

  2. AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

    Date pendingSunghwan Han, Youngtae Han, Youngmin YiFlow-Matching Vision-Language-ActionAction Generation

  3. APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

    Jun 10, 2026Kechun Xu, Zhenjie Zhu, Anzhe Chen +2Pretraining