ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection
Authors: Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, +3 more
Organizations: Harbin Institute of Technology, Shenzhen · Institute for Artificial Intelligence, Great Bay University · Shenzhen Loop Area Institute · Macao Polytechnic University · Shenzhen Technology University · Dongguan Key Laboratory for Intelligence and Information Technology
Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
Figures & tables
Figure 1 : Overview of VLA paradigms and motivation for ATI-VLA . Figures (a)–(c) compare general (non-predictive) VLA, predictive VLA, and our proposed action-centric predictive framework, ATI-VLA . Figures (d)–(f) identify key limitations of existing predictive VLA approaches, including observation–action modality gaps and slow convergence induced by joint optimization, which are responsible for their inferior performance relative to general VLA. By addressing these limitations, ATI-VLA enables the predictive VLA paradigm to more fully realize its potential, achieving state-of-the-art performance with faster convergence.
Figure 2 : Step 1: Actionable Representation Alignment. Left: The current observation and action are encoded and vector-quantized into a shared codebook, then decoded by a unified attention-based decoder to jointly predict the future observation Ot+n and action At+n . Right: t -SNE visualization of observation (main/wrist) and action latents during training: they are initially dispersed and modality-separated, but progressively converge into an aligned shared structure after training.
Figure 3 : Step 2: Action-Centric Adaptive Injection. Predictive observation latents are adaptively transformed via a lightweight side-path and injected into LLM layers to modulate action query tokens. Interval-based injection provides effective predictive guidance while preserving stable, action-centric decoding.
Category
Method
Success Rate( ↑ )
Rank( ↓ )
Spatial
Object
Goal
Long
Average
General VLA
OpenVLA-OFT [RSS’25] [ 20 ]
97.6%
98.4%
97.9%
94.5%
97.1%
2
π0 [RSS’25] [ 2 ]
96.8%
98.8%
95.8%
85.2%
94.2%
7
PD-VLA [IROS’25] † [ 44 ]
95.5%
96.7%
94.9%
91.7%
94.7%
6
Predictive VLA
ATM [RSS’24] [ 50 ]
68.5%
68.0%
77.8%
39.3%
63.4%
11
Seer [ICLR’25] [ 45 ]
-
-
-
87.7%
-
-
Table 1 : Simulation Results on LIBERO. Comparison of task success rates and their ranks.
Table 5
Method
Object P
Markers C
Plate H
T-shirt F
T-shirt F(OOD)
Average
Open
+Pack
C1
+C2
Pick
+Pass
+Place
S1
+S2
+S3
S1
+S2
+S3
Galaxea R1 Lite platform
OpenVLA-OFT†
18/25
17/25
17/25
15/25
18/25
17/25
15/25
17/25
15/25
14/25
10/25
8/25
8/25
55%
UniVLA†
14/25
13/25
13/25
13/25
15/25
14/25
12/25
15/25
15/25
12/25
13/25
10/25
7/25
46%
F1 †
16/25
15/25
19/25
17/25
15/25
14/25
14/25
17/25
15/25
14/25
16/25
12/25
11/25
57%
DreamVLA†
15/25
14/25
17/25
14/25
13/25
11/25
11/25
15/25
14/25
14/25
13/25
12/25
10/25
50%
Table 4 : Real-World Experiments on Different Platforms. The table reports success rates at each stage for five long-horizon tasks, including an out-of-distribution (OOD) variant of T-shirt Folding. Results marked with “†” were reproduced using the same experimental settings as ATI-VLA. The average success rate corresponds to the completion rate of the entire task sequence.
Table 7
Figure 4 : Ablation on Actionable Representation Alignment Strategy. t-SNE of observation/action latents under different alignment strategies: (a) dual decoder, (b) sequential decoding, (c) dual codebook, (d) our method (shared codebook + joint decoding).
Figure 5 : Visualization results of the proposed method. (a) shows future observation prediction results. (b) illustrates step-by-step execution of real-world manipulation tasks.
Table 10
Figure 6 : Convergence Comparison. ATI-VLA converges faster and performs better than ATI-VLA w/o alignment, while Fig. 1(f) further shows faster convergence over prior predictive VLA baselines.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Parameters
FLOPs
LIBERO
RoboTwin 2.0
Real World
Cost (h)
OpenVLA-OFT (Baseline)†
7.71 B
8.45 T
96.6
58.8
55.0
30.2
ATI-VLA
8.04 B (+4.2%)
8.51 T (+0.7%)
97.9 (+1.3)
72.3 (+13.5)
70.0 (+15.0)
16.2 (-46.4%)
Appendix
Table 10 : Trade-off Between Model Complexity and Task Performance.
Method
LIBERO
RoboTwin 2.0
Real-World
Baseline + Direct Future Prediction
95.8%
52.4%
48%
Baseline + injection without alignment
93.4%
51.3%
50%
Baseline + with alignment but without adaptive injection
96.2%
60.7%
62%
ATI-VLA (Ours)
97.9%
72.3%
70%
Appendix
Table 11 : Controlled Same-Backbone Ablation. All variants use the same backbone, the same training data, and the same evaluation protocol.
Method
LIBERO
RoboTwin 2.0
Real-World
Image Prediction (CoT-VLA-like)
95.2%
52.0%
50%
World-Knowledge Prediction (DreamVLA-like)
94.4%
50.3%
46%
ATI-VLA (Ours)
97.9%
72.3%
70%
Appendix
Table 12 : Controlled Comparison over Predicted Modalities. All variants use the same backbone, training data, and evaluation protocol, differing only in the predicted modality and predictive design. Results report average success rates.
LLM
Architectural Prior
Norm
Layers
Best s
LLaMA-2-7B
Dense causal attention
RMSNorm
32
8 (1/4)
Mistral-7B
Sliding-window attention
RMSNorm
32
8 (1/4)
OPT-1.3B
Dense causal attention
LayerNorm
24
8 (1/3)
Appendix
Table 13 : Generalization of Injection Interval Across Backbones. We evaluate LLM backbones with different attention patterns and normalization schemes. The preferred injection interval remains around one-third to one-fourth of the model depth.
Setting
LIBERO
RoboTwin 2.0
Real-World
Effect of Action Information in Stage 1
Observation-only Alignment
95.6%
64.3%
61%
Full Alignment (Ours)
97.9%
72.3%
70%
Effect of Predictive Alignment Objective
Reconstruction-based Alignment
96.0%
62.4%
59%
Predictive Alignment (Ours)
97.9%
72.3%
70%
Appendix
Table 14 : Ablation on Stage-1 Alignment Design. We examine whether action information and predictive learning are necessary for learning action-grounded latent representations.
Method
Semantic
Pose/Control
Transition
Success
w/o Alignment
26
33
49
92
ATI-VLA (Ours)
11
19
41
129
Appendix
Table 15 : Failure-Mode Diagnosis in Long-Horizon Real-World Tasks. We categorize failures on Plate Handover and T-shirt Folding, with 100 trials per task and 200 trials in total for each method. The non-aligned variant fails more often in semantic grounding, pose/control grounding, and transition-phase stages.
Method
Safe Steps ↑
Pre-Failure Steps ↑
Gap ↓
w/o Alignment
0.59
0.37
0.22
ATI-VLA (Ours)
0.80
0.70
0.10
Appendix
Table 17 : Cross-Modal Similarity Before Failure. We compare cross-modal similarity on safe steps and pre-failure steps. The non-aligned variant exhibits a larger alignment drop before failure, while ATI-VLA maintains higher alignment in these high-risk regions.
Method
Avg. Cos(gpred,gact)↑
Neg. Ratio ↓
RoboTwin 2.0 ↑
Real-World ↑
DreamVLA
-0.23
43%
–
50%
Direct Future Prediction
-0.29
39%
52.4%
48%
Direct Future Prediction + PCGrad
-0.09
21%
57.6%
55%
ATI-VLA (Align + Inject)
Avoided by design
Avoided by design
72.3%
70%
Appendix
Table 18 : Gradient Conflict Between Prediction and Action Objectives. We measure the gradient cosine similarity between prediction and action losses on shared parameters. PCGrad partially alleviates the conflict, while ATI-VLA avoids it by design and achieves the best performance.
Figure 7 : Extended Future Observation Prediction from Aligned Action Representations. Additional qualitative examples of future observation prediction on the LIBERO and RoboTwin 2.0 benchmarks. Each triplet shows the current observation, ground-truth future observation, and the predicted future observation from the aligned action representation.
Figure 8 : Extended Step-by-Step Real-World Task Execution. Further visualization of all real-world tasks. Each frame captures a critical stage in the sequential execution.
Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-quality data and is limited by the scarcity of real-world robot action datasets. To facilitate VLA model learning with abundant unlabeled human videos, Latent Action Models (LAM) learn latent action representations from visual dynamics to provide additional supervision for VLA learning. However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. This enables reciprocal benefits where LAMs learn with action trajectories to avoid spurious visual changes, while VLAs are regularized by forward dynamics learned within LAMs to reduce hallucinations of functionally ineffective trajectories. We demonstrate LARA versatility and effectiveness for pre-training, post-training enhancement of pre-trained VLA models, and LAM refinement, achieving an average of ~10%, ~5%, and ~15% improvement over 3 simulation and 1 meticulously designed real-world robotic manipulation benchmarks.
Mengya Liu, Baoxiong Jia, Jiangyong Huang +2
State Key Laboratory of General Artificial Intelligence, BIGAI · Peking University · Delta Intelligence
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of massive tokens leading to high inference latency and increased training cost, and (ii) insufficient utilization of generated actions resulting in potential performance loss. To address these issues, we develop a training framework to finetune VLA models for generating significantly fewer action tokens with high parallelism, effectively reducing inference latency and training cost. Furthermore, we introduce an inference optimization technique with a novel voting-based ensemble strategy to combine current and previous action predictions, improving the utilization of generated actions and overall performance. Our results demonstrate that we achieve superior performance compared with state-of-the-art VLA models, achieving significantly higher success rates and 39× faster inference than OpenVLA with 46 Hz throughput on edge platforms, demonstrating practical deployability. The code is available at https://github.com/LukeLIN-web/VOTE.
Juyi Lin, Amir Taherin, Arash Akbari +11
Northeastern University, Boston, USA · EmbodyX,San Mateo,USA
Vision-Language-Action (VLA) models aim for general robot learning by aligning action as a modality within powerful Vision-Language Models (VLMs). Existing VLAs rely on end-to-end supervision to implicitly enable the action decoding process to learn task-relevant features. However, without explicit guidance, these models often overfit to spurious correlations, such as visual shortcuts or environmental noise, limiting their generalization. In this paper, we introduce GuidedVLA, a framework designed to manually guide the action generation to focus on task-relevant factors. Our core insight is to treat the action decoder not as a monolithic learner, but as an assembly of functional components. Individual attention heads are supervised by manually defined auxiliary signals to capture distinct factors. As an initial study, we instantiate this paradigm with three specialized heads: object grounding, spatial geometry, and temporal skill logic. Across simulation and real-robot experiments, GuidedVLA improves success rates in both in-domain and out-of-domain settings compared to strong VLA baselines. Finally, we show that the quality of these specialized factors correlates positively with task performance and that our mechanism yields decoupled, high-quality features. Our results suggest that explicitly guiding action-decoder learning is a promising direction for building more robust and general VLA models.
Xiaosong Jia, Bowen Yang, Zuhao Ge +17
Institute of Trustworthy Embodied AI (TEAI), Fudan University · 2Shanghai Key Laboratory of Multimodal Embodied AI · Shanghai Jiao Tong University +1