ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection
Authors: Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, +3 more
Organizations: Harbin Institute of Technology, Shenzhen · Institute for Artificial Intelligence, Great Bay University · Shenzhen Loop Area Institute · Macao Polytechnic University · Shenzhen Technology University · Dongguan Key Laboratory for Intelligence and Information Technology
Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
Figures & tables
Figure 1 : Overview of VLA paradigms and motivation for ATI-VLA . Figures (a)–(c) compare general (non-predictive) VLA, predictive VLA, and our proposed action-centric predictive framework, ATI-VLA . Figures (d)–(f) identify key limitations of existing predictive VLA approaches, including observation–action modality gaps and slow convergence induced by joint optimization, which are responsible for their inferior performance relative to general VLA. By addressing these limitations, ATI-VLA enables the predictive VLA paradigm to more fully realize its potential, achieving state-of-the-art performance with faster convergence.
Figure 2 : Step 1: Actionable Representation Alignment. Left: The current observation and action are encoded and vector-quantized into a shared codebook, then decoded by a unified attention-based decoder to jointly predict the future observation Ot+n and action At+n . Right: t -SNE visualization of observation (main/wrist) and action latents during training: they are initially dispersed and modality-separated, but progressively converge into an aligned shared structure after training.
Figure 3 : Step 2: Action-Centric Adaptive Injection. Predictive observation latents are adaptively transformed via a lightweight side-path and injected into LLM layers to modulate action query tokens. Interval-based injection provides effective predictive guidance while preserving stable, action-centric decoding.
Category
Method
Success Rate( ↑ )
Rank( ↓ )
Spatial
Object
Goal
Long
Average
General VLA
OpenVLA-OFT [RSS’25] [ 20 ]
97.6%
98.4%
97.9%
94.5%
97.1%
2
π0 [RSS’25] [ 2 ]
96.8%
98.8%
95.8%
85.2%
94.2%
7
PD-VLA [IROS’25] † [ 44 ]
95.5%
96.7%
94.9%
91.7%
94.7%
6
Predictive VLA
ATM [RSS’24] [ 50 ]
68.5%
68.0%
77.8%
39.3%
63.4%
11
Seer [ICLR’25] [ 45 ]
-
-
-
87.7%
-
-
Table 1 : Simulation Results on LIBERO. Comparison of task success rates and their ranks.
Table 5
Method
Object P
Markers C
Plate H
T-shirt F
T-shirt F(OOD)
Average
Open
+Pack
C1
+C2
Pick
+Pass
+Place
S1
+S2
+S3
S1
+S2
+S3
Galaxea R1 Lite platform
OpenVLA-OFT†
18/25
17/25
17/25
15/25
18/25
17/25
15/25
17/25
15/25
14/25
10/25
8/25
8/25
55%
UniVLA†
14/25
13/25
13/25
13/25
15/25
14/25
12/25
15/25
15/25
12/25
13/25
10/25
7/25
46%
F1 †
16/25
15/25
19/25
17/25
15/25
14/25
14/25
17/25
15/25
14/25
16/25
12/25
11/25
57%
DreamVLA†
15/25
14/25
17/25
14/25
13/25
11/25
11/25
15/25
14/25
14/25
13/25
12/25
10/25
50%
Table 4 : Real-World Experiments on Different Platforms. The table reports success rates at each stage for five long-horizon tasks, including an out-of-distribution (OOD) variant of T-shirt Folding. Results marked with “†” were reproduced using the same experimental settings as ATI-VLA. The average success rate corresponds to the completion rate of the entire task sequence.
Table 7
Figure 4 : Ablation on Actionable Representation Alignment Strategy. t-SNE of observation/action latents under different alignment strategies: (a) dual decoder, (b) sequential decoding, (c) dual codebook, (d) our method (shared codebook + joint decoding).
Figure 5 : Visualization results of the proposed method. (a) shows future observation prediction results. (b) illustrates step-by-step execution of real-world manipulation tasks.
Table 10
Figure 6 : Convergence Comparison. ATI-VLA converges faster and performs better than ATI-VLA w/o alignment, while Fig. 1(f) further shows faster convergence over prior predictive VLA baselines.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Parameters
FLOPs
LIBERO
RoboTwin 2.0
Real World
Cost (h)
OpenVLA-OFT (Baseline)†
7.71 B
8.45 T
96.6
58.8
55.0
30.2
ATI-VLA
8.04 B (+4.2%)
8.51 T (+0.7%)
97.9 (+1.3)
72.3 (+13.5)
70.0 (+15.0)
16.2 (-46.4%)
Appendix
Table 10 : Trade-off Between Model Complexity and Task Performance.
Method
LIBERO
RoboTwin 2.0
Real-World
Baseline + Direct Future Prediction
95.8%
52.4%
48%
Baseline + injection without alignment
93.4%
51.3%
50%
Baseline + with alignment but without adaptive injection
96.2%
60.7%
62%
ATI-VLA (Ours)
97.9%
72.3%
70%
Appendix
Table 11 : Controlled Same-Backbone Ablation. All variants use the same backbone, the same training data, and the same evaluation protocol.
Method
LIBERO
RoboTwin 2.0
Real-World
Image Prediction (CoT-VLA-like)
95.2%
52.0%
50%
World-Knowledge Prediction (DreamVLA-like)
94.4%
50.3%
46%
ATI-VLA (Ours)
97.9%
72.3%
70%
Appendix
Table 12 : Controlled Comparison over Predicted Modalities. All variants use the same backbone, training data, and evaluation protocol, differing only in the predicted modality and predictive design. Results report average success rates.
LLM
Architectural Prior
Norm
Layers
Best s
LLaMA-2-7B
Dense causal attention
RMSNorm
32
8 (1/4)
Mistral-7B
Sliding-window attention
RMSNorm
32
8 (1/4)
OPT-1.3B
Dense causal attention
LayerNorm
24
8 (1/3)
Appendix
Table 13 : Generalization of Injection Interval Across Backbones. We evaluate LLM backbones with different attention patterns and normalization schemes. The preferred injection interval remains around one-third to one-fourth of the model depth.
Setting
LIBERO
RoboTwin 2.0
Real-World
Effect of Action Information in Stage 1
Observation-only Alignment
95.6%
64.3%
61%
Full Alignment (Ours)
97.9%
72.3%
70%
Effect of Predictive Alignment Objective
Reconstruction-based Alignment
96.0%
62.4%
59%
Predictive Alignment (Ours)
97.9%
72.3%
70%
Appendix
Table 14 : Ablation on Stage-1 Alignment Design. We examine whether action information and predictive learning are necessary for learning action-grounded latent representations.
Method
Semantic
Pose/Control
Transition
Success
w/o Alignment
26
33
49
92
ATI-VLA (Ours)
11
19
41
129
Appendix
Table 15 : Failure-Mode Diagnosis in Long-Horizon Real-World Tasks. We categorize failures on Plate Handover and T-shirt Folding, with 100 trials per task and 200 trials in total for each method. The non-aligned variant fails more often in semantic grounding, pose/control grounding, and transition-phase stages.
Method
Safe Steps ↑
Pre-Failure Steps ↑
Gap ↓
w/o Alignment
0.59
0.37
0.22
ATI-VLA (Ours)
0.80
0.70
0.10
Appendix
Table 17 : Cross-Modal Similarity Before Failure. We compare cross-modal similarity on safe steps and pre-failure steps. The non-aligned variant exhibits a larger alignment drop before failure, while ATI-VLA maintains higher alignment in these high-risk regions.
Method
Avg. Cos(gpred,gact)↑
Neg. Ratio ↓
RoboTwin 2.0 ↑
Real-World ↑
DreamVLA
-0.23
43%
–
50%
Direct Future Prediction
-0.29
39%
52.4%
48%
Direct Future Prediction + PCGrad
-0.09
21%
57.6%
55%
ATI-VLA (Align + Inject)
Avoided by design
Avoided by design
72.3%
70%
Appendix
Table 18 : Gradient Conflict Between Prediction and Action Objectives. We measure the gradient cosine similarity between prediction and action losses on shared parameters. PCGrad partially alleviates the conflict, while ATI-VLA avoids it by design and achieves the best performance.
Figure 7 : Extended Future Observation Prediction from Aligned Action Representations. Additional qualitative examples of future observation prediction on the LIBERO and RoboTwin 2.0 benchmarks. Each triplet shows the current observation, ground-truth future observation, and the predicted future observation from the aligned action representation.
Figure 8 : Extended Step-by-Step Real-World Task Execution. Further visualization of all real-world tasks. Each frame captures a critical stage in the sequential execution.