Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter's success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter's average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT's success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.
Figures & tables
Figure 1: Conceptual view of the perception-to-action transition. Standard VLAs require Action Tokens to acquire and organize action-relevant information during decoding. PAIR instead learns perception-derived Bridge Tokens shaped by the structure of expert actions and injects them into the initial Action Tokens, so that the Action Expert begins from an action-ready intermediate state and refines it toward continuous control. The Action Autoencoder is used only during training.
Figure 2: Overview of PAIR. The Masked Action Autoencoder encodes expert actions into Action Latent Tokens, which guide the perception-derived Bridge Tokens through representation alignment. The Bridge Tokens are projected into the Action Token space and injected at the entrance of the Action Expert .
Method
Spatial
Object
Goal
Long
Avg.
SpatialVLA ( Qu et al., 2025 )
88.2
89.9
78.6
55.5
78.1
WorldVLA ( Cen et al., 2025 )
87.6
96.2
83.4
60.0
81.8
π0 -FAST ( Pertsch et al., 2025 )
96.4
96.8
88.6
60.2
85.5
LARA (full) ( Liu et al., 2026b )
88.0
92.0
88.5
86.0
88.6
GR00T N1 ( Bjorck et al., 2025 )
94.4
97.6
93.0
90.6
93.9
π0 ( Black et al., 2025 )
96.8
98.8
95.8
85.2
94.2
Table 1: Simulation results on LIBERO. Success rate (%) is reported for each task suite and their average. Best results are in bold.
Method
Camera
Robot
Language
Light
Background
Noise
Layout
Total
OpenVLA ( Kim et al., 2025b )
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
WorldVLA ( Cen et al., 2025 )
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
NORA ( Hung et al., 2025 )
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
UniVLA ( Bu et al., 2025 )
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
π0 ( Black et al., 2025 )
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
VLA-Adapter ( Wang et al., 2026 )
36.2
37.9
74.6
70.6
76.1
58.0
69.7
59.1
Table 2: Zero-shot LIBERO-Plus results with LIBERO-trained checkpoints. Total follows the benchmark’s overall score, not the seven-category mean. Best results are in bold.
Method
1
2
3
4
5
Avg. Len.
OpenVLA ( Kim et al., 2025b )
91.3
77.8
62.0
52.1
43.5
3.27
RoboVLMs ( Li et al., 2026 )
98.0
93.6
85.4
77.8
70.4
4.25
VPP ( Hu et al., 2025 )
96.5
90.9
86.6
82.0
76.9
4.33
UnifiedVLA ( Wang et al., 2025b )
98.9
94.8
89.0
82.8
75.1
4.41
DreamVLA ( Zhang et al., 2025 )
98.2
94.6
89.5
83.4
78.1
4.44
NIAF ( Liu et al., 2026a )
99.7
95.9
90.6
84.8
76.4
4.47
Table 3: Simulation results on CALVIN ABC → D. We report success rates (%) for completing at least one to five consecutive tasks and average completed sequence length. Best results are in bold.
Figure 3: Seven real-world tasks on the UF850, spanning language grounding, multi-step execution, and precise manipulation.
Figure 4: Real-world success rates over 20 trials per task. PAIR improves OpenVLA-OFT across all seven tasks, with an overall gain of 13.6 percentage points.
Figure 5: Bridge Tokens connect task semantics to action structure. (a) Linear-probe balanced accuracy for task identity and action clusters. (b) Action-chunk R2 for baseline Action Tokens across layers (blue) and Bridge Tokens before the Action Expert (dashed). (c) RSA and top-10 neighbor overlap with ground-truth actions. Error bars show fold standard deviations in (a,b), and task standard deviations for Bridge Tokens or 5th–95th permutation percentiles for shuffled controls in (c) .
Position
Spatial
Object
Goal
Long
Avg.
Start (base)
99.2
99.6
98.6
96.2
98.4
End
98.4
99.0
97.6
94.6
97.4
Middle
98.8
98.9
98.3
95.6
97.9
All
98.7
99.0
97.9
94.8
97.6
Table 4: Injection position ablation on the four LIBERO task suites. Values are success rates (%).
λalign
Spatial
Object
Goal
Long
Avg.
0.1 (base)
99.2
99.6
98.6
96.2
98.4
0.01
98.8
99.2
98.0
95.2
97.8
0
98.2
98.8
97.4
94.4
97.2
Table 5: Alignment loss weight ablation on the four LIBERO task suites. Values are success rates (%) .
Table 7: Injection strength ablation on the four LIBERO task suites. Values are success rates (%).
AE Training
Spatial
Object
Goal
Long
Avg.
Noise (base)
99.2
99.6
98.6
96.2
98.4
Clean
98.4
99.0
97.6
95.4
97.6
Appendix
Table 8: Action AE noise ablation on the four LIBERO task suites. Values are success rates (%).
Figure 6: Task identity and action-cluster decodability across depth. Blue curves show baseline Action Tokens with fold-standard-deviation bands. Green dashed lines indicate the pre-LLM Bridge Token score. Layer 0 precedes the Transformer; layers 1–24 denote completed blocks.
Figure 7: Continuous-action decodability of Action Tokens across depth. Ridge readouts predict the full eight-step ground-truth action chunk from frozen Action Tokens. Scores are 1−MAE after target normalization, pooled over five trajectory-grouped held-out folds on 1,000 shared LIBERO-Spatial examples. PAIR tokens are measured after Bridge Token injection; baseline tokens are measured before and after each Transformer block.
Figure 8: Action-cluster geometry and kernel alignment. (a,b) Cosine similarities between GT-derived action-cluster centroids in the two spaces, using a shared color scale. A0–A7 index unsupervised clusters rather than named behaviors. (c) Within-task linear CKA. Error bars indicate task standard deviation for Bridge Tokens and the 5th–95th percentile interval of the trajectory-permutation reference.
Stage / backbone
Batch
Steps
Learning rate
Schedule
LoRA rank
Action Autoencoder
64
50k
3×10−4
cosine, 1k warmup
–
VLA-Adapter + PAIR
64
150k
2×10−4
cosine
64
OpenVLA-OFT + PAIR
64
50k
5×10−4
constant
32
Appendix
Table 9: Default training hyperparameters. Batch denotes the effective batch size across four GPUs. The OpenVLA-OFT learning rate remains constant because its decay milestone follows the 50k-step run .
Task
Instruction and success criterion
Demos.
Object placement
“Place the object in the Y bowl.” Place the object in the instructed bowl.
51
Button pressing
“First press X , then Y , and finally Z .” Press the three buttons in the specified order.
50
Cube placement
“Put X in the cup.” Select the instructed cube among the candidates and place it in the cup.
51
Cup stacking
“Stack X on top of Y .” Stack the instructed source cup on the target cup.
51
Board wiping
“Wipe the marked circle off the whiteboard.” Grasp the eraser and remove the marked circle.
30
Water pouring
“Pour water from the red cup into the blue cup.” Transfer the water between the specified cups.
30
Appendix
Table 10: Real-world task definitions and dataset sizes. X , Y , and Z denote instance-specific objects or colors.
Figure 9: Demonstration sequences for the seven real-world tasks. Each row shows five chronological frames from one recording, with short labels identifying the manipulation stages.
Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-quality data and is limited by the scarcity of real-world robot action datasets. To facilitate VLA model learning with abundant unlabeled human videos, Latent Action Models (LAM) learn latent action representations from visual dynamics to provide additional supervision for VLA learning. However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. This enables reciprocal benefits where LAMs learn with action trajectories to avoid spurious visual changes, while VLAs are regularized by forward dynamics learned within LAMs to reduce hallucinations of functionally ineffective trajectories. We demonstrate LARA versatility and effectiveness for pre-training, post-training enhancement of pre-trained VLA models, and LAM refinement, achieving an average of ~10%, ~5%, and ~15% improvement over 3 simulation and 1 meticulously designed real-world robotic manipulation benchmarks.
Mengya Liu, Baoxiong Jia, Jiangyong Huang +2
State Key Laboratory of General Artificial Intelligence, BIGAI · Peking University · Delta Intelligence
Vision-Language-Action (VLA) models that couple pretrained Vision-Language Models (VLMs) with continuous action experts have achieved strong manipulation performance, yet generalization to out-of-distribution (OOD) language instructions remains poor. A known challenge is the structural imbalance in VLA data, where language is far less diverse than visual and action content, making policies prone to visual shortcuts. While discrete-action methods mitigate this through vision-language co-training, continuous action experts lack such protection: they start from random initialization and learn entirely from imbalanced data, producing noisy gradients that corrupt the VLM and fail to exploit its language capability. We address this from a Bayesian perspective, factorizing the policy into a language-agnostic Vision-Action (VA) prior and a language-conditioned VLA likelihood, and propose APT, a two-stage training method emphasizing Action expert PreTraining. In Stage 1, the action expert is pretrained as a VA prior on vision-action pairs from a frozen VLM, bypassing the language imbalance. In Stage 2, language tokens are injected through a gated fusion mechanism that integrates VLM features while preserving the learned visuomotor prior. APT applies to mainstream VLA architectures, including the π and GR00T-style architectures. Comprehensive experiments validate that APT achieves consistent gains on unseen instructions and compositional tasks. Project Page: https://xukechun.github.io/papers/APT/
Kechun Xu, Zhenjie Zhu, Anzhe Chen +2
Zhejiang University · Zhejiang Humanoid Robot Innovation Center
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Existing designs either expose only a narrow part of the representation hierarchy or rigidly match each decoder block to one VLM layer, restricting access to complementary task evidence across depths. We introduce LIRA, a local cross-layer action-conditioning mechanism that formulates VLM-to-action conditioning as depth-aware information routing. LIRA operates on task-token features and LIRA Query features derived from intermediate VLM states, then assigns each Parallel Fusion Block a depth-aligned local window centered on its corresponding VLM layer. Parallel Fusion Blocks aggregate neighboring LIRA Query features and integrate them with task-token features and proprioceptive inputs before action prediction. This routing interface leaves the backbone architecture, action decoder, and supervised training recipe unchanged. Across LIBERO, LIBERO-Plus, CALVIN ABC→D, and real-world manipulation, LIRA improves the principal aggregate metrics over the VLA-Adapter baseline under the same 0.5B-parameter configuration. In zero-shot transfer to LIBERO-Plus, LIRA increases average success from 59.1% to 78.0%, an 18.9-point gain indicating improved robustness under controlled distribution shifts.