Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task π0.5 policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.
Figures & tables
Figure 1: Overview of ViDAL. (a) Action VAE supervised only by reconstruction. (b) VLAs supervised with visual prediction signals. (c) ViDAL (ours) grounds the action latent space in future visual dynamics, and uses the Action VAE as a plug-in action representation at deployment. Right column: consistent success-rate gains on simulation benchmarks and realrobot platforms.
Figure 2: Framework of ViDAL. Top: the Action VAE is trained with visual-dynamics supervision from a frozen video encoder. Bottom: the frozen Action VAE then plugs into VLA training (left) and inference (right), serving as the action representation between the policy and the robot.
Method
Object
Spatial
Goal
Long
Average
Action-tokenization VLAs
π0 -FAST [ 8 ]
97.2
96.6
96.0
86.8
94.2
FASTer [ 11 ]
99.4
98.0
98.6
95.4
97.9
Visual-Dynamics Augmented VLAs
GR00T-N1.5 [ 33 ]
97.6
94.4
93.0
90.6
93.9
WorldVLA [ 19 ]
96.2
87.6
83.4
60.0
81.8
Table 1: Success rates (%) on LIBERO. Methods are grouped; gray rows are our +ViDAL variants and deltas are relative to the corresponding backbone.
π0 [ 4 ]
StarVLA-OFT [ 26 ]
StarVLA-OFT +ViDAL
π0.5 [ 5 ]
π0.5 +ViDAL
Task
Clean
Random
Clean
Random
Clean
Random
Clean
Random
Clean
Random
Adjust Bottle
90.0
56.0
96.0
0.0
100.0
5.0
84.0
73.0
99.0
86.0
Place Empty Cup
37.0
11.0
72.0
4.0
66.0
0.0
80.0
56.0
89.0
83.0
Click Alarmclock
63.0
11.0
91.0
14.0
81.0
6.0
56.0
42.0
79.0
79.0
Open Laptop
85.0
46.0
31.0
0.0
79.0
6.0
75.0
48.0
91.0
77.0
Place Burger Fries
80.0
4.0
96.0
6.0
97.0
4.0
68.0
47.0
85.0
77.0
Table 2: Success rates (%) on RoboTwin 2.0 under Clean/Random settings; we show 15 representative tasks and report Average over all 50 tasks. Gray columns are our +ViDAL variants and deltas are relative to the corresponding backbone.
Variant
Ldyna
Lfreq
LKL
Clean
Random
Vanilla Action VAE
✓
57.4
38.3
w/o Ldyna
✓
✓
62.2
40.1
w/o Lfreq
✓
✓
63.1
40.4
w/o LKL
✓
✓
63.2
39.7
Full ViDAL
✓
✓
✓
65.5
43.1
Table 3: Training objective ablation of ViDAL on RoboTwin 2.0 with π0.5 . Clean/Random success rates (%) average over 50 tasks; reconstruction loss is always on.
Figure 3: Action latent design ablation of ViDAL on RoboTwin 2.0 with π0.5 . Each panel sweeps one design choice with the others fixed; stars mark the default.
Figure 4: Component comparison for ViDAL on RoboTwin 2.0 with π0.5 . (a) frozen Video Encoder or Video VAE Encoder for Ldyna ; (b) form of dynamics supervision; (c) policy prediction target and inference decoding branch ( Latent / Raw : predict only the latent or only raw actions; Both, lat./raw : predict both, decode from the named branch).
Figure 5: Scaling analysis of ViDAL on RoboTwin 2.0 with π0.5 . (a) success rate against the number of training tasks ∣T∣ (eval subset varies with ∣T∣ ); (b) the corresponding success rate improvement Δ (+ViDAL- π0.5 ), which is comparable across ∣T∣ ; (c) success rate against demonstrations per task N on a fixed 50-task eval; (d) the corresponding success rate improvement Δ against N .
Figure 6: Real-world deployment of ViDAL on two robot platforms. Top: a single-arm Franka Research 3 with a ZED 2i RGB-D camera on four tabletop tasks. Bottom: a dual-arm ARX with three RealSense D405 cameras on three bimanual tasks. Each panel shows the hardware setup, representative rollouts, and success rates of π0.5 with and without ViDAL over 10 rollouts per task.
Figure 7: Action latent space geometry on 50 RoboTwin 2.0 tasks. Left/middle: 2D PCA projections of raw action chunks and ViDAL latents, colored by task. Right: cumulative explained variance vs. principal-component index.
Figure 8: Future video prediction from the ViDAL action latent. For each episode ( left: real-world Franka; right: real-world ARX), top: ground-truth future frames; bottom: frames decoded via the future visual latent predictor and frozen video VAE decoder.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
LIBERO / Franka
RoboTwin 2.0 / ARX
Action dimension da
7
14
Action chunk length H
32
48
Action latent dimension D
8
16
Observation clip length
33 frames
49 frames
Image resolution
128×128
96×128
Video latent tokens Tv×Ns
9×64
13×48
Appendix
Table 4: Key Action VAE training configurations. We report only the settings that define the action-latent interface and visual-dynamics target.
π0
StarVLA-OFT
StarVLA-OFT +ViDAL
π0.5
π0.5 +ViDAL
Task
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Adjust Bottle
90.0
56.0
96.0
0.0
100.0
5.0
84.0
73.0
99.0
86.0
Beat Block Hammer
43.0
21.0
58.0
1.0
68.0
2.0
68.0
19.0
86.0
28.0
Blocks Ranking RGB
19.0
5.0
45.0
0.0
9.0
0.0
48.0
22.0
71.0
41.0
Blocks Ranking Size
7.0
1.0
27.0
0.0
3.0
0.0
26.0
7.0
37.0
11.0
Click Alarmclock
63.0
11.0
91.0
14.0
81.0
6.0
56.0
42.0
79.0
79.0
Appendix
Table 5: Full RoboTwin 2.0 success rates (%) over all 50 tasks. Gray columns indicate the corresponding +ViDAL variants.
Figure 9: Future visual prediction examples on simulation benchmarks (LIBERO and RoboTwin 2.0). Each example contains a ground-truth row and a prediction row decoded from the future visual latent predicted using the ViDAL action latent.
Figure 10: Future visual prediction examples on real-robot datasets (Franka and ARX). Each example contains a ground-truth row and a prediction row decoded from the future visual latent predicted using the ViDAL action latent.
Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends on large-scale, high-quality data and is limited by the scarcity of real-world robot action datasets. To facilitate VLA model learning with abundant unlabeled human videos, Latent Action Models (LAM) learn latent action representations from visual dynamics to provide additional supervision for VLA learning. However, LAM and VLA are typically trained separately, leaving LAM ungrounded during VLA training and VLA models constrained by frozen LAM representations. To address these issues, we propose Latent Action Representation Alignment (LARA), a plug-and-play framework that jointly optimizes LAM and VLA via representation alignment. This enables reciprocal benefits where LAMs learn with action trajectories to avoid spurious visual changes, while VLAs are regularized by forward dynamics learned within LAMs to reduce hallucinations of functionally ineffective trajectories. We demonstrate LARA versatility and effectiveness for pre-training, post-training enhancement of pre-trained VLA models, and LAM refinement, achieving an average of ~10%, ~5%, and ~15% improvement over 3 simulation and 1 meticulously designed real-world robotic manipulation benchmarks.
Mengya Liu, Baoxiong Jia, Jiangyong Huang +2
State Key Laboratory of General Artificial Intelligence, BIGAI · Peking University · Delta Intelligence
Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified action space across embodiments. However, existing LAMs often rely on discrete quantization encode and decode pipelines, which can lead to trivial frame reconstruction behavior, limited representational capacity, and a lack of physically meaningful structure. We introduce RotVLA, a VLA framework built on a continuous rotational latent action representation. Latent actions are modeled as elements of SO(n), providing continuity, compositionality, and structured geometry aligned with real-world action dynamics. A triplet frame learning framework further enforces meaningful temporal dynamics while avoiding degeneration. RotVLA consists of a VLM backbone and a flow-matching action head, pretrained on large-scale cross-embodiment robotic datasets and human videos with latent-action supervision. For downstream robot control, the flow-matching head is extended into a unified action expert that jointly denoises latent and robot actions. Here, latent actions serve as a latent planner, providing high-level guidance that conditions action generation. With only 1.7B parameters and 1700+ hours of pretraining data, RotVLA achieves 98.2% on LIBERO and 89.6% / 88.5% on RoboTwin2.0 under clean and randomized settings, respectively. It also demonstrates strong real-world performance on manipulation tasks, consistently outperforming existing VLA models.
Qiwei Li, Xicheng Gong, Xinghang Li +5
Wangxuan Institute of Computer Technology, Peking University · Xiaomi Robotics · CASIA
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this limitation by conditioning policies on predicted futures, yet existing approaches typically rely on computationally expensive video generation with substantial pixel-level redundancy. We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. At the core of LaWAM is a latent-action-conditioned Latent World Model (LaWM). We obtain LaWM by training a latent action model in the latent space of a pretrained vision foundation model and repurposing its forward decoder to predict future observation features for scene evolution. LaWAM then conditions action generation on these predicted latent visual subgoals to enable dynamics-aware robot control. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference. LaWAM runs in 187 ms per action-chunk prediction and achieves up to 24x lower wall-clock latency than pixel-space WAMs.
Jialei Chen, Kai Wang, Kang Chen +9
Jilin University · Zhongguancun Academy · Nankai University +4