Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the visual dynamics induced by these actions. We introduce ViDAL, a Visual Dynamics-grounded Action Latent Space that anchors continuous action latents in the future visual dynamics of the scene. Specifically, ViDAL learns action latent space by training an Action Variational Autoencoder (Action VAE) to reconstruct action chunks while aligning its latent with future scene dynamics. When integrated into downstream robot policies, the proposed Action VAE serves as a plug-in action interface compatible with multiple VLA architectures and enables optional future-video prediction as an additional capability. Empirically, ViDAL outperforms competitive baselines on LIBERO with 98.1% average success, improves a multi-task π0.5 policy on RoboTwin 2.0 from 54.3% to 65.5% (Clean) and from 33.2% to 43.1% (Random) success rates over 50 dual-arm tasks, and yields 20.0% and 23.4% absolute success-rate gains on real-world single-arm Franka and dual-arm ARX robot platforms.
Figures & tables
Figure 1: Overview of ViDAL. (a) Action VAE supervised only by reconstruction. (b) VLAs supervised with visual prediction signals. (c) ViDAL (ours) grounds the action latent space in future visual dynamics, and uses the Action VAE as a plug-in action representation at deployment. Right column: consistent success-rate gains on simulation benchmarks and realrobot platforms.
Figure 2: Framework of ViDAL. Top: the Action VAE is trained with visual-dynamics supervision from a frozen video encoder. Bottom: the frozen Action VAE then plugs into VLA training (left) and inference (right), serving as the action representation between the policy and the robot.
Method
Object
Spatial
Goal
Long
Average
Action-tokenization VLAs
π0 -FAST [ 8 ]
97.2
96.6
96.0
86.8
94.2
FASTer [ 11 ]
99.4
98.0
98.6
95.4
97.9
Visual-Dynamics Augmented VLAs
GR00T-N1.5 [ 33 ]
97.6
94.4
93.0
90.6
93.9
WorldVLA [ 19 ]
96.2
87.6
83.4
60.0
81.8
Table 1: Success rates (%) on LIBERO. Methods are grouped; gray rows are our +ViDAL variants and deltas are relative to the corresponding backbone.
π0 [ 4 ]
StarVLA-OFT [ 26 ]
StarVLA-OFT +ViDAL
π0.5 [ 5 ]
π0.5 +ViDAL
Task
Clean
Random
Clean
Random
Clean
Random
Clean
Random
Clean
Random
Adjust Bottle
90.0
56.0
96.0
0.0
100.0
5.0
84.0
73.0
99.0
86.0
Place Empty Cup
37.0
11.0
72.0
4.0
66.0
0.0
80.0
56.0
89.0
83.0
Click Alarmclock
63.0
11.0
91.0
14.0
81.0
6.0
56.0
42.0
79.0
79.0
Open Laptop
85.0
46.0
31.0
0.0
79.0
6.0
75.0
48.0
91.0
77.0
Place Burger Fries
80.0
4.0
96.0
6.0
97.0
4.0
68.0
47.0
85.0
77.0
Table 2: Success rates (%) on RoboTwin 2.0 under Clean/Random settings; we show 15 representative tasks and report Average over all 50 tasks. Gray columns are our +ViDAL variants and deltas are relative to the corresponding backbone.
Variant
Ldyna
Lfreq
LKL
Clean
Random
Vanilla Action VAE
✓
57.4
38.3
w/o Ldyna
✓
✓
62.2
40.1
w/o Lfreq
✓
✓
63.1
40.4
w/o LKL
✓
✓
63.2
39.7
Full ViDAL
✓
✓
✓
65.5
43.1
Table 3: Training objective ablation of ViDAL on RoboTwin 2.0 with π0.5 . Clean/Random success rates (%) average over 50 tasks; reconstruction loss is always on.
Figure 3: Action latent design ablation of ViDAL on RoboTwin 2.0 with π0.5 . Each panel sweeps one design choice with the others fixed; stars mark the default.
Figure 4: Component comparison for ViDAL on RoboTwin 2.0 with π0.5 . (a) frozen Video Encoder or Video VAE Encoder for Ldyna ; (b) form of dynamics supervision; (c) policy prediction target and inference decoding branch ( Latent / Raw : predict only the latent or only raw actions; Both, lat./raw : predict both, decode from the named branch).
Figure 5: Scaling analysis of ViDAL on RoboTwin 2.0 with π0.5 . (a) success rate against the number of training tasks ∣T∣ (eval subset varies with ∣T∣ ); (b) the corresponding success rate improvement Δ (+ViDAL- π0.5 ), which is comparable across ∣T∣ ; (c) success rate against demonstrations per task N on a fixed 50-task eval; (d) the corresponding success rate improvement Δ against N .
Figure 6: Real-world deployment of ViDAL on two robot platforms. Top: a single-arm Franka Research 3 with a ZED 2i RGB-D camera on four tabletop tasks. Bottom: a dual-arm ARX with three RealSense D405 cameras on three bimanual tasks. Each panel shows the hardware setup, representative rollouts, and success rates of π0.5 with and without ViDAL over 10 rollouts per task.
Figure 7: Action latent space geometry on 50 RoboTwin 2.0 tasks. Left/middle: 2D PCA projections of raw action chunks and ViDAL latents, colored by task. Right: cumulative explained variance vs. principal-component index.
Figure 8: Future video prediction from the ViDAL action latent. For each episode ( left: real-world Franka; right: real-world ARX), top: ground-truth future frames; bottom: frames decoded via the future visual latent predictor and frozen video VAE decoder.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
LIBERO / Franka
RoboTwin 2.0 / ARX
Action dimension da
7
14
Action chunk length H
32
48
Action latent dimension D
8
16
Observation clip length
33 frames
49 frames
Image resolution
128×128
96×128
Video latent tokens Tv×Ns
9×64
13×48
Appendix
Table 4: Key Action VAE training configurations. We report only the settings that define the action-latent interface and visual-dynamics target.
π0
StarVLA-OFT
StarVLA-OFT +ViDAL
π0.5
π0.5 +ViDAL
Task
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Adjust Bottle
90.0
56.0
96.0
0.0
100.0
5.0
84.0
73.0
99.0
86.0
Beat Block Hammer
43.0
21.0
58.0
1.0
68.0
2.0
68.0
19.0
86.0
28.0
Blocks Ranking RGB
19.0
5.0
45.0
0.0
9.0
0.0
48.0
22.0
71.0
41.0
Blocks Ranking Size
7.0
1.0
27.0
0.0
3.0
0.0
26.0
7.0
37.0
11.0
Click Alarmclock
63.0
11.0
91.0
14.0
81.0
6.0
56.0
42.0
79.0
79.0
Appendix
Table 5: Full RoboTwin 2.0 success rates (%) over all 50 tasks. Gray columns indicate the corresponding +ViDAL variants.
Figure 9: Future visual prediction examples on simulation benchmarks (LIBERO and RoboTwin 2.0). Each example contains a ground-truth row and a prediction row decoded from the future visual latent predicted using the ViDAL action latent.
Figure 10: Future visual prediction examples on real-robot datasets (Franka and ARX). Each example contains a ground-truth row and a prediction row decoded from the future visual latent predicted using the ViDAL action latent.