SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models
Authors: Tianfu Li, Haoxuan Xu, Wenbo Chen, Haitian Li, Changchuan Yang, Xinhu Zheng, Jun Ma, Yuan Liu, +2 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou). · The Hong Kong University of Science and Technology. · Nanyang Technological University. · Zhejiang University.
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive iterative denoising, while one-step alternatives can underperform their multi-step counterparts. To reconcile efficient future modeling with strong action performance, we present SLIP-VLA, a policy learning framework that equips VLA models with a Single-Step Latent Imagination for future-aware action prediction. SLIP-VLA obtains temporally dense future latent representations with a single denoising update, and we improve the perceptual sufficiency of these representations by aligning intermediate latents with future geometric and semantic features. We further improve their control sufficiency through action-conditioned latent world modeling and inverse dynamics modeling, explicitly coupling latent transitions with robot actions. SLIP-VLA achieves state-of-the-art performance across diverse simulation benchmarks and real-world manipulation tasks, while its single-step latent imagination takes only 12 ms.
Figures & tables
Fig. 2: Pipeline of SLIP-VLA. The VLM backbone first encodes the multimodal context, and the World DiT then performs a single denoising pass to produce temporally dense future latents. (a) For perceptual sufficiency shaping, selected intermediate representations are aligned with future VGGT- Ω geometric features and DINOv3 semantic features. (b) For control sufficiency shaping, deeper latent representations are regularized through forward and inverse latent dynamics to preserve action-dependent transitions. The shaped future representations are injected layer-wise into the Action DiT for action generation, while all auxiliary shaping modules are used only during training.
Method
Spatial
Object
Goal
Long
Avg.
General VLAs
OpenVLA [ 1 ]
84.7
88.4
79.2
53.7
76.5
OpenVLA-OFT [ 35 ]
97.6
98.4
97.9
94.5
97.1
π0 [ 2 ]
98.0
96.8
94.4
88.4
94.4
π0.5 [ 3 ]
98.8
98.2
98.0
92.4
96.9
Future-Aware VLAs
TABLE I: LIBERO success rates (%). Best and second-best results are bolded and underlined , respectively.
Method
Cam.
Robot
Lang.
Light
BG
Noise
Layout
Avg.
General VLAs
OpenVLA [ 1 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
OpenVLA-OFT [ 35 ]
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
π0 [ 2 ]
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 -Fast [ 36 ]
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
Future-Aware VLAs
TABLE II: LIBERO-Plus success rates (%). Best and second-best results are bolded and underlined , respectively.
Method
Clean
Rand.
Avg.
General VLAs
π0 [ 2 ]
65.9
58.4
62.2
π0.5 [ 3 ]
82.7
76.8
79.8
Future-Aware VLAs
WorldVLA [ 4 ]
42.5
32.2
37.4
World Action Models
TABLE III: RoboTwin 2.0 success rates (%). Best and second-best results are bolded and underlined , respectively.
Fig. 3: Qualitative analysis of perceptual and control information in the single-step latent. (a) Compared with the naive single-step baseline, SLIP-VLA recovers more coherent future geometry and more concentrated task-relevant semantic responses. (b) For forward dynamics, the correct action produces a prediction closer to the target latent, while a shuffled action induces clear spatial deviations; the action-response map highlights regions most affected by the action change.
Fig. 4: Effect of denoising steps on future prediction and control performance. The top row shows representative predicted videos with one to four denoising steps, while the curves report PSNR and LIBERO-Plus success rate. Additional denoising provides no consistent improvement in visual prediction quality and does not improve action performance, with the highest success rate achieved using a single step.
Fig. 5: Real-world evaluation of SLIP-VLA. (a) The real-world manipulation platform. (b.1) Our SLIP-VLA successfully completes the in-distribution corn-to-plate task. (b.2) It remains successful under visual distribution shifts in background, object appearance, and object orientation. (b.3) SLIP-VLA further completes both two-step and three-step long-horizon tasks involving sequential interaction with the pot and the target object.
Variant
LIBERO
LIBERO-Plus
Naive Single-Step
97.7
79.0
w/ Perceptual Sufficiency
98.3
81.4
w/ Control Sufficiency
98.5
82.3
w/ Perceptual + Control Sufficiency
98.9
83.2
TABLE IV: Ablation study of perceptual and control sufficiency shaping. Success rates (%) on LIBERO and LIBERO-Plus.
Geometry
Semantics
Variant
Pearson ↑
AbsRel ↓
mIoU ↑
Dice ↑
Naive Single-Step
0.804
0.188
0.388
0.486
SLIP-VLA (Ours)
0.916
0.112
0.590
0.683
TABLE V: Quantitative analysis of the perceptual sufficiency of frozen single-step latent representations.
Evaluation Condition
Distance ↓
Δ vs. Correct ↑
Forward Dynamics
Correct Action
0.012
–
Gripper Shuffled
0.527
0.515
Motion Shuffled
0.772
0.759
Fully Shuffled
0.876
0.864
Inverse Dynamics
TABLE VI: Quantitative analysis of control sufficiency in SLIP-VLA. Forward dynamics measures latent prediction distance under different action perturbations, while inverse dynamics measures action recovery error under different latent-transition perturbations. Lower is better.
Source Initialization
LIBERO
LIBERO-Plus
Gaussian Noise
98.6
82.5
Current Latent
98.8
82.9
Current + Noise (Ours)
98.9
83.2
TABLE VII: Ablation study of source initialization strategies. Success rates (%) on LIBERO and LIBERO-Plus.
Task
π0
SLIP-VLA (Ours)
In-Distribution
Corn → Plate
72
86
Out-of-Distribution
Unseen Background
58
70
Unseen Object Appearance
54
70
Unseen Object Orientation
52
66
TABLE VIII: Real-world success rates (%). Each task condition is evaluated over 50 trials.
Department of Electric Engineering, Korea Advanced Institute of Science and Technology · University Of Science and Technology Beijing · Shanghai Jiao Tong University +1