SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models
Authors: Tianfu Li, Haoxuan Xu, Wenbo Chen, Haitian Li, Changchuan Yang, Xinhu Zheng, Jun Ma, Yuan Liu, +2 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou). · The Hong Kong University of Science and Technology. · Nanyang Technological University. · Zhejiang University.
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive iterative denoising, while one-step alternatives can underperform their multi-step counterparts. To reconcile efficient future modeling with strong action performance, we present SLIP-VLA, a policy learning framework that equips VLA models with a Single-Step Latent Imagination for future-aware action prediction. SLIP-VLA obtains temporally dense future latent representations with a single denoising update, and we improve the perceptual sufficiency of these representations by aligning intermediate latents with future geometric and semantic features. We further improve their control sufficiency through action-conditioned latent world modeling and inverse dynamics modeling, explicitly coupling latent transitions with robot actions. SLIP-VLA achieves state-of-the-art performance across diverse simulation benchmarks and real-world manipulation tasks, while its single-step latent imagination takes only 12 ms.
Figures & tables
Fig. 2: Pipeline of SLIP-VLA. The VLM backbone first encodes the multimodal context, and the World DiT then performs a single denoising pass to produce temporally dense future latents. (a) For perceptual sufficiency shaping, selected intermediate representations are aligned with future VGGT- Ω geometric features and DINOv3 semantic features. (b) For control sufficiency shaping, deeper latent representations are regularized through forward and inverse latent dynamics to preserve action-dependent transitions. The shaped future representations are injected layer-wise into the Action DiT for action generation, while all auxiliary shaping modules are used only during training.
Method
Spatial
Object
Goal
Long
Avg.
General VLAs
OpenVLA [ 1 ]
84.7
88.4
79.2
53.7
76.5
OpenVLA-OFT [ 35 ]
97.6
98.4
97.9
94.5
97.1
π0 [ 2 ]
98.0
96.8
94.4
88.4
94.4
π0.5 [ 3 ]
98.8
98.2
98.0
92.4
96.9
Future-Aware VLAs
TABLE I: LIBERO success rates (%). Best and second-best results are bolded and underlined , respectively.
Method
Cam.
Robot
Lang.
Light
BG
Noise
Layout
Avg.
General VLAs
OpenVLA [ 1 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
OpenVLA-OFT [ 35 ]
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
π0 [ 2 ]
13.8
6.0
58.8
85.0
81.4
79.0
68.9
53.6
π0 -Fast [ 36 ]
65.1
21.6
61.0
73.2
73.2
74.4
68.8
61.6
Future-Aware VLAs
TABLE II: LIBERO-Plus success rates (%). Best and second-best results are bolded and underlined , respectively.
Method
Clean
Rand.
Avg.
General VLAs
π0 [ 2 ]
65.9
58.4
62.2
π0.5 [ 3 ]
82.7
76.8
79.8
Future-Aware VLAs
WorldVLA [ 4 ]
42.5
32.2
37.4
World Action Models
TABLE III: RoboTwin 2.0 success rates (%). Best and second-best results are bolded and underlined , respectively.
Fig. 3: Qualitative analysis of perceptual and control information in the single-step latent. (a) Compared with the naive single-step baseline, SLIP-VLA recovers more coherent future geometry and more concentrated task-relevant semantic responses. (b) For forward dynamics, the correct action produces a prediction closer to the target latent, while a shuffled action induces clear spatial deviations; the action-response map highlights regions most affected by the action change.
Fig. 4: Effect of denoising steps on future prediction and control performance. The top row shows representative predicted videos with one to four denoising steps, while the curves report PSNR and LIBERO-Plus success rate. Additional denoising provides no consistent improvement in visual prediction quality and does not improve action performance, with the highest success rate achieved using a single step.
Fig. 5: Real-world evaluation of SLIP-VLA. (a) The real-world manipulation platform. (b.1) Our SLIP-VLA successfully completes the in-distribution corn-to-plate task. (b.2) It remains successful under visual distribution shifts in background, object appearance, and object orientation. (b.3) SLIP-VLA further completes both two-step and three-step long-horizon tasks involving sequential interaction with the pot and the target object.
Variant
LIBERO
LIBERO-Plus
Naive Single-Step
97.7
79.0
w/ Perceptual Sufficiency
98.3
81.4
w/ Control Sufficiency
98.5
82.3
w/ Perceptual + Control Sufficiency
98.9
83.2
TABLE IV: Ablation study of perceptual and control sufficiency shaping. Success rates (%) on LIBERO and LIBERO-Plus.
Geometry
Semantics
Variant
Pearson ↑
AbsRel ↓
mIoU ↑
Dice ↑
Naive Single-Step
0.804
0.188
0.388
0.486
SLIP-VLA (Ours)
0.916
0.112
0.590
0.683
TABLE V: Quantitative analysis of the perceptual sufficiency of frozen single-step latent representations.
Evaluation Condition
Distance ↓
Δ vs. Correct ↑
Forward Dynamics
Correct Action
0.012
–
Gripper Shuffled
0.527
0.515
Motion Shuffled
0.772
0.759
Fully Shuffled
0.876
0.864
Inverse Dynamics
TABLE VI: Quantitative analysis of control sufficiency in SLIP-VLA. Forward dynamics measures latent prediction distance under different action perturbations, while inverse dynamics measures action recovery error under different latent-transition perturbations. Lower is better.
Source Initialization
LIBERO
LIBERO-Plus
Gaussian Noise
98.6
82.5
Current Latent
98.8
82.9
Current + Noise (Ours)
98.9
83.2
TABLE VII: Ablation study of source initialization strategies. Success rates (%) on LIBERO and LIBERO-Plus.
Task
π0
SLIP-VLA (Ours)
In-Distribution
Corn → Plate
72
86
Out-of-Distribution
Unseen Background
58
70
Unseen Object Appearance
54
70
Unseen Object Orientation
52
66
TABLE VIII: Real-world success rates (%). Each task condition is evaluated over 50 trials.
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. We propose SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy. SLIM learns action-grounded predictive latents that capture both action-conditioned future transitions and the actions that explain observed changes. SLIM learns these representations through self-supervised masked trajectory prediction, combining action reconstruction with future-latent prediction. A compact Mixture-of-Transformers (MoT) backbone models interactions between observation latents and action tokens. The resulting policy is trained with flow matching for language-conditioned action generation. Across simulation benchmarks and real-world evaluation, SLIM matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.
Jingkai Wang, Zihan Tang, Gu Zhang +7
1Fudan University · 2Beijing Academy of Artificial Intelligence · 3Tsinghua University +1
Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Imagine-RL, which augments noise-space VLA post-training with action-conditioned visual-torque imagination. For each candidate action chunk, a frozen visual-torque latent world model (VTLWM) autoregressively predicts compact future representations without pixel reconstruction. A current image-state-action query attends to observed histories and predicted futures, while previous-window prediction residuals provide token-wise confidence priors that suppress unreliable future tokens. By combining current evidence with predicted consequences, the action critic better evaluates candidate actions and supervises the actor, while the VLA and VTLWM remain frozen. Across four real-robot tasks with 50 evaluation trials per task, Imagine-RL uses only 100 RL trajectories and improves the average success rate by (23.6%) over DSRL and by (60%) over VLA baselines.
Kejia Hu, Wentong Zhai, Bo Zhao +1
Department of Electric Engineering, Korea Advanced Institute of Science and Technology · University Of Science and Technology Beijing · Shanghai Jiao Tong University +1
Visual-Language-Action models (VLAs) have advanced generalist robot control by mapping multimodal observations and language instructions directly to actions, but sparse action supervision often encourages shortcut mappings rather than representations of dynamics, contact, and task progress. Recent world-action models introduce future prediction through video rollouts, yet pixel-space prediction is a costly and indirect substrate for control, as it may model visual details irrelevant to action generation and introduces substantial training or inference overhead. We present Being-H0.7, a latent world-action model that brings future-aware reasoning into VLA-style policies without generating future frames. Being-H0.7 inserts learnable latent queries between perception and action as a compact reasoning interface, and trains them with a future-informed dual-branch design: a deployable prior branch infers latent states from the current context, while a training-only posterior branch replaces the queries with embeddings from future observations. Jointly aligning the two branches at the latent reasoning space leads the prior branch to reason future-aware, action-useful structure from current observations alone. At inference, Being-H0.7 discards the posterior branch and performs no visual rollout. Experiments across six simulation benchmarks and diverse real-world tasks show that Being-H0.7 achieves state-of-the-art or comparable performance, combining the predictive benefits of world models with the efficiency and deployability of direct VLA policies.