In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, applying this principle to the learned latent space of WMs introduces a fundamental challenge: because WMs have fully learned state spaces and dynamics inferred from high-dimensional observations, it is unclear how to define latent-space disturbances that faithfully represent uncertainty in the underlying system. Our key idea is to model a latent-space disturbance as a perturbation to the learned latent dynamics that induces pessimistic but plausible transitions. Specifically, we construct a set of plausible latent dynamics by combining a dynamics-aware similarity metric that captures plausible transitions with out-of-distribution detection that excludes implausible latent states. We calibrate this uncertainty set over latent dynamics using conformal prediction, ensuring that WM imaginations induced by the latent disturbance remain plausible without becoming overly pessimistic. We then jointly optimize robust robot actions and the worst-case latent disturbances through game-theoretic optimization. We leverage this latent-space robust optimization to robustify policy steering, considering two paradigms: latent safety filtering and sample-and-verify steering of a generative control policy. Our controlled simulation experiments show that our latent disturbance enables robust decision-making directly in WM latent spaces, and hardware experiments with a Franka manipulator show that modeling latent disturbances enables robust policy steering, reducing failures by 70% in safety filtering and 54% in sampling-based policy steering. Project website: https://junwon.me/LatentDisturbance/.
Figures & tables
Fig. 2: Naughty 3D Dubins’ Car. (a) Environment with the vehicle and a failure set at the center. (b) For an action sequence that turns right, an adverse disturbance can instead drive the vehicle into failure, making the decision non-robust. (c) Driving straight remains safe even in the worst case. However, an implausible imagination can make this robust action appear unsafe.
Naughty
Positional Disturbances
Method
B.Acc. ↑
FPR ↓
FNR ↓
Fail ↓
B.Acc. ↑
FPR ↓
FNR ↓
Fail ↓
Nominal
0.932
0.006
0.130
0.114
0.831
0.000
0.337
0.218
WoN
0.959
0.057
0.026
0.061
0.911
0.162
0.017
0.034
CVaR0.05
0.944
0.017
0.095
0.062
0.943
0.028
0.087
0.084
CVaR0.1
0.935
0.012
0.118
0.061
0.936
0.018
0.109
0.101
CVaR0.2
0.927
0.010
0.135
0.062
0.927
0.011
0.135
0.134
TABLE I: 3D Dubins’ Car World Model: Quantitative Results.
Fig. 3: Qualitative Results: 3D Dubins’ Car. Solid lines represent the ground-truth unsafe-set boundaries, and red regions denote the unsafe sets computed by each method. LUCID closely approximates the ground-truth unsafe sets of robust optimization under two different types of disturbances. In contrast, Nominal is overly optimistic, while the sampling-based baselines such as WoN and CVaR inaccurately approximate unsafe sets.
Uncertainty Set
Naughty
Positional Disturbances
Method
Dyn.
Conf.
OOD
B.Acc. ↑
FPR ↓
FNR ↓
B.Acc. ↑
FPR ↓
FNR ↓
LUCID
✓
✓
✓
0.958
0.037
0.046
0.956
0.023
0.066
w/o OOD
✓
✓
✗
0.809
0.381
0.001
0.621
0.759
0.000
0.5ϵKL
✓
✗
✓
0.927
0.011
0.136
0.864
0.001
0.271
2.0ϵKL
✓
✗
✓
0.926
0.126
0.022
0.630
0.733
0.007
Euclidean
✗
✓
✓
0.605
0.790
0.000
0.875
0.001
0.250
TABLE II: 3D Dubins’ Car: Uncertainty Set Ablation. Dyn., Conf., and OOD represent dynamics-aware similarity, conformal calibration, and in-distribution constraints, respectively.
Fig. 4: Qualitative: Naughty Dubins’ Car with Optimistic Disturbances. Solid lines show the ground-truth unsafe-set boundaries of the disturbance-free system, and dashed lines show the failure set. Without the in-distribution constraint, the disturbance exploits implausible latent states, trivially imagining trajectories that avoid failure from any non-failure state.
Fig. 5: Safety Value Function: Nominal vs. LUCID . Safeguarding the same πtask from the same initial states, the robust safety filter proactively intervenes based on a plausibly pessimistic imagination ( o^28d ) in which πtask can lead to failure, while the nominal imagination ( o^28 ) predicts a non-failure outcome. The resulting robust action prevents failure. In contrast, the nominal latent safety filter overestimates safety and intervenes too late based on an optimistic imagination ( o^50 ), leading to failure under an adverse realization.
Safety Filter
Success ↑
Failure ↓
Safety Gain ↑
Cond. Failure ↓
Robust Rate ↑
No Filter
0.33
0.47
–
–
0.10
Nominal
0.31
0.40
0.17 ± 0.2
0.46 ± 0.5
0.11
WoN
0.35
0.31
-0.04 ± 0.2
0.32 ± 0.4
0.07
CVaR 0.1
0.30
0.31
-0.11 ± 0.2
0.32 ± 0.4
0.10
UnConf.
0.05
0.06
0.10 ± 0.03
0.06 ± 0.23
0.90
LUCID
0.34
0.11
0.21 ± 0.15
0.11 ± 0.31
0.77
TABLE V: Results: Trajectory Replay Across 20 Different Physics.
Fig. 6: Block Pouring: World Model Imaginations. (a) Initial state: the orange block is tilted while πtask remains still. (b) The nominal imagination predicts a non-failure outcome. (c) With the conformalized uncertainty set, the latent disturbance imagines a plausible adverse transition in which the green block slides and falls. (d) Without the conformalized uncertainty set, the latent disturbance induces an implausible transition in which the green block falls in an infeasible manner, resulting in an overly pessimistic imagination.
πtask
Safety Filter
Success ↑
Failure ↓
Safety Gain ↑
Cond. Failure ↓
DreamerV3
No Filter
0.43
0.56
–
–
Nominal
0.60
0.39
0.21 ± 0.24
0.40 ± 0.48
WoN
0.34
0.54
0.10 ± 0.20
0.54 ± 0.49
CVaR 0.1
0.00
0.34
-0.05 ± 0.05
0.34 ± 0.47
UnConf.
0.00
0.00
0.29 ± 0.08
0.00 ± 0.06
LUCID
0.81
0.18
0.27 ± 0.22
0.19 ± 0.39
TABLE VI: Ablation: Deterministic Block-Pouring with Fixed Physics
Fig. 7: Qualitative Results: Preventing Sunny-Side-Down Egg. LUCID preemptively identifies teleoperator actions that may cause the egg to fall or flip and intervenes with robust actions that remain safe under worst-case dynamics. In contrast, the Nominal safety filter relies on overly optimistic safety values and non-robust actions, leading to sunny-side-down outcomes.
Fig. 8: Filtering the Same πtask Across Different Surfaces. ( Left ) Failure rates under each spatula surface condition. ( Right ) Robust rate, measuring the fraction of trajectories that remain safe across all surface conditions. LUCID consistently minimizes failures and achieves a higher robust rate, whereas the Nominal safety filter prevents failures only under certain surface conditions.
Fig. 9: Quantitative Results: Success Rates. ( Left ) Safety filtering over 20 rollouts for Diffusion Policy and π0.5 . ( Right ) Steering π0.5 with sample-and-verify action selection, using world-model imaginations to evaluate candidate action chunks. In both settings, LUCID robustly improves success rates.
Fig. 10: Qualitative: Policy Steering with World Model. (a) Initial observation and sampled candidate actions. (b) For each candidate, Nominal imagines outcomes using fz , LUCID computes worst-case fzπd within the calibrated uncertainty set F , while Unconformalized fzπd^ can exploit out-of-distribution latent states. (c) Imagined outcomes for candidate actions. LUCID selects an action that remains successful under the plausible worst-case, whereas the other schemes lose discriminability among candidate actions.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Gaussian
Categorical
Input
[h,μ,σ]
h
Input dimension
512+32+32
512
Hidden layers
4 layers, 512 units each
Normalization / activation
LayerNorm / ReLU
Output residual
(Δμ,Δσ)
Δϕ
Output dimension
32+32
32×32
Appendix
TABLE VII: Residual parameterization of πd in ( 11 ).
Hyperparameter
Value
Image Dimension
128×128×3
Action Dimension
1 (continuous)
Stochastic Latent
Gaussian
Deterministic Dimension
512
Stochastic Dimension
32
Activation Function
SiLU
Appendix
TABLE VIII: 3D Dubins’ Car Latent World Model Settings.
Parameter
Value
Actor Learning Rate
10−3
Critic Learning Rate
10−4
Disturbance Learning Rate
10−4
Discount Factor
0.85→0.9999
Training Iterations
100,000
Replay Buffer Size
100,000
Appendix
TABLE IX: 3D Dubins’ Car Latent Safety Filter Settings.
Parameter
Randomization Range
Cube Position x
[−0.02,0.02] m
Cube Position y
[−0.02,0.02] m
Mass
[0.8,1.2] kg
Static Friction
[0.2,0.8]
Dynamic Friction
[0.1,0.5]
Restitution
[0.05,0.3]
Appendix
TABLE X: Physical Randomization of Block Pouring in IsaacLab.
Hyperparameter
Value
Image Dimension
2×128×128×3
Proprioception Dimension
7 (eef position/quaternion)
Action Dimension
7 (continuous)
Stochastic Latent
Categorical
Deterministic Dimension
512
Stochastic Dimension
32×32
Appendix
TABLE XI: Block Pouring World Model Training Settings.
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions on the dynamics of a world model. If the world model is incorrect, the resulting value function can inherit its errors and produce overconfident safety estimates. Existing latent safety filters often rely on auxiliary signals such as ensemble disagreement or value-target consistency residuals for adaptation, but these signals can remain small even when the world model's predictions deviate from observations. We propose an adaptive latent safety filter that calibrates safety reasoning using directly observed world-model error. Our method uses Adaptive Conformal Inference to construct online uncertainty sets from discrepancies between predicted and observation-inferred latent states, then evaluates safety pessimistically by minimizing the learned value function over these sets. This allows the filter to remain minimally conservative when the world model is accurate, while becoming more cautious when observations reveal model mismatch. We provide a finite-time coverage guarantee for the adaptive uncertainty radius. Through simulation and hardware experiments, we show that our method significantly reduces failures relative to state-of-the-art latent safety filters while preserving task completion.
John Cao, Somil Bansal
Department of Aeronautics and Astronautics, Stanford University
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this limitation by conditioning policies on predicted futures, yet existing approaches typically rely on computationally expensive video generation with substantial pixel-level redundancy. We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. At the core of LaWAM is a latent-action-conditioned Latent World Model (LaWM). We obtain LaWM by training a latent action model in the latent space of a pretrained vision foundation model and repurposing its forward decoder to predict future observation features for scene evolution. LaWAM then conditions action generation on these predicted latent visual subgoals to enable dynamics-aware robot control. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference. LaWAM runs in 187 ms per action-chunk prediction and achieves up to 24x lower wall-clock latency than pixel-space WAMs.
Jialei Chen, Kai Wang, Kang Chen +9
Jilin University · Zhongguancun Academy · Nankai University +4