Organizations: University of Washington · University of Georgia · Western Washington University · King Abdulaziz City for Science and Technology · HUMAIN
World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.86% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 1.45% and 1.00% on two DreamWAM configurations and to 10.60% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.
Figures & tables
Figure 1: TAPDreamer constructs a patch using a public visual encoder and observations from one source task, then freezes it for deployment across tasks and WAMs within each benchmark. Examples compare clean and patched execution on drawer opening with FastWAM, pot placement with Motus, and bowl placement with DreamWAM.
Figure 2: Overview of TAPDreamer . (a) Pixel updates maximize the latent displacement through a frozen public visual encoder. (b) The resulting patch is frozen and reused at each step across tasks and models within the same benchmark.
Setup
LIBERO
RoboTwin
Spatial
Object
Goal
Long-hor.
Overall
Clean
98.00
±0.00
99.60
±0.00
97.60
±0.00
95.40
±0.00
97.65
±0.00
90.86
±0.00
Random
85.60
-12.40
78.00
-21.60
90.00
-7.60
72.20
-23.20
81.45
-16.20
79.20
-11.66
TAPDreamer
0.00
-98.00
0.00
-99.60
0.00
-97.60
0.00
-95.40
0.00
-97.65
0.00
-90.86
Table 1: Closed-loop FastWAM success rates (percent) on complete task sets. LIBERO and RoboTwin each use one fixed patch, optimized on Spatial t6 and pick diverse bottles , respectively. Small subscripts report the signed success-rate difference from Clean, in percentage points. LIBERO Overall averages the four suites.
Victim
Suite
Clean
Random
TAPDreamer
LIBERO
DW-U
Spatial
96.20
±0.00
82.20
-14.00
0.00
-96.20
DW-U
Object
99.60
±0.00
96.00
-3.60
0.00
-99.60
DW-U
Goal
98.00
±0.00
90.00
-8.00
5.80
-92.20
DW-U
Long-horizon
96.20
±0.00
70.00
-26.20
0.00
-96.20
DW-U
Overall
97.50
±0.00
84.55
-12.95
1.45
-96.05
Table 2: Frozen-patch transfer across models within each benchmark. LIBERO results use the first optimization initialization. Success rates are percentages; small subscripts report signed differences from Clean for the same model and suite, in percentage points. DW-U and DW-J denote DreamWAM-uncond and DreamWAM-joint. LIBERO Overall averages the four suites.
Attention input
Attention output
Final latent
Random
0.007
0.008
0.013
TAPDreamer
0.007
0.708
0.794
Table 3: Fraction of squared representation change outside the projected patch footprint: ∥ΔXO∥F2/∥ΔX∥F2 , where ΔX=Xp−X , X∈{U,H,Z} denotes attention input, attention output, and final latent, respectively, and O denotes positions outside the footprint at the corresponding layer. Values are medians over 48 states.
Figure 4: Spatial spread and attention restoration. (a,b) Per-position final-latent L2 changes with a shared color scale; dashed boxes mark the patch footprint. (c) Mean L2 change outside the footprint over 48 states, with 95% confidence intervals. The dotted line denotes the random patch.
Metric
TAPDreamer
Random
Cross-task cosine similarity
0.971
0.603
Magnitude CV
0.056
0.071
Template cosine similarity
0.981
0.716
Relative L1 residual
0.206
0.809
L1 bound ratio
0.833
0.069
Table 4: Consistency of final-latent changes outside the patch footprint. CV denotes coefficient of variation.
Figure 5: Optimization trajectories for three initializations: (a) latent L1 displacement and (b) source-template agreement. Dotted lines denote random-patch results.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setup
Upper left
Upper right
Lower left
Lower right
(0,0)
(0,143)
(143,0)
(143,143)
Random
81.45
97.55
96.20
93.80
Reoptimized
0.00
-81.45
0.00
-97.55
0.00
-96.20
0.00
-93.80
Moved
0.00
-81.45
93.50
-4.05
94.85
-1.35
95.00
+1.20
Appendix
Table 5: Ablation study of patch position. Reoptimized patches are optimized at the given position; moved patches are copies of the unchanged upper-left patch. Success rates are percentages; subscripts report signed differences from Random at the same position, in percentage points.
Setup
LIBERO
Spatial
Object
Goal
Long-hor.
Overall
Clean
98.00
99.60
97.60
95.40
97.65
Random
85.60
-12.40
78.00
-21.60
90.00
-7.60
72.20
-23.20
81.45
-16.20
UPA-RFAS
84.20
-13.80
76.80
-22.80
77.40
-20.20
67.80
-27.60
76.55
-21.10
UADA
82.00
-16.00
77.00
-22.60
77.20
-20.40
67.20
-28.20
75.85
-21.80
TAPDreamer
0.00
-98.00
0.00
-99.60
0.00
-97.60
0.00
-95.40
0.00
-97.65
Appendix
Table 6: Direct transfer of VLA patches to FastWAM. Values are LIBERO task success rates (percent); subscripts report signed differences from Clean in the same suite or Overall column, in percentage points. Clean, Random, and TAPDreamer use the reference results in Table 1 .
Restored component
Latent retained
Action retained
Random latent retained
Restore clean attention branch
0.3327
0.321
1.000
Weights A restored to clean
0.3364
0.320
1.000
Values V restored to clean
0.3327
0.321
1.000
Appendix
Table 7: Fraction of patch-induced change retained after attention restoration. Latent and action columns report TAPDreamer ; the last column reports random-patch latent retention. Values are medians over 48 states.
Figure 7: Evolution of perturbation structure across observations during optimization. For three initializations, the panels show displacement on source and held-out observations, source-template agreement, the effective rank of 48 perturbation fields, and the ratio of the frozen source-template lower bound to measured L1 displacement. In (a), solid and dashed curves denote source and held-out states, respectively. Dotted lines show results with random patches.
Figure 8: Example of closed-loop behavior under TAPDreamer . Left: observation sequences under clean and perturbed conditions. Right: corresponding three-dimensional end-effector position trajectories. Green dashed and red solid paths denote Clean and TAPDreamer , respectively. The triangle marks the starting position and stars mark the endpoints; t counts environment execution steps. Positions are in meters in the world coordinate frame.
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.
Qi Li, Xingyi Yang, Xinchao Wang
National University of Singapore · The Hong Kong Polytechnic University
World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions. To diagnose this failure, we conduct action-head attention analysis and causal interventions. We find that the action decoder fails to focus on task-relevant interaction regions and remains sensitive to perturbations in task-irrelevant areas. This reveals a representation mismatch: hidden states optimized for visual reconstruction are not inherently organized in a form useful for low-level action control. In this paper, we propose AGRA, an Action-Grounded Representation Alignment objective that regularizes the world-action interface by aligning intermediate video diffusion features with spatially coherent semantic representations from a foundation visual encoder. We evaluate AGRA on real-world manipulation tasks. Experiments show that AGRA makes world model representations more action-grounded: by focusing the action decoder on the correct interaction regions, it improves object localization accuracy and affordance understanding, and makes the policy more robust to perturbations in task-irrelevant regions. As a result, AGRA consistently improves both in-distribution performance and out-of-distribution generalization over the baseline world action model.
Generative visual models offer a foundation for learning representations of physical dynamics, yet their extension to continuous control raises a fundamental question: do visual prediction and action generation require separate computational pathways? Existing approaches usually introduce trainable action heads or separate action experts to bridge low-dimensional states and high-dimensional visual representations. In this work, we explore whether the visual backbone's existing capacity can also support control when actions are expressed in a compatible representation. Thus, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch. This allows a single model to predict both how the robot should move and what the scene may look like afterward. Visual prediction and action generation become parts of the same generative process, without a dedicated action head or separate action expert. Experiments with subsampled training windows show gains over a matched dual-expert control, while benchmark evaluations reach 91.8% success rate on LIBERO-Plus and 96.12% on RoboTwin 2.0 in a full-data setting with additional augmented demonstrations. More broadly, the result suggests that capability need not be added where it can be inherited: the constraint on extending a generative backbone is the interface a new signal is written in, not the capacity to model it.
Tianheng Wang, Zhou Xie, Heng Jia +3
Westlake University · Lanzhou University · Zhejiang University +2