World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but existing PTQ methods such as smoothing and rotation are insufficient to maintain the precision of action generation. To overcome this limitation, we propose Q-WAM, a new 4-bit weight-activation quantization for WAMs that preserves the actions the model generates. Specifically, we introduce the \textit{Action Observability Gramian (AOG)}, which measures how much rounding errors in each weighted combination of a layer's input channels change the final action through all denoising steps. We also develop Action-Subspace Protection (ASP), which keeps the few most action-sensitive channel combinations in a tiny 16-bit low-rank branch and quantizes the complementary weights and activations to 4 bits, both as dense matrix multiplications that run efficiently on GPUs. Finally, to preserve action quality with minimal overhead, we identify the experts that matter most for the generated action by aggregating the AOG-derived action mass across the layers of each expert and apply ASP only to those experts. We evaluate Q-WAM on three WAMs, both in simulation and in real-world deployment. On the RoboTwin 2.0 benchmark, it reaches 89.6--93.0% average success rate, within 1.1 percentage points of the 16-bit models, while reducing the memory of the quantized blocks by 3.1--3.4×. Our method outperforms the strongest baseline, SVDQuant, by 2.5--8.7 percentage points. On a Unitree G1 humanoid and a bimanual UR3 robot, it improves success over SVDQuant by 12.8-17.6 percentage points.
Figures & tables
Figure 1: Removing outliers is not enough. (a) The activation entering a layer has a few outliers. (b) Smoothing and rotation remove them. (c) Quantizing the layer still damages the action, through a few AOG directions. (d) Our proposed ASP protects those directions and the damage is gone.
Figure 2: Overview of Q-WAM. (a) The AOG Gℓ measures how far a rounding error in each weighted combination of layer ℓ ’s input channels moves the action in later denoising steps. (b) ASP keeps the most sensitive of these combinations in a 16-bit branch and quantizes the deflated remainder to 4 bits; the action mass decides which experts receive it.
Figure 3: Deciding what to protect with the AOG. (a) Action error from quantizing a single layer, against the layer’s local error and its AOG score. (b) The first 128 eigenvalues of the rotated AOG for several action-expert layers, which fall quickly; the dashed line marks the rank we keep. (c) The action expert’s share of the action mass and of the parameters in each model.
Model
Method
Precision
BPW
Clean
Randomized
Avg.
Mem (GB)
Fast-WAM ( 2 -expert MoT, video generation backbone)
bf16 (upper bound)
W16A16
16.00
92.28
91.34
91.81
11.85
SmoothQuant
W4A4
4.00
21.98
16.48
19.23
2.98
ViDiT-Q
W4A4
4.41
82.00
80.48
81.24
3.61
SVDQuant
W4A4
4.84
89.12
87.38
88.25
3.60
Q-WAM (Ours)
W4A4
4.62
91.62
89.94
90.78
3.44
Table 1: RoboTwin 2.0 results. BPW is the weight cost; Mem is the targeted-block memory.
Figure 4: Real-world evaluation tasks. Frames progress from left to right within each task.
Unitree G1
UR3 Bimanual
Model
Method
Tool Sort
Item Cls.
Table Clean
Drawer
Cube Stacking
Avg. (%)
bf16
19/25
13/25
17/25
20/25
12/25
64.8
SVDQuant
16/25
10/25
12/25
12/25
4/25
43.2
Fast-WAM
Q-WAM (Ours)
18/25
11/25
14/25
19/25
8/25
56.0
bf16
21/25
17/25
22/25
23/25
21/25
83.2
SVDQuant
15/25
12/25
19/25
22/25
5/25
58.4
Table 2: Real-world results (successes/trials) on two embodiments and five tasks.
Model
Method
BPW
Clean
Randomized
Avg.
Mem (GB)
bf16 (upper bound)
16.00
92.28
91.34
91.81
11.85
per-group W4A4
4.50
82.60
81.80
82.20
3.35
+ smoothing and rotation
4.50
87.78
87.32
87.55
3.35
Fast-WAM
+ ASP (Ours)
4.62
91.62
89.94
90.78
3.44
bf16 (upper bound)
16.00
92.82
93.70
93.26
9.11
per-group W4A4
4.50
21.16
20.14
20.65
2.64
Table 3: Component ablation on RoboTwin 2.0 at a fixed weight group size.
Table 8
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: ImageWAM’s memory comparison for the ASP overhead.
Figure 7: Real-world experiment setup.
model
activation
weight
BPW
clean
randomized
Avg.
Mem (GB)
Fast-WAM
ASP
group INT4
4.62
91.62
89.94
90.78
3.44
ASP
SVDQuant
4.95
91.44
90.54
90.99
3.68
ImageWAM
ASP
group INT4
4.58
93.00
92.94
92.97
2.69
ASP
SVDQuant
4.83
90.56
90.16
90.36
2.83
LingBot-VA
ASP
group INT4
4.76
90.20
88.92
89.56
3.27
ASP
SVDQuant
5.03
90.52
88.88
89.70
3.43
Appendix
Table 5: Q-WAM with SVDQuant as the weight quantizer. ASP is held fixed on the activation side and only the weight side varies, between group-wise INT4 and SVDQuant’s low-rank decomposition.
Figure 8: Share of Fast-WAM’s action mass carried by each linear layer, for the action expert (top) and the video expert (bottom) on one logarithmic color scale. Rows are layer types and columns are blocks. The bars on the right give each layer type’s share of its own expert’s mass. Gray cells have exactly zero mass.
Calibration seed
Clean
Randomized
Avg.
42 (original set)
93.08
92.28
92.68
7
92.62
93.08
92.85
13
91.48
91.98
91.73
21
92.34
92.44
92.39
Mean ± std.
92.38±0.67
92.44±0.46
92.41±0.49
Appendix
Table 6: Sensitivity of ImageWAM to the calibration set. Each row recomputes the AOG, the protected subspaces and the checkpoint from a different set of calibration episodes, one per task, and evaluates it on the full RoboTwin suite.
bf16
SmoothQuant
ViDiT-Q
SVDQuant
Q-WAM
Task
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
adjust_bottle
100
100
47
21
100
100
100
100
100
100
beat_block_hammer
97
98
18
10
94
94
98
94
100
96
blocks_ranking_rgb
100
100
1
0
96
97
100
99
100
100
blocks_ranking_size
95
93
1
0
83
90
89
91
90
92
click_alarmclock
99
100
96
92
100
100
100
100
100
100
Appendix
Table 7: Per-task success rates (%) of Fast-WAM on RoboTwin 2.0 for every method in Table 1 , under the clean (Clean) and domain-randomized (Rand.) conditions. The best quantized average is shown in bold.
bf16
SmoothQuant
ViDiT-Q
SVDQuant
Q-WAM
Task
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
adjust_bottle
100
100
96
97
99
99
100
100
100
100
beat_block_hammer
100
98
42
27
91
97
95
97
98
92
blocks_ranking_rgb
96
99
89
91
73
82
97
98
99
99
blocks_ranking_size
92
97
67
71
60
82
95
93
93
96
click_alarmclock
100
100
100
100
100
99
100
99
99
100
Appendix
Table 8: Per-task success rates (%) of ImageWAM on RoboTwin 2.0 for every method in Table 1 , under the clean (Clean) and domain-randomized (Rand.) conditions. The best quantized average is shown in bold.
bf16
SmoothQuant
ViDiT-Q
SVDQuant
Q-WAM
Task
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
adjust_bottle
96
96
96
90
95
95
100
96
98
94
beat_block_hammer
94
100
90
86
90
83
98
94
100
92
blocks_ranking_rgb
96
94
82
58
76
65
88
88
94
96
blocks_ranking_size
96
82
66
46
81
53
86
62
100
80
click_alarmclock
100
100
100
98
100
100
100
100
100
100
Appendix
Table 9: Per-task success rates (%) of LingBot-VA on RoboTwin 2.0 for every method in Table 1 , under the clean (Clean) and domain-randomized (Rand.) conditions. The best quantized average is shown in bold.
College of Intelligent Robotics and Advanced Manufacturing, Fudan University · School of Data Science and Engineering, East China Normal University · Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University
Robotics Program, Colorado School of Mines, Golden, CO USA. · Department of Civil and Coastal Engineering, University of Florida, Gainesville, FL USA. · University of Southern California - Institute for Creative Technology, Los Angeles, CA USA. +1