World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly different effects on final actions, making numerical accuracy alone insufficient for reliable control. We introduce SteerQuant, a 4-bit quantization framework for WAMs that steers errors toward computations with less influence on final actions. It maps how each stream's quantization errors affect final actions and uses this map to guide shared channel scaling. Activation scaling is further calibrated for each stream and denoising step to accommodate changes in activation ranges and action impact. This adapts quantization to different stream requirements without duplicating weights or increasing bit-widths for selected streams. To reduce the extra kernel launches and memory traffic introduced by scaling, we develop Rudder, a 4-bit inference engine for WAMs that fuses scaling and output compensation into low-bit kernels. Under W4A8 and W4A4, SteerQuant maintains mean LIBERO success within 0.8 percentage points of full precision, while delivering up to 2.23× denoising speedup over BF16 across three WAMs with reduced peak GPU memory usage. On a real dual-arm robot, W4A8 deployment achieves a 1.35× end-to-end inference speedup while maintaining average task success relative to BF16.
Figures & tables
Figure 1: Control effects of matched local errors. (a) Video- and action-stream comparisons at Blocks 14 and 24 at the final denoising step, with local errors matched within each block. (b) Action-stream comparisons across Steps 0–4 at Block 24, with local errors matched across steps.
Figure 2: Overview of SteerQuant. Γℓ,τ is a diagonal matrix of token-wise stream gains. The upper-right gray box depicts offline weight preparation; the other gray boxes denote fused inference kernels.
Figure 3: Rudder’s fused execution. The first kernel combines channel scaling, stream gains, and activation quantization. The second performs low-bit GEMM and applies dequantization, inverse-gain compensation, and bias before writing the output. Only integer activations and row scales pass between kernels. Dashed boxes indicate kernel boundaries.
LIBERO
RoboLab
Cosmos-Policy
FastWAMJoint
Cosmos-Edge
Precision
Method
Spatial ↑
Object ↑
Goal ↑
Long ↑
Avg. ↑
Spatial ↑
Object ↑
Goal ↑
Long ↑
Avg. ↑
SR ↑
FP
Original
98.3
100.0
97.9
98.0
98.5
99.2
99.4
98.6
98.4
98.9
21.83
W4A8
SmoothQuant
98.0
99.7
98.4
95.6
97.9
96.8
98.4
97.2
96.2
97.2
13.67
PTQ4DiT
98.8
99.6
96.1
97.7
98.1
97.8
98.6
98.0
96.0
97.6
13.67
Q-DiT
97.5
99.9
96.8
96.7
97.7
97.4
98.4
97.8
96.6
97.6
12.17
Table 1: Closed-loop task success rates (%) on LIBERO and RoboLab. For LIBERO, Avg. denotes the mean across the four suites.
Cosmos-Policy
FastWAMJoint
Cosmos-Edge
Precision
Method
Lat. ↓
Spd. ↑
Peak ↓
Storage ↓
Lat. ↓
Spd. ↑
Peak ↓
Storage ↓
Lat. ↓
Spd. ↑
Peak ↓
Storage ↓
FP
Original
261.460
1.000×
4.241
3.644
360.450
1.000×
12.116
11.230
881.637
1.000×
8.492
6.276
SmoothQuant
134.967
1.937×
1.607
1.191
263.368
1.369×
3.363
2.983
570.893
1.544×
6.405
4.314
W4A8
SteerQuant
133.455
1.959×
1.642
1.191
260.977
1.381×
3.363
2.983
563.858
1.564×
6.397
4.312
W4A4
QuaRot
122.657
2.132×
1.604
1.189
221.647
1.626×
3.363
2.983
569.943
1.547×
6.396
4.312
Atom
192.997
1.355×
1.679
1.258
309.021
1.166×
3.510
3.188
766.027
1.151×
6.463
4.361
Table 2: Inference efficiency and memory footprint across models. Lat.: denoising latency (ms); Spd.: speedup relative to FP; Peak: peak allocated GPU memory (GB); Storage: packed weight storage (GiB). PTQ4DiT and Q-DiT are omitted because their official implementations do not provide low-bit inference kernels for the evaluated settings.
Variant
Routing
Modulation
Weighting
RMSE ↓
95% CI of Δ RMSE
Base
×
×
Uniform
0.1581
[−0.0430,−0.0357]
Routing only
✓
×
Map-derived
0.1224
[−0.0044,−0.0026]
Modulation only
×
✓
Map-derived
0.1529
[−0.0368,−0.0312]
Both, uniform
✓
✓
Uniform
0.1416
[−0.0262,−0.0052]
SteerQuant
✓
✓
Map-derived
0.1189
–
Table 3: Contribution of individual components and action guidance. Standardized action RMSE is averaged over W4A8 and W4A4 on Cosmos-Policy, using 1,800 observations from 600 trajectories across 40 tasks. Paired, unadjusted 95% CIs are computed by resampling whole trajectories within each task–seed stratum and refer to ΔRMSE=RMSESteerQuant−RMSErow . Negative values favor SteerQuant.
Figure 4: Action-guided quantization error analysis and real-world execution. Left: Cosmos-Policy W4A8 on 1,800 observations. Errors are token-normalized squared Frobenius differences from FP outputs under matched FP inputs, averaged over observations and streams within each group. Axes show percentage changes relative to uniform weighting on symmetric logarithmic scales. Green points indicate lower high-impact and higher low-impact error. Right: Representative SteerQuant executions on the AgileX Cobot Magic dual-arm platform. Rows show battery insertion, object packing, and cup stacking, with time progressing from left to right.
Method
Battery Insertion
Pack Objects
Stack Cups
Avg. score (%) ↑
Latency (ms) ↓
Speedup ↑
FP (BF16)
83.33
69.44
75.00
75.93
462.79
1.00×
SteerQuant (W4A8)
75.00
88.89
83.33
82.41
341.97
1.35×
Table 4: Real-world task performance and inference efficiency. Task scores are reported in percent and averaged equally across tasks. Scoring details are provided in Appendix D . We report the median end-to-end inference latency and speedup, measured from sending a policy request to receiving the predicted actions. Robot motion execution time is excluded.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Comparison
Block
Step
Stream
Action RMSE
Success
FP
–
–
–
–
7/9
Stream–depth
14
4
Video
0.246±0.0292
5/9
Stream–depth
14
4
Action
0.0833±0.00788
9/9
Stream–depth
24
4
Video
0.0500±0.00575
9/9
Stream–depth
24
4
Action
0.361±0.0458
4/9
Step
24
0
Action
0.00330±0.000331
7/9
Appendix
Table 5: Controlled intervention summary. RMSE is mean ± s.d.; success is the number of successful rollouts out of nine. The Block 24, Step 4 action condition is shared by both comparisons.
Figure 5: Relative action impact of semantic streams. Scores aggregate over 320 observations and eight stream-aligned Linear families, additionally averaging over steps in (a) and blocks in (b). Each column is normalized by its maximum and multiplied by 100; colors compare streams within a column. The marker indicates Block 17.
Observations
Spearman correlation
Top-10% overlap
5
0.898 [0.859, 0.958]
0.957 [0.914, 0.965]
10
0.927 [0.864, 0.969]
0.966 [0.929, 0.973]
20
0.947 [0.874, 0.981]
0.973 [0.954, 0.979]
40
0.960 [0.911, 0.987]
0.980 [0.972, 0.984]
80
0.977 [0.938, 0.993]
0.985 [0.981, 0.988]
160
0.993 [0.970, 0.999]
0.989 [0.986, 0.992]
Appendix
Table 6: Calibration-size convergence. Median agreement with the 320-observation map, with 2.5th–97.5th percentile intervals over 200 subsets per size, sampled without replacement within each subset.
World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but existing PTQ methods such as smoothing and rotation are insufficient to maintain the precision of action generation. To overcome this limitation, we propose Q-WAM, a new 4-bit weight-activation quantization for WAMs that preserves the actions the model generates. Specifically, we introduce the \textit{Action Observability Gramian (AOG)}, which measures how much rounding errors in each weighted combination of a layer's input channels change the final action through all denoising steps. We also develop Action-Subspace Protection (ASP), which keeps the few most action-sensitive channel combinations in a tiny 16-bit low-rank branch and quantizes the complementary weights and activations to 4 bits, both as dense matrix multiplications that run efficiently on GPUs. Finally, to preserve action quality with minimal overhead, we identify the experts that matter most for the generated action by aggregating the AOG-derived action mass across the layers of each expert and apply ASP only to those experts. We evaluate Q-WAM on three WAMs, both in simulation and in real-world deployment. On the RoboTwin 2.0 benchmark, it reaches 89.6--93.0% average success rate, within 1.1 percentage points of the 16-bit models, while reducing the memory of the quantized blocks by 3.1--3.4×. Our method outperforms the strongest baseline, SVDQuant, by 2.5--8.7 percentage points. On a Unitree G1 humanoid and a bimanual UR3 robot, it improves success over SVDQuant by 12.8-17.6 percentage points.
Arash Akbari, Arman Akbari, Jingwu Luo +7
Northeastern University · EmbodyX · University of Minnesota - Twin Cities +1
World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-training quantization (PTQ) methods are poorly suited to WAMs because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment. We present QuantWAMs, a PTQ framework that aligns quantization decisions with the calibration context defined by model structure, rollout distribution, and task objective. QuantWAMs introduces three strategies: shared-basis outlier calibration, which pools activation evidence only across coordinate-compatible modules; co-training-objective saliency, which computes empirical-Fisher scores from the joint video--action gradient and assigns weight precision at a calibration-stable layer granularity; and fixed-intervention rollout auditing, which revises denoising-step protection schedules using reachable closed-loop states without changing the precision budget. We evaluate QuantWAMs on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and real-robot manipulation with an AgiBot G2. Under a W4A4-dominant setting, the reported simulation means differ from FP16 by 0.2--0.7 percentage points. Real-robot trials further establish deployment feasibility on three manipulation tasks. For the targeted video and action blocks, QuantWAMs reduces peak weight-and-activation memory to about 29% of FP16 and provides 1.4--1.6× block-level speedups.
Jiacheng Zhou, Jinfan Lv, Ruixuan Li +4
College of Intelligent Robotics and Advanced Manufacturing, Fudan University · School of Data Science and Engineering, East China Normal University · Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present WAMACHINE, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that WAMACHINE achieves 1.47-3.05× speedups in observation-to-action latency and 2.23-3.27× speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Zhinnan Liu, Haozhi Han, Ruge Zhang +8
Xiamen University · Institute for AI Industry Research, Tsinghua University · School of Computer Science, Peking University +2