World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly different effects on final actions, making numerical accuracy alone insufficient for reliable control. We introduce SteerQuant, a 4-bit quantization framework for WAMs that steers errors toward computations with less influence on final actions. It maps how each stream's quantization errors affect final actions and uses this map to guide shared channel scaling. Activation scaling is further calibrated for each stream and denoising step to accommodate changes in activation ranges and action impact. This adapts quantization to different stream requirements without duplicating weights or increasing bit-widths for selected streams. To reduce the extra kernel launches and memory traffic introduced by scaling, we develop Rudder, a 4-bit inference engine for WAMs that fuses scaling and output compensation into low-bit kernels. Under W4A8 and W4A4, SteerQuant maintains mean LIBERO success within 0.8 percentage points of full precision, while delivering up to 2.23× denoising speedup over BF16 across three WAMs with reduced peak GPU memory usage. On a real dual-arm robot, W4A8 deployment achieves a 1.35× end-to-end inference speedup while maintaining average task success relative to BF16.
Figures & tables
Figure 1: Control effects of matched local errors. (a) Video- and action-stream comparisons at Blocks 14 and 24 at the final denoising step, with local errors matched within each block. (b) Action-stream comparisons across Steps 0–4 at Block 24, with local errors matched across steps.
Figure 2: Overview of SteerQuant. Γℓ,τ is a diagonal matrix of token-wise stream gains. The upper-right gray box depicts offline weight preparation; the other gray boxes denote fused inference kernels.
Figure 3: Rudder’s fused execution. The first kernel combines channel scaling, stream gains, and activation quantization. The second performs low-bit GEMM and applies dequantization, inverse-gain compensation, and bias before writing the output. Only integer activations and row scales pass between kernels. Dashed boxes indicate kernel boundaries.
LIBERO
RoboLab
Cosmos-Policy
FastWAMJoint
Cosmos-Edge
Precision
Method
Spatial ↑
Object ↑
Goal ↑
Long ↑
Avg. ↑
Spatial ↑
Object ↑
Goal ↑
Long ↑
Avg. ↑
SR ↑
FP
Original
98.3
100.0
97.9
98.0
98.5
99.2
99.4
98.6
98.4
98.9
21.83
W4A8
SmoothQuant
98.0
99.7
98.4
95.6
97.9
96.8
98.4
97.2
96.2
97.2
13.67
PTQ4DiT
98.8
99.6
96.1
97.7
98.1
97.8
98.6
98.0
96.0
97.6
13.67
Q-DiT
97.5
99.9
96.8
96.7
97.7
97.4
98.4
97.8
96.6
97.6
12.17
Table 1: Closed-loop task success rates (%) on LIBERO and RoboLab. For LIBERO, Avg. denotes the mean across the four suites.
Cosmos-Policy
FastWAMJoint
Cosmos-Edge
Precision
Method
Lat. ↓
Spd. ↑
Peak ↓
Storage ↓
Lat. ↓
Spd. ↑
Peak ↓
Storage ↓
Lat. ↓
Spd. ↑
Peak ↓
Storage ↓
FP
Original
261.460
1.000×
4.241
3.644
360.450
1.000×
12.116
11.230
881.637
1.000×
8.492
6.276
SmoothQuant
134.967
1.937×
1.607
1.191
263.368
1.369×
3.363
2.983
570.893
1.544×
6.405
4.314
W4A8
SteerQuant
133.455
1.959×
1.642
1.191
260.977
1.381×
3.363
2.983
563.858
1.564×
6.397
4.312
W4A4
QuaRot
122.657
2.132×
1.604
1.189
221.647
1.626×
3.363
2.983
569.943
1.547×
6.396
4.312
Atom
192.997
1.355×
1.679
1.258
309.021
1.166×
3.510
3.188
766.027
1.151×
6.463
4.361
Table 2: Inference efficiency and memory footprint across models. Lat.: denoising latency (ms); Spd.: speedup relative to FP; Peak: peak allocated GPU memory (GB); Storage: packed weight storage (GiB). PTQ4DiT and Q-DiT are omitted because their official implementations do not provide low-bit inference kernels for the evaluated settings.
Variant
Routing
Modulation
Weighting
RMSE ↓
95% CI of Δ RMSE
Base
×
×
Uniform
0.1581
[−0.0430,−0.0357]
Routing only
✓
×
Map-derived
0.1224
[−0.0044,−0.0026]
Modulation only
×
✓
Map-derived
0.1529
[−0.0368,−0.0312]
Both, uniform
✓
✓
Uniform
0.1416
[−0.0262,−0.0052]
SteerQuant
✓
✓
Map-derived
0.1189
–
Table 3: Contribution of individual components and action guidance. Standardized action RMSE is averaged over W4A8 and W4A4 on Cosmos-Policy, using 1,800 observations from 600 trajectories across 40 tasks. Paired, unadjusted 95% CIs are computed by resampling whole trajectories within each task–seed stratum and refer to ΔRMSE=RMSESteerQuant−RMSErow . Negative values favor SteerQuant.
Figure 4: Action-guided quantization error analysis and real-world execution. Left: Cosmos-Policy W4A8 on 1,800 observations. Errors are token-normalized squared Frobenius differences from FP outputs under matched FP inputs, averaged over observations and streams within each group. Axes show percentage changes relative to uniform weighting on symmetric logarithmic scales. Green points indicate lower high-impact and higher low-impact error. Right: Representative SteerQuant executions on the AgileX Cobot Magic dual-arm platform. Rows show battery insertion, object packing, and cup stacking, with time progressing from left to right.
Method
Battery Insertion
Pack Objects
Stack Cups
Avg. score (%) ↑
Latency (ms) ↓
Speedup ↑
FP (BF16)
83.33
69.44
75.00
75.93
462.79
1.00×
SteerQuant (W4A8)
75.00
88.89
83.33
82.41
341.97
1.35×
Table 4: Real-world task performance and inference efficiency. Task scores are reported in percent and averaged equally across tasks. Scoring details are provided in Appendix D . We report the median end-to-end inference latency and speedup, measured from sending a policy request to receiving the predicted actions. Robot motion execution time is excluded.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Comparison
Block
Step
Stream
Action RMSE
Success
FP
–
–
–
–
7/9
Stream–depth
14
4
Video
0.246±0.0292
5/9
Stream–depth
14
4
Action
0.0833±0.00788
9/9
Stream–depth
24
4
Video
0.0500±0.00575
9/9
Stream–depth
24
4
Action
0.361±0.0458
4/9
Step
24
0
Action
0.00330±0.000331
7/9
Appendix
Table 5: Controlled intervention summary. RMSE is mean ± s.d.; success is the number of successful rollouts out of nine. The Block 24, Step 4 action condition is shared by both comparisons.
Figure 5: Relative action impact of semantic streams. Scores aggregate over 320 observations and eight stream-aligned Linear families, additionally averaging over steps in (a) and blocks in (b). Each column is normalized by its maximum and multiplied by 100; colors compare streams within a column. The marker indicates Block 17.
Observations
Spearman correlation
Top-10% overlap
5
0.898 [0.859, 0.958]
0.957 [0.914, 0.965]
10
0.927 [0.864, 0.969]
0.966 [0.929, 0.973]
20
0.947 [0.874, 0.981]
0.973 [0.954, 0.979]
40
0.960 [0.911, 0.987]
0.980 [0.972, 0.984]
80
0.977 [0.938, 0.993]
0.985 [0.981, 0.988]
160
0.993 [0.970, 0.999]
0.989 [0.986, 0.992]
Appendix
Table 6: Calibration-size convergence. Median agreement with the 320-observation map, with 2.5th–97.5th percentile intervals over 200 subsets per size, sampled without replacement within each subset.
College of Intelligent Robotics and Advanced Manufacturing, Fudan University · School of Data Science and Engineering, East China Normal University · Shanghai Key Lab of Intelligent Information Processing, College of Computer Science and Artificial Intelligence, Fudan University