cs.ROSep 30, 2026

SteerQuant: Steering Quantization Error with Action-Guided Scaling in World-Action Models

Authors: Yunhan Wang, Haodong Wang, Zhiming Liu, Zicong Hong, Qianli Liu, Xiaoyi Pang, Yangjia Hu, Quanxin Shou, +2 more

Organizations: Tongji · HKUST · HIT · EPFL

Abstract

World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly different effects on final actions, making numerical accuracy alone insufficient for reliable control. We introduce SteerQuant, a 4-bit quantization framework for WAMs that steers errors toward computations with less influence on final actions. It maps how each stream's quantization errors affect final actions and uses this map to guide shared channel scaling. Activation scaling is further calibrated for each stream and denoising step to accommodate changes in activation ranges and action impact. This adapts quantization to different stream requirements without duplicating weights or increasing bit-widths for selected streams. To reduce the extra kernel launches and memory traffic introduced by scaling, we develop Rudder, a 4-bit inference engine for WAMs that fuses scaling and output compensation into low-bit kernels. Under W4A8 and W4A4, SteerQuant maintains mean LIBERO success within 0.8 percentage points of full precision, while delivering up to 2.23×2.23\times denoising speedup over BF16 across three WAMs with reduced peak GPU memory usage. On a real dual-arm robot, W4A8 deployment achieves a 1.35×1.35\times end-to-end inference speedup while maintaining average task success relative to BF16.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection

    Sep 27, 2026Arash Akbari, Arman Akbari, Jingwu Luo +7Faster-WamWorld Models

  2. QuantWAMs: Calibrating at the Right Granularity for World Action Models

    Jul 30, 2026Jiacheng Zhou, Jinfan Lv, Ruixuan Li +4Post-Training QuantizationWorld Models

  3. Efficient World Action Model Inference with Adaptive Intermediate States

    Sep 28, 2026Zhinnan Liu, Haozhi Han, Ruge Zhang +8Efficient World-Action ModelWorld Models