cs.LGOct 5, 2026

Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization

Authors: Yan Scholten, Rachel Lawrence, James Hensman, Stephan Günnemann, Alicia Curth, Riccardo Grazzi

Organizations: Dept. of Computer Science & Munich Data Science Institute, Technical University of Munich · Microsoft Research Cambridge, UK

Abstract

Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the cost of a serial schedule that becomes a bottleneck at scale. As a solution, we propose parallel quantization with activation denoising, which recovers much of the sequential benefit while keeping quantization fully parallel. Rather than re-calibrating layer-by-layer, we take a robustness perspective and model the upstream error as noise, regularizing to be robust to it through a preprocessing step followed by metric-weighted rounding. Applied at every layer, this regularization forms a depth-compounding smoothness penalty that dampens how strongly quantization errors amplify through the model. Unlike orthogonal rotations commonly used in quantization, which must preserve the model's function, we multiply the weights by a more general linear transformation. We find that the two are complementary and their effects compound. Empirically, our robustness regularization recovers a significant part of sequential quantization's benefit in a single parallel pass, at a fraction of its time. Overall, by treating compounding quantization errors as a robustness problem, we offer a principled foundation for more efficient and accurate LLM quantization at scale.

Figures & tables

Appendix figures & tables35 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Why Does Post-Training Quantization Work?

    Sep 10, 2026Yuxiang Chen, Michael Beyer, Jun Zhu +1LLM QuantizationPost-Training Quantization

  2. DynamicPTQ: Mitigating Activation Quantization Collapse via Residual-Stream Dynamics

    Jun 10, 2026Zimo Zhao, Maolin Wang, Bowen Yu +3Mixed-Precision QuantizationLLM Quantization

  3. Theory-optimal Quantization Based on Flatness

    May 11, 2026Xiusheng Huang, Zhe Li, Xuanwu Yin +5LLM QuantizationActivation Quantization