cs.LGSep 28, 2026

TORQUE: Optimizing What (not) to Quantize Before and After Rotation

Authors: Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik

Organizations: University College London · Harvard University · VMware Research by Broadcom

Abstract

Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotations by jointly optimizing how many and which coordinates to preserve at high precision both before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We derive a quantization error upper bound and prove that top-kk pre-rotation retention minimizes it for each kk. This reduces the search over coordinate subsets to an optimization over kk, enabling a fast optimizer that uses offline codebooks and parallel parameter selection for practical implementation. We demonstrate an improved tradeoff between reconstruction accuracy and storage cost through numerical evaluation under the Gaussian model and experiments on nearest-neighbor retrieval, KV-cache compression, and activation compression.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven

    May 7, 2026Ran Ben-Basat, William Kuszmaul, Michael Mitzenmacher +2Randomized Hadamard TransformsResidual Vector Quantization

  2. Output-Aware Rotation for INT2 KV-Cache Quantization

    Aug 3, 2026Vincent-Daniel Yun, Woosang Lim, Minsoo Cheong +4Key-Value Cache QuantizationKv-Cache Management