cs.CVAug 4, 2026

TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

Authors: Seokho HanDongwei WangJinhee KimYiran ChenKang Eun JeonHuanrui YangJong Hwan Ko

Organizations: Department of Electrical and Computer Engineering, Sungkyunkwan University, Korea · Department of Electrical and Computer Engineering, University of Arizona, USA · Department of Electrical and Computer Engineering, Duke University, USA · Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)

Abstract

Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-case arithmetic throughout the denoising trajectory. We introduce Temporal-Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TASQ stores one shared maximum-precision weight buffer and learns a Temporal-Spatial LSB Mask that selects a lower effective precision for each layer and denoising stage by truncating least-significant bits. Storage therefore remains fixed by the worst case, while BitOPs decrease at less sensitive stages without per-stage weight copies or runtime search. A Temporal-Precision Engine maps the learned schedule to bit-serial execution, where cycles scale with effective precision and switching precision has no measured cycle overhead. On PixArt-Sigma, SANA-1.6B, and SDXL-Turbo, TASQ achieves quality comparable to static quantization with less computation. Together with the Temporal-Precision Engine, it reduces execution cycles by 25 to 50 percent over static quantization and by 6.1 to 7.5x over a naive static 8-bit bit-serial execution. Code is available at https://github.com/seokho-han/tasq.

Explore similar work

CardsList