cs.LGSep 27, 2026

Chameleon: Dynamic Format Adapter for Efficient Diffusion

Authors: Arnab Sanyal, Sandeep Chinchali

Organizations: UT SWARM Lab, Department of Electrical & Computer Engineering The University of Texas, Austin TX 78712

Abstract

Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the number format\mathit{number\ format} in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best format depends on the distribution being encoded, and that distribution differs across weight channels, across layers, and along the diffusion timestep, where activation distributions slide from heavy-tailed and noise-dominated to tightly clustered and structured. We propose Chameleon, a PTQ framework that holds the bit-width fixed and treats the format itself as a discrete variable, chosen per weight channel and per (layer, timestep bucket) activation tensor. Activation formats come from {INT8, FP8 E4M3, FP8 E5M2, MXFP8, MXINT8}, selected ahead of time from two cheap statistics (empirical kurtosis and the closed-form diffusion SNR) and stored in a lookup table; weight formats come from {INT8, MXINT8} at 8 bits or {INT4, NF4, FP4 E2M1, MXINT4, MXFP4} at 4 bits, selected offline by reconstruction error. An architectural fork adapts the same selection layer to multi-step UNets, single-step distilled models, and Diffusion Transformers. Across SDXL, SDXL-Turbo, and PixArt-αα on COCO-2014, Chameleon achieves the best FID in all six backbone ×\times bit-width settings, with CLIP within 0.24 of the FP16 reference and the best of all quantized methods at W4A8W_{4}A_{8}.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

    May 10, 2026Yongsen Cheng, Kai Liu, Kaiwen Tao +5QuantizedDiffusion Models

  2. KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

    Jul 23, 2026Yann Bouquet, Alireza Khodamoradi, Kristof Denolf +1Post-Training QuantizationDiffusion Transformers

  3. FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models

    Jul 7, 2026Bowen Xue, Zihan Min, Xingyang Li +8Post-Training QuantizationDiffusion Models