Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.
Figures & tables
Figure 1 : ResidualQuant progressively recovers BF16-level accuracy while reducing KV storage and improving decode throughput. Results on Ouro-1.4B with group size g=32 in every quantized loop. Left: Starting from ① direct INT2 quantization, we successively introduce ② Last-loop residual quantization, ③ least-square scaling, ④ rotation, and ⑤ mixed precision to achieve ResidualQuant . Accuracy increases from 27.8% to 76.0%, matching the BF16 baseline of 75.0%, while KV cache storage is reduced by 80.7%. Right: At 8k context on an RTX 5090, ResidualQuant achieves 3.27× the peak decode throughput and 4× the largest feasible batch size compared to the BF16 counterpart.
Figure 2 : Post-RoPE key magnitudes in Ouro-1.4B. Left: Loop 1, ∣K1∣ . Middle: Loop 2, ∣K2∣ . Right: the residual after Least Square scaling of Loop 1’s keys, ∣K2−α2K1∣ . All panels use the same height and color scales. Visualization details and last-loop anchor comparisons are provided in Appendix C .
Figure 3 : Loop-wise KV statistics and the effects of prediction and reference choices on Ouro-1.4B. Left: Decode throughput with INT2/4 and g=32 on an RTX 5090, comparing the First-/Last-loop and Previous-/Next-loop reference families with OptR-H rotations. The plotted measurements use Last-loop and Next-loop, respectively. For each context length, each method runs at its maximum feasible batch size, with speedups normalized to BF16 FlashAttention at the same batch size. Middle: Mean K and V norms across loops. Right: Key prediction MSE using Last-loop references under uniform INT2 and INT4, with group size g=32 and no rotation. Solid bars show prediction MSE with LS scaling, while hatched portions indicate the additional error without LS scaling ( α=1 ). Analysis details are provided in Appendix B .
Key NMSE ×103↓
Reference
Loop 1
Loop 2
Loop 3
Loop 4
Avg
First-loop
179.39
44.88
51.71
54.30
82.57
Last-loop
53.55
28.17
19.30
178.20
69.80
Table 1 : Reference choice on Ouro-1.4B. Uniform INT2 with group size g=32 . NMSE compares original and reconstructed keys on all 500 MATH500 problems (Appendix B ).
Precision
direct
+ OSCAR
+ OptR-H
Residual
+ OSCAR
+ OptR-H
INT2
54.2
65.8
64.2
67.8
74.8
75.2
INT4
74.2
75.0
75.2
74.2
77.2
74.4
Table 4: Complementary effects of residual quantization and rotation. MATH500 accuracy (%) on Ouro-1.4B with uniform INT2 and INT4, using group sizes g=16 and g=32 , respectively. direct denotes standard uniform quantization, while Residual uses LS scaling with the Last-loop reference policy. OSCAR ( Zhou et al., 2026 ) and OptR-H ( Yun et al., 2026 ) are applied to direct or residual quantization. BF16 accuracy is 75.0%.
direct quantization
Residual Quantization
Config
g
beff
w/o Rotation
w/ Rotation
beff
w/o Rotation
w/ Rotation
Uniform
[4,4,4,4]
–
4.5
74.20
75.20
4.6
74.20
74.40
[2,2,2,2]
32
2.5
27.80
55.40
2.6
60.60
71.60
[2,2,2,2]
16
3.0
54.20
64.20
3.1
67.80
75.20
Descending
Table 5 : Loop-wise precision allocation ablations. MATH500 accuracy (%) on Ouro-1.4B and effective KV bitwidth beff . We compare direct quantization and residual quantization with and without OptR-H rotation. Residual quantization uses LS scaling and the Last-loop policy. Uniform applies the same precision to all loops, Descending uses [4,2,2,2] , and Ascending uses [2,2,2,4] , allocating the higher precision to the first and last loop, respectively. Group size g shown in the table applies only to INT2, while INT4 uses group size 32. Blue shading marks our final INT2/4 precision configuration. BF16 accuracy is 75.0%.
Figure 4 : Decode throughput across batch sizes and context lengths. ResidualQuant with group size g=32 in every loop and BF16 on Ouro-1.4B using an RTX 5090, at 2k, 4k, 8k, and 16k context. Vertical and horizontal arrows indicate gains in peak throughput and the largest feasible tested power-of-two batch, respectively.
Method
GSM8K
MATH500
HumanEval
MBPP
Ouro-1.4B (4 loops)
BF16
78.77
75.00
71.95
74.87
Naive
49.51
53.00
62.20
65.08
+ Residual Quantization
56.86 (+7.35)
61.00 (+8.00)
62.20 (+0.00)
68.52 (+3.44)
FlatQuant
68.31
67.40
67.68
69.58
+ Residual Quantization
70.20 (+1.89)
68.80 (+1.40)
70.12 (+2.44)
71.69 (+2.11)
Table 6 : Extension of residual quantization to W4A4 activation quantization. Accuracy or pass@1 (%) when applying residual quantization with LS scaling to W4A4 activation quantization, with KV caches kept in BF16. We consider three baselines: Naive round-to-nearest quantization, FlatQuant ( Sun et al., 2025 ) , and LoopQ ( Fang et al., 2026 ) . For each baseline, colored parentheses report the absolute accuracy change from applying residual quantization. Details are provided in Appendix H .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Benchmark
Shots
Prompt format
Max tokens
Ouro-1.4B
GSM8K
3
GSM8K CoT
1,024
MATH500
0
CoT, boxed answer
2,048
HumanEval
0
EvalPlus-style
1,024
MBPP
0
EvalPlus-style
2,048
Huginn-3.5B
GSM8K
8
GSM8K CoT
1,024
MATH500
4
Adapted Minerva
2,048
Appendix
Table 7 : Prompt formats and maximum generated tokens. Ouro-1.4B uses plain-text prompts, while Huginn-3.5B uses its tokenizer’s chat template.
Figure 5 : Key distributions relative to the last-loop anchor. Rows correspond to loops 1, 2, and 3. Columns show ∣Kr∣ , ∣K4∣ , and ∣Kr−K4∣ on shared scales.
Schedule
g
beff
w/o OptR-H
w/ OptR-H
BF16
–
16
42.84
8×[2,2,2,4]
16
3.5
41.70
42.38
32
3.1
43.14
42.61
4×[2×7,4]
16
3.3
41.55
42.53
32
2.9
41.93
42.15
2×[2×15,4]
16
3.3
41.93
42.68
Appendix
Table 8 : Huginn-3.5B GSM8K accuracy (%, 1,319 questions) and effective KV bitwidth beff across anchor intervals. Each group uses one final INT4 anchor. 2×n denotes n consecutive INT2 loops. Group size g applies to INT2, while INT4 uses g=32 . Both residual quantization variants use LS scaling and the Last-loop policy. Effective bitwidth includes quantized values, scales, offsets, and LS coefficients, following Appendix A.3 , and is the same with and without OptR-H.
No rotation
+ OSCAR
+ OptR-H
Reference / Method
Direct
Residual
Norm-ratio
LS scaling
Direct
LS scaling
Direct
LS scaling
INT2 , g=16
Direct
54.20
–
–
–
65.80
–
64.20
–
Previous-loop
–
69.20
73.00
71.20
–
74.60
–
72.20
Next-loop
–
71.60
67.60
73.20
–
73.60
–
72.80
First-loop
–
46.00
69.00
69.80
–
71.20
–
71.40
Appendix
Table 9 : Scaling coefficient, reference, and rotation ablations. MATH500 accuracy (%) on Ouro-1.4B with uniform INT2, uniform INT4, and mixed INT2/4. Uniform INT2 uses group size g=16 , while uniform INT4 and mixed INT2/4 use g=32 in every loop. For uniform precision, direct quantization is independent of the reference policy. For mixed INT2/4, Previous-loop and First-loop use [4,2,2,2] , while Next-loop and Last-loop use [2,2,2,4] . The direct and rotation-only baselines are listed separately for each mixed-precision schedule. Residual variants compare different scaling coefficients and reference choices, with OSCAR ( Zhou et al., 2026 ) or OptR-H ( Yun et al., 2026 ) applied to residual quantization with LS scaling. Blue row shading marks the selected Last-loop policy, blue column shading marks residual quantization with LS scaling and OptR-H, and the darker intersections indicate our final configuration. BF16 accuracy is 75.0%.
Figure 6 : Decode throughput with INT2 group size g=16 . Ouro-1.4B on an RTX 5090 at 2k, 4k, 8k, and 16k context. ResidualQuant uses [2,2,2,4] , Last-loop references, LS scaling, and shared OptR-H rotations, with group size g=16 for INT2 residuals and g=32 for the INT4 anchor. Each run generates 128 tokens, with throughput measured over the 127 decode steps after the first output token and averaged over three trials following one warm-up. Arrows show peak-throughput gains and ratios of the largest plotted batch sizes relative to BF16.
Average Bitwidth
Method
BF16 KV tokens
Prefill
Decode
Accuracy (%)
BF16 + sparse
All
9.40
16.00
76.0
FlashLoop
Recent 64–127
2.64
4.50
72.8
ResidualQuant ( g=16 ) + sparse
Recent 64
2.18
3.47
74.4
ResidualQuant ( g=32 ) + sparse
Recent 64
2.01
3.09
73.4
Appendix
Table 10 : Combining ResidualQuant with sparse KV updates and attention. MATH500 accuracy on Ouro-1.4B and average KV bitwidth during prefill and decode. All methods use the same sparse KV update and attention mechanisms from FlashLoop. For ResidualQuant , g denotes the INT2 residual group size, while INT4 anchors use g=32 . Average bitwidth includes quantized values, scales, offsets, and LS coefficients, but excludes recent BF16 states.
Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64× end-to-end speedup and up to 6× KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
Wanqi Yang, Shiwei Liu
ELLIS Institute Tübingen Max Planck Institute for Intelligent Systems Tübingen AI Center
Looped language models (LoopLMs) improve parameter efficiency by recursively reusing Transformer blocks, enabling deeper computation under a fixed model size. However, this reuse makes LoopLMs more fragile under post-training quantization (PTQ). We present the first systematic study of quantization in LoopLMs and identify three challenges: distribution shift across roles, state reuse across loop transitions, and recursive error accumulation. To address these challenges, we propose LoopQ, a loop-aware PTQ framework that preserves a shared quantized backbone while introducing lightweight adaptations. LoopQ combines activation scaling, selective transformation, cross-loop state alignment, and trajectory-aware optimization to reduce distributional mismatch within loops and error accumulation across loops. Experiments across seven benchmarks show that, under W4A4 quantization, LoopQ improves average downstream accuracy by 68.8% and reduces average perplexity by 87.7% compared with the strongest static PTQ baseline.
Large-scale transformer training and deployment are increasingly constrained by the transfer of activations, gradients, and optimizer states across accelerators. Low-bit quantization offers a natural remedy, but transformer activations are often heavy-tailed and outlier-dominated, making simple quantization highly lossy. We show that this difficulty is not only a property of the quantizer, but also of the architecture. Specifically, residual connections can drive transformer activations away from Gaussianity during training. Using controlled comparisons between residual and residual-free transformers, we demonstrate that this effect leads to substantially higher quantization error and accuracy degradation at low precision in residual models. We explain the phenomenon through an excess kurtosis analysis, showing that residual mixing can amplify non-Gaussianity, whereas dense mixing in residual-free contracts non-Gaussianity. We then show that residual-free transformers can be made trainable using orthogonal initialization, spectral or second-order optimization, and depth-aware scaling of attention temperature. In language tasks, while there is a small drop in full precision performance, these models retain near-Gaussian activations and exhibit significantly improved robustness to low-bit quantization. Our results identify an accuracy--compressibility trade-off in transformer design and motivate architecture-level approaches to quantization-friendly foundation models.
Yiping Ji, Mahalakshmi Sabanayagam, Peyman Moghadam +2
Australian Institute for Machine Learning, Adelaide University · DATA61, CSIRO