Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.
Figures & tables
Figure 1 : ResidualQuant progressively recovers BF16-level accuracy while reducing KV storage and improving decode throughput. Results on Ouro-1.4B with group size g=32 in every quantized loop. Left: Starting from ① direct INT2 quantization, we successively introduce ② Last-loop residual quantization, ③ least-square scaling, ④ rotation, and ⑤ mixed precision to achieve ResidualQuant . Accuracy increases from 27.8% to 76.0%, matching the BF16 baseline of 75.0%, while KV cache storage is reduced by 80.7%. Right: At 8k context on an RTX 5090, ResidualQuant achieves 3.27× the peak decode throughput and 4× the largest feasible batch size compared to the BF16 counterpart.
Figure 2 : Post-RoPE key magnitudes in Ouro-1.4B. Left: Loop 1, ∣K1∣ . Middle: Loop 2, ∣K2∣ . Right: the residual after Least Square scaling of Loop 1’s keys, ∣K2−α2K1∣ . All panels use the same height and color scales. Visualization details and last-loop anchor comparisons are provided in Appendix C .
Figure 3 : Loop-wise KV statistics and the effects of prediction and reference choices on Ouro-1.4B. Left: Decode throughput with INT2/4 and g=32 on an RTX 5090, comparing the First-/Last-loop and Previous-/Next-loop reference families with OptR-H rotations. The plotted measurements use Last-loop and Next-loop, respectively. For each context length, each method runs at its maximum feasible batch size, with speedups normalized to BF16 FlashAttention at the same batch size. Middle: Mean K and V norms across loops. Right: Key prediction MSE using Last-loop references under uniform INT2 and INT4, with group size g=32 and no rotation. Solid bars show prediction MSE with LS scaling, while hatched portions indicate the additional error without LS scaling ( α=1 ). Analysis details are provided in Appendix B .
Key NMSE ×103↓
Reference
Loop 1
Loop 2
Loop 3
Loop 4
Avg
First-loop
179.39
44.88
51.71
54.30
82.57
Last-loop
53.55
28.17
19.30
178.20
69.80
Table 1 : Reference choice on Ouro-1.4B. Uniform INT2 with group size g=32 . NMSE compares original and reconstructed keys on all 500 MATH500 problems (Appendix B ).
Precision
direct
+ OSCAR
+ OptR-H
Residual
+ OSCAR
+ OptR-H
INT2
54.2
65.8
64.2
67.8
74.8
75.2
INT4
74.2
75.0
75.2
74.2
77.2
74.4
Table 4: Complementary effects of residual quantization and rotation. MATH500 accuracy (%) on Ouro-1.4B with uniform INT2 and INT4, using group sizes g=16 and g=32 , respectively. direct denotes standard uniform quantization, while Residual uses LS scaling with the Last-loop reference policy. OSCAR ( Zhou et al., 2026 ) and OptR-H ( Yun et al., 2026 ) are applied to direct or residual quantization. BF16 accuracy is 75.0%.
direct quantization
Residual Quantization
Config
g
beff
w/o Rotation
w/ Rotation
beff
w/o Rotation
w/ Rotation
Uniform
[4,4,4,4]
–
4.5
74.20
75.20
4.6
74.20
74.40
[2,2,2,2]
32
2.5
27.80
55.40
2.6
60.60
71.60
[2,2,2,2]
16
3.0
54.20
64.20
3.1
67.80
75.20
Descending
Table 5 : Loop-wise precision allocation ablations. MATH500 accuracy (%) on Ouro-1.4B and effective KV bitwidth beff . We compare direct quantization and residual quantization with and without OptR-H rotation. Residual quantization uses LS scaling and the Last-loop policy. Uniform applies the same precision to all loops, Descending uses [4,2,2,2] , and Ascending uses [2,2,2,4] , allocating the higher precision to the first and last loop, respectively. Group size g shown in the table applies only to INT2, while INT4 uses group size 32. Blue shading marks our final INT2/4 precision configuration. BF16 accuracy is 75.0%.
Figure 4 : Decode throughput across batch sizes and context lengths. ResidualQuant with group size g=32 in every loop and BF16 on Ouro-1.4B using an RTX 5090, at 2k, 4k, 8k, and 16k context. Vertical and horizontal arrows indicate gains in peak throughput and the largest feasible tested power-of-two batch, respectively.
Method
GSM8K
MATH500
HumanEval
MBPP
Ouro-1.4B (4 loops)
BF16
78.77
75.00
71.95
74.87
Naive
49.51
53.00
62.20
65.08
+ Residual Quantization
56.86 (+7.35)
61.00 (+8.00)
62.20 (+0.00)
68.52 (+3.44)
FlatQuant
68.31
67.40
67.68
69.58
+ Residual Quantization
70.20 (+1.89)
68.80 (+1.40)
70.12 (+2.44)
71.69 (+2.11)
Table 6 : Extension of residual quantization to W4A4 activation quantization. Accuracy or pass@1 (%) when applying residual quantization with LS scaling to W4A4 activation quantization, with KV caches kept in BF16. We consider three baselines: Naive round-to-nearest quantization, FlatQuant ( Sun et al., 2025 ) , and LoopQ ( Fang et al., 2026 ) . For each baseline, colored parentheses report the absolute accuracy change from applying residual quantization. Details are provided in Appendix H .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Benchmark
Shots
Prompt format
Max tokens
Ouro-1.4B
GSM8K
3
GSM8K CoT
1,024
MATH500
0
CoT, boxed answer
2,048
HumanEval
0
EvalPlus-style
1,024
MBPP
0
EvalPlus-style
2,048
Huginn-3.5B
GSM8K
8
GSM8K CoT
1,024
MATH500
4
Adapted Minerva
2,048
Appendix
Table 7 : Prompt formats and maximum generated tokens. Ouro-1.4B uses plain-text prompts, while Huginn-3.5B uses its tokenizer’s chat template.
Figure 5 : Key distributions relative to the last-loop anchor. Rows correspond to loops 1, 2, and 3. Columns show ∣Kr∣ , ∣K4∣ , and ∣Kr−K4∣ on shared scales.
Schedule
g
beff
w/o OptR-H
w/ OptR-H
BF16
–
16
42.84
8×[2,2,2,4]
16
3.5
41.70
42.38
32
3.1
43.14
42.61
4×[2×7,4]
16
3.3
41.55
42.53
32
2.9
41.93
42.15
2×[2×15,4]
16
3.3
41.93
42.68
Appendix
Table 8 : Huginn-3.5B GSM8K accuracy (%, 1,319 questions) and effective KV bitwidth beff across anchor intervals. Each group uses one final INT4 anchor. 2×n denotes n consecutive INT2 loops. Group size g applies to INT2, while INT4 uses g=32 . Both residual quantization variants use LS scaling and the Last-loop policy. Effective bitwidth includes quantized values, scales, offsets, and LS coefficients, following Appendix A.3 , and is the same with and without OptR-H.
No rotation
+ OSCAR
+ OptR-H
Reference / Method
Direct
Residual
Norm-ratio
LS scaling
Direct
LS scaling
Direct
LS scaling
INT2 , g=16
Direct
54.20
–
–
–
65.80
–
64.20
–
Previous-loop
–
69.20
73.00
71.20
–
74.60
–
72.20
Next-loop
–
71.60
67.60
73.20
–
73.60
–
72.80
First-loop
–
46.00
69.00
69.80
–
71.20
–
71.40
Appendix
Table 9 : Scaling coefficient, reference, and rotation ablations. MATH500 accuracy (%) on Ouro-1.4B with uniform INT2, uniform INT4, and mixed INT2/4. Uniform INT2 uses group size g=16 , while uniform INT4 and mixed INT2/4 use g=32 in every loop. For uniform precision, direct quantization is independent of the reference policy. For mixed INT2/4, Previous-loop and First-loop use [4,2,2,2] , while Next-loop and Last-loop use [2,2,2,4] . The direct and rotation-only baselines are listed separately for each mixed-precision schedule. Residual variants compare different scaling coefficients and reference choices, with OSCAR ( Zhou et al., 2026 ) or OptR-H ( Yun et al., 2026 ) applied to residual quantization with LS scaling. Blue row shading marks the selected Last-loop policy, blue column shading marks residual quantization with LS scaling and OptR-H, and the darker intersections indicate our final configuration. BF16 accuracy is 75.0%.
Figure 6 : Decode throughput with INT2 group size g=16 . Ouro-1.4B on an RTX 5090 at 2k, 4k, 8k, and 16k context. ResidualQuant uses [2,2,2,4] , Last-loop references, LS scaling, and shared OptR-H rotations, with group size g=16 for INT2 residuals and g=32 for the INT4 anchor. Each run generates 128 tokens, with throughput measured over the 127 decode steps after the first output token and averaged over three trials following one warm-up. Arrows show peak-throughput gains and ratios of the largest plotted batch sizes relative to BF16.
Average Bitwidth
Method
BF16 KV tokens
Prefill
Decode
Accuracy (%)
BF16 + sparse
All
9.40
16.00
76.0
FlashLoop
Recent 64–127
2.64
4.50
72.8
ResidualQuant ( g=16 ) + sparse
Recent 64
2.18
3.47
74.4
ResidualQuant ( g=32 ) + sparse
Recent 64
2.01
3.09
73.4
Appendix
Table 10 : Combining ResidualQuant with sparse KV updates and attention. MATH500 accuracy on Ouro-1.4B and average KV bitwidth during prefill and decode. All methods use the same sparse KV update and attention mechanisms from FlashLoop. For ResidualQuant , g denotes the INT2 residual group size, while INT4 anchors use g=32 . Average bitwidth includes quantized values, scales, offsets, and LS coefficients, but excludes recent BF16 states.