Looped language models provide a parameter-efficient way to scale iterative test-time computation by repeatedly executing a shared recurrent core. Post-training quantization (PTQ) can reduce the memory footprint and inference cost of looped language models, but errors introduced by a quantized shared core affect subsequent cores. Among PTQ methods, channel scaling and orthogonal rotations preserve the floating-point computation while producing representations with different quantization quality. We find that quantization configuration candidate rankings can change with recurrent depth, motivating configuration selection at the target deployment depth. However, evaluating every candidate over the full calibration set at this depth is costly. We therefore propose Loopy, a PTQ framework that formulates shared-core quantization through a recurrent-depth-aware objective, selecting shared low-bit representations by their final prediction loss at the target deployment depth. Channel scaling and orthogonal rotations parameterize the candidate representations. To approximately solve this selection problem efficiently, Loopy progressively allocates calibration windows to promising candidates while preserving complete target-depth execution, using only forward evaluations. Across eight settings, Loopy achieves the state-of-the-art results among different baselines. On Ouro-1.4B under W4A4, Loopy reduces LAMBADA perplexity by 36.5% relative to SpinQuant. Our code is available at https://github.com/Shameless0817/Loopy-review.git.
Figures & tables
Figure 1: Looped language model.
Figure 2: Loopy method overview. Top: channel scaling and rotation define quantized representations shared across recurrent steps. Bottom: progressive selection allocates more calibration windows to surviving candidates while preserving the target depth T .
Figure 3: Rotation improves activation-side quantization on Ouro-1.4B. (a) Attention and feed-forward network (FFN) inputs of physical block 22 at recurrent step 4, before and after rotation. Appendix F provides all 24 blocks. (b) Activation normalized mean-squared error (NMSE) distributions without and with rotation, on a logarithmic scale.
Figure 4: Accuracy across recurrent depths. WinoGrande accuracy (%) for 20 configurations of Ouro-1.4B W4A4. Columns identify scaling power p and rotation seed s and are ordered by decreasing accuracy at T=4 . Rows show recurrent depth. A shared color scale preserves absolute differences across depths. Appendix H shows the corresponding within-depth rankings.
Figure 5: Operator-wise quantization transforms . Channel scaling and rotation in attention and FFN, with weight-side transforms merged before quantization and input transforms applied online.
Ouro-2.6B
Ouro-1.4B
Huginn-0125
Recurrent-Llama-3.2-1B
Method
WT2
LMB
Avg.
WT2
LMB
Avg.
WT2
LMB
Avg.
WT2
LMB
Avg.
FP ref.
10.51
4.23
69.65
12.03
5.29
64.66
13.21
7.30
50.42
20.72
17.23
38.63
W4A8
AWQ
11.46
5.22
66.45
13.64
7.68
60.51
14.14
8.85
47.37
28.16
50.70
35.10
SmoothQuant
19.44
15.20
56.15
53.27
134.60
37.78
—
—
25.82
243.77
—
26.40
QuaRot
11.76
5.30
66.76
14.25
8.29
59.09
—
—
25.61
26.03
31.15
36.60
Table 1: W4A8 and W4A4 results on looped language models. WT2 and LMB denote WikiText-2 and LAMBADA perplexity ( ↓ ); Avg. is mean task accuracy (%, ↑ ). Detailed accuracy results are shown in Tables 5 and 6 . Bold indicates the best low-bit result per model and precision.
WT2 PPL ↓
Avg. Acc. (%) ↑
Model
LoopQ †
Loopy
LoopQ †
Loopy
Ouro-1.4B
15.81
14.73
57.08
57.75
Ouro-2.6B
15.89
12.35
58.13
64.37
Table 2: W4A4 comparison with LoopQ. † : published results ( Fang et al., 2026 ) ; Avg. Acc. averages the five tasks in Table 1 .
Figure 6: Selection quality–calibration cost trade-off. WikiText-2 perplexity versus calibration time for (a) Ouro-1.4B W4A4 ( T=4 ) and (b) Huginn-0125 W4A8 ( T=32 ). Black stars denote exhaustive-search references. Lower perplexity is better.
W/A
KV setting
PPL ↓
Δ PPL (%)
Avg. Acc. (%) ↑
W4A4
BF16 KV
11.4174
—
54.5266
BF16 KV+ H
11.4139
−0.03
54.4818
KV8
11.4109
−0.06
54.1423
KV4
12.1660
+6.56
51.9853
KV4+ H
11.8062
+3.41
53.3883
W4A8
BF16 KV
9.3425
—
61.7154
Table 3: KV quantization results. Δ PPL is relative to BF16 KV within each W/A setting; Avg. Acc. averages four tasks. H denotes the Q/K Hadamard transform.
W4A4
W4A8
Method
B1/128
B4/2048
B1/128
B4/2048
QuaRot
9.08∗
35.14∗
8.80
34.75
SpinQuant
9.16∗
35.85∗
9.14
35.78
Loopy (online)
10.01
40.93
10.22
41.17
Loopy (offline) †
11.21
45.45
11.37
45.83
Table 4: Decoding throughput on NVIDIA H20 (tokens/s). B/input denotes batch size and input length. Protocol details and additional configurations appear in Appendix D .
Table 6: Complete W4A4 results. Avg. is the arithmetic mean of the five displayed task accuracies (%). Large perplexities are shown in scientific notation.
W/A
KV setting
ARC-C
WinoGrande
LAMBADA
HellaSwag
W4A4
BF16 KV
44.20
62.51
45.97
65.43
BF16 KV+ H
44.97
60.85
46.59
65.51
KV8
45.65
58.96
46.48
65.48
KV4
40.78
58.64
44.34
64.17
KV4+ H
40.96
61.88
45.68
65.04
W4A8
BF16 KV
51.02
64.56
61.48
69.80
Appendix
Table 7: Per-task accuracy (%) for KV quantization. H denotes the Q/K Hadamard transform.
Model
BF16
W4A4
W3A4
W3A3
Ouro-1.4B
12.03
14.73
23.77
1653.88
Ouro-2.6B
10.51
12.35
17.65
361.85
Appendix
Table 8: WikiText-2 perplexity under lower-bit quantization. All quantized results use Loopy. Lower is better.
Figure 7: Calibration depth at fixed deployment depth. Token-normalized WikiText-2 perplexity with 2048-token windows for (a) Ouro-2.6B W4A8, (b) Ouro-1.4B W4A4, and (c) Huginn-0125 W4A8. Orange markers indicate matched calibration and deployment depths. Each configuration has one run; no error bars are shown. Vertical scales differ across panels.
W4A4
W4A8
Method
B1/128
B4/2048
B1/128
B4/2048
QuaRot
9.08∗
35.14∗
8.80
34.75
SpinQuant
9.16∗
35.85∗
9.14
35.78
Loopy (online)
10.01
40.93
10.22
41.17
Loopy (offline)
11.21
45.45
11.37
45.83
+ down rotation
11.17
45.56
11.37
45.60
Appendix
Table 9: Complete H20 decoding throughput (tokens/s, ↑ ). The last two rows add online rotations to Loopy (offline).
Figure 8: Activation distributions in blocks 1–4 of Ouro-1.4B at recurrent step 4. Each row shows attention input, rotated attention input, FFN input, and rotated FFN input, from left to right.
Figure 9: Activation distributions in blocks 5–8 of Ouro-1.4B at recurrent step 4. Each row shows attention input, rotated attention input, FFN input, and rotated FFN input, from left to right.
Figure 10: Activation distributions in blocks 9–12 of Ouro-1.4B at recurrent step 4. Each row shows attention input, rotated attention input, FFN input, and rotated FFN input, from left to right.
Figure 11: Activation distributions in blocks 13–16 of Ouro-1.4B at recurrent step 4. Each row shows attention input, rotated attention input, FFN input, and rotated FFN input, from left to right.
Figure 12: Activation distributions in blocks 17–20 of Ouro-1.4B at recurrent step 4. Each row shows attention input, rotated attention input, FFN input, and rotated FFN input, from left to right.
Figure 13: Activation distributions in blocks 21–24 of Ouro-1.4B at recurrent step 4. Each row shows attention input, rotated attention input, FFN input, and rotated FFN input, from left to right.
Figure 14: Activation distributions across recurrent steps. Input activations to the query projection in physical block 12 of Ouro-1.4B for the same input window. Panels labeled Step 0–3 correspond to the first through fourth recurrent calls. The overall activation structure remains similar, while local peaks and maximum magnitudes vary across calls despite shared weights.
Figure 15: Outlier-channel stability at the block-12 query-projection input. Left to right, stepwise Top-1%/5% channel overlap, maximum absolute activation for fixed channels (logarithmic color scale), and token exceedance frequency. The frequency threshold is 1.84375; the four-step common Max Top-1% fraction is 81.0%.
Figure 17: Within-depth configuration rankings. WinoGrande accuracy ranks for Ouro-1.4B W4A4, using the same candidate pool and column order as Figure 4 . Darker cells indicate better ranks (1 is best); stars mark the displayed top-ranked configuration at each depth.