Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.
Figures & tables
Figure 1: LoopCD lets an earlier loop guide the final one: more accuracy at full depth, less compute at half. (a) LoopCD reads an earlier state through the same output layers and uses it to guide the final prediction. (b) Accuracy at full depth without LoopCD (hatched) and the gain with it (solid), for eight model and benchmark pairs; Huginn at R=32 uses hidden-state guidance. (c) Seven-benchmark mean against forward compute, on a log axis broken into one segment per model: the unguided model at full depth (hatched) and LoopCD at half the iterations (filled), with the compute saved; all six settings are in Figure 4 .
Figure 2: Looped Transformers differ in which part of the network loops. Ouro repeats its full stack; Huginn, Parcae, and Looped-Qwen3 repeat a shared block (blue) between fixed prelude and coda layers (gray). Layer and iteration counts are for Ouro-1.4B, Huginn-0125, Parcae-1.3B, and Looped-Qwen3 (Table A1 ).
Figure 3: LoopCD applies the contrast before or after the coda and language modeling head. LoopCD-Hidden combines h1 and hR (Eq. (2) ) and evaluates the coda and lm_head once; LoopCD-Logits evaluates both states through them and combines the logits (Eq. (3) ), in two passes. Rust marks the reference, sage the final state, steel the guided result.
AIME 2024
AIME 2025
OlympiadBench
Mean Δ
Model
Pass@1
Pass@10
Pass@1
Pass@10
Pass@1
Pass@10
Pass@1
Pass@10
Ouro-1.4B-Thinking
50.83
80.33
38.96
68.09
61.27
76.69
LoopCD, fixed
56.88 +6.05
84.82 +4.49
46.25 +7.29
69.23 +1.14
63.75 +2.48
78.66 +1.97
+5.27
+2.53
LoopCD, adaptive
59.17 +8.34
88.19 +7.86
47.08 +8.12
70.30 +2.21
63.48 +2.21
78.99 +2.30
+6.22
+4.12
Ouro-2.6B-Thinking
61.88
82.65
49.58
75.95
64.05
78.22
LoopCD, fixed
70.62 +8.74
88.33 +5.68
58.33 +8.75
81.21 +5.26
66.66 +2.61
79.52 +1.30
+6.70
+4.08
Table 1: LoopCD improves mathematical reasoning and code generation at full recurrent depth. Panel (a): Ouro-Thinking, pass@1 and pass@10 from sixteen samples per problem; subscripts give the change from the baseline. Panel (b): five configurations, base and extended tests; a guided score is red below its baseline.
Model
ARC-C
ARC-E
SciQ
MMLU
HellaSwag
WinoGrande
PIQA
Mean Δ
Ouro-1.4B
60.41
84.09
94.70
44.45
74.48
70.80
78.51
LoopCD, fixed
62.29 +1.88
84.39 +0.30
95.30 +0.60
45.35 +0.90
75.30 +0.82
69.85 −0.95
79.05 +0.54
+0.58
LoopCD, adaptive
61.69 +1.28
84.26 +0.17
95.20 +0.50
45.67 +1.22
75.57 +1.09
68.35 −2.45
78.73 +0.22
+0.29
Ouro-2.6B
66.21
88.01
94.20
51.09
79.52
77.66
80.25
LoopCD, fixed
68.60 +2.39
89.10 +1.09
94.60 +0.40
51.82 +0.73
81.07 +1.55
76.87 −0.79
80.74 +0.49
+0.84
LoopCD, adaptive
69.71 +3.50
89.02 +1.01
94.30 +0.10
52.30 +1.21
81.66 +2.14
75.93 −1.73
80.69 +0.44
+0.95
Table 2: Every logit configuration improves its seven-benchmark mean under LoopCD. Seven multiple-choice benchmarks, each option scored by likelihood. Baselines in gray, then LoopCD at fixed and adaptive strength; subscripts give the change from the baseline, and bold marks the better rule’s mean (Appendix B.2 ).
Huginn-0125 R=32
Huginn-0125 R=16
Parcae-1.3B R=8
Benchmark
Baseline
LoopCD
Baseline
LoopCD
Baseline
LoopCD
ARC-C
37.97
38.74 +0.77
37.37
37.97 +0.60
40.36
42.66 +2.30
ARC-E
64.18
64.52 +0.34
64.14
64.35 +0.21
64.60
65.36 +0.76
HellaSwag
66.56
68.86 +2.30
66.32
68.16 +1.84
55.99
58.19 +2.20
WinoGrande
61.17
62.35 +1.18
61.01
62.83 +1.82
58.96
59.12 +0.16
PIQA
75.52
75.57 +0.05
75.19
75.08 −0.11
72.63
71.11 −1.52
Table 3: LoopCD-Hidden raises every configuration’s mean at no added output pass. LoopCD-Hidden at fixed strength. Baselines in gray; subscripts give the change from the baseline.
Figure 4: With LoopCD, half the recurrent iterations match or beat full depth for less compute. (a) Change in the seven-benchmark mean from the unguided model at full depth: the unguided model at half depth (hatched) and LoopCD at half depth (filled). (b) Forward FLOPs at half depth (solid) as a share of the unguided full-depth pass (hatched); the lighter segment is the readout of logit guidance.
Figure 5: A Looped Transformer’s first recurrent iteration is much weaker than its final one. (a) After the first iteration (open) against the final one (filled), seven-benchmark mean accuracy, six models. (b) Ouro after each recurrent iteration, ARC-Challenge, 25-shot. (c) Huginn after each of its 32 recurrent iterations, ARC-Challenge, zero-shot.
Figure 6: LoopCD gains most where the weak prediction disagrees most with the final prediction. (a) Jensen–Shannon divergence (bits) between Ouro’s prediction at each recurrent iteration and its final one, ARC-Challenge. (b) Kullback–Leibler divergence (nats) between Huginn’s prediction at each recurrent iteration and its final one. (c) The gain of LoopCD-Logits at ω=0.5 against the share of questions on which the reference alone picks a different option from the final prediction, four models, ARC-Challenge (filled) and HellaSwag (open).
Figure 7: LoopCD gains most from the first recurrent iteration as its weak prediction. (a) The gain of LoopCD-Logits at ω=0.5 on ARC-Challenge by reference iteration k , grouped by model, the first iteration in full colour. (b) Huginn at R=32 , LoopCD-Hidden at ω=0.5 : seven-benchmark mean change by reference state, burn-in shaded.
Figure 8: The re-ranking carries the gain, and the gain lands where the model is uncertain. (a) An illustrative contrast zR−z1 splits into a part parallel to zR , which rescales the logits like a temperature, and an orthogonal part, which re-ranks tokens. (b) HellaSwag, the two Ouro models: accuracy change as the strength grows, re-ranking part alone (thick) and full update (thin). (c) ARC-Challenge questions in five equal groups by the unguided prediction’s confidence, least confident first: accuracy change per group, LoopCD-Logits at ω=0.5 , four models.
Figure 9: Multiple-choice scoring tolerates about twice the fixed strength that generation does. (a) Seven-benchmark mean change against the fixed logit strength, six models; the band spans the strengths at which they peak, and the dashed line marks the selected ω=0.5 . (b) Generation mean change, five models; the band marks the selected 0.2 to 0.3 , and beyond ω≈0.6 the curves fall steeply.
Figure 10: The adaptive rule keeps its gain over a wider range of strength than a fixed value. (a) Seven-benchmark mean change against the adaptive cap ωmax , six models; the band spans the caps at which they peak. (b) Mean over the six models, the fixed rule against its strength (dashed) and the adaptive rule against its cap (solid); the band is where the adaptive mean is still above the baseline and the fixed mean is not.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Model
# Params
#P
#R
#C
R
Leff
Ouro-1.4B
1.4B
0
24
0
4
96
Ouro-2.6B
2.6B
0
48
0
4
192
Huginn-0125 ( R = 32)
3.5B
2
4
2
32
132
Huginn-0125 ( R = 16)
3.5B
2
4
2
16
68
Parcae-370M
0.37B
4
4
4
8
40
Parcae-1.3B
1.3B
8
8
8
8
80
Appendix
Table A1: Layer structure of the evaluated models. #P , #R , and #C count prelude, recurrent-block, and coda layers; R is the iteration count, and Leff=#P+R#R+#C counts layer applications. The Ouro-Thinking checkpoints share the Ouro architectures listed; Looped-Qwen3 is built from the dense Qwen3-4B. Layer counts do not give FLOPs, as layers differ in width (Table A7 ).
Model
ARC-C
ARC-E
HellaSwag
WinoGrande
PIQA
SciQ
MMLU
LoopCD-Logits
Ouro-1.4B / 2.6B
25
8
10
5
0
0
5
Huginn-0125 R=32 / R=16
0
0
0
0
0
0
0
Parcae-370M / 1.3B
25
0
0
0
0
0
5
Looped-Qwen3
25
0
0
0
0
0
5
LoopCD-Hidden
Appendix
Table A2: Demonstrations per prompt on the multiple-choice benchmarks. 0 means zero-shot. A guided run always uses its baseline’s count, so every change in Tables 2 and 3 is measured at a fixed prompt. Where the two forms differ, as for Looped-Qwen3, compare their changes rather than their absolute scores.
Model
GSM8K
MMLU-Pro
HumanEval(+)
MBPP(+)
LoopCD-Logits
Ouro-1.4B / 2.6B
3
5
0
0
Huginn-0125 R=32 / R=16
8 †
5
0
0
Looped-Qwen3
3
5
0
0
LoopCD-Hidden
Huginn-0125 R=32 / R=16
–
–
0
0
Appendix
Table A3: Prompting configurations across generation benchmarks. † denotes GSM8K-CoT. Code benchmarks use no worked examples. HumanEval(+) and MBPP(+) include both base and extended tests.
Configuration
Fixed ω
Cap ωmax
Multiple-choice scoring, LoopCD-Logits
Ouro-1.4B
0.5
1.0
Ouro-2.6B
0.5
1.0
Huginn-0125 R=32
0.5
0.5
Huginn-0125 R=16
0.5
0.5
Parcae-370M
0.5
1.0
Appendix
Table A4: Selected hyperparameters for standard-budget experiments. ω is fixed strength, and ωmax is the adaptive cap. LoopCD-Hidden rows specify the reference state hb to avoid initialization noise; LoopCD-Logits rows use h1 . Dashes denote untested modes. Reduced-depth configurations follow Table A8 .
GSM8K
MMLU-Pro
Model
Baseline
LoopCD, fixed
LoopCD, adaptive
Baseline
LoopCD, fixed
LoopCD, adaptive
Ouro-1.4B
75.82
77.48 +1.66
77.86 +2.04
48.87
49.14 +0.27
48.88 +0.01
Ouro-2.6B
81.27
82.34 +1.07
80.44 −0.83
56.08
57.50 +1.42
56.99 +0.91
Huginn-0125 R=32
42.91
41.62 −1.29
43.06 +0.15
13.59
14.50 +0.91
14.23 +0.64
Huginn-0125 R=16
38.21
39.80 +1.59
38.44 +0.23
13.66
14.26 +0.60
14.18 +0.52
Looped-Qwen3
85.29
84.15 −1.14
84.23 −1.06
58.86
58.85 −0.01
59.32 +0.46
Appendix
Table A5: LoopCD-Logits holds or raises MMLU-Pro in every configuration; GSM8K changes are small and mixed. Five configurations at full depth. Baselines in gray; subscripts give the change from the baseline (Appendix B.2 ).
AIME 2024
AIME 2025
OlympiadBench
Mean Δ
Model
pass@1
pass@10
pass@1
pass@10
pass@1
pass@10
pass@1
pass@10
Looped-Qwen3
64.79
82.91
55.00
75.68
66.52
78.91
–
–
LoopCD, fixed
61.88 −2.91
86.17 +3.26
54.58 −0.42
77.42 +1.74
66.52 +0.00
79.60 +0.69
−1.11
+1.90
LoopCD, adaptive
61.04 −3.75
86.10 +3.19
53.96 −1.04
81.28 +5.60
66.29 −0.23
79.65 +0.74
−1.67
+3.18
Appendix
Table A6: On mathematical reasoning, LoopCD lowers Looped-Qwen3’s pass@1 and raises its pass@10. Both metrics from sixteen samples per problem, as in Table 1 (a).
LoopCD-Logits
Model
Forward TFLOPs
FLOPs ×
head (%)
coda or tail (%)
Layers after the loop
LoopCD-Hidden ×
Ouro-1.4B
5.360
1.019
1.92
0.00
0
1.000
Ouro-2.6B
10.617
1.010
0.97
0.00
0
1.000
Huginn-0125 R=32
54.526
1.022
0.65
1.51
2
1.000
Huginn-0125 R=16
28.261
1.042
1.25
2.90
2
1.000
Parcae-370M
0.593
1.152
5.80
9.42
4
1.000
Appendix
Table A7: LoopCD-Logits adds computational overhead based on post-loop layers, whereas LoopCD-Hidden adds negligible cost. Forward TFLOPs represent one unguided 512-token prefill. The logit multiplier is relative to this baseline, with the extra work divided between the head and the coda/frozen tail applied to the reference state. LoopCD-Hidden relies solely on vector operations before a single output pass, yielding a 1.000× multiplier.
Model
Guidance
Full → reduced
Reference
Fixed ω
Huginn-0125
Logit
32 → 16
h5
0.5
Huginn-0125
Hidden-state
32 → 16
h6
0.5
Parcae-370M
Logit
8 → 4
h1
0.75
Parcae-1.3B
Logit
8 → 4
h1
0.5
Parcae-1.3B
Hidden-state
8 → 4
h1
0.75
Looped-Qwen3
Hidden-state
8 → 4
h1
1.0
Appendix
Table A8: Settings of the six half-depth comparisons in Figure 4 . Each guided model runs half the recurrent iterations of its full-depth baseline, on the multiple-choice benchmarks; Looped-Qwen3 exits after four of its eight damped substeps and keeps the coefficient 1/8 . h1 is the state after the first iteration. The strength is fixed across benchmarks within each setting.
Figure A1: Half depth saves the loop’s share of the forward pass, and LoopCD-Logits pays for the layers after it. One bar per configuration, scaled to its unguided full-depth pass: prelude, recurrent iterations, then coda and head. Dashed blocks are the iterations a half-depth run skips; the outlined block at the right is LoopCD-Logits’ second coda and head. LoopCD-Hidden adds no visible block. Analytic FLOPs at a 512-token prefill.
Setting
Iterations
Forward FLOPs at half depth, guidance included (fraction of the full pass)
Removed (%)
Huginn-0125, logits
32 → 16
0.54
46.0%
Huginn-0125, hidden states
32 → 16
0.52
48.2%
Parcae-370M, logits
8 → 4
0.78
22.5%
Parcae-1.3B, logits
8 → 4
0.73
27.3%
Parcae-1.3B, hidden states
8 → 4
0.61
39.2%
Looped-Qwen3, hidden states
8 → 4
0.76
23.6%
Appendix
Table A9: Each half-depth setting spends about half to three quarters of the full-depth forward pass, guidance included. Analytic forward FLOPs at a 512-token prefill, as a fraction of the same checkpoint’s unguided full-depth pass; logit settings include their second readout.
Figure A2: LoopCD-Hidden keeps most of the logit form’s gain behind a shallow coda and exceeds it in code generation. Mean change at the selected settings of Tables 2 and 3 , LoopCD-Logits (solid) against LoopCD-Hidden (outlined), each against its own unguided model. (a) Seven-benchmark mean, three configurations in order of coda depth, with the share of the logit gain the hidden-state form keeps. (b) Four-column code-generation mean, Huginn at both depths.
Figure A3: On most benchmarks the hidden-state form matches or exceeds the logit form; Parcae-1.3B is the exception. Change from the unguided model per benchmark at the selected settings of Tables 2 and 3 , LoopCD-Logits (filled) and LoopCD-Hidden (open). Rows are benchmarks; columns are the configurations with a coda, in order of coda depth.
ω=0.5
ω=1
Model
Reference
hidden states
logits
hidden states
logits
Ouro-1.4B
h1
0.058
0.050
0.19
0.13
Ouro-2.6B
h1
0.052
0.038
0.17
0.11
Huginn-0125 R=32
h1
0.048
0.083
0.15
0.24
Huginn-0125 R=32
h6
0.013
0.016
0.044
0.058
Parcae-1.3B
h1
0.021
0.030
0.046
0.105
Appendix
Table A10: At the same strength, LoopCD-Hidden moves the output more than LoopCD-Logits in Ouro and less behind a coda. Mean per-position Jensen–Shannon divergence (bits) from the unguided final prediction, 300 HellaSwag documents. Parcae’s values sit above a 0.006-bit floor that two identical passes already show.
Figure A4: How far guidance moves the output depends on the coda. Mean per-position Jensen–Shannon divergence (bits) from the unguided final prediction against the strength, 300 HellaSwag documents; LoopCD-Logits solid with filled markers, LoopCD-Hidden dashed with open markers. (a, b) Ouro, no coda: the hidden-state form moves the output more. (c, d) Huginn at R=32 , references h1 and h6 . (e) Parcae-1.3B, deep coda: the hidden-state form moves it less. Panels of one family share a y-axis.
Reference k
Alone (%)
Disagrees (%)
Gain at ω=0.5
Huginn-0125 , final prediction 37.54%
1
22.78
51.3
3.07
2
22.10
46.8
1.79
4
30.12
35.8
0.94
6
32.25
29.5
1.11
8
33.45
20.0
0.60
Appendix
Table A11: Guidance gain depends on disagreement rather than standalone accuracy. LoopCD-Logits at ω=0.5 with reference iteration k . Each block’s header gives the unguided model’s accuracy; “alone” is the reference’s own accuracy. Shot counts as in Table A2 .
Figure A5: Each Ouro recurrent iteration reworks the state; Looped-Qwen3’s damped substeps move it gradually. Standardized cosine similarity to the final state across the unrolled layers, ARC-Challenge, with a band for the spread over documents. Alternate shading marks Ouro’s recurrent iterations and Looped-Qwen3’s substeps; dots mark the ends of Ouro’s recurrent iterations, the ring the reference h1 of Tables 1 and 2 , the star the final state.
Figure A6: Across ten configurations, Ouro’s similarity drops at each recurrent-iteration boundary and Looped-Qwen3’s climbs more smoothly. Standardized cosine similarity of each intermediate state to the final state against normalized depth. Bands give the spread over documents; the star marks the final state.
Figure A7: Predictions inside a recurrent iteration stay far from the final one, and Huginn’s entropy settles by its sixteenth iteration. (a) Mean Jensen–Shannon divergence (bits) between the prediction at each unrolled layer and the final prediction, first answer position, ARC-Challenge; shading marks Ouro’s four recurrent iterations, dots their ends. (b) Next-token entropy at answer positions against the recurrent iteration, Huginn at R=32 : median and interquartile band, with a dashed line at 2 nats.
Gain on each fifth, closest to farthest (points)
Answers changed
Net gain
Model
Benchmark
1
2
3
4
5
Ouro-1.4B
ARC-Challenge
+6.41
+3.85
+0.43
0.00
0.00
6.6%
+2.1
Ouro-2.6B
ARC-Challenge
+7.26
+3.85
+1.28
+0.43
0.00
5.8%
+2.6
Huginn-0125
ARC-Challenge
+7.26
+4.70
+2.13
+0.85
+0.43
14.9%
+3.1
Parcae-1.3B
ARC-Challenge
+13.25
+5.56
-0.43
-0.85
0.00
8.8%
+3.5
Ouro-1.4B
HellaSwag
+3.69
+0.10
0.00
0.00
0.00
3.9%
+0.76
Appendix
Table A12: The gain comes from the least confident questions and fades toward the most confident. LoopCD-Logits at ω=0.5 , reference h1 , questions in five equal groups (fifths) by the unguided prediction’s confidence, least confident first; the last two columns give the share of answers changed and the net change in accuracy. ARC-Challenge and HellaSwag, shot counts as in Table A2 .
Option
After the first pass, alone
Final prediction
Guided, ω=0.5
a flood
-1.500
-0.770
-0.934
a tornado (correct)
-1.113
-0.209
-0.170
a hurricane
-0.969
-0.258
-0.247
an earthquake
-0.525
-0.201
-0.199
Appendix
Table A13: Guidance flips a close decision on one ARC-Challenge question. Ouro-1.4B, 25-shot: per-character log-likelihood of each option after the first pass, at the final pass, and under LoopCD-Logits ( ω=0.5 , reference h1 ); bold marks the leading option.
Figure A8: LoopCD-Hidden follows the logit re-ranking, in its answers and in its direction. HellaSwag, ω=0.5 , reference h1 . (a) Questions, of 10,042, on which LoopCD-Hidden’s answer differs from the full logit update (light) and from its re-ranking part alone (solid). (b) Cosine similarity in mean-centred logit space between the LoopCD-Hidden update and the re-ranking part (solid) or the full update (dashed), mean over 300 documents.
Figure A9: Every generation benchmark gains only over a narrow range of strength. Change from the unguided model against the fixed logit strength, five models, on GSM8K, MMLU-Pro, HumanEval, and MBPP; a symmetric-log axis keeps the fall beyond ω≈0.6 in frame. Huginn’s GSM8K runs use chain of thought.
Answers changed
Least confident fifth
Most confident fifth
Model
ω=0.25
0.5
1
ω=0.5
1
ω=0.5
Ouro-1.4B
3.8%
6.6%
9.0%
+6.4
+3.9
0.0
Huginn-0125
9.1%
14.9%
20.7%
+7.3
+5.1
+0.4
Parcae-1.3B
–
8.8%
–
+13.3
–
0.0
Appendix
Table A14: A larger strength changes more answers, yet gains less on the least confident questions. LoopCD-Logits on ARC-Challenge at three fixed strengths ω : the share of answers changed, and the accuracy change (points) in the least and the most confident of five equal groups (fifths) by the unguided prediction’s confidence. A dash marks a strength not run for that model.