Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/
Figures & tables
Intervention
Parameters
Layer 14
Layer 26
LoRA
66K
96.5
54.5
Projection LoRA
606K
90.0
52.5
FLAS flow, 3 steps
66K
94.0
54.0
FLAS block, 1 step
168M
95.5
56.0
Table 1: Different edits share an early placement preference. Qwen3-8B exact accuracy (%) on two-chain, 24-line programs; 200 programs per cell with matched training. This comparison evaluation is separate from the headline LoRA result. Step counts are the same at both layers.
Held-out model
Placement bracket
Cutoff
Llama-3.2-1B
7–8
8.0 ✓
Qwen3-4B
20–22
22.0 ✓
Gemma-3-12B
24–27
23.0 ✓
OLMo-3-32B
22–26
19.0
Table 2: The frozen cutoff locates three held-out limits. Brackets join the last working and next tested layer. Checks allow one layer beyond either edge, as preregistered. All entries are layer indices; all nine models and baseline rules appear in Appendix Table 7 .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
LoRA
Learned scale s
Qwen3-8B, layer 14
1.0006
Qwen3-8B, longer-trained LoRA
1.008
Ouro-1.4B, layer 6, every loop
1.004
Ouro-1.4B, longer-trained LoRA
1.002
Ouro-1.4B, first loop only
1.039
Huginn, trained up to 12 lines
1.06
Appendix
Table 3: Learned scales distinguish the LoRA’s scalar and low-rank components. Each value is the learned s in M(h)=(sI+BA)h . Longer-trained LoRAs use the training lengths described below; the Qwen3-8B re-entry LoRA accompanies additional passes through layers 14–22.
Model
Layers
Choice accuracy by chain length (chance 1/3 )
Reach
Cutoff
Value copy
Relay
Unlocked
1
2
3
4
5
6
(lines)
(rel. depth)
(rel. depth)
(line)
(lines)
Llama-3.2-1B
16
0.98
0.50
0.41
0.29
0.30
0.36
1.4
0.50
0.94
3
6
Qwen3-0.6B
28
0.99
0.52
0.34
0.35
0.25
0.39
1.4
0.36
0.92
3
13
Qwen3-1.7B
28
0.99
0.76
0.40
0.35
0.27
0.32
1.8
0.64
0.92
3
22
Llama-3.2-3B
28
0.98
0.78
0.47
0.28
0.28
0.38
1.9
0.45
0.83
3
≥ 24
Qwen3-4B
36
0.99
0.77
0.51
0.36
0.34
0.35
1.8
0.61
0.92
4
≥ 24
Appendix
Table 4: Frozen reach stays short across thirteen standard models. Choice accuracy uses three chains and 200 programs per cell. Cutoff and value-copy layers are relative to depth; value copy is the first layer at which the final token holds half the root-value effect, averaged over one- to three-line chains. Relay is the last consecutively readable line in two-chain, eight-line programs. Unlocked reach uses exact accuracy with a rank-8 LoRA at the best tested layer: six to thirteen placements per model, three for Llama-3.2-3B, and two for Qwen3-14B. A dash denotes an untrained LoRA.
Chains, order
Model
4 lines
8
12
16
20
24
2, level
frozen
57.5
44.5
25.5
15.5
14.5
15.5
2, level
LoRA
100.0
100.0
99.5
98.5
98.0
99.0
2, interleaved
frozen
60.0
44.0
33.0
21.5
18.0
17.0
2, interleaved
LoRA
99.5
95.0
89.5
82.0
85.5
76.5
3, level
frozen
48.5
15.5
8.0
8.0
–
–
3, level
LoRA
97.0
83.5
71.5
58.0
–
–
Appendix
Table 5: The Qwen3-8B LoRA transfers across assignment orders and chain counts. Exact accuracy (%), 200 programs per cell. Training uses two chains in level order up to twenty lines; dashes denote unevaluated cells.
Change
Where
Parameters
Text KL
8 lines
16 lines
24 lines
none (frozen)
0
0
36.0
19.5
17.5
LoRA: h←sh+BAh
rank 8, layer 14
66K
0.005
100.0
99.0
96.5
the same LoRA at three layers
layers 6, 10, 14, shared
66K
0.006
100.0
100.0
94.5
low-rank flow (FLAS-style), 3 steps
rank 8, layer 14
66K
0.003
100.0
99.5
94.0
low-rank flow, 1 step
rank 8, layer 14
66K
0.002
98.0
95.0
79.0
FLAS flow block, 3 steps
MLP + time, layer 14
168M
0.012
99.5
99.5
90.5
Appendix
Table 6: Several forms of early intervention extend reference chains. Exact accuracy (%) on two-chain programs, 200 per length, with 50% root-choice chance. Parameter counts denote trained parameters. Text KL is the divergence from frozen WikiText predictions per token, averaged over the final 200 training steps. Layer-26 interventions are after the working region.
Model
Layers
Limit bracket
Cutoff
45%
Copy −11.5
Qwen3-8B
36
20–21
20.5 ✓
16.2
20.2 ✓
OLMo-3-7B
32
12–15
15.0 ✓
14.4 ✓
12.7 ✓
Llama-3.1-8B
32
13–15
12.5 ✓
14.4 ✓
13.4 ✓
Qwen3-1.7B
28
12–14
18.0
12.6 ✓
13.5 ✓
Qwen3-0.6B
28
12–14
10.0
12.6 ✓
13.7 ✓
Llama-3.2-1B
16
7–8
8.0 ✓
7.2 ✓
2.8
Appendix
Table 7: The cutoff predicts three of four held-out placement brackets within the preregistered tolerance. All entries are layers. The first five models informed the rules; the last four were held out. “Cutoff” is the frozen measurement, “Copy” the value-copy rule, and checks mark predictions within the observed bracket expanded by one layer on either side.
Model
Before / after
Seed 0 (before)
Seed 0 (after)
Seed 1 (before)
Seed 1 (after)
Bracket stability
Qwen3-8B
20 / 21
20.50
5.17
21.76
6.49
100.0%
OLMo-3-7B
12 / 15
21.47
6.62
20.93
7.66
100.0%
Llama-3.1-8B
13 / 15
16.59
2.93
20.00
3.11
100.0%
Qwen3-1.7B
12 / 14
14.86
6.74
–
–
100.0%
Qwen3-0.6B
12 / 14
8.85
4.47
–
–
100.0%
Llama-3.2-1B
7 / 8
4.59
3.68
4.52
3.39
85.5%
Appendix
Table 8: Independent seeds preserve the observed placement contrasts. Reach is measured in lines at 80% exact accuracy. Bracket stability is the fraction of 2,000 binomial resamples retaining the bracket; it conditions on the evaluated grid and seed.
Model
Intervention
Seeds (EM)
EM
Gain [95% CI]
Qwen3-8B
frozen
–
52.9
–
LoRA at 6
–
64.3
+11.4 [+8.6, +14.3]
LoRA at 10
–
62.1
+9.2 [+6.1, +12.2]
LoRA at 20
–
57.7
+4.8 [+1.9, +7.8]
LoRA at 26
–
53.2
+0.3 [-2.2, +3.0]
LoRA at 30
–
51.3
-1.6 [-4.2, +1.2]
Appendix
Table 9: MuSiQue comparisons for Qwen3-8B with fixed prompts. EM and gains average seeds per question; intervals are paired 95% question-bootstrap intervals on 900 gold-context development questions. Ranges are inclusive. Projection LoRA learning rate is 3⋅10−4 unless stated.
Model
Intervention
Seeds (EM)
EM
Gain [95% CI]
OLMo-3-7B
frozen
–
54.4
–
LoRA at 4
–
63.9
+9.4 [+6.2, +12.6]
LoRA at 8
–
60.7
+6.2 [+3.1, +9.3]
LoRA at 12
–
60.3
+5.9 [+2.6, +9.2]
LoRA at 16
–
53.3
-1.1 [-4.2, +2.1]
LoRA at 20
–
52.4
-2.0 [-4.8, +0.8]
Appendix
Table 10: MuSiQue comparisons for OLMo-3-7B and Llama-3.1-8B. Same 900-question protocol as Table 9 ; projection LoRA learning rate is 10−4 .
Ouro-1.4B
T=1
T=2
T=3
T=4
Ouro-1.4B frozen
0.0 [0.0, 0.0]
1.9 [1.7, 2.1]
2.3 [1.9, 2.6]
2.5 [2.3, 2.8]
steering vector
–
2.9 [2.7, 3.2]
6.1 [5.6, 6.5]
7.1 [6.8, 7.7]
LoRA, every loop (seed 0)
1.9 [1.7, 2.2]
7.1 [6.7, 7.7]
17.0 [16.2, 17.5]
26.7 [26.1, 27.6]
LoRA, every loop (seed 1)
1.8 [1.6, 2.1]
7.3 [7.0, 7.8]
20.0 [18.5, 21.1]
≥ 24.0 [24.0, 24.0]
LoRA, every loop (seed 2)
2.3 [2.1, 2.5]
7.0 [6.8, 7.4]
17.9 [17.2, 18.8]
≥ 24.0 [24.0, 24.0]
LoRA, first loop only
–
9.5 [8.9, 10.1]
21.5 [20.0, 23.2]
26.3 [25.3, 27.4]
Appendix
Table 11: Ouro reach over the first four loops. Lines at 80% choice accuracy, with 95% parametric-bootstrap intervals; the same trained LoRAs continue in Table 12 .
Ouro-1.4B
T=6
T=8
T=12
Ouro-1.4B frozen
2.3 [2.1, 2.5]
2.2 [1.9, 2.4]
–
steering vector
6.8 [5.9, 7.6]
6.3 [5.6, 6.9]
–
LoRA, every loop (seed 0)
37.4 [36.1, 39.2]
40.9 [39.1, 43.9]
–
LoRA, every loop (seed 1)
≥ 24.0 [24.0, 24.0]
≥ 24.0 [24.0, 24.0]
–
LoRA, every loop (seed 2)
≥ 24.0 [23.0, 24.0]
≥ 24.0 [24.0, 24.0]
–
LoRA, first loop only
24.5 [23.0, 25.7]
24.0 [19.5, 25.7]
–
Appendix
Table 12: Ouro reach with additional inference loops. Same estimates and conventions as Table 11 .
Huginn-0125
r=2
r=4
r=8
Huginn-0125 frozen
0.0 [0.0, 0.0]
0.0 [0.0, 0.0]
1.6 [1.4, 2.0]
LoRA (up to 12 lines)
0.0 [0.0, 0.0]
3.4 [2.7, 4.3]
16.8 [15.6, 17.6]
longer-trained LoRA (up to 24 lines)
< 2
4.6 [3.6, 5.4]
30.5 [27.8, 34.8]
Huginn-0125
r=16
r=32
r=64
Huginn-0125 frozen
1.4 [0.0, 2.0]
1.5 [1.1, 2.0]
–
LoRA (up to 12 lines)
≥ 24.0 [24.0, 24.0]
≥ 24.0 [23.1, 24.0]
–
Appendix
Table 13: Huginn reach by recurrence count. The standard LoRA trains up to twelve lines and the longer-trained LoRA up to 24; intervals and censoring follow Table 11 .
We introduce training-free looped transformers, in which a lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional fine-tuning, continued training, or architectural changes. Unlike prior looped transformer methods that train with the looped structure end-to-end, we retrofit recurrence onto pretrained models at test time. We show that naive block reapplication usually degrades performance, highlighting the importance of the loop application strategy. Motivated by viewing a pre-norm transformer block as a forward Euler step on an ODE, we instead treat looping as a refinement of the same approximation, replacing one large update with smaller damped sub-steps. Across seven dense, sparse MoE, and MLA+MoE model families, our method improves Qwen3-4B-Instruct by +2.64 pp on MMLU-Pro, Qwen3-30B-A3B-Instruct by +1.14 pp on CommonsenseQA, and Moonlight-16B-A3B-Instruct by +1.20 pp on OpenBookQA.
We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. With a novel post-training method, Loopie develops strong reasoning abilities and achieves frontier-level reasoning performance.
Chain-of-Thought (CoT) prompting substantially improves the sample efficiency of transformers, reducing the complexity of tasks like parity learning from exponential to polynomial in the input length. However, generating explicit reasoning steps at inference is computationally expensive. Implicit Chain-of-Thought (ICoT) has emerged as a promising empirical remedy that trains models to internalize intermediate steps within their hidden states, but its theoretical foundations remain poorly understood. We give the first theoretical analysis of ICoT, proving that an L-layer transformer trained under our proposed Log-ICoT curriculum learns k-parity with poly(n) samples and L=log2k training stages. This matches the sample efficiency of explicit CoT while eliminating its inference overhead, and extends prior one-layer parity guarantees to multi-layer architectures. Compared to standard ICoT, which removes thinking tokens one at a time, Log-ICoT removes them in geometric chunks, reducing the number of stages from linear in k to logarithmic. Experiments on multi-layer transformers confirm the theory and visualize how reasoning is progressively absorbed into deeper layers.