Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, we establish practical training recipes for looped language models, with three main results. (1) We develop a compute-efficient from-scratch pipeline that reduces the training budget from 7.7T tokens in Ouro to 310B tokens while retaining strong reasoning performance. Pretraining followed by high-quality mid-training, together with learning-rate warmup and stronger exit-gate regularization, enables stable recurrent training without prior multi-stage schedules. (2) Under controlled comparisons, our 1.4B LoopLM outperforms a parameter-matched dense model trained on the same data and token budget on all 12 evaluated benchmarks, including +14 points on GSM8K, +10 on MATH, and +22 on DROP. At matched inference compute, it approaches a 3.9B dense model on mathematical reasoning and reading comprehension while using only 36% as many parameters. (3) We introduce a minimal recipe for converting pretrained dense models into looped ones: a single learned input-mixing scalar and a smoothed exit loss, with no step-specific parameters. Applied to Qwen3-1.7B-Base, Looped Qwen improves over an identically continued dense baseline on every evaluated benchmark across two data regimes, with statistically clear gains on GSM8K, MATH, and MMLU-Pro on the curated mixture. Together, these results make looped language models substantially cheaper to train from scratch and practical to introduce into existing pretrained checkpoints, while isolating the gains due to recurrence itself.
Figures & tables
Figure 1: Looped language models trained from scratch or converted from a pretrained dense model, each compared with dense models trained on the same data. (a) A 1.4B LoopLM (ours) trained from scratch on 310B tokens beats its parameter-matched dense twin on every benchmark and stays close to its inference-compute-matched 3.9B twin on math and reading comprehension; LoopLM is shown at four iterations, the 3.9B twin matches three. (b) Looped Qwen, converted from Qwen3-1.7B-Base with one learned scalar and a smoothed exit loss, improves on that checkpoint given the same continued training. (c) Mean score per group at each recurrent step: mathematics (GSM8K, MATH, MATH500), code execution (CRUXEval) and knowledge and commonsense (the six multiple-choice benchmarks of Table 2 ).
Figure 2: Why T=4 . (a) Per-loop LM loss of a T=8 run ( β=0.1 , 124k steps) and the gain of each loop over the previous one: loops five to eight change the loss by about 0.01 nats or less. (b) Exit probability per loop in the same run; the gate spreads its mass evenly over loops four to eight. (c) The same for the main T=4 run. Values are means over the last 5k steps.
Figure 3: Gate collapse and what prevents it. (a) Probability of exiting after the first loop over the first 6k steps for β∈{0.05,0.15} , with a 1k-step warmup and with Ouro’s constant schedule; the constant-schedule β=0.15 run is separate from the cosine-schedule run of Appendix B . (b) LM loss over the same steps: the collapsed run is indistinguishable by loss alone.
Pretraining
Mid-training
Initialization
From scratch
Pretraining checkpoint
Peak learning rate
3×10−4
1×10−4
Final learning rate
3×10−5
1×10−5
LR scheduler
Warmup + cosine decay
LR warmup steps
1,000
Weight decay
0.1
Table 1: Our two-stage training recipe. Both stages share the architecture, recurrent depth, and optimizer; the main difference is the data and learning-rate schedule.
Figure 4: Final Looped Qwen architecture. The pretrained decoder is shared across steps; optional LoRA adapters are step-specific. Later steps mix the original embeddings with the preceding hidden state, while a shared exit gate accumulates stopping probability.
Benchmark
dense-ouro-1.4B
LoopLM (ours)
dense-ouro-3.9B
1 iter.
2 iter.
3 iter.
4 iter.
Math and reading comprehension
GSM8K
69.14±1.27
60.65±1.35
76.50±1.17
81.58±1.07
83.32±1.03
82.79±1.04
MATH
54.10±0.67
44.18±0.67
60.90±0.65
63.64±0.64
63.80±0.64
64.54±0.64
MATH500
55.00±2.22
45.00±2.22
61.40±2.18
63.60±2.15
63.20±2.16
63.60±2.15
DROP (F1)
38.59±0.48
22.20±0.41
45.02±0.49
57.40±0.49
60.57±0.48
58.57±0.49
Table 2: Benchmark scores after mid-training (Section 3.4 ). LoopLM is evaluated at a fixed exit after k iterations; dense-ouro-3.9B matches the inference compute of three iterations.
Model
Tokens
GSM8K
MATH
MATH500
DROP
MMLU
MMLU-Pro
ARC-C
ARC-E
HellaSwag
Winogrande
LoopLM (ours), 4 iter.
0.31
83.3
63.8
63.2
60.6
56.8
24.5
51.0
78.0
59.3
63.9
LoopLM (ours), 3 iter.
0.31
81.6
63.6
63.6
57.4
55.9
24.3
50.5
77.8
58.9
62.9
Ouro-1.4B
7.7
78.9
70.9
82.4
49.7
67.4
48.6
60.9
84.0
74.3
72.3
Huginn-3.5B
0.8
32.6
12.6
13.2
17.8
31.4
—
38.2
69.9
65.2
59.4
Gemma-3-1B
2
2.1
3.7
41.0
42.4
39.9
11.3
38.4
73.0
62.3
58.2
Llama-3.2-1B
9
7.1
3.3
7.4
28.0
32.2
11.8
32.8
—
59.4
62.8
Table 3: LoopLM after mid-training against published models between 1B and 4B parameters. Bold marks the best score in each column; underline marks the second best.
Benchmark
Qwen3-1.7B-Base †
Looped Qwen, shared
Looped Qwen, LoRA-32
Qwen3-4B-Base †
GSM8K
66.64±1.30
70.96±1.25
69.90±1.26
82.87±1.04
MATH
51.14±0.66
53.34±0.66
54.46±0.67
65.92±0.63
MMLU
60.62±0.39
61.05±0.39
60.99±0.39
71.83±0.36
MMLU-Pro
30.76±0.41
34.52±0.42
36.24±0.43
49.52±0.44
ARC-Challenge
53.84±1.46
55.46±1.45
54.95±1.45
64.42±1.40
ARC-Easy
79.97±0.82
80.13±0.82
80.51±0.81
86.32±0.71
Table 4: Conversion on the curated instruction mixture, 60B tokens, evaluated at a fixed exit after two of three recurrent steps. ± is the standard error; † denotes continued training on the same mixture for the same number of steps; bold marks Looped Qwen ahead of Qwen3-1.7B-Base † .
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Final normalization
Decoder norm
RMS
FRMS
Identity
Ouro-norm (baseline)
1.98
–
2.11
Pre-norm
2.03
2.07
2.11
Magic-norm
–
2.04
2.07
Post-norm
–
–
7.37
Sandwich-norm
–
–
7.40
Appendix
Table 5: LM loss after 19K training steps with a 500K-token batch.
Figure 5: Training LM loss (15-step moving average) for representative pre-norm-family variants and for post-norm and sandwich-norm. The pre-norm-family curves nearly overlap, while post-norm and sandwich-norm plateau early.
Figure 6: Gradient norm by decoder-layer depth at recurrent iteration t=2 . Pre-norm preserves gradients throughout the block, whereas post-norm and sandwich-norm concentrate almost all gradient in the final decoder layer.
Figure 7: Post-norm training LM loss (15-step moving average) as a function of recurrent depth T , using a 500K-token batch and a 10K-step warmup. Lower recurrent depth mitigates the failure, but remains worse than the pre-norm-family configurations.
Decoder norm
Final norm
Loss
#Spikes
Max grad norm
Ouro-norm
RMS
1.98
0
42.3
Ouro-norm
Identity
2.11
5
1.7×103
Pre-norm
RMS
2.03
0
15.7
Pre-norm
FRMS
2.07
0
28.4
Pre-norm
Identity
2.11
11
5.2×107
Magic-norm
FRMS
2.04
0
30.0
Appendix
Table 6: Effect of final normalization on loss and training stability. Loss is measured at 19K steps; spikes and maximum gradient norm are measured over the first 25K steps.
β
Schedule
Steps
Peak p0
Late p0
Loss at 2K
Outcome
0.05
Constant
2,267
1.00
1.00 ∗
2.38
No recovery in 2.3K steps
0.10
Constant
54,558
1.00
0.11
2.36
Recovers
0.15
Cosine
24,579
0.99
0.14
2.39
Recovers
0.50
Constant
27,204
0.66
0.20
2.59
Never collapses
∗ Mean over the whole run, which was stopped at step 2,267.
Appendix
Table 7: No-warmup β sweep. Peak p0 is the maximum first-step exit probability after the first 50 steps; late p0 is its mean over steps 5K to 6K. Loss is the LM loss at step 2K, the latest step all four runs reached. The β=0.05 run is the collapsed run shown in Figure 3 .
Figure 8: First-step exit probability p0 during the first 8K training steps without warmup. The shaded region marks p0>0.95 . The β=0.05 run stays collapsed until it was stopped at step 2,267; β=0.10 and 0.15 recover, and β=0.50 never enters the shaded region after its first 50 steps.
MATH target; mid-trained checkpoint; solve rate ≈0.77
MATH, 0-shot (target)
55.68±0.68
–
55.18 ( −0.50 )
Appendix
Table 8: RLTT results at learning rate 10−5 after 140 optimizer steps. The reported uncertainty is the lm-evaluation-harness standard error and depends on the benchmark rather than the checkpoint. GSM8K in the MATH run is a guard metric, not an RL target.
GSM8K run
MATH run
KL, first 20 steps
0.013
<0.001
KL, last 20 steps
0.084
0.003
Gradient norm, first 20 steps
0.58
0.12
Gradient norm, last 20 steps
0.64
0.11
Appendix
Table 9: Training telemetry for the two 10−5 runs, averaged over the first and last 20 optimizer steps. KL is measured against the reference policy and therefore starts at zero by construction.
Parameter
Value
Initialization
Random, std. 0.02
Decoder layers
24 , shared across recurrent steps
Hidden / FFN dimension
2048 / 5632
Attention heads
16 query, 16 key–value; no GQA
Attention head dimension
128
Attention pattern
Full causal attention; no sliding window
Appendix
Table 10: Backbone of the from-scratch 1.4B LoopLM, matching Ouro-1.4B ( Zhu et al., 2025 ) .
Parameter
Value
Recurrent steps T
4 applications of the shared decoder
Initial state
Token embeddings e
Later-step input
Previous hidden state ht−1
Embedding re-injection
None
Exit gate
Shared linear projection 2048→1 with sigmoid
Exit distribution
Learned per token; final step receives residual mass
Appendix
Table 11: Recurrence, exit gate, and training objective. The configuration is shared by both training stages.
Source
Weight
Nemotron-CC (high quality)
66.5%
Nemotron-CC-Math v1
15.0%
MegaMath (high quality)
4.6%
Nemotron Cascade SFT Stage 1
9.6%
Nemotron Post-Training (code)
3.8%
OpenCoder annealing corpus
0.5%
Appendix
Table 12: Pretraining mixture using mosaic-random sampling.
Source
Weight
FLAN
18.0%
TaskSource
9.0%
Natural Reasoning
7.0%
WebInstruct-verified
4.0%
Platypus
2.0%
No Robots
2.0%
Appendix
Table 13: Mid-training mixture based on HRM-Text sources. The first block preserves general knowledge; the second emphasizes mathematics and reasoning.
Projection
s0
Loss
Δ vs. none
None (baseline)
—
1.78
—
Per-step
truncated-normal
1.90
+0.12
Shared
truncated-normal
1.91
+0.12
Shared
learned zero
1.91
+0.12
Per-step
learned zero
1.91
+0.13
Appendix
Table 14: Training LM loss with and without concatenation embedding injection, at a matched budget of 40K steps.
Parameter
Value
Initialization
Qwen3-1.7B-Base checkpoint
Trainable backbone
Decoder, embeddings, and LM head
Decoder layers
28
Hidden / FFN dimension
2048 / 6144
Attention heads
16 query, 8 key–value
Attention head dimension
128
Appendix
Table 15: Backbone configuration of the converted Qwen3-1.7B model.
Parameter
Value
Recurrence
Recurrent steps
T=3 applications of the shared decoder
Step 0 input
Original token embeddings e
Later-step input
zt=we+(1−w)ht−1
Injection parameter
One trainable scalar shared across steps
Initialization
a=4.6 , w=σ(a)=0.99005
Appendix
Table 16: Recurrent modules and objective used for Looped Qwen.
Parameter
Value
Adapter layout
One independent bank per recurrent step
Rank
32
Attention targets
q , k , v , output projections
MLP targets
Gate, up, down projections
Parameterization
Unscaled BAx
Initialization
Kaiming-uniform A , zero B
Appendix
Table 17: Step-specific LoRA configuration.
Parameter
Value
Token budget
60B tokens
Optimizer updates
28,611
Sequence length
4096
Optimizer
Fused AdamW
AdamW coefficients
β1=0.9 , β2=0.95 , ϵ=10−8
Peak / minimum LR
3×10−5 / 3×10−6
Appendix
Table 18: Optimization and batch configuration for the 60B-token conversion run.
Source
Weight (%)
Nemotron Post-Training Dataset v1
30.00
Nemotron Post-Training Dataset v1 (math)
27.90
Nemotron Post-Training Dataset v1 (code)
23.10
Nemotron SFT Instruction Following Chat v2
11.90
Nemotron Instruction Following Chat v1
3.86
Nemotron Science v1
3.24
Appendix
Table 19: Curated instruction mixture used for the 60B-token Looped Qwen run.
Figure 9: Conversion training logs, one seed per configuration. (a) Naive recurrence without input re-injection and (b) scalar injection w=σ(a) with a0=3.5 , ϵ=0 and no LoRA: loss of steps 2 and 3 minus the loss of step 1 on the same batch. Shading marks where the exit probability of step 1 is above 0.9; before about 4k steps the gap is above the plotted range. (c) Per-step LM loss of the run with a learned per-step projection. (d) Step-2 loss at a0=4.6 with rank-8 LoRA, without and with exit-loss smoothing; the unsmoothed run was stopped at 4.9k steps. (e) Learned injection weight w on a logit axis for a0=2.2 , 3.5 and 4.0 ( ϵ=0 , no LoRA) and a0=4.6 ( ϵ=0.05 , rank-8 LoRA). (f) The gap of (a) and (b) for the instruction-mixture runs with a0=4.6 and ϵ=0.05 , without adapters (solid) and with rank-32 LoRA (dashed), in thousandths of a nat. Plotted points are logged means over 5 to 10 optimizer steps (every step in the unsmoothed run of (d)); in (a), (b) and (f) the thick line is a 31-point moving average of these points.
Benchmark
Qwen3-1.7B-Base †
Looped Qwen, shared
Looped Qwen, LoRA-32
Qwen3-4B-Base †
GSM8K
60.12±1.35
61.87±1.34
63.68±1.32
76.50±1.17
MATH
38.00±0.65
38.08±0.65
39.22±0.65
47.12±0.66
MMLU
61.18±0.40
62.73±0.39
62.15±0.39
71.04±0.36
MMLU-Pro
32.43±0.42
32.96±0.42
33.88±0.42
46.28±0.44
ARC-Challenge
55.29±1.45
55.38±1.45
54.61±1.45
61.35±1.40
ARC-Easy
80.81±0.81
81.27±0.80
80.81±0.80
85.48±0.72
Appendix
Table 20: Conversion on the 84B-token web-scale mixture. Looped models are evaluated after two of three recurrent steps. ± denotes standard error; † continued on the same mixture for the same number of steps; bold marks scores above Qwen3-1.7B-Base † .
Model
Inference
Training
Total training
(GFLOPs/token)
(GFLOPs/token)
(FLOPs)
LoopLM, exit k=3 / k=4
7.60 / 10.07
32.0
9.9×1021
dense-ouro-1.4B
2.67
8.0
2.5×1021
dense-ouro-3.9B
7.60
22.8
7.1×1021
Ouro-1.4B schedule (estimate)
10.07
32.0 / 64.0
3.4×1023
Looped Qwen, exit k=2 / k=3
6.26 / 9.08
31.0
1.9×1021 (60B)
Appendix
Table 21: Compute per token, with the usual approximation of 2N FLOPs for a forward pass and 6N for a training step, where N counts the parameters applied to a token: the non-embedding layers times the recurrent steps, plus the LM head. Looped models train with all recurrent steps and apply the LM head at every step. Attention-score FLOPs are left out; they are equal within each matched pair. The Ouro-1.4B total is an estimate from its reported stage budgets and recurrent depths on the same architecture.
Modern LLMs are trained to "think" primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training and under-leverages pre-training data. We present and open-source Ouro, named after the recursive Ouroboros, a family of pre-trained Looped Language Models (LoopLM) that instead build reasoning into the pre-training phase through (i) iterative computation in latent space, (ii) an entropy-regularized objective for learned depth allocation, and (iii) scaling to 7.7T tokens. Ouro 1.4B and 2.6B models enjoy superior performance that match the results of up to 12B SOTA LLMs across a wide range of benchmarks. Through controlled experiments, we show this advantage stems not from increased knowledge capacity, but from superior knowledge manipulation capabilities. We also show that LoopLM yields reasoning traces more aligned with final outputs than explicit CoT. We hope our results show the potential of LoopLM as a novel scaling direction in the reasoning era. Our model is available here: http://ouro-llm.github.io.
Rui-Jie Zhu, Zixuan Wang, Kai Hua +30
1ByteDance Seed · 2UC Santa Cruz · 3Princeton University +8
Looped computation shows promise in improving the reasoning-oriented performance of LLMs by scaling test-time compute. However, existing approaches typically require either training recurrent models from scratch or applying disruptive retrofits, which involve substantial computational costs and may compromise pretrained capabilities. To address these limitations, we introduce \textbf{Looped Depth Up-Scaling} (LoopUS), a post-training framework that converts a standard pretrained LLM into a looped architecture. As a key technical contribution, LoopUS recasts the pretrained LLM into an encoder, a looped reasoning block, and a decoder. It operationalizes this latent-refinement architecture through four core components: (1) block decomposition, guided by staged representation dynamics; (2) an input-dependent selective gate to mitigate hidden-state drift; (3) random deep supervision for memory-efficient learning over long recursive horizons; and (4) a confidence head for adaptive early exiting. Collectively, these mechanisms transform a standard non-looped model into a looped form while stabilizing it against both computational bottlenecks and representation collapse. Through stable latent looping, LoopUS improves reasoning-oriented performance without extending the generated traces or requiring recurrent training from scratch. For more details, see https://thrillcrazyer.github.io/LoopUS
Taekhyun Park, Yongjae Lee, Dohee Kim +1
Department of Data Science Pusan National University Busan, Republic of Korea · Department of Industrial Engineering Pusan National University Busan, Republic of Korea · Department of Artificial Intelligence Engineering Changwon National University Changwon, Republic of Korea
Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two design choices limit these gains. Each iteration sees only the previous output and cannot directly access earlier computations. Moreover, a fixed loop count wastes depth on easy inputs while leaving hard ones with too little computation. We introduce RecurTrace, which addresses both limitations using the loop's own trajectory. Specifically, Loop Memory Attention lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone. A halting head then reads the loop state and predicts whether to continue, with supervision from an oracle that identifies when additional depth still reduces loss. In a controlled MathQA comparison on the same looped backbone, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, exceeding the best fixed loop depth by 2.2 points at matched compute. By comparison, ACT and PonderNet collapse to one loop, and CALM reaches only 54.1% with 5.6 loops, while the stronger LoopUS-Conf and TaH-Mismatch baselines reach 55.3% at 3.2 loops and 55.7% at 2.1 loops. Finally, RecurTrace improves generation accuracy over same-budget fine-tuned baselines at 0.6B, 1.7B, 4B, and 8B, with the gain growing with model size from 0.6 to 3.4 points.
Yuxiang Wang, Kunyu Feng, Yingda Shen +3
The Chinese University of Hong Kong, Shenzhen · The Chinese University of Hong Kong · Tianjin University +2