Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong performance on iterative reasoning and length-generalization tasks suggests an appealing explanation: recurrence may provide an inductive bias that lets the model reuse a learned algorithm across loops. However, weight sharing alone does not imply that every loop performs the same operation. This raises a basic question: is each loop actually repeating the same computation, and if not, what routes the shared parameters to different operations? We study this question using graph walks as a test case. In the model's native trajectories, decoded predictions can advance by different numbers of graph steps or remain at a reached target, showing that recurrent progress need not follow a fixed one-loop-one-step pattern. We then show that a frozen loop can be steered toward different transitions by modifying its entering hidden state: a learned linear layer J selects the desired transition without changing the shared Transformer layers. To test how this steering works, we use activation patching and find that attention patterns can recover its effects and switch the selected transition. Across five matched pairs of graph models, changing intermediate supervision during backbone training changes which transitions J can induce. This suggests that J selects computations learned by the backbone rather than creating new algorithms. Together, these results show that the hidden state can control shared computation, with attention routing as a causal pathway.
Figures & tables
Figure 1: Overview of looped computation and state steering. Phenomenon: intermediate readouts reveal different patterns of progress across weight-shared loops. Control: a learned linear layer changes the state entering the next loop to select a different transition while the backbone remains frozen. Mechanism: patching attention patterns from a steered run into an unsteered run transfers routing while preserving the receiving run’s values, recovering steering effects. Flames mark trainable components; snowflakes mark frozen components.
Figure 2: Color shows prediction frequency. Red dashed lines mark the training loop L=8 or L=6 ; white dots mark the unique most frequent prediction at each loop.
(a) from h6
(b) from F(Jone(h6))
(c) from F(Jtwo(h6))
Figure 4: Controller reuse on 512 new held-out graphs with D8L6 seed 6. The horizontal axis counts additional controlled loops after h6 ; exact match compares the predicted node with the cumulative target at each call. Curves average two map fits on the same 3,175 examples; mixed averages 32 fixed sequences. Shading shows 95% graph-bootstrap intervals.
Figure 5: Attention routing and state steering. (a) Attention computation. (b) Attention-pattern and output patches between graphs. (c) Attention-pattern patches between steered ( J ) and unsteered ( 0 ) runs. D8L6 bars are averaged over backbone means; markers indicate seeds 6 (circle), 10 (diamond), and 13 (triangle). Ouro error bars show 95% Wilson intervals.
Figure 6: Pattern patching between one-hop and two-hop runs. Arrows indicate the patch direction. “Raw pattern” uses the run’s own attention pattern, while “Patched pattern” uses the second-layer pattern from the other run. Points show means across backbones, and horizontal bars show their ranges.
Figure 7: Matched graph supervision. Colors mark answers relative to u=fG8(s) . Bars average two map fits per backbone and five paired backbone seeds.
Figure 8: Ouro antonym cancellation. (a,b) Readout accuracy (%) as a function of loop and deletion depth. (c) Accuracy using four loops and a fitted map.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Execution beyond map-training ranges. (a) Parity seed 2 at t=n : dark curves are smoothed; faint points are individual measurements. (b) Graph continuation: lines average twelve backbones and two maps each; shading shows 95% bootstrap intervals over backbone means. Gray regions mark map-training ranges.
Figure 10: Parity readout timing for seed 2. (a,b) Accuracy by input length n and loop t , without and with J ; insets enlarge lengths 54–60. The diagonal marks t=n ; dots mark the best nearby integer loop. (c) Readout offset t−n and fitted trends.
Figure 11: Native D8L8 readouts for seeds 0–5.
Figure 12: Native D8L8 readouts for seeds 6–11.
Figure 13: Native D8L6 readouts for seeds 3–7.
Experiment
Scored count
Inclusion and intervention
Native trajectories
5,120 per backbone
512 ten-node cycles; all starts; no prediction filtering.
Target selection
4,110
Distinct current, one-hop, and two-hop labels; no prediction filtering; one additional F .
Pattern vs. output
3,966 / 3,963
Distinct semantic answers and correct clean readouts in both runs; selected second-layer head H2.
Steering-pattern patching, Figure 5 c
4,044
Distinct current, one-hop, and two-hop labels; no prediction filtering; all L2 heads at the answer position.
Output restoration
4,413 / 4,417
Eligible clean transitions broken by query replacement; restore H2 or control H3.
Target switching
4,116
Separate 512-graph cohort; distinct current, one-hop, and two-hop labels; all second-layer heads and token positions.
Appendix
Table 3: D8L6 evaluation populations. Counts are graph–start instances or paired runs, as appropriate. All map fits within a row use the stated graph population; correctness-based eligibility can differ by fit.
Figure 14: Native readouts of the D8L6 backbones used in the attention experiments.
Figure 15: Paired backbone comparisons for (a) one-hop and (b) two-hop targets. Lines connect the two supervision regimes for the same backbone seed. Points average two map fits; horizontal offsets separate overlapping points.
Supervision
Intervention
Stay
One hop
Two hops
Final-only
None
100.0±0.0
0.0±0.0
0.0±0.0
Stepwise
None
0.0±0.0
100.0±0.0
0.0±0.0
Final-only
J
100.0±0.0
66.5±10.0
50.5±15.2
Stepwise
J
100.0±0.0
100.0±0.0
13.2±6.8
Appendix
Table 13: Matched ten-node D8L8 supervision comparison. Entries are mean target accuracies ± sample standard deviation across five backbone seeds (%). We first average the two map fits within each backbone. Under J , each column uses its own target-specific map.
One hop
Two hops
Seed
Supervision
Fit 1
Fit 2
Fit 1
Fit 2
0
Final-only
78.75
76.46
62.05
56.80
0
Stepwise
100.00
100.00
25.57
15.54
1
Final-only
74.09
72.71
66.76
65.55
1
Stepwise
100.00
100.00
11.58
12.79
2
Final-only
56.42
55.43
57.77
56.90
Appendix
Table 14: Matched D8L8 map accuracy (%) on 4,137 common test instances. Each pair of columns reports the two independently fitted maps for that target.
Experiment
Backbone
Map
Placement
D8L6 control
20,000 updates; seed 6 selected at 16,000
Rank 48; 8,000 updates; two fits per target
All tokens before the next F ; maps reused in compositions.
Matched graph supervision
20,000 updates; five paired final checkpoints
Rank 48; 8,000 updates; two fits per target
At h8 , followed by one additional F .
Ouro letter-walk
Checkpoint at update 200
Dense affine; checkpoint at update 500
Before loops 2–4; all tokens.
Ouro cancellation
500 updates per training strategy
Residual rank 128; 500 updates
After each loop’s RMSNorm, including loop 4 before the output head.
Appendix
Table 15: Training budgets and intervention sites for the main-text experiments. Backbone updates identify the selected checkpoint or the training budget as specified; map updates give the fitting budget.
Task or request range
J+F
F only
J only
Stop
Successor
100.00
75.00
71.09
70.31
Weekday
100.00
61.72
78.12
52.34
Doubling
100.00
3.91
4.69
3.12
Fibonacci
100.00
0.00
0.00
0.00
Collatz
48.44
2.34
2.34
2.34
All long requests ( k=5 –8)
89.69
28.59
31.25
25.62
Appendix
Table 16: Qwen continuation accuracy (%). All branches share the controlled computation through h4 ; column headings specify what happens afterward. Each task row contains 128 long requests ( k=5 –8). The pooled short and long rows each contain 640 prompts.
Figure 16: State steering in synthetic KG relation composition. Each point uses 1,024 queries, paired between the two conditions, with exactly n backbone calls for a length- n query. Gray shading marks backbone training lengths 1–3; blue shading marks map training lengths 4–16. The dashed vertical line marks the end of map training coverage.
Relation length
Queries
F only (%)
With affine J (%)
1–3
3,072
99.87
100.00
4–16
13,312
2.53
99.92
17–24
8,192
1.54
42.87
25–32
8,192
1.56
1.65
Appendix
Table 17: KG target accuracy aggregated over equally sized per-length evaluation sets. Every row compares the same queries and the same number of frozen-backbone calls.
Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across T iterations, they attain the effective depth and reasoning capabilities of larger models at a fixed parameter count. Yet existing approaches suffer from latent overthinking and undifferentiated computation, largely because intermediate representations receive no guidance across loops. Multi-token prediction (MTP) supplies exactly the dense, forward-looking supervision the loop is missing. We propose \textsc{LoopMTP}, which links the two through a structural correspondence in latent space: a model that loops T times can anticipate T future tokens. \textsc{LoopMTP} realizes this by softly aligning the hidden state of loop t with the embedding of the token t steps ahead, while a lightweight gate preserves useful information across iterations. \textsc{LoopMTP} improves average accuracy by up to 8.1% (relative) over the non-looped baseline, with training remaining stable for up to 15 loops.
Behzad Shomali, Markus Frey, David Berghaus +2
Lamarr Institute · University of Bonn · Fraunhofer IAIS
When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from controlled populations on group word problems. (1) The budget law: free training installs a linear computation frontier, a mechanism that solves v positions per loop, whose speed is priced by the training contract: v ~ n_train/T_train (exponent 0.98 +/- 0.04, R^2=0.99), exactly unity under T=n training. SGD selects a frontier matching the minimum the contract demands; granting more test-time loops than ever trained rescues late positions at fixed input length, yielding a principled halting rule T* = ceil(n / v-hat). (2) Architecture prior, not expressivity, picks the algorithm: standard-depth transformers learn parallel scans on this family; weight tying flips the selection to the serial frontier, even when positional addressing for a log-depth scan is supplied. At matched depth and parameters, untied models extrapolate worst and fail to learn A5 at all. (3) The walls are not where circuit complexity says: NC1-completeness costs nothing (A5 generalizes fully), while group order does (S5's 120x120 operator deadlocks joint learning) -- and an operator-first curriculum dissolves the wall in every seed. (4) Mechanisms are portable, not mandatable: warm-starting across budget contracts transfers the algorithm in every seed, re-pricing its speed, while imposing seriality through the input schedule fails where free training succeeds. These results are invisible to standard instruments, which provably saturate at the fixed points trained loops converge to. We introduce a head instrument, the convergence-time scaling tau(n,i), validate it causally via damage cones whose slope reproduces v, and show in-distribution head measurements predict out-of-distribution fate where tail metrics do not. Results replicate on the public easy-to-hard benchmark.
We introduce training-free looped transformers, in which a lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional fine-tuning, continued training, or architectural changes. Unlike prior looped transformer methods that train with the looped structure end-to-end, we retrofit recurrence onto pretrained models at test time. We show that naive block reapplication usually degrades performance, highlighting the importance of the loop application strategy. Motivated by viewing a pre-norm transformer block as a forward Euler step on an ODE, we instead treat looping as a refinement of the same approximation, replacing one large update with smaller damped sub-steps. Across seven dense, sparse MoE, and MLA+MoE model families, our method improves Qwen3-4B-Instruct by +2.64 pp on MMLU-Pro, Qwen3-30B-A3B-Instruct by +1.14 pp on CommonsenseQA, and Moonlight-16B-A3B-Instruct by +1.20 pp on OpenBookQA.