Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.
Figures & tables
Figure 1: TaH2 improves both test-time and iteration-depth scaling. Mean AIME24–26 accuracy over 32 samples for 1.7B models post-trained from the same checkpoint. (a) Output-token cutoffs sweep 4K–16K; lines are fitted trends. (b) Accuracy at a 16K cutoff as the depth ceiling M grows.
Figure 2: TaH2’s architecture and training scheme. Left: the decider decides whether to continue after each iteration. Right: the updater reinjects token embeddings at each iteration, and stopping probabilities weight per-iteration predictions. All modules share parameters across iterations. The decider receives online depth supervision after each iteration; the no-gradient lookahead supplies the target at the stopping point and is omitted at inference.
Domain
Benchmark
Std.
Huginn
Ouro
TaH2-fixed
TaH2
M=1
M=3
M=7
M=2
M=4
M=2
M=4
M=2
M=4
M=8
math
AIME24
11.0
13.9
10.6
10.5
14.5
15.7
14.3
15.7
16.3
16.3
AIME25
13.0
14.2
14.0
13.9
13.4
15.0
15.4
16.5
16.9
17.9
AIME26
11.9
11.8
13.3
11.0
10.3
13.2
13.8
14.3
14.2
14.5
AMC23
45.3
45.8
49.0
45.9
45.9
51.1
52.2
50.7
50.9
52.8
MATH500
75.2
74.4
75.2
75.1
73.5
76.1
78.5
78.1
78.6
79.3
Table 1: Accuracy (%) of Qwen3-1.7B models across ten benchmarks. Olympiad: OlympiadBench; HE: HumanEval. Bold/underline indicate the best/second-best accuracy per row; subscripts show gains over Standard in points. FLOPs/token is total decoding FLOPs divided by total generated tokens across benchmarks, relative to Standard.
Figure 4: Depth and test-time scaling at 1.7B. (a) Each marker denotes a model trained with the labelled depth ceiling M ; lower is better. (b) Dark segments show output-token cutoffs within 16K; light segments extend beyond it to 32K.
Table 5Table 6
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Hyper-parameter
Value
global batch (samples)
128
sequence length
16384
learning rate
4×10−5
max gradient norm
1.0
epochs
3
warmup ratio
0.03
Appendix
Table 5: Training hyper-parameters shared by all variants at a given scale.
Figure 7: Extended duo-causal attention in TaH2. Each cell denotes a (token id, iteration depth) pair; green arrows and cells show the keys visible to token 3. (a, b) Executed depths at inference and training; dashed cells are no-gradient lookahead iterations used only for depth supervision. (c, d) The corresponding attention masks over concatenated per-iteration KV caches, where shaded cells are visible. Lookahead queries cannot attend to other tokens’ lookahead states; all other attention follows the duo-causal rule.
Model
Loop scope
Depth allocation
Added module
Added params
Standard
none ( M=1 )
–
–
0
TaH2
all L layers
per-token adaptive
updater + decider
46.2M (2.6%)
TaH2-fixed
all L layers
fixed
updater
21.0M (1.2%)
Ouro (fixed-depth)
all L layers
fixed
–
0
Ouro (adaptive)
all L layers
per-token adaptive
exit gate
2.0K ( < 0.001%)
Huginn
layers 8–21
fixed
input adapter
4.2M (0.24%)
Appendix
Table 6: Architectures used in the 1.7B comparison. Added parameters are reported as counts and fractions of total model parameters.
Scale
Backbone
Updater
Decider
Total added
Added (%)
1.7B
1720.57M
20.98M
25.18M
46.16M
2.61%
4B
4022.47M
32.78M
39.33M
72.11M
1.76%
8B
8190.74M
83.90M
100.68M
184.59M
2.20%
Appendix
Table 7: Updater and decider parameters across backbone scales. Counts include projections, MLPs and RMSNorm scales and are rounded independently. Percentages are relative to total model parameters.
Benchmark
Domain
#Problems
Metric
Subset
AIME24 / AIME25 / AIME26
math
30 / 30 / 30
avg@32
–
AMC23
math
40
avg@32
–
MATH500 ( Lightman et al., 2023 )
math
500
avg@4
–
OlympiadBench ( He et al., 2024 )
math
675
avg@4
–
IMO-AnswerBench
math
400
avg@4
–
GPQA ( Rein et al., 2023 )
QA
198
avg@8
Diamond
Appendix
Table 8: Benchmarks used in this paper. The default output-token limit is 32K.
Table 10: Training FLOPs ( 1018 ) of the 1.7B post-training runs, split into the terms of Equation 21 . Fwd. + bwd. reports 3FLOPsforward .
Teacher
AIME24
AIME25
AIME26
Mean
Qwen3-8B
11.04
13.02
11.87
11.98
Qwen3-32B
9.69
11.98
9.38
10.35
Qwen3-235B-A22B
10.42
9.90
10.21
10.18
Appendix
Table 11: AIME accuracy (%) for Qwen3-1.7B-Base students trained on responses from different teachers. All evaluations use avg@32 and a 32K output-token limit.
Model
Inference ceiling
AIME24
AIME25
AIME26
Avg.
Mean depth
Standard
1
11.0
13.0
11.9
12.0
1.00
TaH2 ( M=2 )
2
15.7
16.5
14.3
15.5
1.23
TaH2 ( M=4 )
4
16.3
16.9
14.2
15.8
1.49
TaH2 ( M=8 )
4
13.6
15.1
11.4
13.4
1.55
TaH2 ( M=8 )
8
16.3
17.9
14.5
16.2
2.22
TaH2 ( M=8 )
12
16.9
17.2
14.5
16.2
2.72
Appendix
Table 12: Effects of changing the inference depth ceiling. M denotes the training ceiling; the upper group provides reference models evaluated at their training ceilings. Accuracy (%) uses avg@32; Avg. is the mean across AIME24–26.
Figure 8: Validation loss during post-training at 1.7B. Insets show final loss differences from Standard; negative values indicate improvements.
Figure 9: Per-benchmark test-time scaling on the six math benchmarks at 1.7B. Each panel compares Standard with TaH2 at M∈{2,4,8} using the same nine output-token cutoffs as Figure 4(b) . Dark segments extend through 16K and light segments through 32K.
Figure 10: Parallel test-time scaling: mean AIME24–26 cons@ n versus decoding FLOPs per problem for n=1,…,32 , one model family per panel with Standard for reference.
Figure 11: Token-level iteration depth in sampled correct responses from OlympiadBench, HumanEval and GPQA. Darker shading indicates more iterations; mean depths refer to the displayed spans.
Figure 12: Attention mass by key iteration for three representative heads at query iteration 8. Bars show means over 100 sequences; error bars indicate one sample-level standard deviation. Queries are averaged within each sequence first.
Figure 13: Attention maps for the same heads, with columns indexing heads and rows indexing key iterations. Each query position is averaged over sequences that reach iteration 8. All panels share a logarithmic colour scale; grey marks positions with no executed queries. Prompt keys are omitted without renormalising attention.
Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same Transformer blocks, trading additional computation for improved performance without increasing parameter count or context length. Because the number of loop iterations can be adjusted at inference, it also provides a natural mechanism for balancing performance and test-time compute. However, Looped Transformer still suffers from training instability when the number of loop iterations increases. Our analysis reveals that this instability stems from two sources: gradient oscillation and residual explosion. To address these two problems, we propose the Fully Looped Transformer, which introduces two parameter-free modifications: (1) Fully Looped Architecture, which distributes inter-loop signals across all layers to mitigate residual explosion; (2) Attention Injection, which reuses the existing attention block to suppress gradient oscillation. These modifications stabilize training dynamics, enabling the Fully Looped Transformer to be trained stably up to 12 loop iterations, whereas other baseline looped models collapse in this regime. In milder settings where Looped Transformer does not collapse, Fully Looped Transformer still improves average downstream-task performance by up to 13.2%. Overall, our experiments demonstrate that Fully Looped Transformer improves training stability, enhances downstream performance, and provides preliminary adaptability under different test-time compute budgets by varying loop iterations at inference.
Looped Transformers (LT) have emerged as a powerful architecture by iterating their layers multiple times before decoding the final token. However, pairing them with full attention retains quadratic complexity, making them computationally expensive and slow. We introduce LT2 (Linear-Time Looped Transformers), a family of looped architectures that replace quadratic softmax attention with subquadratic, linear-time attention. We study two variants: LT2-linear with linear attention and LT2-sparse with sparse attention. We find that looping uniquely synergizes with these variants: it enables iterative memory refinement in linear attention and progressively expands the effective receptive field in sparse attention. We formalize these benefits theoretically and demonstrate consistent empirical gains across controlled recall, state-tracking, and language modeling tasks. We then explore LT2-hybrid, which combines different attention variants in a looped setting. Two variants are especially promising: LT2-hybrid (GDN+DSA), which interleaves linear and sparse attention to maximize efficiency and matches the standard looped transformer's quality at fully linear-time cost; and LT2-hybrid (Full+GDN), which interleaves GDN with a small fraction of full attention layers to maximize quality, surpassing the standard looped transformer in both performance and efficiency. We also show how to convert a pre-trained LT into an LT2-hybrid model. With about 1B tokens of training, our converted model, Ouro-hybrid-1.4B, outperforms industry-level 1B models and is competitive with industry-level 4B models while retaining the speed benefits of linear-time attention. Together, these results show a clear path toward making looped transformers more scalable and advancing efficient, capable small language models.
Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64× end-to-end speedup and up to 6× KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
Wanqi Yang, Shiwei Liu
ELLIS Institute Tübingen Max Planck Institute for Intelligent Systems Tübingen AI Center