Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64× end-to-end speedup and up to 6× KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
Figures & tables
Figure 1: Comparison of performance-matched looped and non-looped models at an 8K context length 2 2 2 Accuracy results for these models are reported in Appendix A . . FlashLoop substantially reduces both computation and KV-cache memory for looped Transformers, achieving 5.8 × and 1.8 × reductions in KV cache memory and 8K prefill FLOPs.
Figure 2: Recurrent refinement becomes structurally redundant along three complementary axes. (a) Token-level cross-loop redundancy : hidden-state change becomes increasingly concentrated at a subset of tokens. (b) Attention-column redundancy : later-loop attention outputs can be reconstructed from a small subset of key columns identified by the preceding loop. (c) Cross-loop representation redundancy : at the same quantization setting, recursively quantizing adjacent-loop KV residuals gives lower reconstruction error than independently quantizing complete KV states.
Figure 3: Additional evidence for cross-loop redundancy across different looped models. (a) Different looped transformers have similar cross-loop redundancy: decode attention concentrates on a small set of attention columns, (b) whose importance becomes increasingly stable across loops, (c–d) while adjacent-loop K/V residuals decrease with loop depth motivating a residual-based quantization for cross-loop KV states.
Model
Method
MATH-500
ARC-C
GSM8K
HellaSwag
WinoGrande
Avg.
Δ Avg.
Speedup ↑
KV Memory ↑
Ouro-1.4B
Original
65.80
59.98
78.77
74.24
71.74
70.11
–
1.00×
1.00×
FlashLoop
65.40
60.67
79.61
74.33
72.38
70.48
+0.37
1.59×
5.85×
Ouro-1.4B-Thinking
Original
46.80
62.03
81.65
72.61
72.45
67.11
–
1.00×
1.00×
FlashLoop
46.20
61.69
80.39
72.38
70.51
66.23
−0.88
1.59×
5.85×
Ouro-2.6B
Original
52.20
66.13
81.80
79.54
76.40
71.21
–
1.00×
1.00×
FlashLoop
52.80
65.10
82.56
79.58
75.37
71.08
−0.13
1.64×
6.06×
Table 1: Task performance and efficiency of Ouro and Huginn with FlashLoop. We report accuracy (%) on five benchmarks. Δ Avg. denotes the average-score difference from the original model in percentage points. KV Memory means the reduction of peak KV cache memory relative to original models.
Method
MATH-500
ARC-C
GSM8K
Avg.
Original
65.80
59.98
78.77
68.18
Loop-aware Sparse Attention
69.00
60.24
77.48
68.91
+ Cross-Loop Token-Sparse Updates
66.60
60.07
79.23
68.63
FlashLoop w/ Per-loop KIVI4
64.40
59.87
77.10
67.12
FlashLoop
65.40
60.67
79.61
68.56
Table 2: Component ablation of FlashLoop on Ouro-1.4B. FlashLoop w/ Per-loop KIVI4 means FlashLoop with full KV cache quantization along the same quantization axis. All values are accuracy (%).
Method
MATH-500
ARC-C
GSM8K
Avg.
Original
65.80
59.98
78.77
68.18
Loop-aware Sparse Attention
69.00
60.24
77.48
68.91
+ Cross-Loop Token-Sparse Updates
66.60
60.07
79.23
68.63
FlashLoop w/ Per-loop KIVI4
64.40
59.87
77.10
67.12
FlashLoop
65.40
60.67
79.61
68.56
Table 2: Component ablation of FlashLoop on Ouro-1.4B. FlashLoop w/ Per-loop KIVI4 means FlashLoop with full KV cache quantization along the same quantization axis. All values are accuracy (%).
Model
Original
H 2 O
Last-step
FlashLoop
( −75% KV)
( −75% KV)
Ouro-1.4B
70.11
62.81
68.33
70.48
Ouro-1.4B-T
67.11
62.68
65.09
66.23
Ouro-2.6B
71.21
67.35
69.96
71.08
Ouro-2.6B-T
72.60
68.01
70.78
72.24
Huginn-3.5B
42.60
40.68
32.18
42.77
Table 3: Comparison with alternative KV-cache strategies. We report average accuracy (%) across five benchmarks. "T" represents thinking version.
Figure 4: Efficiency analysis of each component. We cumulatively enable cross-loop token-sparse updates, loop-aware sparse attention, and 4-bit KV residual quantization, reporting KV-cache memory, prefill latency, and decode latency.
Figure 5: Accuracy–efficiency trade-offs of FlashLoop hyperparameters on Ouro-1.4B. We vary the retained attention-key ratio (left), Loop-3/4 token-update ratios (middle), and KV residual quantization precision (right), while keeping the remaining components fixed.
Figure 6: Scaling with context length. FlashLoop delivers increasingly substantial memory savings and speedups as the context length grows.
Figure 7: Scaling with loops. As the number of loops increases, FlashLoop progressively reduces the accumulated prefill FLOPs and KV-cache memory of both Ouro-1.4B and Huginn-3.5B.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
ARC-C
HellaSwag
WinoGrande
GSM8K
MATH-500
Ouro-2.6B R4
66.13
79.54
76.40
81.80
52.20
Qwen3-8B
66.10
79.60
76.80
83.09
62.30
Gemma3-12B
72.44
83.68
77.74
77.18
63.20
Llama-3.1-8B
60.75
81.97
77.11
78.17
52.90
Appendix
Table 4: Benchmark results of open-sourced performance-matched looped and non-looped models. All models are base models. The best score in each column is bolded , and the second-best is underlined .
Figure 8: Cross-loop redundancy in Huginn-3.5B. We record the same metrics of Huginn-3.5B as in Figure 3 , which show similar redundancy as in Ouro family.
Figure 9: Prefill attention maps for a random sampled MATH-500 prompt. There are obvious important key-column patterns across different layers and loops, while the important columns are stable across loops.
Model family
Loop(s)
Tokens
Columns
Mode
Ouro-1.4B
1–2
100%
100%
Dense warm-up
3
25%
10%
Sparse
4
10%
10%
Sparse
Ouro-1.4B-Thinking
1–2
100%
100%
Dense warm-up
3
25%
12%
Sparse
4
10%
12%
Sparse
Appendix
Table 5: Loop-wise sparsity configurations used in the main experiments. Token retention is the fraction of prompt-token rows recomputed during prefill. Column retention is the fraction of attention columns recomputed during decode.
Setting
WikiText-2 PPL ↓
Needle-in-a-Haystack (%) ↑
8K
16K
32K
Single-needle
Multi-needles
Ouro-1.4B
10.526
4.231
5.321
83.0
9.2
Ouro-1.4B + FlashLoop
10.513
4.266
5.589
84.5
7.5
Appendix
Table 6: Long-context quality evaluation. WikiText-2 perplexity is measured on 128 continuation tokens (lower is better). Needle-in-a-Haystack accuracy is measured over full single-needle and multi-needles samples.
We introduce training-free looped transformers, in which a lightweight inference-time wrapper loops a contiguous mid-stack block of layers of a frozen checkpoint without additional fine-tuning, continued training, or architectural changes. Unlike prior looped transformer methods that train with the looped structure end-to-end, we retrofit recurrence onto pretrained models at test time. We show that naive block reapplication usually degrades performance, highlighting the importance of the loop application strategy. Motivated by viewing a pre-norm transformer block as a forward Euler step on an ODE, we instead treat looping as a refinement of the same approximation, replacing one large update with smaller damped sub-steps. Across seven dense, sparse MoE, and MLA+MoE model families, our method improves Qwen3-4B-Instruct by +2.64 pp on MMLU-Pro, Qwen3-30B-A3B-Instruct by +1.14 pp on CommonsenseQA, and Moonlight-16B-A3B-Instruct by +1.20 pp on OpenBookQA.
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.
Looped Transformers have recently demonstrated strong performance in both reasoning and language tasks by reusing a shared set of parameters across multiple iterations, achieving parameter efficiency without sacrificing representational power. Besides, looped Transformers perform inference directly in the latent space (latent reasoning) to reduce the number of tokens consumed during inference, thereby achieving improved sample efficiency. However, these models typically apply a fixed recursion depth uniformly to every token, leading to suboptimal compute allocation and leaving significant efficiency gains on the table. In this work, we propose \textbf{dynamic token-choice routing} for looped transformers, enabling each token to adaptively determine its own number of loop iterations based on its hidden state. We use a dynamic router to decide whether a token should continue recursing or exit early, allowing simple tokens to bypass unnecessary computation while hard tokens receive deeper processing. To ensure that this adaptive mechanism does not compromise decoding efficiency, we further introduce recursion-wise KV caching, which maintains an independent key-value cache for each recursion loop. This design ensures that tokens at different depths only attend to their corresponding cached states, effectively eliminating redundant computations for exited tokens and enabling fast autoregressive decoding. Extensive experiments show that T-LoopFormer reaches the sota performance under the same parameters on PPL and 10 zero-shot reasoning tasks, even surpassing the base model at 24x FLOPs and our model could reach the lowest inference latency, which validate the effectiveness of token-choice router and recursion-wise KV cache. Code: https://github.com/YuMingQian1234/T-LoopFormer.
Mingqian Yu, Wenpeng Zhang, Peilin Zhao
Institute of Automation Chinese Academy of Sciences Beijing, China · Independent Reasearcher · School of Artificial Intelligence Shanghai Jiao Tong University Shanghai, China