Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64× end-to-end speedup and up to 6× KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
Figures & tables
Figure 1: Comparison of performance-matched looped and non-looped models at an 8K context length 2 2 2 Accuracy results for these models are reported in Appendix A . . FlashLoop substantially reduces both computation and KV-cache memory for looped Transformers, achieving 5.8 × and 1.8 × reductions in KV cache memory and 8K prefill FLOPs.
Figure 2: Recurrent refinement becomes structurally redundant along three complementary axes. (a) Token-level cross-loop redundancy : hidden-state change becomes increasingly concentrated at a subset of tokens. (b) Attention-column redundancy : later-loop attention outputs can be reconstructed from a small subset of key columns identified by the preceding loop. (c) Cross-loop representation redundancy : at the same quantization setting, recursively quantizing adjacent-loop KV residuals gives lower reconstruction error than independently quantizing complete KV states.
Figure 3: Additional evidence for cross-loop redundancy across different looped models. (a) Different looped transformers have similar cross-loop redundancy: decode attention concentrates on a small set of attention columns, (b) whose importance becomes increasingly stable across loops, (c–d) while adjacent-loop K/V residuals decrease with loop depth motivating a residual-based quantization for cross-loop KV states.
Model
Method
MATH-500
ARC-C
GSM8K
HellaSwag
WinoGrande
Avg.
Δ Avg.
Speedup ↑
KV Memory ↑
Ouro-1.4B
Original
65.80
59.98
78.77
74.24
71.74
70.11
–
1.00×
1.00×
FlashLoop
65.40
60.67
79.61
74.33
72.38
70.48
+0.37
1.59×
5.85×
Ouro-1.4B-Thinking
Original
46.80
62.03
81.65
72.61
72.45
67.11
–
1.00×
1.00×
FlashLoop
46.20
61.69
80.39
72.38
70.51
66.23
−0.88
1.59×
5.85×
Ouro-2.6B
Original
52.20
66.13
81.80
79.54
76.40
71.21
–
1.00×
1.00×
FlashLoop
52.80
65.10
82.56
79.58
75.37
71.08
−0.13
1.64×
6.06×
Table 1: Task performance and efficiency of Ouro and Huginn with FlashLoop. We report accuracy (%) on five benchmarks. Δ Avg. denotes the average-score difference from the original model in percentage points. KV Memory means the reduction of peak KV cache memory relative to original models.
Method
MATH-500
ARC-C
GSM8K
Avg.
Original
65.80
59.98
78.77
68.18
Loop-aware Sparse Attention
69.00
60.24
77.48
68.91
+ Cross-Loop Token-Sparse Updates
66.60
60.07
79.23
68.63
FlashLoop w/ Per-loop KIVI4
64.40
59.87
77.10
67.12
FlashLoop
65.40
60.67
79.61
68.56
Table 2: Component ablation of FlashLoop on Ouro-1.4B. FlashLoop w/ Per-loop KIVI4 means FlashLoop with full KV cache quantization along the same quantization axis. All values are accuracy (%).
Method
MATH-500
ARC-C
GSM8K
Avg.
Original
65.80
59.98
78.77
68.18
Loop-aware Sparse Attention
69.00
60.24
77.48
68.91
+ Cross-Loop Token-Sparse Updates
66.60
60.07
79.23
68.63
FlashLoop w/ Per-loop KIVI4
64.40
59.87
77.10
67.12
FlashLoop
65.40
60.67
79.61
68.56
Table 2: Component ablation of FlashLoop on Ouro-1.4B. FlashLoop w/ Per-loop KIVI4 means FlashLoop with full KV cache quantization along the same quantization axis. All values are accuracy (%).
Model
Original
H 2 O
Last-step
FlashLoop
( −75% KV)
( −75% KV)
Ouro-1.4B
70.11
62.81
68.33
70.48
Ouro-1.4B-T
67.11
62.68
65.09
66.23
Ouro-2.6B
71.21
67.35
69.96
71.08
Ouro-2.6B-T
72.60
68.01
70.78
72.24
Huginn-3.5B
42.60
40.68
32.18
42.77
Table 3: Comparison with alternative KV-cache strategies. We report average accuracy (%) across five benchmarks. "T" represents thinking version.
Figure 4: Efficiency analysis of each component. We cumulatively enable cross-loop token-sparse updates, loop-aware sparse attention, and 4-bit KV residual quantization, reporting KV-cache memory, prefill latency, and decode latency.
Figure 5: Accuracy–efficiency trade-offs of FlashLoop hyperparameters on Ouro-1.4B. We vary the retained attention-key ratio (left), Loop-3/4 token-update ratios (middle), and KV residual quantization precision (right), while keeping the remaining components fixed.
Figure 6: Scaling with context length. FlashLoop delivers increasingly substantial memory savings and speedups as the context length grows.
Figure 7: Scaling with loops. As the number of loops increases, FlashLoop progressively reduces the accumulated prefill FLOPs and KV-cache memory of both Ouro-1.4B and Huginn-3.5B.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
ARC-C
HellaSwag
WinoGrande
GSM8K
MATH-500
Ouro-2.6B R4
66.13
79.54
76.40
81.80
52.20
Qwen3-8B
66.10
79.60
76.80
83.09
62.30
Gemma3-12B
72.44
83.68
77.74
77.18
63.20
Llama-3.1-8B
60.75
81.97
77.11
78.17
52.90
Appendix
Table 4: Benchmark results of open-sourced performance-matched looped and non-looped models. All models are base models. The best score in each column is bolded , and the second-best is underlined .
Figure 8: Cross-loop redundancy in Huginn-3.5B. We record the same metrics of Huginn-3.5B as in Figure 3 , which show similar redundancy as in Ouro family.
Figure 9: Prefill attention maps for a random sampled MATH-500 prompt. There are obvious important key-column patterns across different layers and loops, while the important columns are stable across loops.
Model family
Loop(s)
Tokens
Columns
Mode
Ouro-1.4B
1–2
100%
100%
Dense warm-up
3
25%
10%
Sparse
4
10%
10%
Sparse
Ouro-1.4B-Thinking
1–2
100%
100%
Dense warm-up
3
25%
12%
Sparse
4
10%
12%
Sparse
Appendix
Table 5: Loop-wise sparsity configurations used in the main experiments. Token retention is the fraction of prompt-token rows recomputed during prefill. Column retention is the fraction of attention columns recomputed during decode.
Setting
WikiText-2 PPL ↓
Needle-in-a-Haystack (%) ↑
8K
16K
32K
Single-needle
Multi-needles
Ouro-1.4B
10.526
4.231
5.321
83.0
9.2
Ouro-1.4B + FlashLoop
10.513
4.266
5.589
84.5
7.5
Appendix
Table 6: Long-context quality evaluation. WikiText-2 perplexity is measured on 128 continuation tokens (lower is better). Needle-in-a-Haystack accuracy is measured over full single-needle and multi-needles samples.
Institute of Automation Chinese Academy of Sciences Beijing, China · Independent Reasearcher · School of Artificial Intelligence Shanghai Jiao Tong University Shanghai, China