Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, which uses online labels indicating whether further iteration improves the prediction. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy-compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched test-time compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline continues to grow from +2.8 points at depth 2 to +3.9 points at depth 8. Our code is available at https://github.com/thu-nics/TaH.
Figures & tables
Figure 1: TaH2 improves both test-time and iteration-depth scaling. Mean AIME24–26 accuracy over 32 samples for 1.7B models post-trained from the same checkpoint. (a) Output-token cutoffs sweep 4K–16K; lines are fitted trends. (b) Accuracy at a 16K cutoff as the depth ceiling M grows.
Figure 2: TaH2’s architecture and training scheme. Left: the decider decides whether to continue after each iteration. Right: the updater reinjects token embeddings at each iteration, and stopping probabilities weight per-iteration predictions. All modules share parameters across iterations. The decider receives online depth supervision after each iteration; the no-gradient lookahead supplies the target at the stopping point and is omitted at inference.
Domain
Benchmark
Std.
Huginn
Ouro
TaH2-fixed
TaH2
M=1
M=3
M=7
M=2
M=4
M=2
M=4
M=2
M=4
M=8
math
AIME24
11.0
13.9
10.6
10.5
14.5
15.7
14.3
15.7
16.3
16.3
AIME25
13.0
14.2
14.0
13.9
13.4
15.0
15.4
16.5
16.9
17.9
AIME26
11.9
11.8
13.3
11.0
10.3
13.2
13.8
14.3
14.2
14.5
AMC23
45.3
45.8
49.0
45.9
45.9
51.1
52.2
50.7
50.9
52.8
MATH500
75.2
74.4
75.2
75.1
73.5
76.1
78.5
78.1
78.6
79.3
Table 1: Accuracy (%) of Qwen3-1.7B models across ten benchmarks. Olympiad: OlympiadBench; HE: HumanEval. Bold/underline indicate the best/second-best accuracy per row; subscripts show gains over Standard in points. FLOPs/token is total decoding FLOPs divided by total generated tokens across benchmarks, relative to Standard.
Figure 4: Depth and test-time scaling at 1.7B. (a) Each marker denotes a model trained with the labelled depth ceiling M ; lower is better. (b) Dark segments show output-token cutoffs within 16K; light segments extend beyond it to 32K.
Table 5Table 6
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Hyper-parameter
Value
global batch (samples)
128
sequence length
16384
learning rate
4×10−5
max gradient norm
1.0
epochs
3
warmup ratio
0.03
Appendix
Table 5: Training hyper-parameters shared by all variants at a given scale.
Figure 7: Extended duo-causal attention in TaH2. Each cell denotes a (token id, iteration depth) pair; green arrows and cells show the keys visible to token 3. (a, b) Executed depths at inference and training; dashed cells are no-gradient lookahead iterations used only for depth supervision. (c, d) The corresponding attention masks over concatenated per-iteration KV caches, where shaded cells are visible. Lookahead queries cannot attend to other tokens’ lookahead states; all other attention follows the duo-causal rule.
Model
Loop scope
Depth allocation
Added module
Added params
Standard
none ( M=1 )
–
–
0
TaH2
all L layers
per-token adaptive
updater + decider
46.2M (2.6%)
TaH2-fixed
all L layers
fixed
updater
21.0M (1.2%)
Ouro (fixed-depth)
all L layers
fixed
–
0
Ouro (adaptive)
all L layers
per-token adaptive
exit gate
2.0K ( < 0.001%)
Huginn
layers 8–21
fixed
input adapter
4.2M (0.24%)
Appendix
Table 6: Architectures used in the 1.7B comparison. Added parameters are reported as counts and fractions of total model parameters.
Scale
Backbone
Updater
Decider
Total added
Added (%)
1.7B
1720.57M
20.98M
25.18M
46.16M
2.61%
4B
4022.47M
32.78M
39.33M
72.11M
1.76%
8B
8190.74M
83.90M
100.68M
184.59M
2.20%
Appendix
Table 7: Updater and decider parameters across backbone scales. Counts include projections, MLPs and RMSNorm scales and are rounded independently. Percentages are relative to total model parameters.
Benchmark
Domain
#Problems
Metric
Subset
AIME24 / AIME25 / AIME26
math
30 / 30 / 30
avg@32
–
AMC23
math
40
avg@32
–
MATH500 ( Lightman et al., 2023 )
math
500
avg@4
–
OlympiadBench ( He et al., 2024 )
math
675
avg@4
–
IMO-AnswerBench
math
400
avg@4
–
GPQA ( Rein et al., 2023 )
QA
198
avg@8
Diamond
Appendix
Table 8: Benchmarks used in this paper. The default output-token limit is 32K.
Table 10: Training FLOPs ( 1018 ) of the 1.7B post-training runs, split into the terms of Equation 21 . Fwd. + bwd. reports 3FLOPsforward .
Teacher
AIME24
AIME25
AIME26
Mean
Qwen3-8B
11.04
13.02
11.87
11.98
Qwen3-32B
9.69
11.98
9.38
10.35
Qwen3-235B-A22B
10.42
9.90
10.21
10.18
Appendix
Table 11: AIME accuracy (%) for Qwen3-1.7B-Base students trained on responses from different teachers. All evaluations use avg@32 and a 32K output-token limit.
Model
Inference ceiling
AIME24
AIME25
AIME26
Avg.
Mean depth
Standard
1
11.0
13.0
11.9
12.0
1.00
TaH2 ( M=2 )
2
15.7
16.5
14.3
15.5
1.23
TaH2 ( M=4 )
4
16.3
16.9
14.2
15.8
1.49
TaH2 ( M=8 )
4
13.6
15.1
11.4
13.4
1.55
TaH2 ( M=8 )
8
16.3
17.9
14.5
16.2
2.22
TaH2 ( M=8 )
12
16.9
17.2
14.5
16.2
2.72
Appendix
Table 12: Effects of changing the inference depth ceiling. M denotes the training ceiling; the upper group provides reference models evaluated at their training ceilings. Accuracy (%) uses avg@32; Avg. is the mean across AIME24–26.
Figure 8: Validation loss during post-training at 1.7B. Insets show final loss differences from Standard; negative values indicate improvements.
Figure 9: Per-benchmark test-time scaling on the six math benchmarks at 1.7B. Each panel compares Standard with TaH2 at M∈{2,4,8} using the same nine output-token cutoffs as Figure 4(b) . Dark segments extend through 16K and light segments through 32K.
Figure 10: Parallel test-time scaling: mean AIME24–26 cons@ n versus decoding FLOPs per problem for n=1,…,32 , one model family per panel with Standard for reference.
Figure 11: Token-level iteration depth in sampled correct responses from OlympiadBench, HumanEval and GPQA. Darker shading indicates more iterations; mean depths refer to the displayed spans.
Figure 12: Attention mass by key iteration for three representative heads at query iteration 8. Bars show means over 100 sequences; error bars indicate one sample-level standard deviation. Queries are averaged within each sequence first.
Figure 13: Attention maps for the same heads, with columns indexing heads and rows indexing key iterations. Each query position is averaged over sequences that reach iteration 8. All panels share a logarithmic colour scale; grey marks positions with no executed queries. Prompt keys are omitted without renormalising attention.