We present TraceRelay, an attention-aligned recurrent architecture that distributes persistent representations over a rolling sequence of low-dimensional traces. Local right looking attention forms increments from lower-layer representations; delivery is delayed until all attended inputs are in the causal past. A fixed additive phase recurrence accumulates the delayed increments, and left-looking attention reads the resulting residual augmented stream. A stride-wise prefix sum supports parallel prefill and bounded-buffer continuation. We study 36 small-model runs on Equal Repeats, bounded Dyck closing-type prediction, and causal Most-Freq generation, using three seeds per setting. At trained length 256, Equal Repeats models with recurrent phase inheritance reach 98.81-99.69% accuracy versus 50.73-51.63% for separately trained variants without inheritance, despite the latter receiving more updates. Accuracy drops sharply at lengths beyond the training range. At the longest evaluated lengths, models with more dimensions in the middle layer's recurrent traces perform better on Dyck (76.34% versus 55.49% close accuracy at length 4096), whereas models with fewer trace dimensions perform better on five-symbol Most-Freq (70.74% versus 55.60% exact generation at length 1024). These contrasting cases motivate further study of how the size of recurrent representations should be chosen for different tasks, without establishing a general rule across tasks or model configurations.
Figures & tables
Figure 1: Reference computation graph. Right attention observes lower-layer positions i through i+R ; its increment is delivered at i+W , with W>R . The stride-wise phase scan and cos/sin readout precede the residual merge and Left attention. The displayed readout applies to initialized lineages; until a valid delayed update arrives, the readout is zero. Left Q/K/V consume the merged stream. The orange branches pass the merged representation mt to the attention residual and the post-attention state ht to the FFN residual. Right does not read this layer’s phase. Normalization is implicit, as in the equations.
Figure 2: Equal Repeats: with and without phase inheritance. Accuracy on 10,000 common examples per length, mean ± sample SD over three training seeds. Shading marks lengths beyond the 32–256 training range. Models without inheritance are separately trained with the same task and geometry; all use 20,000 updates, while models with inheritance stop earlier. The strong separation within the trained range largely disappears in OOD evaluation. The dotted line is the 1/3 chance reference.
Figure 3: Bounded Dyck closing-type prediction. Three-seed mean ± sample SD for k=8,m=10 , using 2,000 common sequences per length. (a) Close accuracy versus length; shaded lengths exceed the training maximum 256. (b,c) Matching-distance and pre-close-depth slices at T=4096 , with 4,096,000 close targets in total. The starred distance bin 1025–2048 has only 43 targets; the empty >2048 bin is omitted. The 513–1024 bin has 1,639 targets. Dotted horizontal lines in (b,c) mark 12.5% chance. Full slice counts are in Table 3 .
Figure 4: Most-Freq causal generation. Five symbols, separated counts with minimum gap four, training lengths 128–256, and 1,024 common evaluation examples per length. Curves show mean ± sample SD over three training seeds; shading denotes length OOD. (a) Exact greedy ranking including predicted EOS, alongside same-input shortcut baselines. (b) Top-1 accuracy. Suffix frequency (32) ranks symbols present in the last 32 input tokens by descending frequency there, breaking ties by first occurrence within that suffix. First appearance lists symbols seen in the full input in order of first occurrence, ignoring their counts. Fixed symbol order always lists 0,1,2,3,4 . All three heuristics append EOS and use the same evaluation inputs. No gold output is fed back during model generation. Narrower centers have higher OOD means in this sweep, with unequal early-stopped training exposure.
Task
Phase inheritance
dc
ID: 256
OOD: 1024
Equal Repeats
On
64
98.81±0.58
33.76±1.82
Equal Repeats
On
32
99.24±0.56
35.63±2.99
Equal Repeats
On
16
99.69±0.13
32.84±1.66
Equal Repeats
Off
64
51.03±0.94
37.48±0.14
Equal Repeats
Off
32
51.63±1.24
37.65±0.23
Equal Repeats
Off
16
50.73±0.52
37.47±0.09
Table 1: Best-checkpoint test results, mean ± sample SD over seeds 42–44 (percent). ID is T=256 and OOD is T=1024 for every task. Metrics are sequence classification (Equal Repeats), token-weighted close type (Dyck), and exact generated ranking including EOS (Most-Freq).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Equal Repeats
Dyck
Most-Freq
Phase inheritance
On and off; 18 runs
On; 9 runs
On; 9 runs
Task parameters
Binary 0a1b0c ; three relation classes
k=8 bracket types; maximum depth m=10
V=5 ; decreasing count, first-appearance ties
Sampler
Inclusive repeated length r∈[1,⌊T/2⌋] ; equal-all maps to class 0
Completion-count DP over bounded paths; uniform opening types
Separated positive counts, gap ≥4 ; randomized symbol assignment and order
Online source length
Uniform integer T∈[32,256]
Uniform even T∈[32,256]
Uniform integer T∈[128,256]
ID validation lengths
32, 64, 128, 256
32, 64, 128, 256
128, 192, 256
Final ID lengths
32, 64, 128, 256
32, 64, 128, 256
128, 192, 256
Appendix
Table 2: Executed protocols. All tasks sweep dc∈{64,32,16} with three training seeds per variant and width. ID validation occurs every 250 updates; two consecutive mean ID scores ≥99% trigger stopping. OOD results never control training. Table 5 gives individual stopping/selected steps and token exposure.
Slice
Closes
dc=64
dc=32
dc=16
1–8
3,194,038
87.27±2.62
72.57±10.69
62.65±14.31
9–16
360,329
48.41±6.04
38.07±1.98
34.52±2.93
17–32
257,308
30.90±3.45
27.24±1.32
27.23±3.41
33–64
159,952
28.85±4.23
26.04±2.23
26.83±3.21
65–128
80,458
31.20±4.76
26.90±2.19
27.89±3.08
129–256
32,358
32.93±5.24
26.64±2.36
27.94±3.24
Appendix
Table 3: Dyck diagnostics at T=4096 : percent close accuracy, mean ± sample SD. Counts are targets in the common 2,000-sequence evaluation set, not independent replications across models. An empty bin is unavailable, not zero accuracy.
Metric
dc=64
dc=32
dc=16
Full exact
55.60±12.82
65.40±2.39
70.74±4.95
Top-1
86.82±2.05
89.10±3.00
90.76±7.91
Top-2 prefix
74.41±7.43
78.71±2.03
84.05±6.15
Set exact
81.84±7.26
87.96±3.97
89.55±4.67
Termination
82.91±7.32
96.22±5.36
95.64±5.61
Appendix
Table 4: Most-Freq diagnostics at T=1024 (percent; three-seed mean ± SD). Top-2 requires the first two outputs in the correct order. Set exact requires each source symbol exactly once but ignores rank; termination records a predicted EOS within the decode limit.
Figure 5: Optimization histories. Every logged ID validation mean for all 36 runs; colors indicate width and line styles indicate seed. Stars mark the checkpoint selected by ID accuracy with a cross-entropy tie-break. The dashed horizontal line is 99%; stopping requires two consecutive passes. Curves end at their actual stopping updates. Nine Equal Repeats runs without phase inheritance and one Dyck run exhaust the budget. Metrics and horizontal scales are task-specific, and the curves are not smoothed. Table 5 gives steps, tokens, and wall time.
Task / inheritance
dc
Seed
Best
Final
Tokens (M)
Min.
Params
ER / On
64
42
2,500
4,250
39.10
6.7
295,875
ER / On
64
43
5,750
7,000
65.10
10.7
295,875
ER / On
64
44
7,750
7,750
71.52
12.3
295,875
ER / On
32
42
6,000
9,250
85.51
14.6
268,227
ER / On
32
43
8,250
8,250
76.43
13.0
268,227
ER / On
32
44
10,250
10,250
94.80
15.9
268,227
Appendix
Table 5: Actual training exposure for all 36 runs. Best is the selected checkpoint update; final is the stopping update. Tokens are millions of physical training tokens; Most-Freq includes separator and teacher-forced output tokens. Minutes are logged training-loop wall time, including periodic ID validation but excluding final evaluation. On/off indicates phase inheritance; † denotes budget exhaustion without satisfying the two-pass threshold.
Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two design choices limit these gains. Each iteration sees only the previous output and cannot directly access earlier computations. Moreover, a fixed loop count wastes depth on easy inputs while leaving hard ones with too little computation. We introduce RecurTrace, which addresses both limitations using the loop's own trajectory. Specifically, Loop Memory Attention lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone. A halting head then reads the loop state and predicts whether to continue, with supervision from an oracle that identifies when additional depth still reduces loss. In a controlled MathQA comparison on the same looped backbone, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, exceeding the best fixed loop depth by 2.2 points at matched compute. By comparison, ACT and PonderNet collapse to one loop, and CALM reaches only 54.1% with 5.6 loops, while the stronger LoopUS-Conf and TaH-Mismatch baselines reach 55.3% at 3.2 loops and 55.7% at 2.1 loops. Finally, RecurTrace improves generation accuracy over same-budget fine-tuned baselines at 0.6B, 1.7B, 4B, and 8B, with the gain growing with model size from 0.6 to 3.4 points.
Yuxiang Wang, Kunyu Feng, Yingda Shen +3
The Chinese University of Hong Kong, Shenzhen · The Chinese University of Hong Kong · Tianjin University +2
In streaming tasks, recurrent models can carry latent computation across time, allowing each update to build on representations produced earlier. This raises a basic question: once temporal recurrence provides sequential computation across steps, how much depth is still needed within each step? Prior work has shown that recurrence can make shallow models competitive. We instead study this question as a compute-allocation problem, varying within-step depth, expert width, and the number of parallel experts per layer across several compute budgets. For each budget, we compare the best observed recurrent and non-recurrent allocations and the performance they achieve under approximately matched per-step computation. Across Sokoban and autoregressive FineWeb language modeling, we find that temporal recurrence shifts the best observed compute allocation toward substantially fewer layers, with comparable or better performance.
Ivan Anokhin, Johan Obando-Ceron, Irina Rish +1
Mila – Qu´ebec AI Institute · Universit´e de Montr´eal · Sakana AI
Recurrent-depth language models, such as looped Transformers, repeatedly apply shared network blocks to refine latent representations without generating explicit intermediate reasoning tokens. However, each step recomputes full attention over the entire context, repeating costly global routing. We study how attention routing evolves across recurrent depth and find a consistent separation in convergence timescales: attention support and distributions stabilize substantially earlier than hidden states and attention outputs. This suggests two stages of recurrent inference: early discovery of a sparse working set, followed by representation refinement over largely stable routing support. Motivated by this finding, we introduce WISE (Working-set Inference with Support Exploitation), a training-free method that uses unrestricted attention during early recurrent steps to discover a block-structured working set, then reuses its support in later steps while keeping attention weights and recurrent refinement dynamic. Controlled interventions show that multi-step discovery yields more effective working sets than first-step selection, and that support reuse better preserves model behavior than more restrictive forms of attention reuse. Across multi-hop QA benchmarks, WISE largely preserves full-attention performance. Matched context-scaling experiments reveal an increasingly favorable quality-efficiency tradeoff as routing support becomes sparser with longer contexts. A sparse-attention implementation achieves up to a 1.76x late-step attention speedup over native FlashAttention at 4K context. Code: https://github.com/tbn5pj/WISE_code.
Ke Wan, Chen Chen
Department of Computer Science University of Virginia Charlottesville, VA, USA