cs.LGOct 8, 2026

TraceRelay: Attention-Aligned Recurrence over Rolling Traces

Authors: Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung

Organizations: College of Pharmacy, Chungnam National University · Department of Computer Science & Engineering, Chungnam National University

Abstract

We present TraceRelay, an attention-aligned recurrent architecture that distributes persistent representations over a rolling sequence of low-dimensional traces. Local right looking attention forms increments from lower-layer representations; delivery is delayed until all attended inputs are in the causal past. A fixed additive phase recurrence accumulates the delayed increments, and left-looking attention reads the resulting residual augmented stream. A stride-wise prefix sum supports parallel prefill and bounded-buffer continuation. We study 36 small-model runs on Equal Repeats, bounded Dyck closing-type prediction, and causal Most-Freq generation, using three seeds per setting. At trained length 256, Equal Repeats models with recurrent phase inheritance reach 98.81-99.69% accuracy versus 50.73-51.63% for separately trained variants without inheritance, despite the latter receiving more updates. Accuracy drops sharply at lengths beyond the training range. At the longest evaluated lengths, models with more dimensions in the middle layer's recurrent traces perform better on Dyck (76.34% versus 55.49% close accuracy at length 4096), whereas models with fewer trace dimensions perform better on five-symbol Most-Freq (70.74% versus 55.60% exact generation at length 1024). These contrasting cases motivate further study of how the size of recurrent representations should be chosen for different tasks, without establishing a general rule across tasks or model configurations.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

    Sep 3, 2026Yuxiang Wang, Kunyu Feng, Yingda Shen +3Memory-Augmented Language ModelsRecurrent Transformers

  2. Temporal Recurrence Favors Fewer Layers

    Sep 14, 2026Ivan Anokhin, Johan Obando-Ceron, Irina Rish +1Recurrent Neural Networks

  3. Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

    Sep 23, 2026Ke Wan, Chen ChenLong-Context Language Model InferenceRecurrent Transformers