cs.CLOct 6, 2026

Recurrent Looped Transformer

Authors: Yifan Zhang, Jichen Feng, Shihan Qin

Organizations: Princeton University · University of Pennsylvania

Abstract

State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based S5S_5 permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based S5S_5 to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based S5S_5 from 100% to 20%.

Figures & tables

Explore similar work

Apr 23, 2026cs.LG

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding

Transformers process tokens in parallel but are temporally shallow: at position tt, each layer attends to key-value pairs computed based on the previous layer, yielding a depth capped by the number of layers. Recurrent models offer unbounded temporal depth but suffer from optimization instability and historically underutilize modern accelerators. We introduce the Recurrent Transformer, a simple architectural change where each layer attends to key-value pairs computed off its own activations, yielding layerwise recurrent memory while preserving standard autoregressive decoding cost. We show that the architecture can emulate both (i) a conventional Transformer and (ii) token-to-token recurrent updates under mild assumptions, while avoiding optimization instability. Naively, prefill/training appears bandwidth-bound with effective arithmetic intensity near 11 because keys and values are revealed sequentially; we give an exact tiling-based algorithm that preserves the mathematical computation while reducing HBM traffic from Θ(N2)Θ(N^2) to Θ(Nlog⁡N)Θ(N\log N), increasing effective arithmetic intensity to Θ(N/log⁡N)Θ(N/\log N) for sequence length NN. On 150M and 300M parameter C4 pretraining, Recurrent Transformers improve cross-entropy over a parameter-matched Transformer baseline and achieve the improvement with fewer layers (fixed parameters), suggesting that recurrence can trade depth for width, thus reducing KV cache memory footprint and inference latency.
Sep 24, 2026cs.LG

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

Looped Transformers have attracted substantial attention as a parameter-efficient approach to increasing computational depth through repeated application of shared Transformer blocks. However, their practical advantages over conventional Transformers remain under debate: each additional loop incurs another Transformer pass and requires caching another set of KV states, causing inference FLOPs and KV-cache memory to grow continuously with loop depth. This overhead becomes particularly severe at large loop counts and long context, preventing the parameter efficiency of Looped Transformers from translating into practical inference efficiency. In this paper, we find that much of the additional computation and storage introduced by looping is redundant. As recurrence proceeds, state changes become increasingly concentrated on a small subset of tokens; attention-output differences are dominated by a sparse and stable subset of key columns; and KV residuals between adjacent loops become progressively more amenable to low-bit quantization. Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization. Across several Looped Transformers models, FlashLoop delivers lossless accuracy while achieving up to 1.64×\times end-to-end speedup and up to 6×\times KV-cache memory reduction, substantially improving the practicality of scaling Looped Transformers to greater computational depths and longer context.
May 30, 2026cs.LG

Refit the Probe: Single-Direction Ablation Is Not a Necessity Test

Probes are routinely paired with an intervention: ablate the direction the probe found, run the model, and read the change in task accuracy, taking a large drop as evidence that the computation depends on what the probe read and a near-zero drop as evidence that it does not. Either inference requires that the ablation have removed the target from the layer. We find that the ablation does not remove what it targets. A probe refitted on the ablated activations recovers its original accuracy in every cell we test, and keeps recovering when the probe's entire row space is deleted rather than a single axis, because the quantity survives in the orthogonal complement. Because a refitted probe recovers, neither a large task drop nor a near-zero one establishes whether the model needed the target, and one probe fit detects this. Replacing the ablation with iterative nullspace projection, scored against random subspaces of matched dimension, reverses the conclusion: representations that looked causally inert carry most of the task. The correction also separates where a variable is most readable from where deleting it does most damage, and those are not the same layer in any pretrained model we study. The erasure is defined by a linear probe family, so removing a nonlinearly encoded quantity remains open.