We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. \ours{} consists of two coupled models: a prefiller
Q, which leverages full attention\footnote{In practice, we use interleaved full and sliding-window attention for
Q, as this yields stronger performance. The essential requirement is that
Q be more expressive than
P, with access to the full history.} to produce memory targets
mt′, and a decoder
P, which uses only sliding-window attention and recurrent K/V injection to produce decoder memories
mt for next-token prediction. We train \ours{} with a memory consistency loss that aligns
mt with
mt′, allowing inference to use
P alone. Empirically, \ours{} improves validation loss and downstream pretraining benchmarks over sliding-window and latent recurrent transformer baselines. Moreover, sharing parameters between
P and
Q reduces parameter memory while preserving most of the gains.