Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
Figures & tables
Figure 1: Joint and layerwise optimization of nonlinear recurrent memory. (a) TTT jointly optimizes both layers through a shared reconstruction objective. (b) DeltaTTT assigns a local reconstruction objective to each layer and performs layerwise delta-rule updates. (c) Both methods use the nonlinear readout after the writes.
Update rule
Loss ↓
Single ↑
32k ↑
Serial
1.99
28.25
11.13
Parallel
1.98
37.87
30.87
Table 1: Serial and parallel TTT with SwishGLU inner models, 0.5B parameters, and 92.5B training tokens. Single denotes mean single-needle retrieval accuracy across S1/S2/S3; 32k denotes the mean accuracy across S1/S2/S3 at 32k context length.
Figure 2: Fitting fixed Full-Attention K/V pairs at different replay budgets.
Method
Gradient/write evaluated at
State recurrence
Output computed from
Nonlinear
Efficient training
Token-wise TTT ( Sun et al., 2024 )
Previous token
Token-wise
Current token
✓
×
Mini-batch TTT ( Sun et al., 2024 )
Batch start
Token-wise
Current token
✓
×
Chunk-TTT ( Zhang et al., 2025 )
Chunk start
Chunk-wise
Chunk start
✓
✓
E 2 -TTT ( Zhong et al., 2026 )
Chunk start
Token-wise
Chunk start
✓
✓
Parallel TTT
Initial state
Chunk-wise
Chunk start
✓
✓
DeltaNet ( Yang et al., 2024 )
Previous token
Token-wise
Current token
×
✓
Table 2: Comparison of memory update schedules. Gradient/write evaluated at denotes the fast-weight state used to compute each update; state recurrence denotes the granularity of state evolution; output computed from denotes the state used for query readout. TTT rows assume nonlinear inner models. Parallel TTT adapts fixed-initialization batch-gradient TTT to our chunkwise, read-before-write baseline (Eq. 4 ).
Model
Params (M)
Train loss ↓
Single ↑
S1 ↑
S2 ↑
S3 ↑
4K
8K
16K
32K
Avg.
4K
8K
16K
32K
Avg.
4K
8K
16K
32K
Avg.
Full attention
536.95
1.918769
72.42
100.0
100.0
99.2
95.2
98.60
95.4
91.2
95.6
88.2
92.60
50.0
27.8
22.2
4.2
26.05
DeltaNet backbone
DeltaNet
537.64
1.955473
44.38
99.8
98.4
97.2
91.8
96.80
83.0
32.6
10.8
2.4
32.20
7.6
4.2
2.2
2.6
4.15
Gated DeltaNet
562.80
1.967150
47.60
100.0
100.0
100.0
98.8
99.70
76.0
43.6
19.4
5.0
36.00
17.6
7.2
2.2
1.4
7.10
DeltaTTT
541.17
1.950079
48.78
100.0
100.0
100.0
99.8
99.95
93.4
44.0
14.2
4.6
39.05
21.4
4.6
1.6
1.8
7.35
Table 3: Results grouped by backbone at the 92.5B-token budget. All recurrent configurations use SWA with window 512. Single denotes the mean over S1/S2/S3.
Figure 3: Average Position-wise loss on 43 held-out documents (19 books, 24 papers), each at least 32K tokens long.
Model
Throughput/GPU ↑
Peak memory/GPU ↓
Full attention
1.00×
1.00×
DeltaNet
1.42×
0.53×
LaCT Serial (SwiGLU)
0.83×
0.68×
LaCT Parallel (SwiGLU)
1.07×
0.86×
DeltaTTT(ours)
1.37×
0.69×
Table 4: Training efficiency relative to full attention ( 1.00× ).
Model
Configuration
Train loss ↓
Single ↑
(a) Write address
DeltaTTT-Pre (selected)
Address from At−1
1.962162
54.32
DeltaTTT-Post
Address from At
1.960538
52.22
(b) Progressive component removal
DeltaTTT (selected)
Full model
1.952344
50.47
DeltaTTT
−B online update
1.967580
36.02
Table 5: Ablations at 0.5B model and 92.5B training tokens, Removals in (b) are cumulative; Single is the mean single-needle retrieval accuracy across S1/S2/S3.
Figure 4: Learned A0 and B0 at layer 13, head 1. Top: main configuration; bottom: the variant with both matrices initialized to zero.
Model
Train tokens
Initial MSE
Final MSE
Reduction
DeltaNet
92.5B
0.00408555
0.00118968
70.88%
DeltaNet(learned init)
92.5B
0.00466836
0.00109973
76.44%
DeltaTTT
92.5B
0.00468315
0.00084878
81.88%
Table 6: Initial and final full-sequence reconstruction errors on the same 32K validation sequence. Each checkpoint uses its self-generated K/V features. Reduction denotes the relative decrease in MSE.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Training data
Long-Data-Collections, 92.5B tokens
Model dimensions
24 layers; hidden width 1,024; FFN width 4,096
Vocabulary and embeddings
131,072 tokens; tied input/output embeddings
Sequence length
32,768 tokens
Optimizer
AdamW, (β1,β2)=(0.9,0.95) , ϵ=10−8
Weight decay / gradient clipping
0.1 / 1.0
Appendix
Table 7: Shared training configuration
Model
Initial states
Train loss ↓
Single ↑
DeltaTTT
A0=B0=0 (learnable)
1.95
47.70
DeltaTTT
(A0)ij,(B0)ij∼N(0,0.022)
1.95
48.00
DeltaTTT
(A0)ij∼N(0,0.022),B0=LR+0.5I
1.95
48.78
Appendix
Table 8: Initial-state ablations for DeltaTTT on the DeltaNet backbone. Single denotes average single-needle accuracy (%).
Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorly: it is sequential in time, limiting parallelism, and suffers from vanishing or exploding gradients, making long-range associations difficult to learn. We propose Supervised Memory Training (SMT), a method for training nonlinear RNNs that sidesteps recurrent credit propagation entirely by reducing RNN training to supervised learning on one-step memory transition labels (mt,xt+1)→mt+1. SMT acquires these memory labels by training a Transformer-based encoder on a predictive state objective--retaining only information from the past necessary to predict the future. By decoupling what to remember from how to update memory, SMT enables time-parallel RNN training with a stable O(1) length gradient path between any two tokens--without ever unrolling the RNN. We find that SMT outperforms BPTT when pretraining various RNN architectures on tasks like language modeling and pixel sequence modeling. SMT enables nonlinear RNNs to better capture long-range dependencies and train in parallel, potentially unlocking the scaling of models that build temporal abstractions of past experience.
To address the increasing long-context compute limitations of softmax attention, several subquadratic recurrent operators have been developed. This work includes models such as Mamba-2, DeltaNet, Gated DeltaNet (GDN), and Kimi Delta Attention (KDA). As the space of recurrences grows, a parallel line of work has arisen to taxonomize them. One compelling view is the test-time regression (TTR) framework, which interprets recurrences as performing online least squares updates that learn a linear map from the keys to values. Existing delta-rule recurrences can be seen as first-order approximations to this objective, but notably ignore the curvature of the least-squares loss during optimization. In this work, we address this by introducing preconditioning to these recurrences. Starting from the theory of online least squares, we derive equivalences between linear attention and the delta rule in the exactly preconditioned case. Next, we realize this theory in practice by proposing a diagonal approximation: this enables us to introduce preconditioned variants of DeltaNet, GDN, and KDA alongside efficient chunkwise parallel algorithms for computing them. Empirically, we find that our preconditioned delta-rule recurrences yield consistent performance improvements across synthetic recall benchmarks and language modeling at the 340M and 1B scale.
Linear Attention (LA) offers a promising paradigm for scaling large language models (LLMs) to long sequences by avoiding the quadratic complexity of self-attention. Recent LA models such as Mamba2 and GDN interpret linear recurrences as closed-form online stochastic gradient descent (SGD), but naive SGD updates suffer from rapid information decay and suboptimal convergence in optimization. While momentum-based optimizers provide a natural remedy, they pose challenges in simultaneously achieving training efficiency and effectiveness. To address this, we develop a chunkwise parallel algorithm for LA with a stepwise momentum rule by geometrically reordering the update coefficients. Further, from a dynamical systems perspective, we analyze the momentum-based recurrence as a second-order system that introduces complex conjugate eigenvalues. This analysis guides the design of stable gating constraints. The resulting model, Momentum DeltaNet (MDN), leverages Triton kernels to achieve comparable training throughput with competitive linear models such as Mamba2 and KDA. Extensive experiments on the 400M and 1.3B parameter models demonstrate consistent performance improvements over strong baselines, including Transformers, Mamba2 and GDN, across diverse downstream evaluation benchmarks. Code: https://github.com/HuuYuLong/MomentumDeltaNet .
Yulong Huang, Xiang Liu, Hongxiang Huang +5
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China