Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
Figures & tables
Figure 1: Joint and layerwise optimization of nonlinear recurrent memory. (a) TTT jointly optimizes both layers through a shared reconstruction objective. (b) DeltaTTT assigns a local reconstruction objective to each layer and performs layerwise delta-rule updates. (c) Both methods use the nonlinear readout after the writes.
Update rule
Loss ↓
Single ↑
32k ↑
Serial
1.99
28.25
11.13
Parallel
1.98
37.87
30.87
Table 1: Serial and parallel TTT with SwishGLU inner models, 0.5B parameters, and 92.5B training tokens. Single denotes mean single-needle retrieval accuracy across S1/S2/S3; 32k denotes the mean accuracy across S1/S2/S3 at 32k context length.
Figure 2: Fitting fixed Full-Attention K/V pairs at different replay budgets.
Method
Gradient/write evaluated at
State recurrence
Output computed from
Nonlinear
Efficient training
Token-wise TTT ( Sun et al., 2024 )
Previous token
Token-wise
Current token
✓
×
Mini-batch TTT ( Sun et al., 2024 )
Batch start
Token-wise
Current token
✓
×
Chunk-TTT ( Zhang et al., 2025 )
Chunk start
Chunk-wise
Chunk start
✓
✓
E 2 -TTT ( Zhong et al., 2026 )
Chunk start
Token-wise
Chunk start
✓
✓
Parallel TTT
Initial state
Chunk-wise
Chunk start
✓
✓
DeltaNet ( Yang et al., 2024 )
Previous token
Token-wise
Current token
×
✓
Table 2: Comparison of memory update schedules. Gradient/write evaluated at denotes the fast-weight state used to compute each update; state recurrence denotes the granularity of state evolution; output computed from denotes the state used for query readout. TTT rows assume nonlinear inner models. Parallel TTT adapts fixed-initialization batch-gradient TTT to our chunkwise, read-before-write baseline (Eq. 4 ).
Model
Params (M)
Train loss ↓
Single ↑
S1 ↑
S2 ↑
S3 ↑
4K
8K
16K
32K
Avg.
4K
8K
16K
32K
Avg.
4K
8K
16K
32K
Avg.
Full attention
536.95
1.918769
72.42
100.0
100.0
99.2
95.2
98.60
95.4
91.2
95.6
88.2
92.60
50.0
27.8
22.2
4.2
26.05
DeltaNet backbone
DeltaNet
537.64
1.955473
44.38
99.8
98.4
97.2
91.8
96.80
83.0
32.6
10.8
2.4
32.20
7.6
4.2
2.2
2.6
4.15
Gated DeltaNet
562.80
1.967150
47.60
100.0
100.0
100.0
98.8
99.70
76.0
43.6
19.4
5.0
36.00
17.6
7.2
2.2
1.4
7.10
DeltaTTT
541.17
1.950079
48.78
100.0
100.0
100.0
99.8
99.95
93.4
44.0
14.2
4.6
39.05
21.4
4.6
1.6
1.8
7.35
Table 3: Results grouped by backbone at the 92.5B-token budget. All recurrent configurations use SWA with window 512. Single denotes the mean over S1/S2/S3.
Figure 3: Average Position-wise loss on 43 held-out documents (19 books, 24 papers), each at least 32K tokens long.
Model
Throughput/GPU ↑
Peak memory/GPU ↓
Full attention
1.00×
1.00×
DeltaNet
1.42×
0.53×
LaCT Serial (SwiGLU)
0.83×
0.68×
LaCT Parallel (SwiGLU)
1.07×
0.86×
DeltaTTT(ours)
1.37×
0.69×
Table 4: Training efficiency relative to full attention ( 1.00× ).
Model
Configuration
Train loss ↓
Single ↑
(a) Write address
DeltaTTT-Pre (selected)
Address from At−1
1.962162
54.32
DeltaTTT-Post
Address from At
1.960538
52.22
(b) Progressive component removal
DeltaTTT (selected)
Full model
1.952344
50.47
DeltaTTT
−B online update
1.967580
36.02
Table 5: Ablations at 0.5B model and 92.5B training tokens, Removals in (b) are cumulative; Single is the mean single-needle retrieval accuracy across S1/S2/S3.
Figure 4: Learned A0 and B0 at layer 13, head 1. Top: main configuration; bottom: the variant with both matrices initialized to zero.
Model
Train tokens
Initial MSE
Final MSE
Reduction
DeltaNet
92.5B
0.00408555
0.00118968
70.88%
DeltaNet(learned init)
92.5B
0.00466836
0.00109973
76.44%
DeltaTTT
92.5B
0.00468315
0.00084878
81.88%
Table 6: Initial and final full-sequence reconstruction errors on the same 32K validation sequence. Each checkpoint uses its self-generated K/V features. Reduction denotes the relative decrease in MSE.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Training data
Long-Data-Collections, 92.5B tokens
Model dimensions
24 layers; hidden width 1,024; FFN width 4,096
Vocabulary and embeddings
131,072 tokens; tied input/output embeddings
Sequence length
32,768 tokens
Optimizer
AdamW, (β1,β2)=(0.9,0.95) , ϵ=10−8
Weight decay / gradient clipping
0.1 / 1.0
Appendix
Table 7: Shared training configuration
Model
Initial states
Train loss ↓
Single ↑
DeltaTTT
A0=B0=0 (learnable)
1.95
47.70
DeltaTTT
(A0)ij,(B0)ij∼N(0,0.022)
1.95
48.00
DeltaTTT
(A0)ij∼N(0,0.022),B0=LR+0.5I
1.95
48.78
Appendix
Table 8: Initial-state ablations for DeltaTTT on the DeltaNet backbone. Single denotes average single-needle accuracy (%).