Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.
Figures & tables
Design
p
E
∣We∣
∣Ωℓ∣
Reachable
Active
Memories
TTT-Dense (LaCT)
L
LH
S/H
H
S
S/H
L
Dense, fixed compute
p
pH
S/H
H
S
S/H
p
Dense, fixed state
p
pH
LS/(pH)
H
LS/p
LS/(pH)
p
uTTT-Dense
1
H
LS/H
H
LS
LS/H
1
TTT-MoE
L
E
u
E/L
Eu/L
Ku
L
Routed, p groups
p
E
u
E/p
Eu/p
Ku
p
Table 1: Designs along the group count p . H counts fast-weight heads and S is one private dense layer’s full state. E counts SwiGLU triples of size ∣We∣ ; layer ℓ reaches ∣Ωℓ∣∣We∣ , while Active counts the state applied by one token head. The fixed-compute dense row stores pS at unchanged per-token fast-weight compute; the other dense rows preserve total state LS , and routed rows preserve Eu . The uTTT-Dense row is size-matched. A write-enabled chunk ends in one update per group, ( ).
Figure 2: State-read forward and backward chains. Layer-private TTT (top) and universal TTT (bottom). (a) Forward reads use the chunk-start state. Inner-loop writes prepare the state for later reads during both training and inference (Section ). Bands trace top-down (1) and bottom-up (2) routes. (b) The restricted read-output Jacobian traces backward through historical inner updates to write inputs; it is a factor of outer task-loss backpropagation. Each update includes the inner-loss gradient computation. Queries and predicted coefficient producers are held fixed here; full training retains those additional paths. Hollow nodes lack this restricted state-chain path; writers below layer ℓ in earlier chunks can still reach the read through the residual stream. Colors identify the forward writer-to-reader relation; arrows point backward. Routing determines which paths are active (Appendix ).
Figure 3: Loss–retrieval Pareto plot , at (a) 124M and (b) 760M parameters. One point per model (Section ); the horizontal axis is mean per-token loss beyond 30K in 32K-token sequences, and the vertical axis is mean RULER accuracy over 4K–32K. Arrows run through the fast-weight models from LaCT to uTTT-MoE. The shaded band in (b) spans the four fast-weight models, whose losses differ by 0.0047 nat. The vertical-axis break in (b) marks a compressed empty stretch.
124M \CT@row@color
760M
Model
PTL ↓
RULER ↑
4K
8K
16K
32K \CT@row@color
PTL ↓
RULER ↑
4K
8K
16K
32K
Baselines
Transformer †
3.0869
16.52
23.82
20.24
13.10
8.93 \CT@row@color
2.5978
34.16
41.70
37.95
33.36
23.62
Transformer-SWA
3.1243
14.50
26.93
14.73
9.94
6.41 \CT@row@color
2.6325
22.27
45.23
22.64
12.67
8.57
Gated DeltaNet-SWA
3.0994
12.32
23.78
12.93
8.30
4.26 \CT@row@color
2.6075
21.48
40.40
23.12
13.66
8.76
DeltaNet-SWA
3.1336
11.69
22.52
11.23
7.63
5.36 \CT@row@color
2.6086
20.77
43.43
20.56
11.79
7.28
Table 2: Long-context language modeling at 32K. PTL is the mean next-token loss at positions ≥30 K on Books3, and the five columns beside it are RULER accuracy (%): the mean over 4K, 8K, 16K and 32K, then each of those lengths. Bold marks the best value in each column among the bounded-state models; the full-attention Transformer ( † ) keeps a key–value cache that grows with context and is excluded from that comparison. The lower block holds the fast-weight models, of which uTTT-MoE is the model this paper proposes.
Figure 4: State-read backward and forward dependence. LLM diagnostics use 124M models and individual Books3 documents’ first 32K tokens. (a) Squared-gradient shares from a layer-1, chunk-8 read-output Jacobian probe, normalized per document and averaged over 300 documents. This probes a local factor of outer task-loss backpropagation. Colors are logarithmic; gray is zero. Boxed counts require nonzero gradients in every document; circles mark the read. (b) On 1,024 documents, bars show loss increases after freezing state writes or reading only own writes, relative to each model’s normal predictions over chunks 2–8; dots show unmodified tail PTL at positions ≥30,000 , measured on these unpacked documents and therefore not on the scale of Table . Full uses full uTTT-MoE backpropagation; Cut stops cross-layer read-to-write gradients during training; TTT is TTT-MoE. (c) PSNR drop after removing one writer–reader route in uTTT-MoE-e64-p1, averaged over 1,019 GSO objects with 23 target views each. Diagonal routes are included. Colors are linear, with drops below 0.01 dB white; the outline marks the maximum. Panels (b,c) use fixed trained checkpoints.
Figure 5: Measured training efficiency. (a) Training throughput of the three ownership regimes on one node of eight A100 or H100 GPUs (32K context, batch 1 per GPU), in thousands of tokens per second per GPU. (b) Where a 760M training step’s time goes (one A100, batch 1); attention and MLP are identical across the three models, and other is the backward pass and the optimizer. (c) Training loss of the three 760M models against H100 GPU-hours.
Figure 6: Sharing (row A) and compute scaling (row B) on Objaverse/GSO : PSNR against the target-view index, i.e. the number of posed views written into the pool; the number at the right end of each curve is its value at view 23. The e , p , w , and a in model names are defined in Section . Row A holds per-layer fast-weight compute fixed ( a=1 , w=1 ) and widens sharing from p=8 to p=1 ; row B holds total fast-weight state fixed (64 experts, width 8) and raises per-layer compute from one unit to eight. The MoE panel of row B also shows the three baselines at the same protocol: the full-attention reference (black dashed) and the linear-attention Gated DeltaNet and DeltaNet (grey dotted). Full metrics and 95% bootstrap intervals are in Table .
Model
PSNR ↑
View 23 PSNR ↑
SSIM ↑
LPIPS ↓
MoE A: e×p=64 , top-1
uTTT-MoE-e64-p1
24.68 [24.51, 24.86]
25.48 [25.27, 25.68]
0.860 [0.855, 0.864]
0.173 [0.168, 0.177]
uTTT-MoE-e32-p2
24.20 [24.03, 24.37]
24.98 [24.78, 25.18]
0.854 [0.850, 0.859]
0.181 [0.177, 0.186]
uTTT-MoE-e16-p4
24.11 [23.94, 24.29]
24.89 [24.69, 25.10]
0.854 [0.849, 0.858]
0.182 [0.177, 0.186]
uTTT-MoE-e8-p8 (TTT-MoE-e8)
23.90 [23.73, 24.08]
24.56 [24.36, 24.77]
0.850 [0.845, 0.855]
0.186 [0.181, 0.191]
Dense A: w=1
Table 3: The configurations of Figure on Objaverse/GSO ( n=1,019 objects). Cells: mean over objects of each object’s mean over target views 1–23, with its 95% bootstrap interval over objects (Appendix ); the second PSNR column is view 23 alone. Bold: best per column among fast-weight models; † full-attention reference, not bolded. The tinted row is the best configuration in the table.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Asset
Use in this paper
License or access
Objaverse ( Deitke et al., 2023 )
NVS training
ODC-By v1.0; per-object CC licenses, some non-commercial
Google Scanned Objects ( Downs et al., 2022 )
NVS evaluation
CC BY 4.0
DL3DV-10K ( Ling et al., 2024 )
NVS training and evaluation
CC BY-NC 4.0; gated access
Long-Data-Collections ( Together AI, 2023 )
LM training
Mixed sources, each under its own license
Books3 from The Pile ( Gao et al., 2020 )
LM evaluation (Section ; Appendices , and )
Copyrighted books; withdrawn from its original host in 2023
RULER ( Hsieh et al., 2024 )
Retrieval evaluation, via the LM Evaluation Harness ( Gao et al., 2024 )
Apache-2.0 (harness: MIT); some needle tasks use Paul Graham essays as filler, the QA tasks use the two datasets below, the rest are synthetic
Appendix
Table B1: External datasets, with license or access terms.
124M
760M
Width d / layers L / attention heads
768 / 12 / 12
1536 / 24 / 24
Attention head dimension
64
FFN (SwiGLU ( Shazeer, 2020 ) ) hidden width
2,048
4,096
Normalization
pre-norm RMSNorm ( Zhang & Sennrich, 2019 )
Vocabulary
32,000; untied embeddings
RoPE ( Su et al., 2024 ) θ
106
Appendix
Table D1: Settings shared by every language model. Where one value spans both columns it holds at both scales.
Non-emb. params
Model
FFN width
State / layer
124M
760M
Attention references
Transformer ( Vaswani et al., 2017 )
2,048 / 4,096
KV cache
84.95M
679.6M
Transformer-SWA ( Beltagy et al., 2020 )
2,048 / 4,096
—
84.95M
679.6M
Linear attention, matched in parameters
DeltaNet-SWA ( Yang et al., 2024 )
2,048 / 4,096
64d
85.10M
680.6M
Appendix
Table D2: What distinguishes each language model. Every mixer but full attention runs beside the 4K window of Table and, except Mamba-2’s, reuses its q,k,v projections; the linear-attention mixers apply SiLU and L2 to q,k and use no short convolution. State is what a layer carries across chunks; uTTT-MoE pools it, 4L experts shared by all layers. Parameters exclude the input embedding and output head (24.6M / 49.2M each).
LaCT, TTT-Dense
TTT-MoE
uTTT-MoE
Heads × width
4×d/4
Function
SwiGLU; initial W0,W2 : rank 32 +0.5I
q,k,v
window attention’s, pre-RoPE; SiLU; L2, then RoPE on q,k
Step size
per token, head, and matrix; base 10−3
Update
momentum, Muon ( Jordan et al., 2024 ) (5 steps), row renormalization
Momentum coefficient
per head
per chunk
per chunk
Appendix
Table D3: The fast-weight block ( d : model width; two values are 124M / 760M). LaCT and TTT-Dense compute the same block in a different order: LaCT runs the whole sequence through one layer at a time, TTT-Dense one chunk at a time through all layers.
Model
FFN width
State / layer
Params
References
Transformer
1,536
KV cache ( +1,024 tokens per view)
27.76M
DeltaNet
1,584
8×642 , delta rule
37.04M
Gated DeltaNet
1,408
8×642 , gated delta rule
37.01M
Fast-weight models of Section
64 total experts or total width 8
1,536
8×3×642 per width unit
37.04–37.05M
Appendix
Table D4: What distinguishes each novel-view model. The fast-weight models share Table ; the references change only the memory sublayer and the FFN width. Parameters count every trainable weight, including learnable initial states; the fast-weight row is for 64 total experts or total width 8 (36.25M outside the state and router, 0.79M initial state, and a router of at most 4,096 weights).
Width / layers / attention heads
512 / 8 / 8
Block
attention → fast-weight read → SwiGLU FFN (1,536)
Attention
within one view
Input
2562 , patch 8: 1,024 tokens per view
Views
training 12 (11 posed + 11 target chunks); evaluation 24
Fast-weight heads
8×64
Unit of state
expert: 3×642 weights, shared by heads; width unit: 8 experts
Appendix
Table D5: Settings shared by every novel-view-synthesis fast-weight model of Section ; 36.25M parameters outside the fast-weight state and router.
Figure E1: Retrieval by task and loss by position , at 124M and 760M. (a) RULER accuracy at 4K–32K on the second single-needle task (S-NIAH-2), the multi-value task (MV-NIAH), and, in the last row, the unweighted mean over all thirteen RULER tasks, which is the RULER axis of Figure . (b) Per-token validation loss on Books3 by position within a 32K-token sequence, on a log-scaled axis with positions below 2K omitted as warm-up.
Figure E2: Training seeds and every balancing run of both routed families , on the axes of Figure . (a) Six training seeds of the 124M uTTT-MoE configuration of Table ; the filled star is seed 42, the reported run. (b) Every run of Table (uTTT-MoE, circles) and Table (TTT-MoE; filled squares E/L=4 , hollow E/L=8 ) at 124M and 760M. Colour is the balancing mechanism and the scope over which utilization is measured; the layer-private control has a single scope and takes the layer-scope colours. The ringed circle is the configuration reported as uTTT-MoE at that scale.
Table 17
γ=1
γ=2
γ=4
Balancing
PTL
RULER
PTL
RULER
PTL
RULER
Aux. loss 10−4
—
—
2.5744
23.16
—
—
Aux. loss 5×10−4
—
—
2.5780
23.96
2.5805
24.19
Aux. loss 10−3
2.5783
21.27
2.5783
25.50
2.5791
26.07
Aux. loss 2×10−3
—
—
2.5793
20.57
2.5794
25.56
Aux. loss 5×10−3
—
—
2.5771
25.11
—
—
Appendix
Table E3: The routing-weight scale at 760M. Pool scope and a per-layer router throughout; γ scales the gates as in ( ) and is 1 everywhere else in this paper. Each cell is PTL and RULER (%); a dash marks a combination that was not trained. The γ=1 column is the pool-scope part of the 760M block of Table .
Figure E3: Every run of Table , resolved by task. Cards (a), (b) and (d) sweep the loss-free bias-update rate at 124M pool scope, 124M layer scope and 760M pool scope; cards (c) and (e) draw the unbalanced and auxiliary-loss runs at 124M and 760M. Each card holds the two tasks of Figure (a) and the unweighted mean over all thirteen RULER tasks, whose four points are the 4K–32K columns of Table ; the cards share their rows, so a task reads across the figure. Rates are ordinal and drawn light to dark; the pink line is the configuration reported at that scale, present only in the two pool-scope sweeps. Each panel carries its own accuracy range.
Model
State (M)
PTL ↓
RULER ↑
4K
8K
16K
32K
uTTT-Dense, r=1 (FLOPs-matched)
0.44
3.1290
10.64
20.27
12.23
6.46
3.61
uTTT-Dense, r=2
0.88
3.1232
13.20
23.57
14.66
9.79
4.78
uTTT-Dense, r=4
1.77
3.1249
13.41
21.65
15.12
10.69
6.20
uTTT-Dense, r=6
2.65
3.1165
13.63
23.60
16.01
9.92
5.01
TTT-Dense (Table )
5.31
3.1219
12.65
26.31
13.58
6.00
4.72
Appendix
Table E4: A dense shared state at 124M. The shared state is r times the width of one layer-private state, so r=1 matches TTT-Dense in per-token fast-weight compute and r=12 , the depth of the model, would match it in size; columns as in Table ; the TTT-Dense row repeats Table .
Fixed-target ablation: visible prefix
Reset segments
Model
4K
8K
16K
32K
48K
64K
120K
Seg. 1
Seg. 2
Seg. 3
Seg. 4
Transformer-SWA (32K)
3.222
3.222
3.222
3.222 †
3.222 †
3.222 †
3.222 †
3.224
3.224
3.218
3.214
Transformer-SWA (64K)
3.220
3.220
3.220
3.220
3.220
3.220 †
3.220 †
3.223
3.222
3.216
3.212
Transformer (32K)
3.211
3.198
3.188
3.329 †
4.920 †
5.692 †
5.838 †
3.205
3.201
3.194
3.190
Transformer (64K)
3.226
3.212
3.201
3.193
3.194
3.208 †
5.563 †
3.218
3.215
3.208
3.204
LaCT
3.215
3.214
3.213
3.213 †
3.216 †
3.219 †
3.230 †
3.216
3.216
3.209
3.205
Appendix
Table F1: Fixed-target ablation and reset segments , loss in nats. Left: the same 8,192 target tokens as the visible prefix grows, each prefix and its targets forming one sequence; † : sequence beyond the model’s trained length; bold marks each row’s minimum before rounding. Right: the four consecutive 32 K pieces of the same 128 K. Transformer-SWA has a 4 K attention window.
Figure F1: Per-token loss along 3,000 Books3 documents (positions 4 K– 128 K) , smoothed over 6,144 tokens. (a) Loss of the seven models. (b) Loss relative to the 32K-trained Transformer-SWA with a 4 K attention window. Shaded bands show successive loss differences in the order SWA (64K), LaCT, TTT-MoE, uTTT-MoE, starting from zero; hatching marks a reversal of the ordering. Full-attention models are lines only, clipped at the top. Dotted lines mark the 32 K and 64 K training lengths; LaCT and both expert pools are 32K-trained.
Model
4K
8K
16K
32K
64K
Transformer-SWA, 4K window (32K)
29.1
15.6
9.8
6.0
3.4 †
Transformer-SWA, 4K window (64K)
28.5
16.0
9.5
6.0
3.8
Transformer (32K)
30.9
26.4
23.0
13.8
1.4 †
Transformer (64K)
20.5
10.7
11.4
8.4
5.8
LaCT
23.6
14.8
9.3
5.2
3.3 †
TTT-MoE
28.7
13.3
6.4
4.2
3.3 †
Appendix
Table F2: RULER accuracy (%) by evaluation length , 500 samples per task per length, averaged over the 13 tasks of Figure . Bold marks column maxima at the displayed precision. † : beyond the trained length. Model configurations are given in Appendix .
Figure G1: Efficiency beyond Figure . (a) GPU busy time at 124M against batch per GPU, on one node of eight A100s (top) or H100s (bottom). (b) uTTT-MoE training throughput (top) and peak memory (bottom) against shared-pool size at 124M on one A100, with TTT-Dense and TTT-MoE in the same setting dashed. (c) 760M inference against context length: batch-1 prefill throughput (top, one A100, 4K–32K contexts) and per-token decoding latency (bottom, one A100; TTT-MoE and uTTT-MoE coincide).
124M
760M
Kernels switched off
tok/s
speed-up
tok/s
speed-up
none
12,071
–
5,542
–
all
10,533
1.15
4,710
1.18
routing permutation of k , v and learning rates
11,282
1.07
5,088
1.09
grouped GEMMs of the outer-gradient backward
11,816
1.02
5,158
1.07
pointwise second-order backward
11,950
1.01
5,162
1.07
Appendix
Table G1: Fused-kernel ablation of uTTT-MoE training at batch 1×32 K on one A100. Training tokens per second per GPU, median of three repeats, and the speed-up the group gives; peak memory is the same in every row.
Figure H1: The four modules of Figure on DL3DV ; conventions as there.
Figure H2: The remaining modules on Objaverse/GSO ; conventions as Figure . MoE A ′ and B ′ repeat rows A and B at 32 experts in total, with TTT-MoE-e4 as the p=8 endpoint. C sweeps the pool size under full sharing ( p=1 ): per-layer compute is fixed for the routed pool ( a=1 ) and scales with w for the dense state.
Figure H3: The modules of Figure on DL3DV ; conventions as Figure .
Model
PSNR ↑
View 23 PSNR ↑
SSIM ↑
LPIPS ↓
Layer-private pools ( p=8 )
uTTT-Dense-w1-p8 (TTT-Dense)
23.99 [23.81, 24.16]
24.49 [24.29, 24.69]
0.852 [0.847, 0.856]
0.185 [0.180, 0.189]
uTTT-MoE-e4-p8 (TTT-MoE-e4)
23.80 [23.62, 23.97]
24.44 [24.24, 24.64]
0.849 [0.845, 0.854]
0.188 [0.183, 0.193]
uTTT-MoE-e8-p8 (TTT-MoE-e8)
23.90 [23.73, 24.08]
24.56 [24.36, 24.77]
0.850 [0.845, 0.855]
0.186 [0.181, 0.191]
MoE, universal pool, top-1 (C)
uTTT-MoE-e8-p1
24.72 [24.54, 24.90]
25.55 [25.34, 25.77]
0.861 [0.856, 0.865]
0.172 [0.168, 0.177]
Appendix
Table H1: Every fast-weight configuration on Objaverse/GSO ( n=1,019 objects). Cells: mean over objects of each object’s mean over target views 1–23, with its 95% bootstrap interval over objects; the second PSNR column is view 23 alone. Bold: best per column among fast-weight models; † full-attention reference, not bolded.
Model
PSNR ↑
View 23 PSNR ↑
SSIM ↑
LPIPS ↓
Layer-private pools ( p=8 )
uTTT-Dense-w1-p8 (TTT-Dense)
16.14 [15.91, 16.39]
16.69 [16.35, 17.03]
0.383 [0.365, 0.402]
0.681 [0.673, 0.690]
uTTT-MoE-e4-p8 (TTT-MoE-e4)
15.98 [15.74, 16.22]
16.47 [16.14, 16.81]
0.381 [0.362, 0.399]
0.691 [0.682, 0.699]
uTTT-MoE-e8-p8 (TTT-MoE-e8)
16.02 [15.78, 16.27]
16.54 [16.20, 16.89]
0.379 [0.361, 0.397]
0.688 [0.679, 0.696]
MoE, universal pool, top-1 (C)
uTTT-MoE-e8-p1
16.02 [15.78, 16.27]
16.63 [16.30, 16.98]
0.382 [0.364, 0.401]
0.686 [0.677, 0.694]
Appendix
Table H2: Every fast-weight configuration on DL3DV ( n=140 test scenes); format as Table , with means and bootstrap intervals computed over scenes.
Schedule
GSO V23 PSNR ↑
DL3DV V23 PSNR ↑
k patch tok/s/GPU ↑
min/1k steps ↓
Peak GiB ↓
Sequential schedule
25.58
16.80
100.2
119.9
40.75
Aggregated schedule
25.48
16.79
102.3 (+2.1%)
117.5 (-2.1%)
30.04 (-26.3%)
Appendix
Table H3: Measured cost and final quality of the two write schedules at 64 experts, a=1 . Aggregated-schedule quality is taken from Tables and . View-23 PSNR is the final target view; min/1k steps is wall-clock minutes per thousand optimizer steps. Cost measurements use one H100 80GB with the same 256×256 shape, 8 layers, 12 views, batch 32, and no activation checkpointing. Bold marks the best value per column.
Figure I1: Quantitative reuse of shared experts by redundant tokens (GSO, 1,019 objects, 8.6 M token pairs). All panels compare the write routing of an earlier-view token at one layer with that of a later-view token at a deeper layer, over 28 layer pairs and 8 heads. (a) Same-expert rate versus 3D Euclidean distance for e64-p1; objects are scaled to a longest bounding-box side of 1.8. Filled: geometry-matched pairs; open: controls on the object surface from the same later view, at distance ≥0.08 , grouped by distance. Dashed: pooled control rate; dotted: uniform routing ( 1/64 ). (b) Same-surface/control ratio for each layer pair in e64-p1, maximum over the 8 heads. Colors are logarithmic; only ratios ≥2 are labeled. (c) Same-surface and control rates for four sharing partitions, with their ratios. Rates use all 28 layer pairs, including pairs in different pools that cannot share a physical expert.
Figure I2: A redundant token pair reuses one physical expert across depths (uTTT-MoE-e64-p1, one GSO object). (a) Two source views of the object; the red squares mark two tokens that see the same surface point (mean reprojection error 0.035 source pixels). (b) The same two tokens at the two depths compared: the write of the view-11 token at layer 2 and that of the view-18 token at layer 4. Red arrows mark these writes; both are routed to the same expert of the 64-expert shared pool. Compared at the same layer and head and averaged over the 8 layers and 8 heads, their keys and values have cosine similarities of 0.992 and 0.987 , a mean representation similarity of 0.99 .
Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive-level train-view forward, train-view backward, and causal query-view rules into the full graph-level TTT computation, including the fast-weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning-rate initialization, weight decay, and a single-layer nonlinearity improve performance, while MSE and inner-product losses perform similarly. Deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections and gating provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.
Bohao Tang, Zhen Qin, Yuqi Pan +3
Shanghai Jiao Tong University · Shanghai Innovation Institute · ByteDance Seed
Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While recent in-place TTT methods make fast-weight adaptation possible for pretrained LLMs without redesigning the backbone, they leave a central question unresolved: what should each test-time write store? Existing recipes train the fast weight to match a learned local value proxy but they are not directly tied to the self-supervised next-token prediction signal. We introduce Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state. This makes each local write follow the same causal computation that supports next-token prediction: the value target is a pointwise linear projection of a single next-position contextual state. On RULER Full-13, averaged over 4k to 32k contexts, TTT-NTP is the only method that consistently improves the released backbone across four models spanning three families and a 0.6-8B size range, by 3.9 points on Llama-3.1-8B, 3.0 on Mistral-7B-v0.3, 4.1 on Qwen3-4B, and 2.9 on Qwen3-0.6B. On the real-world LongBench-v2 long-document QA benchmark, TTT-NTP improves over the base model by 5.6 points on Llama-3.1-8B and 3.7 on Mistral-7B-v0.3, while preserving commonsense and knowledge performance. Our code is publicly available at https://github.com/yancyou/TTT-NTP.
Test-time training (TTT) adapts an LLM during generation by reading and updating request-owned state, such as fast weights, low-rank deltas, or streaming learner state. This breaks batched LLM serving, which assumes shared static weights: serial execution is correct but slow, while naive batching can corrupt request state. We formulate this problem as read-write TTT serving and present RW-TTT , which tags each decode step with its owner, version, and READ/WRITE effect, batches only compatible phases, and commits updates only to the owner. On one GPU with eight fast-weight InPlace-TTT streams, RW-TTT reaches 274.61 aggregate tok/s, 9.31x over sequential serving and 3.44x over per-stream replicas under the same memory budget. It preserves behavior on RULER, a long-context benchmark, and passes owner/version checks.
Jian Yang, Zhizhuo Kou, Yao Tian +4
1The Hong Kong University of Science and Technology · 2The Chinese University of Hong Kong · 3National University of Singapore