Universal Test-Time Training
Organizations: University of Wisconsin–Madison · University of Michigan–Ann Arbor · Adobe
Abstract
Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.
Figures & tables
| Design | Reachable | Active | Memories | ||||
|---|---|---|---|---|---|---|---|
| TTT-Dense (LaCT) | |||||||
| Dense, fixed compute | |||||||
| Dense, fixed state | |||||||
| uTTT-Dense | |||||||
| TTT-MoE | |||||||
| Routed, groups |
| 124M \CT@row@color | 760M | |||||||||||
| Model | PTL | RULER | 4K | 8K | 16K | 32K \CT@row@color | PTL | RULER | 4K | 8K | 16K | 32K |
| Baselines | ||||||||||||
| Transformer † | 3.0869 | 16.52 | 23.82 | 20.24 | 13.10 | 8.93 \CT@row@color | 2.5978 | 34.16 | 41.70 | 37.95 | 33.36 | 23.62 |
| Transformer-SWA | 3.1243 | 14.50 | 26.93 | 14.73 | 9.94 | 6.41 \CT@row@color | 2.6325 | 22.27 | 45.23 | 22.64 | 12.67 | 8.57 |
| Gated DeltaNet-SWA | 3.0994 | 12.32 | 23.78 | 12.93 | 8.30 | 4.26 \CT@row@color | 2.6075 | 21.48 | 40.40 | 23.12 | 13.66 | 8.76 |
| DeltaNet-SWA | 3.1336 | 11.69 | 22.52 | 11.23 | 7.63 | 5.36 \CT@row@color | 2.6086 | 20.77 | 43.43 | 20.56 | 11.79 | 7.28 |
| Model | PSNR | View 23 PSNR | SSIM | LPIPS |
|---|---|---|---|---|
| MoE A: , top-1 | ||||
| uTTT-MoE-e64-p1 | 24.68 [24.51, 24.86] | 25.48 [25.27, 25.68] | 0.860 [0.855, 0.864] | 0.173 [0.168, 0.177] |
| uTTT-MoE-e32-p2 | 24.20 [24.03, 24.37] | 24.98 [24.78, 25.18] | 0.854 [0.850, 0.859] | 0.181 [0.177, 0.186] |
| uTTT-MoE-e16-p4 | 24.11 [23.94, 24.29] | 24.89 [24.69, 25.10] | 0.854 [0.849, 0.858] | 0.182 [0.177, 0.186] |
| uTTT-MoE-e8-p8 (TTT-MoE-e8) | 23.90 [23.73, 24.08] | 24.56 [24.36, 24.77] | 0.850 [0.845, 0.855] | 0.186 [0.181, 0.191] |
| Dense A: | ||||
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Asset | Use in this paper | License or access |
|---|---|---|
| Objaverse ( Deitke et al., 2023 ) | NVS training | ODC-By v1.0; per-object CC licenses, some non-commercial |
| Google Scanned Objects ( Downs et al., 2022 ) | NVS evaluation | CC BY 4.0 |
| DL3DV-10K ( Ling et al., 2024 ) | NVS training and evaluation | CC BY-NC 4.0; gated access |
| Long-Data-Collections ( Together AI, 2023 ) | LM training | Mixed sources, each under its own license |
| Books3 from The Pile ( Gao et al., 2020 ) | LM evaluation (Section ; Appendices , and ) | Copyrighted books; withdrawn from its original host in 2023 |
| RULER ( Hsieh et al., 2024 ) | Retrieval evaluation, via the LM Evaluation Harness ( Gao et al., 2024 ) | Apache-2.0 (harness: MIT); some needle tasks use Paul Graham essays as filler, the QA tasks use the two datasets below, the rest are synthetic |
| 124M | 760M | |
| Width / layers / attention heads | 768 / 12 / 12 | 1536 / 24 / 24 |
| Attention head dimension | 64 | |
| FFN (SwiGLU ( Shazeer, 2020 ) ) hidden width | 2,048 | 4,096 |
| Normalization | pre-norm RMSNorm ( Zhang & Sennrich, 2019 ) | |
| Vocabulary | 32,000; untied embeddings | |
| RoPE ( Su et al., 2024 ) | ||
| Non-emb. params | ||||
| Model | FFN width | State / layer | 124M | 760M |
| Attention references | ||||
| Transformer ( Vaswani et al., 2017 ) | 2,048 / 4,096 | KV cache | 84.95M | 679.6M |
| Transformer-SWA ( Beltagy et al., 2020 ) | 2,048 / 4,096 | — | 84.95M | 679.6M |
| Linear attention, matched in parameters | ||||
| DeltaNet-SWA ( Yang et al., 2024 ) | 2,048 / 4,096 | 85.10M | 680.6M | |
| LaCT, TTT-Dense | TTT-MoE | uTTT-MoE | |
| Heads width | |||
| Function | SwiGLU; initial : rank 32 | ||
| window attention’s, pre-RoPE; SiLU; L2, then RoPE on | |||
| Step size | per token, head, and matrix; base | ||
| Update | momentum, Muon ( Jordan et al., 2024 ) (5 steps), row renormalization | ||
| Momentum coefficient | per head | per chunk | per chunk |
| Model | FFN width | State / layer | Params |
| References | |||
| Transformer | 1,536 | KV cache ( tokens per view) | 27.76M |
| DeltaNet | 1,584 | , delta rule | 37.04M |
| Gated DeltaNet | 1,408 | , gated delta rule | 37.01M |
| Fast-weight models of Section | |||
| 64 total experts or total width 8 | 1,536 | per width unit | 37.04–37.05M |
| Width / layers / attention heads | 512 / 8 / 8 |
|---|---|
| Block | attention fast-weight read SwiGLU FFN (1,536) |
| Attention | within one view |
| Input | , patch 8: 1,024 tokens per view |
| Views | training 12 (11 posed + 11 target chunks); evaluation 24 |
| Fast-weight heads | |
| Unit of state | expert: weights, shared by heads; width unit: 8 experts |
| Balancing | PTL | RULER | PTL | RULER | PTL | RULER |
|---|---|---|---|---|---|---|
| Aux. loss | — | — | 2.5744 | 23.16 | — | — |
| Aux. loss | — | — | 2.5780 | 23.96 | 2.5805 | 24.19 |
| Aux. loss | 2.5783 | 21.27 | 2.5783 | 25.50 | 2.5791 | 26.07 |
| Aux. loss | — | — | 2.5793 | 20.57 | 2.5794 | 25.56 |
| Aux. loss | — | — | 2.5771 | 25.11 | — | — |
| Model | State (M) | PTL | RULER | 4K | 8K | 16K | 32K |
|---|---|---|---|---|---|---|---|
| uTTT-Dense, (FLOPs-matched) | 0.44 | 3.1290 | 10.64 | 20.27 | 12.23 | 6.46 | 3.61 |
| uTTT-Dense, | 0.88 | 3.1232 | 13.20 | 23.57 | 14.66 | 9.79 | 4.78 |
| uTTT-Dense, | 1.77 | 3.1249 | 13.41 | 21.65 | 15.12 | 10.69 | 6.20 |
| uTTT-Dense, | 2.65 | 3.1165 | 13.63 | 23.60 | 16.01 | 9.92 | 5.01 |
| TTT-Dense (Table ) | 5.31 | 3.1219 | 12.65 | 26.31 | 13.58 | 6.00 | 4.72 |
| Fixed-target ablation: visible prefix | Reset segments | ||||||||||
| Model | 4K | 8K | 16K | 32K | 48K | 64K | 120K | Seg. 1 | Seg. 2 | Seg. 3 | Seg. 4 |
| Transformer-SWA (32K) | 3.222 | 3.222 | 3.222 | 3.222 † | 3.222 † | 3.222 † | 3.222 † | 3.224 | 3.224 | 3.218 | 3.214 |
| Transformer-SWA (64K) | 3.220 | 3.220 | 3.220 | 3.220 | 3.220 | 3.220 † | 3.220 † | 3.223 | 3.222 | 3.216 | 3.212 |
| Transformer (32K) | 3.211 | 3.198 | 3.188 | 3.329 † | 4.920 † | 5.692 † | 5.838 † | 3.205 | 3.201 | 3.194 | 3.190 |
| Transformer (64K) | 3.226 | 3.212 | 3.201 | 3.193 | 3.194 | 3.208 † | 5.563 † | 3.218 | 3.215 | 3.208 | 3.204 |
| LaCT | 3.215 | 3.214 | 3.213 | 3.213 † | 3.216 † | 3.219 † | 3.230 † | 3.216 | 3.216 | 3.209 | 3.205 |
| Model | 4K | 8K | 16K | 32K | 64K |
|---|---|---|---|---|---|
| Transformer-SWA, 4K window (32K) | 29.1 | 15.6 | 9.8 | 6.0 | 3.4 † |
| Transformer-SWA, 4K window (64K) | 28.5 | 16.0 | 9.5 | 6.0 | 3.8 |
| Transformer (32K) | 30.9 | 26.4 | 23.0 | 13.8 | 1.4 † |
| Transformer (64K) | 20.5 | 10.7 | 11.4 | 8.4 | 5.8 |
| LaCT | 23.6 | 14.8 | 9.3 | 5.2 | 3.3 † |
| TTT-MoE | 28.7 | 13.3 | 6.4 | 4.2 | 3.3 † |
| 124M | 760M | |||
|---|---|---|---|---|
| Kernels switched off | tok/s | speed-up | tok/s | speed-up |
| none | 12,071 | – | 5,542 | – |
| all | 10,533 | 1.15 | 4,710 | 1.18 |
| routing permutation of , and learning rates | 11,282 | 1.07 | 5,088 | 1.09 |
| grouped GEMMs of the outer-gradient backward | 11,816 | 1.02 | 5,158 | 1.07 |
| pointwise second-order backward | 11,950 | 1.01 | 5,162 | 1.07 |
| Model | PSNR | View 23 PSNR | SSIM | LPIPS |
|---|---|---|---|---|
| Layer-private pools ( ) | ||||
| uTTT-Dense-w1-p8 (TTT-Dense) | 23.99 [23.81, 24.16] | 24.49 [24.29, 24.69] | 0.852 [0.847, 0.856] | 0.185 [0.180, 0.189] |
| uTTT-MoE-e4-p8 (TTT-MoE-e4) | 23.80 [23.62, 23.97] | 24.44 [24.24, 24.64] | 0.849 [0.845, 0.854] | 0.188 [0.183, 0.193] |
| uTTT-MoE-e8-p8 (TTT-MoE-e8) | 23.90 [23.73, 24.08] | 24.56 [24.36, 24.77] | 0.850 [0.845, 0.855] | 0.186 [0.181, 0.191] |
| MoE, universal pool, top-1 (C) | ||||
| uTTT-MoE-e8-p1 | 24.72 [24.54, 24.90] | 25.55 [25.34, 25.77] | 0.861 [0.856, 0.865] | 0.172 [0.168, 0.177] |
| Model | PSNR | View 23 PSNR | SSIM | LPIPS |
|---|---|---|---|---|
| Layer-private pools ( ) | ||||
| uTTT-Dense-w1-p8 (TTT-Dense) | 16.14 [15.91, 16.39] | 16.69 [16.35, 17.03] | 0.383 [0.365, 0.402] | 0.681 [0.673, 0.690] |
| uTTT-MoE-e4-p8 (TTT-MoE-e4) | 15.98 [15.74, 16.22] | 16.47 [16.14, 16.81] | 0.381 [0.362, 0.399] | 0.691 [0.682, 0.699] |
| uTTT-MoE-e8-p8 (TTT-MoE-e8) | 16.02 [15.78, 16.27] | 16.54 [16.20, 16.89] | 0.379 [0.361, 0.397] | 0.688 [0.679, 0.696] |
| MoE, universal pool, top-1 (C) | ||||
| uTTT-MoE-e8-p1 | 16.02 [15.78, 16.27] | 16.63 [16.30, 16.98] | 0.382 [0.364, 0.401] | 0.686 [0.677, 0.694] |
| Schedule | GSO V23 PSNR | DL3DV V23 PSNR | k patch tok/s/GPU | min/1k steps | Peak GiB |
|---|---|---|---|---|---|
| Sequential schedule | 25.58 | 16.80 | 100.2 | 119.9 | 40.75 |
| Aggregated schedule | 25.48 | 16.79 | 102.3 (+2.1%) | 117.5 (-2.1%) | 30.04 (-26.3%) |