cs.LGOct 4, 2026

Universal Test-Time Training

Authors: Zefan Cai, Qinzhe Hu, Ziqiao Ma, Hao Tan, Junjie Hu

Organizations: University of Wisconsin–Madison · University of Michigan–Ann Arbor · Adobe

Abstract

Recent Test-Time Training (TTT) architectures compress context into fast weights that are updated online and queried as memory. Existing TTT designs keep this memory private to each layer: it recurs only over time, and depth merely indexes L separate memories. We argue that memory ownership need not be tied to depth, and introduce Universal Test-Time Training (uTTT), in which all layers read and write one shared memory while retaining layer-specific backbone parameters. The shared memory thus recurs over two dimensions, time and depth, with chunks and layers as their units: a write by a deep layer in one chunk can be read by a shallow layer in the next. We instantiate this idea as uTTT-MoE and uTTT-Dense. uTTT-MoE routes each token head to a few experts in a pool shared by all layers; uTTT-Dense applies the whole shared memory at every layer without routing. In language modeling, uTTT-MoE reaches 15.5 and 27.9 RULER accuracy at 124M and 760M, 2.6 and 2.1 points above its layer-private counterpart at equal state and active compute, the highest among tested bounded-state models, with per-token loss matching or beating full attention. In novel view synthesis, sharing at fixed per-layer compute gains 0.92 dB in view-23 object PSNR in routed models and 0.76 dB in dense models.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Modular TTT: Rethinking Test-Time Training as Composable Modules

    Aug 7, 2026Bohao Tang, Zhen Qin, Yuqi Pan +3Test-Time TrainingSequence Modeling

  2. Test-Time Training with Next-Token Prediction

    Jun 19, 2026Xuan Ouyang, Zefan Cai, Junjie HuTest-Time TrainingNext-Token Prediction

  3. RW-TTT: Batched Serving for Request-Owned Test-Time Training State

    May 27, 2026Jian Yang, Zhizhuo Kou, Yao Tian +4Large Language Model ServingTime-To-First-Token