cs.LGSep 28, 2026

GradLev: Token-Parallel Test-Time Training Via Costate Prediction

Authors: Bo Liu, Qiang Liu

Organizations: The University of Texas at Austin

Abstract

Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Modular TTT: Rethinking Test-Time Training as Composable Modules

    Aug 7, 2026Bohao Tang, Zhen Qin, Yuqi Pan +3Test-Time TrainingSequence Modeling

  2. Test-Time Training with Next-Token Prediction

    Jun 19, 2026Xuan Ouyang, Zefan Cai, Junjie HuTest-Time TrainingNext-Token Prediction

  3. RW-TTT: Batched Serving for Request-Owned Test-Time Training State

    May 27, 2026Jian Yang, Zhizhuo Kou, Yao Tian +4Large Language Model ServingTime-To-First-Token