cs.AIOct 1, 2026

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

Authors: Jude Waide, Robert Lieck

Organizations: Department of Computer Science, University of Durham, Durham, United Kingdom.

Abstract

Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning What to Remember: Test-Time Training via Context Distillation

    Aug 3, 2026Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4Efficient Long-Context InferenceTest-Time Training

  2. Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer

    May 7, 2026Yongyi Wang, Hanyu Liu, Lingfeng Li +6Transformer ArchitecturesSequence Modeling

  3. Modular TTT: Rethinking Test-Time Training as Composable Modules

    Aug 7, 2026Bohao Tang, Zhen Qin, Yuqi Pan +3Test-Time TrainingSequence Modeling