Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.
Figures & tables
Figure 1 : The Decision Titan architecture. A trajectory of returns-to-go, states, and actions is projected into a shared embedding space before having positional embeddings added. This sequence is then chunked into subsequences of k timesteps and n persistent memory tokens are appended to the front of each chunk. Following this, each chunk is then passed through L Titans MAL blocks, consisting of a TTT layer that is updated every b timesteps, followed by self-attention with a window size of 3k+n tokens.
Embedding
Learnt
Information
Global
No
Global position
Chunked
No
Local position
Axial Positional
Yes
Global and local position
Table 1 : Embedding Strategies
Figure 2 : The X-Maze environment. This example shows an instruction phase length of 5, a waiting phase length of 5, and a potential instruction count of 10 (an instruction can be any number from 0 to 9). Time flows from left to right in the diagram. An episode is broken into 3 phases: instruction, waiting, and repeat. There are three different possible environment encodings, which provide different information over the course of an episode.
Encoding
Description
Obs Dim
Single Signal
A single binary bit is appended to the observation, set to 1 only at the final timestep of the waiting phase to indicate the start of the repeat phase.
+1
One-hot Signal
A one-hot vector of dimension I is appended, encoding which instruction the agent should next repeat. This is provided at the last step of the waiting phase and every step of the repeat phase.
+I
One-hot Repeat
Identical to One-hot Signal, except the one-hot instruction id is additionally provided during each step of the instruction phase.
+I
Table 2 : X-Maze Encodings
Encoding
Chunked Embedding
Global Embedding
Axial Embedding
Single Signal
−5.6±1.81
−5.72±2.11
−5.76±1.53
One-hot Signal
−0.16±0.46
−0.20±0.49
−0.72±1.28
One-hot Repeat
−0.00±0.00
−0.00±0.00
−0.00±0.00
Table 3 : Mean Return of Embeddings/Encoding Combinations
Heads
[11,15]
[46,50]
[96,100]
1
0.57
0.61
1.00
2
0.42
0.38
0.81
4
0.35
0.37
0.71
Table 4 : Mean First 0 Episode Return
Figure 3 : Mean episode return of 5 DTi runs against training steps. Each run is evaluated every 250 batches by sampling 5 episodes in the same environment used to gather training data. In this case, the environment is X-Maze with a [11,15] waiting phase range with DTi using the axial embedding strategy and a single head in M . All runs rapidly achieve a mean return equivalent to random chance, before at varying points rapidly increasing to 0 once the long-term dependencies are learnt.
Embedding
[11,15]
[46,50]
[96,100]
Chunked Embedding
5
5
2
Global Embedding
5
5
1
Axial Embedding
5
5
5
Table 5 : Number of Successful Runs
Figure 4 : Average episode return of the repeat phase (last 10 environment steps) against waiting phase length, comparing different DTi versions in the X-Maze environment. Trained was performed on 3 different waiting phase length ranges of X-Maze: [11,15] , [46,50] , and [96,100] . 3 different embeddings are tested for each length, chunked embeddings (ce) , global embeddings (te) , and axial positional embeddings (ae) . The results show that DTi can generalise to lengths of at least 1.7x what it was trained on and that chunked embeddings consistently generalise to greater lengths than the other embedding strategies. Interestingly, a model succeeding at a given length does not mean it can succeed at a shorter length. Each data point is the average of at least 5 episodes.
Figure 5 : Visualisation of a successful run in X-Maze with a waiting phase of length 100. L refers to the loss L at time t , eta to the write strength ηt , alpha to the forgetting gate αt , s to the momentum st , and gamma to the momentum decay γt . L is [0,1] normalised, H X refers to head number X .
Figure 6 : Mean return across different waiting phase lengths. The red run is without momentum (nm) and blue is with. 2 runs are averaged in the no momentum scenario, whilst with momentum is the average of 5 runs, and each run is evaluated on 5 different seeds for each point. One of the runs without momentum was suboptimal due to under-training, explaining the slightly lower peak of the no-momentum case.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
MuJoCo
X-Maze
Training
num_batches
40001
-
batch_size
2∗∗
16
learning_rate
1×10−4
1×10−4
Model Architecture
num_layers
3
3
Appendix
Table 6 : Hyperparameters
Figure 7 : Mean episode return against training steps in T-Maze. The waiting phase length is 5, creating a dependency with a range of 6 timesteps. Tested models have a context length of 5, meaning they need a long-term memory in order to solve the environment. DTi solves the environment whilst DT does not, demonstrating that TTT memory works in RL.
Env
Dataset
DT
DTi (no mem)
DTi
Hopper
E
3899±383
3872±391
3880±479
M
2248±862
3379±491
3355±511
Halfcheetah
E
7900±1099
7072±152
7028±124.6
M
7380±1647
6713±207
6608±134
Walker2D
E
6015.8±472.5
5750±233
5848±228
M
5947±73
5899±136
5851±123.4
Appendix
Table 7 : Mean Return on Control Tasks
Figure 8 : Visualisation of a failed run with momentum in X-Maze with a waiting phase of length 100. No head learns to maintain momentum or to minimise the the forget gate.
Figure 9 : Visualisation of a successful run in X-Maze with a waiting phase of length 50. The same observations as the waiting phase length 100 case apply.
Figure 10 : Visualisation of a successful run of DTi without momentum in X-Maze with a waiting phase of length 50. The update rule does not contain st or γt . One head has learnt a small value for the forgetting gate.
Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruction or online adaptation objectives without considering the future utility of retained information. In this work, we propose \textbf{T}est-\textbf{T}ime \textbf{C}ontext \textbf{D}istillation (TTCD), a TTT framework that introduces a self-supervised objective for allocating limited memory capacity for future use. Specifically, TTCD uses a long-window teacher to supervise the fast weights of a short-window student, where the hidden-state discrepancy between them offers a dense, self-supervised signal guiding the model to memorize the contextual information crucial for future token predictions. We focus on an in-place variant: In-Place TTCD (IP-TTCD), which uses the existing MLP parameters as the fast weights. Experiments on long-context language modeling tasks show IP-TTCD consistently outperforms DeltaNet, Gated DeltaNet, sliding-window attention, and TTT when pre-trained from scratch. Furthermore, IP-TTCD allows pre-trained transformer models to adapt their parameters during inference through continual pre-training, gaining long-context capabilities with only a lightweight architectural augmentation. Our results position TTCD as a step toward architectural continual learning.
Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4
Princeton University · UC Santa Cruz · Carnegie Mellon University +1
Decision Transformer (DT) formulates offline reinforcement learning as autoregressive sequence modeling, achieving promising results by predicting actions from a sequence of Return-to-Go (RTG), state, and action tokens. However, RTG is a scalar that summarizes future rewards, containing far less information than typical state or action vectors, yet it consumes the same computational budget per token. Worse, the self-attention cost of Transformers grows quadratically with sequence length, so including RTG as a separate token adds unnecessary overhead. We propose SlimDT, which removes RTG from the autoregressive sequence. Instead, we inject RTG information into the state representations before the sequential modeling step, allowing the Transformer to process only a compact (state, action) sequence. This reduces the sequence length by one-third, directly improving inference efficiency. On the D4RL benchmark, SlimDT surpasses standard DT across various tasks and achieves performance comparable to existing state-of-the-art methods. Decoupling a sparse conditioning signal from an information-rich sequence thus yields both computational gains and higher task performance.
Yongyi Wang, Hanyu Liu, Lingfeng Li +6
School of Computer Science Peking University Beijing, China 100871
Test-time training (TTT) views sequence modeling as an online learning problem in which fast weights are updated by an internal learning rule. Despite the growing number of TTT variants, existing approaches typically hard-code each variant separately, which makes it difficult to design new TTT methods and to isolate the role of each component. To address this, we propose Modular TTT, a framework that represents the inner learner as a directed acyclic graph and exposes the fast-weight network, loss function, learning rate, weight decay, and normalization as explicit design dimensions. Modular TTT automatically composes primitive-level train-view forward, train-view backward, and causal query-view rules into the full graph-level TTT computation, including the fast-weight state transition. Using Modular TTT, we systematically ablate the components of TTT and find that small learning-rate initialization, weight decay, and a single-layer nonlinearity improve performance, while MSE and inner-product losses perform similarly. Deeper fast-weight networks and normalization tend to hurt performance because they induce excessively large activations, while residual connections and gating provide little measurable benefit. Guided by these findings, we train the best resulting variant as 410M- and 1.45B-parameter models on 100B tokens, and observe training loss and benchmark performance comparable to Gated DeltaNet.
Bohao Tang, Zhen Qin, Yuqi Pan +3
Shanghai Jiao Tong University · Shanghai Innovation Institute · ByteDance Seed