Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.
Figures & tables
Environment
AD FLOPs / Action
RAD FLOPs / Action
Compute Ratio ↓
Bandit
59.8 M (short)
43.0 M
0.72 ×
263.4 M (long)
0.16 ×
Darkroom
144.8 M
27.8 M
0.19 ×
Dark Key-to-Door
203.4 M
85.9 M
0.42 ×
Meta-World
1.84 G
130.6 M
0.07 ×
Table 1: Inference FLOPs per action for each sequence.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Lpre=Ext−k:t∼D[∥dψ(qω(xt−k:t))−y∥F2]
Appendix
Algorithm 2 RAD Training
Hyperparameter
Delayed-Bandit
DR
DKTD
Meta-World
AD baseline
Policy model dimension
64
64
64
64
Policy Transformer layers
4
4
4
4
Policy attention heads
4
4
4
8
Policy feed-forward dimension
256
256
256
256
Attention / residual dropout
0.1
0.1
0.1
0.1
Appendix
Table 2: AD and RAD hyperparameters. RAD uses the same policy architecture as AD. Context windows and capacities are measured in tokens unless otherwise specified (1 transition = 3 tokens).
Hyperparameter
DR
DKTD
Meta-World
Parallel streams per learner
100
100
100
Rollout steps per stream
80
50
100
Optimization minibatch size
40
100
200
Epochs per rollout update
20
10
20
Total Timesteps
100,000
100,000
1,000,000
Learning rate
3×10−4
3×10−4
3×10−4
Appendix
Table 3: PPO source-algorithm configuration.
Hyperparameter
Delayed-Bandit
DR
DKTD
Meta-World
AD policy training
Training steps
50k / 100k a
50k
50k
50k
Configured batch size
64
512
512
256
Peak learning rate
3×10−4
3×10−4
3×10−4
3×10−4
Warmup steps
500
1,000
1,000
5,000
RAD compression pretraining
Appendix
Table 4: AD/RAD training hyperparameters
Environment
Start step
Max. c
Short
Medium
Long
Very long
DR
0
1
0.60
0.35
0.05
0.00
DR
30k
3
0.35
0.40
0.20
0.05
DR
50k
6
0.25
0.30
0.30
0.15
DR
75k
∞
0.25
0.25
0.25
0.25
DKTD
0
2
0.30
0.55
0.15
0.00
DKTD
25k
4
0.20
0.45
0.30
0.05
Appendix
Table 5: RAD curriculum. Probabilities correspond to short, medium, long, and very-long compression-count categories.
Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent advancements have focused on reward engineering or data synthesis, few studies exploit the model's intrinsic representation characteristics to guide the training process. In this paper, we first observe the presence of high-magnitude activations within the query and key vectors when processing long contexts. Drawing inspiration from model quantization -- which establishes the criticality of such high-magnitude activations -- and the insight that long-context reasoning inherently exhibits a sparse structure, we hypothesize that these weights serve as the pivotal drivers for effective model optimization. Based on this insight, we propose LongAct, a strategy that shifts from uniform to saliency-guided sparse updates. By selectively updating only the weights associated with these significant activations, LongAct achieves an approximate 8% improvement on LongBench v2 and enhances generalization on the RULER benchmark. Furthermore, our method exhibits remarkable universality, consistently boosting performance across diverse RL algorithms such as GRPO and DAPO. Extensive ablation studies suggest that focusing on these salient features is key to unlocking long-context potential.
Bowen Ping, Zijun Chen, Tingfeng Hui +4
1Peking University · 2Shanghai Jiao Tong University · 3Beijing University of Posts and Telecommunications
Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruction or online adaptation objectives without considering the future utility of retained information. In this work, we propose \textbf{T}est-\textbf{T}ime \textbf{C}ontext \textbf{D}istillation (TTCD), a TTT framework that introduces a self-supervised objective for allocating limited memory capacity for future use. Specifically, TTCD uses a long-window teacher to supervise the fast weights of a short-window student, where the hidden-state discrepancy between them offers a dense, self-supervised signal guiding the model to memorize the contextual information crucial for future token predictions. We focus on an in-place variant: In-Place TTCD (IP-TTCD), which uses the existing MLP parameters as the fast weights. Experiments on long-context language modeling tasks show IP-TTCD consistently outperforms DeltaNet, Gated DeltaNet, sliding-window attention, and TTT when pre-trained from scratch. Furthermore, IP-TTCD allows pre-trained transformer models to adapt their parameters during inference through continual pre-training, gaining long-context capabilities with only a lightweight architectural augmentation. Our results position TTCD as a step toward architectural continual learning.
Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4
Princeton University · UC Santa Cruz · Carnegie Mellon University +1
Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the history retained between model invocations. Learning what to retain typically relies on proximal policy optimization (PPO) with final task rewards, but sparse rewards provide little guidance for individual memory updates. This limitation motivates on-policy distillation (OPD), which supplies dense teacher supervision on student rollouts. For such supervision to be valid, the teacher must evaluate each sampled action under the same state in which it was generated. However, the context rewriting performed during memory compression can break this alignment. When sampled responses are retained and re-encoded for later invocations, flattening the interaction into a persistent history may cause the teacher to score the action under a state that the student never visited during rollout. The action therefore remains on-policy by provenance, but not necessarily by state. We therefore propose Memory-Aligned On-Policy Distillation (MemOPD). MemOPD records the inputs and sampled outputs of each model invocation, restores its original token positions and causal visibility, and packs the reconstructed invocations for efficient teacher scoring. The teacher provides full-vocabulary supervision at the sampled action positions, while PPO preserves the final task objective. Experiments verify state alignment across several context updates and show that it improves F1 by 7.0% over persistent-history teacher scoring in a matched control. Overall, MemOPD-3B improves F1 over PPO by up to 416.2%, while packing yields up to a 1.63x speedup in actor computation during training. The code for this work is publicly available at: https://github.com/TPssp/MemOPD.
Zhiyuan Liu, Tinghong Ye, Chenghao Liu +2
School of Advanced Manufacturing and Robotics, Peking University · Zhejiang University · Harbin Institute of Technology