Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.
Figures & tables
Environment
AD FLOPs / Action
RAD FLOPs / Action
Compute Ratio ↓
Bandit
59.8 M (short)
43.0 M
0.72 ×
263.4 M (long)
0.16 ×
Darkroom
144.8 M
27.8 M
0.19 ×
Dark Key-to-Door
203.4 M
85.9 M
0.42 ×
Meta-World
1.84 G
130.6 M
0.07 ×
Table 1: Inference FLOPs per action for each sequence.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Lpre=Ext−k:t∼D[∥dψ(qω(xt−k:t))−y∥F2]
Appendix
Algorithm 2 RAD Training
Hyperparameter
Delayed-Bandit
DR
DKTD
Meta-World
AD baseline
Policy model dimension
64
64
64
64
Policy Transformer layers
4
4
4
4
Policy attention heads
4
4
4
8
Policy feed-forward dimension
256
256
256
256
Attention / residual dropout
0.1
0.1
0.1
0.1
Appendix
Table 2: AD and RAD hyperparameters. RAD uses the same policy architecture as AD. Context windows and capacities are measured in tokens unless otherwise specified (1 transition = 3 tokens).
Hyperparameter
DR
DKTD
Meta-World
Parallel streams per learner
100
100
100
Rollout steps per stream
80
50
100
Optimization minibatch size
40
100
200
Epochs per rollout update
20
10
20
Total Timesteps
100,000
100,000
1,000,000
Learning rate
3×10−4
3×10−4
3×10−4
Appendix
Table 3: PPO source-algorithm configuration.
Hyperparameter
Delayed-Bandit
DR
DKTD
Meta-World
AD policy training
Training steps
50k / 100k a
50k
50k
50k
Configured batch size
64
512
512
256
Peak learning rate
3×10−4
3×10−4
3×10−4
3×10−4
Warmup steps
500
1,000
1,000
5,000
RAD compression pretraining
Appendix
Table 4: AD/RAD training hyperparameters
Environment
Start step
Max. c
Short
Medium
Long
Very long
DR
0
1
0.60
0.35
0.05
0.00
DR
30k
3
0.35
0.40
0.20
0.05
DR
50k
6
0.25
0.30
0.30
0.15
DR
75k
∞
0.25
0.25
0.25
0.25
DKTD
0
2
0.30
0.55
0.15
0.00
DKTD
25k
4
0.20
0.45
0.30
0.05
Appendix
Table 5: RAD curriculum. Probabilities correspond to short, medium, long, and very-long compression-count categories.