cs.LGSep 28, 2026

Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation

Authors: Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao

Organizations: The Chinese University of Hong Kong, Shenzhen

Abstract

Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scalability in complex, long-horizon tasks. To address this bottleneck, we propose Recurrent Algorithm Distillation (RAD). RAD employs a dual-component architecture: a Compression Transformer that distills extended interaction histories into compact latent tokens, and an AD Transformer that auto-regressively generates actions using a hybrid context of these compressed memories and recent transitions. By maintaining a fixed-size latent buffer, RAD decouples the effective history length from computational complexity, functionally providing the model with a long-horizon memory. Empirical evaluations across diverse environments demonstrate that RAD matches the asymptotic performance of standard AD with significantly reduced context window sizes, offering a scalable solution for efficient in-context decision-making.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

    Apr 16, 2026Bowen Ping, Zijun Chen, Tingfeng Hui +4Efficient Long-Context InferenceLarge Language Model Reinforcement Learning

  2. Learning What to Remember: Test-Time Training via Context Distillation

    Aug 3, 2026Zixuan Wang, Xingyu Dang, Rui-Jie Zhu +4Efficient Long-Context InferenceTest-Time Training

  3. MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents

    Aug 7, 2026Zhiyuan Liu, Tinghong Ye, Chenghao Liu +2Token-Level SupervisionLong-Horizon Task Planning