cs.LGOct 15, 2024

Learning in the Recurrent State: Gradient Descent with Linear Recurrent Networks

Authors: Yudou Tian, Neeraj Mohan Sushma, Harshvardhan Mestha, Nicolo Colombo, David Kappel, Anand Subramoney

Organizations: Center for Cognitive Interaction Technology CITEC, Universität Bielefeld, Germany · Department of Electrical and Electronics Engineering, Birla Institute of Technology and Science Pilani, India · Department of Computer Science, Royal Holloway, University of London, United Kingdom

Abstract

In-context learning lets a sequence model adapt to a new task from examples in its input. A prominent line of work shows how self-attention can be constructed to implement gradient descent on a linear predictor fit to the in-context examples during the forward pass. State-space models (SSMs) and other linear recurrent networks (LRNNs) model sequences at linear time cost, but it is unclear how their recurrent update could carry out the same in-context gradient descent. We introduce Gradient-based Recurrent In-context Learner (GRIL), a diagonal LRNN that factorizes a supervised gradient step into a short-window cross-product write and a multiplicative readout of the next query. For linear regression, this construction accumulates the context gradient in a matrix state and applies it in a single forward pass, with O(f2)O(f^2) learned degrees of freedom. The same design extends to multi-step updates and cross-entropy classification, with a limited MLP-based extension to non-linear regression. We show empirically that trained GRILs recover the behavior and parameters analytically predicted by the construction on synthetic ICL tasks. Furthermore, the same architecture can be extended and trained on general-purpose benchmarks, including Long Range Arena, language modeling and associative recall. Together, these results establish windowed cross-product self-attention as a concrete inductive bias that lets LRNNs learn in context through gradient-descent-like updates, while remaining trainable on general-purpose tasks.

Explore similar work

CardsList
  1. In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective

    May 25, 2026Mingchen Li, Jiatan Huang, Chuxu Zhang +2In-Context LearningSearch-Augmented Reasoning

  2. In-Context Learning as Implicit Policy Gradient

    Jul 25, 2026Masahiro Kaneko, Timothy BaldwinIn-Context LearningIn-Context