cs.LGOct 1, 2026

Universal interpolation for deep residual self-attention networks

Authors: Sibylle Marcotte, Joan Bruna

Organizations: Department of Computer Science New York University

Abstract

Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across layers, motivated by recent models such as the Looped Transformers. More specifically, we ask whether one can find a predefined finite set of parameters, each defining an attention block, such that the resulting finite set of transformations can map any collection of NN sequences of nn tokens to any other collection of NN sequences of nn tokens. Crucially, these transformations are \emph{fixed independently of the input and output} collections: only the order in which the blocks are applied, their signs, and their durations depend on the particular interpolation task. Our main result establishes it for residual softmax attention using only two frozen single-head blocks with Gaussian-initialized projection matrices. The result holds at both continuous and finite depth. We also characterize the restrictions imposed by causal masking and establish corresponding universal interpolation guarantees.

Explore similar work

CardsList
  1. Delta Attention Residuals

    May 13, 2026Cheng Luo, Zefan Cai, Junjie HuLayer-WiseCross-Layer Interactions

  2. Progressive Approximation in Deep Residual Networks: Theory and Validation

    Apr 27, 2026Wei Wang, Xiao-Yong Wei, Qing LiResidual NetworksDeep Learning

  3. Depth-Attention: Cross-Layer Value Mixing for Language Models

    Jun 3, 2026Boyi Zeng, Yiqin Hao, Zitong Wang +7Layer-WiseCross-Layer Interactions