stat.MLSep 24, 2026

Low-Rank Friction for Memory-Efficient Transformer Pretraining

Authors: Rajit Rajpal, Benedict Leimkuhler

Organizations: School of Mathematics University of Edinburgh Edinburgh, UK EH9 3FD

Abstract

iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor ξ∈Rm×nξ\in\mathbb{R}^{m\times n} carries the same O(mn)\mathcal{O}(mn) memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor ξξ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). This reduces the friction memory footprint from O(mn)\mathcal{O}(mn) to O(m+n)\mathcal{O}(m+n) per layer, which approximately halves iKFAD's total optimiser state. Despite this reduction, R-iKFAD maintains parity in performance with iKFAD: experiments on GPT2-Nano, TinyViT, DistilBERT and GPT2-S confirm that it matches or exceeds iKFAD while nearly halving the memory footprint and remaining comparably robust to hyperparameters. We analyse the continuous-time dynamics in two damping regimes. For linear damping (γ>0γ>0) we prove exponential convergence under strong convexity. For γ=0γ=0, the preferred option in our experiments, the friction is generated entirely from past momentum and switches off as the momentum vanishes, so geometric convergence cannot be shown. We nonetheless prove convergence to the minimiser, together with matching upper and lower bounds on the energy: of order t−1t^{-1} when the regularisation scale εstabε_{\mathrm{stab}} is zero, and of order t−1/2t^{-1/2} when it is positive. To our knowledge this is the first convergence rate for a rank-1 factored optimiser in continuous time, and the first such result that does not require positive damping.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PowerStep: Memory-Efficient Adaptive Optimization via ℓp\ell_p-Norm Steepest Descent

    May 11, 2026Yao Lu, Dengdong Fan, Shixun Zhang +1AdamNeural Network Training

  2. AdaPreLoRA: Adafactor Preconditioned Low-Rank Adaptation

    May 9, 2026Ziyun Liu, Fengmiao Bian, Jian-Feng CaiSpectral PreconditioningLow-Rank Structure

  3. Gefen: Optimized Stochastic Optimizer

    Jun 11, 2026Nadav Benedek, Tomer Koren, Ohad FriedScaling Laws