cs.LGOct 7, 2026

Lightweight and Versatile Learned Optimization by Recombination of Gradient History

Authors: Minyoung Choi, Dalta Imam Maulana, Wanyeong Jung

Organizations: KAIST Daejeon, Republic of Korea

Abstract

This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans. The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters. Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible. A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Gefen: Optimized Stochastic Optimizer

    Jun 11, 2026Nadav Benedek, Tomer Koren, Ohad FriedScaling Laws

  2. Efficient Long-Horizon Learning for Learned Optimization

    Jul 7, 2026Xiaolong Huang, Benjamin Thérien, James Harrison +1Meta-LearningAdaptive Optimizers

  3. FOGO: Forgetting-aware Orthogonalization Optimizer

    Jun 9, 2026Toan Nguyen, Yang Liu, Trung Le +2Adaptive OptimizersContinual Learning