cs.LGJul 26, 2026

A Trust-region Framework for Moment Estimation

Authors: Oluwasegun A. Somefun

Abstract

In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc{Adam}, in stochastic gradient optimization. Specifically, the magnitude of the update step associated with each individual parameter is constrained by a finite-order pp-moment trust-region, with p1p\ge1. The resulting derivation leads to a family of learning-rate mechanisms based on second-moment estimation and normalized pp-th-moment estimation. For p=4p=4, this involves kurtosis estimation. Subsequent derivations provide a unified interpretation of moment-estimation-based normalization, learning-rate scheduling, momentum as a spectral first-order lowpass regularization, and operator-level spectral-norm normalization within a common trust-region framework. Preliminary experiments on GPT2-124M trained on FineWeb-Edu and TinyStories suggest that the fourth-moment realization provides its greatest benefit when trust-region constraints are weak. As progressively stronger trust-region controls are introduced, the second-moment realization becomes increasingly competitive, often achieving slightly lower validation loss than its corresponding fourth-moment realization.

Explore similar work

CardsList
  1. Adaptive Momentum and Nonlinear Damping for Neural Network Training

    Jan 30, 2026Aikaterini Karoni, Rajit Rajpal, Benedict Leimkuhler +1Adam