cs.LGOct 7, 2026

Phase-HDC: Replacing Optimizer History with Gradient Thresholds in Discrete Phase Learning

Authors: Ahmed Nebli

Organizations: Mathalyse Research

Abstract

Training a compact model often needs far more memory than storing it, because the optimizer keeps its own records of past gradients. For a hyperdimensional classifier whose learned parameters are low-bit angles, which we call a \emph{phase memory}, these records take several times more memory than the model itself. We ask whether such a model can be trained while storing nothing but the model. The proposed method, Phase-HDC, turns each stored angle by at most one step per update, against the sign of its current gradient, and only when that gradient is large enough. We show that this simple rule is the exact solution of a first-order loss model in which every changed parameter pays a fixed cost. When everything except the update rule is held fixed, Phase-HDC matches the accuracy of Adam with 6-bit moments while storing three times less. Across eleven image, tabular, and text datasets, it stores 16--23×\times less than standard float32 Adam and 4--6×\times less than 8-bit Adam. The price is an average loss of about five accuracy points against float32 Adam, while Phase-HDC is more accurate than 8-bit Adam on six of the eleven datasets, including byte-level text prediction, where 8-bit Adam collapses. Instrumented training runs explain these outcomes. Once parameters must sit on a discrete grid, Adam's moments mainly decide whether a parameter moves at all, a decision that a threshold on the current gradient can make without memory, and coarse quantization of the moments breaks this decision for inputs that the data rarely contain. The storage savings are logical state rather than measured hardware memory.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 11, 2026cs.LG

Gefen: Optimized Stochastic Optimizer

AdamW is a default optimizer for deep learning, but its moment states add two parameter-sized buffers to training memory, increasing the cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook. Gefen reduces AdamW's optimizer memory footprint by up to 8x while maintaining performance, saving 6.5 GiB per billion parameters. Prior work shares second moments across parameters grouped along the Hessian's block-diagonal structure, but relies on hand-specified architectural rules and leaves unexplained why such grouping works. We prove that large mixed Hessian entries constrain the ratio of squared gradients toward one, explaining why shared second moments are accurate when the squared gradients they pool are similar. The Hessian need not be computed: its block structure is inherited by squared gradients, allowing blocks to be found directly. Gefen therefore infers block structure from initial squared gradients, requiring no architecture-specific metadata or user-tuned hyperparameters beyond AdamW defaults. Gefen learns an exact histogram-based dynamic-programming quantization codebook and reuses the blocks for first-moment scaling. Across diverse pretraining experiments, Gefen achieves the lowest peak optimizer memory among compared methods that maintain AdamW-level performance. In single-machine and distributed training, the reduced footprint enables larger microbatches and substantially improves throughput over AdamW, making Gefen a drop-in replacement that can train larger models or use larger global batch sizes. We provide the complete Python implementation, including fused CUDA kernels at https://github.com/ndvbd/Gefen
Aug 23, 2026cs.LG

Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers

Optimizer-state quantization is commonly designed for Adam's dense, parameter-aligned first- and second-moment arrays. This abstraction breaks for memory-efficient optimizers, whose states may be factored, confidence-modulated, or maintained in a projected space, so similar reconstruction error can produce different update error. We formulate optimizer-state quantization as a joint problem over representation, topology, and update semantics. We then introduce Adaptive Log-Space (AL) quantization for non-negative states. AL fits each block's observed nonzero logarithmic interval and reserves a separate code for exact zero, enforcing q=0⇔x=0q = 0 \Leftrightarrow x = 0; signed momentum and state precision remain independently selectable. Controlled probes show that adaptive ranges reduce update error and temporal drift, exact-zero reservation preserves dormant states, and state topology constrains useful block granularity. End-to-end language-model training evaluates the resulting policy across dense, factored, confidence, and projected optimizer states. On TinyLlama-1.1B, AL8 with uniform 8-bit momentum reaches 72.90 perplexity versus 73.54 for bitsandbytes 8-bit AdamW, with comparable optimizer-state storage and higher throughput. CAME matches reference-level final perplexity across three seeds when its non-negative states use AL16, while a semantic grouping-and-protection policy closes most of quantized Adafactor's 100K-step late-loss gap. These results make state topology and update semantics first-class design constraints for optimizer quantization.
May 11, 2026cs.LG

PowerStep: Memory-Efficient Adaptive Optimization via ℓp\ell_p-Norm Steepest Descent

Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial memory overhead. We introduce PowerStep, a memory-efficient optimizer that achieves coordinate-wise adaptivity without storing second-moment statistics. Motivated by ℓp\ell_p-norm steepest descent, PowerStep applies a signed-power transform directly to one momentum buffer. We establish a finite-horizon stationarity bound for exact, unregularized updates, with an O(1/T)O(1/\sqrt{T}) term and a noise-dependent residual. Experiments on Transformers from 124M to 235B parameters show competitive validation quality while halving fp32\texttt{fp32} optimizer-state memory relative to AdamW. Combined with uniform int8\texttt{int8} quantization, PowerStep remains numerically stable and reduces optimizer-state memory by ∼8×\sim8\times compared to fp32\texttt{fp32} AdamW. PowerStep thus provides a simple, memory-efficient alternative for large-scale training.