cs.LGOct 3, 2026

Clean: Second-order LLM Training at Linear Memory Cost via Nyström Sketching

Authors: Beheshteh T. Rakhshan, Sahar Rajabi, Maziar Sargordi Shikai Fang, Guillaume Rabusseau, Sirisha Rambhatla

Organizations: Mila & DIRO, Université de Montréal · Critical ML, University of Waterloo · Independent Researcher · Zhejiang University

Abstract

Training large language models (LLMs) entails a fundamental trade-off: memory-efficient optimizers such as Adam discard cross-parameter curvature, whereas full-curvature methods such as SOAP can accelerate convergence at prohibitive memory costs. We introduce Clean, a memory-efficient and full-curvature optimizer designed to resolve this bottleneck. Clean leverages the randomized Nystrom method to accurately approximate the left and right preconditioners in SOAP, and to reduce the optimizer's memory complexity from quadratic to linear in terms of model dimensions. We subsequently reintegrate the off-subspace components to capture curvature information beyond the low-rank approximation, preserving rich curvature at minimal memory cost. We further propose Q-Clean, a low-precision variant that aggressively compresses optimizer states. Q-Clean reduces optimizer memory consumption by \textbf{over 50%} compared to Muon when pre-training a LLaMA-1.3B architecture, all while maintaining strong and competitive predictive performance. Notably, Clean operates with a smaller optimizer-state footprint than standard AdamW while reaching AdamW's final performance \textbf{26% faster} in wall-clock time. Furthermore, our methods uniquely enable the pre-training of a 13B-parameter model on a single 80GB GPU, providing a scalable, efficient, and accessible approach to large-scale model optimization.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers

    May 9, 2026Aditya RanganathDeep Learning OptimizationMemory-Efficient Optimization

  2. Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling

    May 30, 2026Qiao Xiao, Boqian Wu, Patrik Okanovic +6Language Model PretrainingEfficient Language Model Training

  3. SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces

    Aug 16, 2026Ziming Yu, Shuyao Xiao, Xingyu Zhao +6Memory-Efficient Fine-TuningZeroth-Order Optimization