cs.LGAug 16, 2026

SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces

Authors: Ziming Yu, Shuyao Xiao, Xingyu Zhao, Sike Wang, Pan Zhou, Peiyu Zang, Xiangda Yan, Yongjie Yang, +1 more

Organizations: Beijing Normal University · Singapore Management University · Xiaomi Inc. · Beijing Key Laboratory of Artificial Intelligence for Education · Engineering Research Center of Intelligent Technology and Educational Application (MOE)

Abstract

Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization through a carefully designed dual low-dimensionality strategy: (i) multi-query forward-difference gradient estimation in periodically refreshed random subspaces to mitigate noise amplification in moment buffers, and (ii) Adam updates with periodic restarts performed directly in low-dimensional space rather than full-parameter space. In experiments, this dual design retains memory overhead comparable to momentum-free ZO methods while achieving stronger optimization performance than the evaluated ZO baselines. Theoretically, in the exact-directional limit, KK-query averaging preserves conditional unbiasedness, while the coefficient estimator's covariance and mean-squared error, as well as query-induced second-moment inflation, scale exactly as 1/K1/K. Extensive experiments across SuperGLUE with models from 1.3B to 32B parameters under both full fine-tuning and LoRA schemes demonstrate consistent improvements over competing ZO methods. SubZero+ significantly narrows the performance gap with first-order optimization while preserving ZO's inference-time memory efficiency.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On Adaptivity in Zeroth-Order Optimization

    May 5, 2026Hassan Dbouk, Nidham Gazagnadou, Matthias Reisser +1Zeroth-Order OptimizationStochastic Gradient Descent

  2. ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces

    Jul 1, 2026Xun Dong, Yibo Xu, Naigang Wang +3Zeroth-Order OptimizationZeroth-Order

  3. Universally Empowering Zeroth-Order Optimization via Adaptive Layer-wise Sampling

    Apr 20, 2026Fei Wang, Li Shen, Liang Ding +3Zeroth-Order OptimizationZeroth-Order