SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces
Organizations: Beijing Normal University · Singapore Management University · Xiaomi Inc. · Beijing Key Laboratory of Artificial Intelligence for Education · Engineering Research Center of Intelligent Technology and Educational Application (MOE)
Abstract
Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization through a carefully designed dual low-dimensionality strategy: (i) multi-query forward-difference gradient estimation in periodically refreshed random subspaces to mitigate noise amplification in moment buffers, and (ii) Adam updates with periodic restarts performed directly in low-dimensional space rather than full-parameter space. In experiments, this dual design retains memory overhead comparable to momentum-free ZO methods while achieving stronger optimization performance than the evaluated ZO baselines. Theoretically, in the exact-directional limit, -query averaging preserves conditional unbiasedness, while the coefficient estimator's covariance and mean-squared error, as well as query-induced second-moment inflation, scale exactly as . Extensive experiments across SuperGLUE with models from 1.3B to 32B parameters under both full fine-tuning and LoRA schemes demonstrate consistent improvements over competing ZO methods. SubZero+ significantly narrows the performance gap with first-order optimization while preserving ZO's inference-time memory efficiency.
Figures & tables
| Model | Method | SST-2 | RTE | BoolQ | WIC | MultiRC | COPA | ReCoRD | SQuAD | DROP | Avg. |
| Task type | —— Classification —— | — Multi-choice — | — Generation — | ||||||||
| OPT-1.3B | Zero-shot | 53.5 | 53.4 | 45.5 | 57.5 | 45.4 | 75.0 | 70.5 | 27.2 | 11.1 | 48.8 |
| FO (Adam) | 92.0 | 78.0 | 75.8 | 66.0 | 70.9 | 77.0 | 71.1 | 81.7 | 31.1 | 71.5 | |
| MeZO | 90.2 | 64.3 | 65.0 | 58.3 | 59.7 | 74.0 | 71.0 | 76.4 | 22.5 | 64.6 | |
| SubZero | 90.9 | 67.5 | 65.4 | 58.1 | 59.5 | 75.0 | 70.9 | 76.4 | 25.4 | 65.5 | |
| LOZO | 91.1 | 67.2 | 66.7 | 58.9 | 61.4 | 75.0 | 72.1 | 74.5 | 25.5 | 65.8 | |
| Model | Method | SST-2 | RTE | BoolQ | WIC | MultiRC | COPA | ReCoRD | SQuAD | DROP | Avg. |
| Task type | —— Classification —— | — Multi-choice — | — Generation — | ||||||||
| OPT-13B | MeZO | 93.6 | 62.1 | 69.1 | 56.9 | 57.8 | 86.0 | 81.4 | 83.0 | 28.7 | 68.7 |
| SubZero | 92.2 | 62.5 | 66.7 | 56.7 | 58.3 | 87.0 | 81.8 | 83.7 | 28.3 | 68.6 | |
| LOZO | 92.8 | 61.4 | 64.7 | 55.5 | 60.9 | 85.0 | 82.1 | 83.7 | 29.5 | 68.4 | |
| TeZO | 92.1 | 65.3 | 66.0 | 59.4 | 59.5 | 85.0 | 81.0 | 80.7 | 29.0 | 68.7 | |
| ZO-Muon | 92.9 | 65.3 | 65.9 | 54.9 | 60.0 | 84.0 | 81.3 | 79.4 | 27.6 | 67.9 | |
| Method | SST-2 | RTE | BoolQ | WIC | SQuAD | Avg. |
| Zero-shot | 61.5 | 87.0 | 84.3 | 61.8 | 85.5 | 76.0 |
| MeZO | 87.5 | 90.3 | 86.7 | 65.7 | 87.9 | 83.6 |
| MeZO-LoRA | 62.0 | 87.4 | 84.8 | 64.7 | 84.7 | 76.7 |
| SubZero | 89.1 | 90.3 | 87.7 | 68.8 | 89.6 | 85.1 |
| LOZO | 89.4 | 91.0 | 87.0 | 69.3 | 90.5 | 85.4 |
| ZO-Muon | 90.5 | 91.3 | 88.6 | 71.0 | 90.5 | 86.4 |
| FT | LoRA | |||||||
| Task | MeZO | SubZero | SubZero+ | % | MeZO | SubZero | SubZero+ | % |
| RTE | 30.4 | 30.5 | 30.7 | +1.0 | 30.4 | 30.4 | 30.4 | +0.0 |
| BoolQ | 34.1 | 34.2 | 34.5 | +1.2 | 34.1 | 34.1 | 34.1 | 0.0 |
| SQuAD | 30.8 | 30.9 | 31.1 | +1.0 | 30.8 | 30.8 | 30.8 | +0.0 |
| DROP | 50.4 | 50.5 | 50.8 | +0.8 | 50.4 | 50.4 | 50.4 | 0.0 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | 1 | 5 | 10 | 20 | 30 | 50 | 70 | 80 | 100 | 150 | 200 | |
| SST-2 | 50.9 | 52.6 | 56.7 | 69.8 | 88.3 | 90.9 | 91.5 | 91.6 | 92.2 | 89.9 | 89.7 | +41.3 |
| RTE | 47.7 | 52.7 | 49.8 | 57.0 | 62.8 | 63.9 | 58.1 | 60.3 | 57.0 | 61.4 | 61.0 | +9.3 |
| MultiRC | 55.5 | 55.5 | 52.3 | 59.8 | 57.6 | 60.0 | 57.6 | 59.3 | 59.6 | 58.7 | 61.1 | +4.1 |
| Avg. | 51.4 | 53.6 | 52.9 | 62.2 | 69.6 | 71.6 | 69.3 | 70.4 | 69.6 | 70.0 | 70.6 | +18.2 |
| Method | Configuration |
| MeZO | learning rate ; 20,000 steps (2 forward passes/step) |
| SubZero | learning rate ; rank (OPT-1.3B) or (others); update frequency ; 20,000 steps (2 forward passes/step) |
| LOZO | learning rate ; rank ; reuse interval ; 20,000 steps (2 forward passes/step) |
| TeZO | learning rate ; fixed rank ; 20,000 steps (2 forward passes/step) |
| ZO-Muon | learning rate ; projection rank ; ; projection interval 100; 5 Newton–Schulz iterations; 8,000 steps (4 perturbation queries + 1 baseline forward pass/step) |
| SubZero+ | learning rate ; rank (OPT-1.3B) or (others); update frequency ; ; 400 steps (99 perturbation queries + 1 baseline forward pass/step); Adam and |
| Method | Configuration |
| MeZO | learning rate ; 20,000 steps (2 forward passes/step) |
| SubZero | learning rate ; rank ; update frequency ; 20,000 steps (2 forward passes/step) |
| LOZO | learning rate ; rank ; reuse interval ; 20,000 steps (2 forward passes/step) |
| TeZO | learning rate ; fixed rank ; 20,000 steps (2 forward passes/step) |
| ZO-Muon | learning rate ; projection rank (selected: 16); ; projection interval 100; 5 Newton–Schulz iterations; 8,000 steps (4 perturbation queries + 1 baseline forward pass/step) |
| SubZero+ | learning rate ; rank ; update frequency ; ; 400 steps (99 perturbation queries + 1 baseline forward pass/step); Adam and |