AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning
Authors: Yue Xie, Zhi Zheng, Yunpeng Ba, Xuyang Wu, Xialiang Tong, Zhichao Lu, Tao Zhong, Zhenkun Wang
Organizations: Southern University of Science and Technology · National University of Singapore · Huawei Technologies Ltd. · City University of Hong Kong
Zeroth-order (ZO) optimization offers a memory-efficient alternative for LLM fine-tuning by estimating updates only from forward evaluations of perturbed parameters, without backpropagation or activation storage. However, in billion-parameter LLMs, isotropic perturbations often waste many forward evaluations on weakly informative directions. To make these evaluations more informative, existing ZO methods restrict perturbations to low-dimensional subspaces. Yet the quality of these subspaces is critical: overly compressed or poorly maintained spaces can miss useful update directions. To obtain a high-quality subspace for ZO updates, this paper proposes AIM-ZO, a ZO fine-tuning method based on Activation-Informed Subspace Maintenance. AIM-ZO uses forward activations as local directional information and continuously integrates them into a broad, evolving subspace over training. To access broader gradient-relevant structure while keeping individual perturbations low-dimensional, AIM-ZO activates only a smaller set of shared and sampled directions, decoupling the maintained width from the active width. We evaluate AIM-ZO across 5 LLMs and 11 downstream tasks under matched forward-evaluation budgets; its six-task average exceeds the strongest fully evaluated ZO baseline by 1.26 percentage points on OPT-2.7B and MeZO by 2.85 percentage points on OPT-30B. Our code is available at https://github.com/EkkoXy/AIM-ZO
Figures & tables
Figure 1: Perturbation spaces. MeZO explores the full parameter space; AGZO constructs a subspace from current activations. AIM-ZO maintains a width- K subspace across training and activates a width- k subspace with h shared and k−h sampled directions per perturbation.
Method
RTE
BoolQ
SST-2
WiC
WSC
SQuAD
Avg.
OPT-2.7B
Zero-shot
55.23
52.87
56.65
54.86
36.54
26.92
47.18
MeZO
65.13±1.67
66.10±2.26
92.50±0.54
58.53±0.41
54.81±1.52
80.99±1.42
69.68
CurvZO
64.53±4.83
68.01±0.62
92.96±0.31
57.96±2.07
46.47±7.28
58.71±2.47
64.77
HiZOO
60.29±2.94
67.31±1.34
91.77±0.67
58.88±0.80
53.53±2.42
73.81±0.90
67.60
AGZO
64.40±3.10
66.26±1.46
92.59±0.51
57.02±1.73
47.76±2.22
29.82±5.54
59.64
Table 1: Multi-method comparison on OPT-2.7B and OPT-13B. Scores are accuracy (%) except SQuAD (F1). Avg. is the unweighted mean over six tasks. Best mean scores are bold; subscripts report available sample standard deviations. Evaluation details are in Appendix H .
OPT-30B
Method
RTE
BoolQ
SST-2
WiC
WSC
SQuAD
Avg.
Zero-shot
53.79
40.90
58.83
52.98
37.50
46.48
48.41
MeZO
63.25±1.79
70.74±1.17
91.08±0.36
55.96±1.81
58.08±2.85
80.22±0.99
69.89
AIM-ZO (Ours)
68.16±2.66
74.90±1.53
93.47±0.56
57.71±0.91
59.23±2.77
82.96±0.69
72.74
Qwen3-8B-Base
Method
RTE
BoolQ
SST-2
WiC
MultiRC
SQuAD
Avg.
Table 2: Scaling and model-family comparisons on OPT-30B and Qwen3-8B (%). SQuAD uses F1; other tasks use accuracy. Avg. is the unweighted mean over six tasks. Best mean scores are bold; subscripts report standard deviations. Qwen3-8B entries use three seeds and development-selected checkpoints.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Learning rate η
Initial scale ϵ0
OPT-2.7B
1.5×10−3
8×10−5
OPT-13B
1×10−3
8×10−5
OPT-30B
5×10−4
6×10−5
Qwen3-0.6B-Base
1×10−3
8×10−5
Qwen3-8B-Base
1×10−3
8×10−5
Appendix
Table 3: AIM-ZO learning rates and initial perturbation scales.
Method
Calls/update
Learning rate
Perturbation scale
MeZO
2
(1–5)×10−7
10−3
CurvZO
2
(1–8)×10−7
5×10−4–10−3
AGZO
3
5×10−7
10−4–10−3
HiZOO
3
5×10−7
5×10−4–10−3
LoZO
2
10−7–10−6
10−3
ZO-Muon
5
10−2
10−3
Appendix
Table 4: Baseline hyperparameter search ranges and calls per update for OPT-13B and Qwen3-0.6B, with OPT-2.7B MeZO also included.
Model
Layer
r=8
r=32
r=64
r=128
Qwen3-4B
Early
.622
.820
.879
.923
Middle
.911
.964
.980
.990
Late
.992
.994
.995
.995
OPT-2.7B
Early
.443
.678
.784
.868
Middle
.954
.980
.987
.992
Late
.923
.966
.981
.991
Appendix
Table 5: Temporal held-out shared-space capture. The rank- r space is fitted on eight steps and evaluated on the next eight.
Table 7: Capture of the finite-training-set mean gradient over 16 steps at K=128 ; the current-batch reference uses exact SVD.
Wide subspace ( K=128 )
Active space ( k=64 )
Model
Layer
Difference (pp)
Win rate
Difference (pp)
Win rate
Qwen3-4B
Early
-11.06
3.9%
-10.31
11.7%
Middle
+4.71
98.8%
+4.06
94.5%
Late
+1.02
49.6%
+2.38
64.5%
OPT-2.7B
Early
+8.95
96.1%
+6.29
83.2%
Middle
+1.54
80.9%
+2.19
83.6%
Appendix
Table 8: Oja minus PI-5 current-batch capture in percentage points (pp), and Oja win rates. Pairs share checkpoints and are not independent training runs.
Model
Layer
ηq=.03
.1
.3
.6
Qwen3-4B
Early
(.943,.140)
(.926,.315)
(.892,.414)
(.837,.549)
Middle
(.527,.898)
(.330,.901)
(.338,.894)
(.410,.886)
Late
(.533,.924)
(.618,.923)
(.657,.922)
(.670,.921)
OPT-2.7B
Early
(.161,.713)
(.263,.677)
(.349,.646)
(.384,.635)
Middle
(.349,.943)
(.417,.940)
(.559,.935)
(.617,.933)
Late
(.440,.704)
(.572,.662)
(.663,.648)
(.697,.644)
Appendix
Table 9: Stationary replay: (normalized Φ , gradient capture). Qwen and OPT use steps 500 and 2000, respectively.
Width
Qwen3-4B
OPT-2.7B
8
.7022 / .0020 / .002272
.8137 / .0026 / .002855
16
.7758 / .0068 / .001615
.8458 / .0043 / .002052
32
.8243 / .0121 / .001072
.8736 / .0082 / .001374
64
.8624 / .0277 / .000829
.8975 / .0183 / .000982
128
.8866 / .0510 / .000564
.9166 / .0416 / .000752
Appendix
Table 10: Full active-space width scan using independent Gaussian coefficient matrices. Each cell gives learned capture / random capture / mean learned single-direction first-order cosine.
Model
Space
OP MSE
TP MSE
Reduction
Cosine gap
Qwen3-4B
Learned
50,357
94,002
46.4%
1.192×10−3
Random
1,042
1,680
38.0%
1.911×10−4
Dense
1,825,752
3,487,915
47.7%
1.647×10−4
OPT-2.7B
Learned
30,755
65,777
53.2%
1.129×10−3
Random
764
1,369
44.2%
2.530×10−4
Dense
1,115,083
1,998,325
44.2%
2.280×10−4
Appendix
Table 11: Matched normalized Gaussian control against a fixed batch gradient. MSE is normalized by gradient energy; cosine gap is OP minus TP.
Model
Space
OP MSE
TP MSE
Reduction
Cosine gap
Qwen3-4B
Learned
1,995,835
3,743,201
46.7%
+7.47×10−5
Random
45,462
91,242
50.2%
+3.06×10−5
Dense
96,707,409
192,548,225
49.8%
−3.72×10−5
OPT-2.7B
Learned
1,374,688
2,154,535
36.2%
+8.44×10−5
Random
31,731
53,723
40.9%
+7.29×10−5
Dense
63,415,059
116,824,528
45.7%
−2.92×10−5
Appendix
Table 12: Finite-training-set reference at ϵ=8×10−5 , K=64 , and 16 loss evaluations.
Model
ϵ
Learned
Random
Dense
Qwen3-4B
6×10−5
46.7 / +7.64
50.2 / +2.99
49.8 / -3.69
7.43×10−5
46.7 / +7.48
50.2 / +3.04
49.9 / -3.72
10−3
39.1 / -1.17
50.3 / +3.07
50.6 / -3.70
OPT-2.7B
2.57×10−5
36.5 / +7.66
41.1 / +7.49
45.7 / -2.68
6×10−5
36.3 / +8.51
41.2 / +7.28
45.7 / -2.98
10−3
33.9 / +2.55
40.9 / +7.21
45.7 / -2.82
Appendix
Table 13: Perturbation-scale control against the exact training gradient. Entries are OP MSE reduction (%) / OP–TP cosine gap in units of 10−5 .
Model
Space
OP
TP
Difference
95% interval
Qwen3-4B
Learned
38.72
32.93
5.791
[1.96,9.63]
Dense
2.774
2.443
.3307
[-.514,1.18]
OPT-2.7B
Learned
40.48
32.60
7.882
[4.17,11.6]
Dense
5.482
4.246
1.236
[.536,1.94]
Appendix
Table 14: BF16 implementation-level cosine. All numeric entries are in units of 10−5 ; intervals are paired 95% intervals conditional on the checkpoint and batch.
Variant
RTE
BoolQ
SST-2
SQuAD
MeZO
65.13±1.67
66.10±2.26
92.50±.54
80.99±1.42
Common active basis
66.14±.78
66.84±.63
92.48±.78
81.06±1.02
AIM-ZO
67.51±3.18
67.22±1.67
92.87±.47
81.19±.49
Appendix
Table 15: Common active basis versus per-perturbation resampling on OPT-2.7B. Classification scores are accuracy; SQuAD uses F1 (%).
Maintenance rule
Accuracy
Fixed subspace
68.40±2.04
Oja every ten steps
68.92±3.38
Oja every step
70.20±1.95
Appendix
Table 16: OPT-2.7B RTE maintenance-frequency ablation (five seeds; final development accuracy, %).
Model
K=64,k=h=64
K=128,k=h=64
AIM-ZO
OPT-2.7B
64.26±3.82
66.55±1.10
67.51±3.18
Qwen3-0.6B
77.38±1.71
77.38±.42
77.62±.92
Appendix
Table 17: RTE accuracy (%) with fixed-prefix activation and the AIM-ZO main-result reference.
Maintained width K
Accuracy
64
65.70±1.25
128
65.46±3.62
256
67.15±0.72
Appendix
Table 18: OPT-2.7B RTE maintained-width scan at k=64 and h=48 (three seeds; official-validation accuracy, %).
Shared width h
Sampled tail k−h
Accuracy
0
64
60.77±1.99
48
16
65.46±3.62
64
0
66.55±1.10
Appendix
Table 19: OPT-2.7B RTE shared-width scan at K=128 and k=64 (three seeds; official-validation accuracy, %).
Table 21: Coefficient-rank ablation on OPT-2.7B RTE over three seeds. Accuracy is official-validation accuracy (%); time is mean per-run median non-evaluation step time; memory is peak CUDA allocated memory.
Estimator
Best checkpoint
Final update
OP15-RLOO
71.40±3.02
71.12±3.29
Raw TP8
70.04±4.38
69.28±4.62
Appendix
Table 22: Online OPT-2.7B RTE estimator ablation over five matched seeds (development accuracy, %).
Qwen3-4B
OPT-2.7B
s
Capture
Δcos
Capture
Δcos
0
.0046
.0795
.0020
.0668
.10
.0857
.365
.0931
.448
.25
.2072
.567
.2304
.705
.50
.4094
.798
.4598
.996
.75
.6116
.975
.6894
1.220
Appendix
Table 23: Fixed-width space-quality control with an exact linear oracle. Alignment gains are reported in units of 10−3 .
Model
ϵ
Correlation
Monotone
OP MSE wins
Qwen3-4B
6×10−5
.952
12/12
84/84
10−3
.952
12/12
84/84
OPT-2.7B
2.57×10−5
.927
12/12
84/84
10−3
.926
12/12
84/84
Appendix
Table 24: Real-loss space-quality sweep. Correlation relates measured capture to the OP–TP alignment gain; monotone counts refer to checkpoint/batch settings.
Model
Task
AIM-ZO
MeZO
Ratio
Qwen3-0.6B
RTE
4.339
1.871
2.319
OPT-2.7B
RTE
7.363
3.508
2.099
OPT-13B
BoolQ
26.603
21.919
1.214
Appendix
Table 25: Training-step hours at 40,000 forward evaluations. Ratio is AIM-ZO time divided by MeZO time.
Model
Method
Evaluations/step
s/step
s/eval.
GiB
Qwen3-8B
MeZO
2
.653
.327
19.929
AIM-ZO
16
7.396
.462
19.937
AGZO
3
1.196
.399
18.772
OPT-30B
MeZO
2
1.810
.905
58.552
AIM-ZO
16
14.052
.878
60.034
AGZO
3
2.727
.909
58.087
Appendix
Table 26: Large-model short-run time and memory.
Cache
s/step
Allocated GiB
Reserved GiB
None
9.164
10.251
19.111
Factor restoration
7.273
10.287
19.143
Factor restoration + full noise
7.248
10.576
20.516
Appendix
Table 27: Cache ablation for AIM-ZO . Time averages all ten steps; allocated and reserved memory are reported separately.
Dataset
Metric
MeZO
AIM-ZO
W/T/L
RTE
Accuracy
65.13±1.67
67.51±3.18
4/1/0
BoolQ
Accuracy
66.10±2.26
67.22±1.67
5/0/0
SST-2
Accuracy
92.50±0.54
92.87±0.47
5/0/0
WiC
Accuracy
58.53±0.41
60.13±1.70
5/0/0
MultiRC
Accuracy
61.96±2.04
63.94±1.59
4/0/1
MultiRC
F1a
34.71±10.81
48.91±4.09
5/0/0
Appendix
Table 28: Five-seed official-validation results for OPT-2.7B under 40,000 matched training forwards. Higher is better. W/T/L counts paired-seed outcomes for AIM-ZO relative to MeZO. Scores are percentages.
Method
RTE
BoolQ
SST-2
WiC
MultiRC
SQuAD
Avg.
Zero-shot
58.48
62.20
58.26
51.25
59.08
56.22
57.58
MeZO
77.62±1.11
73.87±0.57
88.42±1.35
55.36±2.44
72.78±1.49
83.23±0.76
75.21
CurvZO
77.14±1.04
68.28±0.80
82.87±1.15
49.74±1.37
71.72±2.86
71.29±3.46
70.17
HiZOO
77.86±0.55
72.76±0.84
88.80±0.88
53.92±1.96
76.11±0.93
80.42±0.89
74.98
AGZO
69.68±4.39
68.40±0.26
84.75±3.04
52.77±3.70
73.89±0.38
73.17±1.43
70.44
ZO-Muon
75.69±0.21
72.73±0.79
88.76±0.70
57.79±2.05
75.10±0.63
81.10±1.36
75.20
Appendix
Table 29: Multi-method comparison on Qwen3-0.6B-Base (%). SQuAD uses F1; other tasks use accuracy. Avg. is the unweighted mean over six tasks; subscripts report available sample standard deviations.
Metric
Zero-shot
MeZO
AIM-ZO
MultiRC F1a
54.70
83.95±0.25
84.48±0.37
MultiRC question EM
28.23
54.60±1.37
56.31±0.44
SQuAD EM
72.40
78.20±2.23
84.63±0.42
Appendix
Table 30: Additional Qwen3-8B-Base evaluation metrics (%; three seeds for trained methods).
Method
Subspace perturbations
Information- guided
History- informed
Width decoupling
MeZO
×
×
×
×
HiZOO
×
✓
✓
×
CurvZO
✓
✓
✓
×
LoZO
✓
×
×
×
SubZero
✓
×
×
×
ZO-Muon
✓
×
×
×
Appendix
Table 31: Design properties of representative ZO methods. Sparse coordinate spaces count as subspace perturbations; guidance and history refer to information used in perturbation construction.