Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family of LLMs with ternary {-1,0,1} weights and 8-bit activations, fine-tuning requires updating full-precision latent weights, and thus the memory footprint of MeZO no longer matches that of inference. A promising solution is to finetune only a subset of the latent weights, but existing sparse zeroth-order (ZO) methods either ignore the ternary structure or require first-order gradient information to build a sparse mask, which is at odds with the purpose of ZO fine-tuning. We propose TerMeZO, a sparse MeZO scheme that exploits the geometry of the ternary quantizer itself to identify the latent weights that are more likely to change values during fine-tuning, at no additional data or memory cost. Our convergence analysis shows that TerMeZO can converge faster than full-parameter MeZO, owing to its optimized reduction of the fine-tuning effective dimension. We run extensive experiments on BitNet models ranging from 1B to 3B parameters, spanning classification, instruction-following, and mathematical reasoning tasks. TerMeZO matches or exceeds the performance of full-parameter MeZO while substantially reducing the fine-tuning memory footprint.
Figures & tables
Method
Near inference memory of BitNet
Matmul-free (BitNet like)
No FO overhead
Sparsity
Quantization- Aware Mask
PEFT — FO fine-tuning
LoRA ( Hu et al., 2022 )
✗
✗
✗
✗
✗
QLoRA ( Dettmers et al., 2023 )
✗
✗
✗
✗
✗
ZO fine-tuning
MeZO ( Malladi et al., 2023 )
✗
✓
✓
✗
✗
Me ZO (LoRA) ( Malladi et al., 2023 )
✗
✗
✓
✗
✗
Table 1 : Comparison of fine-tuning methods for BitNet-style models (Matmul = matrix multiplication).
BP
MeZO
TerMeZO
TerMeZO
Inference
( ρ0=0.05 )
( ρ0=0.01 )
Falcon-E-1B
16.08
3.61
1.33 (est. 1.31)
0.94 (est. 0.93)
0.89
Falcon-E-3B
28.50
6.28
2.08 (est. 2.04)
1.36 (est. 1.34)
1.24
Table 2 : Memory footprint (GB) during fine-tuning of Falcon-E-1B and Falcon-E-3B on an H100 GPU. For TerMeZO, we report both measured and estimated values using ( 13 ). The TerMeZO and inference numbers assume 2 -bit storage of the ternary weights. Results here are measured usi SFT with the context length fixed to 512.
SST2
RTE
CoPA
CB
BoolQ
MultiRC
BP
94.8
84.1
83
88.0
86.1
83.4
ICL
85.4
75.4
80
78.5
76.8
73.1
MeZO
93.2
81.9
78
87.9
82.3
82.2
LoRA-MeZO
55.0
71.6
73
55.5
54.5
66.6
FIM-S-MeZO
92.3
82.4
75
83.8
79.5
81.6
QZO
91.7
79.7
77
82.8
82.9
81.8
Table 3 : Accuracy (%) of TerMeZO and baselines on classification tasks using the Microsoft/bitnet-b1.58-2B-4T BitNet model ( Ma et al., 2025 ) . The fraction of trainable latent weights is fixed to ρ0=0.05 for TerMeZO.
Falcon-E-1B-Base
Falcon-E-3B-Base
GSM8K
MagiCoder
GSM8K
MagiCoder
Base Model
14.0
17.1
20.0
37.8
BP
37.7
32.9
57.4
49.4
MeZO
32.6
25.6
45.3
42.7
FIM-S-MeZO
32.5
25.0
44.5
43.3
S-MeZO (min)
27.8
15.8
26.8
41.5
Table 4 : Performance of TerMeZO with ρ0=0.05 and benchmarks on GSM8K and MagiCoder using two BitNet models, Falcon-E-1B-Base and Falcon-E-3B-Base. We report accuracy (%) on GSM8K and pass@1 (%) on MagiCoder.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
ZO Methods
BP
Batch size
{16,32}
{8,16}
Learning rate
{1e−6,5e−6,1e−5}
{1e−5,5e−5,1e−4}
Learning rate scheduler
Linear
Cosine
Appendix
Table 5 : Hyperparameters for ZO methods and BP on the GLUE/SuperGLUE datasets.
ZO
BP
Batch size
16
8
Learning rate
{1e−5,5e−5,1e−4}
{1e−5,5e−5,1e−4}
Learning rate scheduler
Linear
Cosine
Appendix
Table 6 : Hyperparameters for the instruction-following and reasoning experiments. Shared across all ZO runs: n=5 Gaussian perturbations, ϵ=10−3 , linear learning rate decay, and ρ0=0.05 .
Method
Test loss
GSM8K (%)
Base Model
1.13
14.0
TerMeZO
0.56
31.84
QuZO
0.99
14.71
Appendix
Table 7 : Accuracy after finetuning of TerMeZO and QuZO on GSM8K for Falcon-E-1B-Base . For each method, we report the best result over its learning-rate grid. We found QuZO diverges at the higher learning rates used for TerMeZO and the other ZO baselines and we searched a lower range for it to ensure convergence: η∈{10−6,3×10−7,10−7} .