cs.LGSep 27, 2026

TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge

Authors: Houssem Sifaou, Prabodh Katti, Bipin Rajendran, Osvaldo Simeone

Organizations: Institute for Intelligent Networked Systems, Northeastern University London, UK

Abstract

Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family of LLMs with ternary {-1,0,1} weights and 8-bit activations, fine-tuning requires updating full-precision latent weights, and thus the memory footprint of MeZO no longer matches that of inference. A promising solution is to finetune only a subset of the latent weights, but existing sparse zeroth-order (ZO) methods either ignore the ternary structure or require first-order gradient information to build a sparse mask, which is at odds with the purpose of ZO fine-tuning. We propose TerMeZO, a sparse MeZO scheme that exploits the geometry of the ternary quantizer itself to identify the latent weights that are more likely to change values during fine-tuning, at no additional data or memory cost. Our convergence analysis shows that TerMeZO can converge faster than full-parameter MeZO, owing to its optimized reduction of the fine-tuning effective dimension. We run extensive experiments on BitNet models ranging from 1B to 3B parameters, spanning classification, instruction-following, and mathematical reasoning tasks. TerMeZO matches or exceeds the performance of full-parameter MeZO while substantially reducing the fine-tuning memory footprint.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces

    Jul 1, 2026Xun Dong, Yibo Xu, Naigang Wang +3Zeroth-Order OptimizationZeroth-Order

  2. GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning

    Jun 1, 2026Liyan Tan, Yequan Zhao, Yifan Yang +3Zeroth-Order OptimizationZeroth-Order

  3. Breaking the 1.58-bit Barrier for Ternary LLMs

    Sep 14, 2026Evangelos Georganas, Alexander Heinecke, Pradeep DubeyLarge Language Model QuantizationModel Weights