cs.LGOct 1, 2026

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Authors: Cristian McGee, El Houcine Bergou, Aritra Dutta

Organizations: University of Central Florida Orlando, FL, USA · Mohammed VI Polytechnic University Ben Guerir, Morocco

Abstract

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces

    Aug 16, 2026Ziming Yu, Shuyao Xiao, Xingyu Zhao +6Zeroth-OrderLarge Language Model Fine-Tuning

  2. On Adaptivity in Zeroth-Order Optimization

    May 5, 2026Hassan Dbouk, Nidham Gazagnadou, Matthias Reisser +1Zeroth-Order OptimizationStochastic Gradient Descent

  3. Universally Empowering Zeroth-Order Optimization via Adaptive Layer-wise Sampling

    Apr 20, 2026Fei Wang, Li Shen, Liang Ding +3Zeroth-Order OptimizationZeroth-Order