Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Organizations: University of Central Florida Orlando, FL, USA · Mohammed VI Polytechnic University Ben Guerir, Morocco
Abstract
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.
Figures & tables
| Model | Dataset | Baseline | ZFO Methods (This Paper) | ||||
| AdamW (FO) | MeZO (ZO) | Taylor 2 | Taylor 3 | Padé 2 | Padé 3 | ||
| Qwen-2.5-Math-1.5B | GSM8K | 78.24 0.59 | 0.00 | 81.00 0.11 | 75.56 0.24 | 80.95 0.49 | 80.97 0.73 |
| MATH | 31.77 1.63 | 22.01 3.93 | 40.10 1.19 | 20.18 3.03 | 37.50 1.79 | 39.58 0.45 | |
| SVAMP | 88.11 0.69 | 68.67 10.33 | 90.78 0.96 | 89.11 0.51 | 90.89 1.02 | 90.33 2.19 | |
| AsDiv | 91.45 0.36 | 64.72 3.51 | 94.09 0.29 | 88.27 0.35 | 93.04 0.22 | 93.46 0.18 | |
| OpenBookQA | 26.07 2.53 | 8.00 10.05 | 37.80 17.85 | 39.73 20.99 | 60.67 0.95 | 64.07 1.22 | |
| Method | ARC-Challenge | StrategyQA | FOLIO | CODAH |
|---|---|---|---|---|
| AdamW (FO) | 49.29 16.08 | 68.56 0.91 | 34.48 1.13 | 57.91 3.24 |
| Taylor 2 | 58.10 1.33 | 69.63 1.82 | 36.6 3.26 | 62.83 3.14 |
| Taylor 3 | 54.24 0.94 | 61.67 1.46 | 34.48 1.13 | 55.28 1.66 |
| Padé 2 | 53.58 2.77 | 65.55 4.43 | 36.27 2.73 | 53.96 17.72 |
| Padé 3 | 57.05 1.24 | 70.06 0.61 | 36.44 4.53 | 41.85 20.58 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Study | Model | Model Parameters | Dataset | Dataset Size (Train / Eval) |
| Main Study | Qwen-2.5-Math-1.5B | 1.54B | GSM8K | 7,473 / 1,319 |
| MATH | 512 / 256 | |||
| SVAMP | 700 / 300 | |||
| AsDiv | 1,844 / 461 | |||
| OpenBookQA | 4,957 / 500 | |||
| Phi-2 | 2.70B | SVAMP | 700 / 300 |
| Model | Dataset | Configuration | ||||||
| FO LR | ZO LR | BT2 | BT3 | BP2 | BP3 | Config Type | ||
| Qwen-2.5-Math-1.5B | GSM8K | 5e-6 | ALL | 10 | 5 | 10 | 10 | a |
| MATH | 5e-6 | 1e-6 | 5 | 5 | 5 | 5 | c | |
| SVAMP | 1e-5 | 5e-7 | 3 | 5 | 3 | 3 | a | |
| AsDiv | 5e-6 | 5e-7 | 5 | 5 | 5 | 5 | c | |
| OpenBookQA | 1e-5 | 1e-7 | 10 | 5 | 5 | 5 | a | |
| Model | Dataset | Configuration | |||||
| Learning Rate | BT2 | BT3 | BP2 | BP3 | Config Type | ||
| Qwen-2.5-Math-1.5B | ARC-Challenge | 1e-5 | 3 | 5 | 3 | 3 | c |
| StrategyQA | 1e-5 | 5 | 5 | 3 | 5 | c | |
| FOLIO | 1e-5 | 3 | 3 | 3 | 3 | c | |
| CODAH | 1e-5 | 5 | 5 | 5 | 5 | c | |
| Problem | Objective | Initial position |
|---|---|---|
| Linear least squares | ||
| Nonlinear least squares | ||
| Logistic regression | ||
| Low-rank matrix factorization | ||
| Rosenbrock | ||
| Beale |
| Method | Main Results Wins | Qwen Sub-Study Wins | Win Ratio |
|---|---|---|---|
| Taylor 2 | 10/12 | 4/4 | 14/16 |
| Taylor 3 | 6/12 | 1/4 | 7/16 |
| Padé 2 | 11/12 | 2/4 | 13/16 |
| Padé 3 | 11/12 | 3/4 | 14/16 |
| Model | Dataset | AdamW + Cosine | Prodigy | Worst ZFO | Best ZFO |
|---|---|---|---|---|---|
| Qwen-2.5-Math-1.5B | GSM8K | 72.88 0.54 | 64.32 0.90 | 75.56 0.24 (Taylor 3 ) | 81.00 0.11 (Taylor 2 ) |
| SVAMP | 86.00 0.33 | 83.22 0.19 | 89.11 0.51 (Taylor 3 ) | 90.89 1.02 (Padé 2 ) | |
| AsDiv | 81.98 0.52 | 80.82 0.30 | 88.27 0.35 (Taylor 3 ) | 94.09 0.29 (Taylor 2 ) | |
| OpenBookQA | 59.60 0.87 | 38.73 0.31 | 37.80 17.85 (Taylor 2 ) | 64.07 1.22 (Padé 3 ) | |
| ARC-Challenge | 55.15 1.11 | 30.80 0.15 | 53.58 2.77 (Padé 2 ) | 58.10 1.33 (Taylor 2 ) | |
| StrategyQA | 66.28 0.51 | 60.80 0.95 | 61.67 1.46 (Taylor 3 ) | 70.06 0.61 (Padé 3 ) |
| Dataset | ZFO-Seq | Armijo | BB | Polyak | PLS |
|---|---|---|---|---|---|
| CODAH | 55.47 | 32.81 | 52.73 | 30.47 | 31.25 |
| ARC-Challenge | 55.86 | 57.81 | 54.69 | 35.16 | 35.16 |
| SVAMP | 83.59 | 85.55 | 85.55 | 82.81 | 76.17 |
| Dataset | Base LR | AdamW (FO) | Taylor 2 | Taylor 3 | Padé 2 | Padé 3 |
|---|---|---|---|---|---|---|
| ARC-Challenge | 31.71 8.11 | 56.77 0.61 | 56.51 0.36 | 55.49 0.98 | 49.86 0.27 | |
| StrategyQA | 59.24 10.14 | 70.69 5.10 | 74.87 0.37 | 67.25 13.64 | 48.18 0.44 | |
| ASDiv | 82.08 0.08 | 81.82 0.26 | 78.05 1.41 | 81.87 0.26 | 81.92 0.46 |
| Method | Selected-step frequency (%) | Median | Fallback frequency (%) | ||||
|---|---|---|---|---|---|---|---|
| Taylor 2 | 1.37 | 40.97 | 3.76 | 57.66 | 1.00000 | 0.61 | |
| Taylor 3 | 0.86 | 65.83 | 23.88 | 33.31 | 0.00156 | 0.67 | |
| Padé 2 | 43.41 | 0.00 | 0.00 | 56.59 | 1.00000 | 0.76 | |
| Padé 3 | 76.27 | 3.45 | 2.55 | 20.27 | 0.00000 | 12.49 | |
| Taylor 2 | 5.04 | 30.37 | 15.87 | 64.59 | 1.00000 | ||