Skip to content
AI.info

Research

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Overview Research area: Machine learning optimization, specifically step-size selection for large language model (LLM) fine-tuning; it sits at the intersection of first-order (gradient-based) and zero

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
arXiv
2610.02190
Published
2026-10-01
Authors
Cristian McGee, El Houcine Bergou, Aritra Dutta

AI summary

Overview

Research area: Machine learning optimization, specifically step-size selection for large language model (LLM) fine-tuning; it sits at the intersection of first-order (gradient-based) and zeroth-order (function-evaluation-only) optimization.

Technical level: Intermediate. The empirical sections are readable with basic fine-tuning familiarity, but the theoretical section relies on smoothness assumptions, sub-Gaussian concentration and convergence-rate notation.

Scope: The paper proposes a framework (ZFO) that keeps a standard first-order optimizer in charge of the update direction and uses two extra function evaluations to choose the step size along that direction, then supports it with theory and experiments across LLM fine-tuning and classical optimization problems.

What This Paper Is About

Gradient-based optimizers such as AdamW decide which way to move in parameter space, but the distance moved (the learning rate or step size) is usually set by a schedule or by hyperparameter tuning, and a bad choice can slow training or destabilize it. The authors ask whether cheap zeroth-order evaluations — forward passes that only measure the objective value — can be used to pick a better step along a direction the optimizer already trusts, rather than to estimate a gradient. Their goal is a lightweight wrapper that improves step size without the cost of a full line search and without discarding the reliable directional information that first-order gradients provide.

Key Contributions

  1. A general ZFO optimization framework. The authors decouple direction selection (handled by a first-order optimizer such as AdamW or Muon) from step selection (handled by zeroth-order evaluations along a one-dimensional line). Within this abstraction they develop four concrete instances built from Taylor models (second- and third-order) and Padé models (second- and third-order). The related method GeN is positioned as one particular fixed-budget directional step-selection rule that fits inside the framework.

  2. Theoretical justification of adaptive step selection. They prove that shared-batch finite-difference curvature estimates with common random numbers (CRNs) concentrate around their population counterparts with high probability, that maximizing the resulting local model yields a step whose objective value is near-optimal within a bounded search region, and that the ZFO update sequence converges to a neighborhood of a stationary point whose size is controlled by accumulated step-selection error.

  3. Empirical validation across scales. They evaluate ZFO on controlled low-dimensional classical optimization problems and on LLM fine-tuning tasks, reporting that ZFO is robust to the initial learning rate and matches or outperforms standard first-order and zeroth-order baselines on most benchmarks.

  4. A low-overhead practical mechanism. Because the extra probes are forward passes on an already-constructed minibatch, the authors measure the added cost as a small fraction of iteration time and a moderate increase in peak GPU memory.

Main Findings

  • ZFO improves on the first-order baseline in nearly every LLM setting tested. In Table 1, at least one ZFO variant improves on first-order fine-tuning for every model–dataset pair. Padé 3 is highlighted as a strong representative instance, improving over first-order in 14 of the 16 LLM settings summarized in Table 9.

  • Gains are largest where the baseline is weak or variable. The paper reports the largest gains on SVAMP for Phi-2 and Gemma-2-2B, and on AsDiv for Llama-3.2-1B, plus MATH and OpenBookQA for Qwen-2.5-Math-1.5B. When the first-order baseline is already strong, such as Qwen-2.5-Math-1.5B on SVAMP, improvements are smaller but still positive in mean performance.

  • Which local model wins is task-dependent. The best-performing ZFO variant changes across tasks and models, so the paper treats the preferred local model as objective-dependent rather than universally fixed.

  • Selected example results (final evaluation, mean over 3 seeds). For Qwen-2.5-Math-1.5B on GSM8K: AdamW 78.24 ± 0.59, MeZO 0.00, Taylor 2 81.00 ± 0.11, Taylor 3 75.56 ± 0.24, Padé 2 80.95 ± 0.49, Padé 3 80.97 ± 0.73. On MATH: AdamW 31.77 ± 1.63, MeZO 22.01 ± 3.93, Taylor 2 40.10 ± 1.19, Taylor 3 20.18 ± 3.03, Padé 2 37.50 ± 1.79, Padé 3 39.58 ± 0.45. On AsDiv for Phi-2: AdamW 56.51 ± 18.76, MeZO 10.15 ± 4.88, Padé 3 69.27 ± 1.92. On AsDiv for Llama-3.2-1B: AdamW 16.60 ± 28.12, MeZO 3.33 ± 0.48, Taylor 3 44.77 ± 1.42, Padé 3 52.16 ± 4.75.

  • Qwen sub-study on additional reasoning benchmarks. At least one ZFO method outperforms the fine-tuning baseline on every dataset. The largest improvements are on ARC-Challenge (↑8.81 over first-order fine-tuning) and CODAH (↑4.92). Example means: AdamW on ARC-Challenge 49.29 ± 16.08 versus Taylor 2 58.10 ± 1.33; AdamW on CODAH 57.91 ± 3.24 versus Taylor 2 62.83 ± 3.14.

  • Classical optimization problems show the same pattern. After 100 updates, at least one ZFO variant achieves substantially lower median objective than its corresponding first-order baseline on every displayed problem. Taylor 3 with AdamW is strongest on nonlinear least squares and Beale, while Padé 3 with AdamW is best on Rosenbrock. Figure 2 plots the median log₁₀ objective over three perturbed initializations, and the paper notes GeN has a similar computational budget since both methods perform two additional function evaluations per iteration.

  • The theoretical bounds have two interpretable parts. The CRN curvature concentration result (Theorem 1) bounds the deviation of the finite-difference estimates of the second and third directional derivatives by a deterministic finite-difference bias term controlled by the perturbation scale ε, plus a sampling term that decays at rate 1/√n. Theorem 2 bounds the gap between the ZFO-selected step's objective value and the best step in the interval by an interval-scaled Taylor approximation error (MR³/3) and a finite-difference curvature error (Bε²R²/12). Theorem 3 bounds the minimum expected squared gradient norm by a standard optimization term decaying at rate 1/T plus a term proportional to the average step-selection error Δ̃_T.

  • Convergence is to a neighborhood, not exactly to a stationary point. The global guarantee shows ZFO converges to a stationary neighborhood whose size is governed by the accumulated step-selection error; exact convergence would require average step-selection error to vanish.

  • Computational overhead is modest but real. ZFO leaves the rollout call, reward evaluation, forward pass and backward pass untouched and adds only two forward-pass probes. In the reported configuration the paper measures approximately 5.2% relative wall-clock overhead for Qwen-2.5-MATH-1.5B across five benchmarks, and peak GPU memory overhead ranging from 11.0% (Qwen/GSM8K) to 19.1% (Phi-2/OpenBookQA). Training was done on an NVIDIA H100 80GB GPU.

  • Learning-rate gains are not simply recoverable by retuning. An ablation comparing ZFO against AdamW across fixed learning rates indicates the Table 1 gains cannot generally be recovered by only changing the first-order learning rate. A further ablation on the interaction between learning rate α and search bound β finds that increasing β is generally beneficial at smaller learning rates but can become unstable once α is large, so the two parameters are complementary but not interchangeable.

  • The four ZFO instances use the search interval differently. Table 13 reports that step choices range from primarily interior steps for Taylor 3 to mostly endpoint decisions for Padé 2.

  • A safeguard for rational models. If a Padé model has a pole inside the search interval or its coefficients are ill-defined, ZFO rejects the Padé step and falls back to the quadratic Taylor step to avoid unstable rational extrapolation.

Methodology in Plain English

The starting point is that an ordinary fine-tuning step computes a gradient and applies it with a learning rate. The authors keep the gradient direction but normalize it to a unit vector, so the only remaining question is how far to travel along that unit direction. They define a search interval whose upper end is the baseline step multiplied by a factor β, so β = 1 corresponds exactly to the standard update and larger β allows longer steps in the same direction.

Within that interval they build a cheap one-dimensional picture of the objective. The current forward pass gives the objective value at step zero; the backward pass that the optimizer already performs gives the directional derivative; and two extra forward passes at +ε and −ε along the same direction supply enough information to estimate second- and third-order curvature by finite differences. From those few numbers they fit a small local model — a quadratic or cubic Taylor polynomial, or a rational Padé function — and pick the step that maximizes that model inside the interval, falling back to the quadratic Taylor step if a Padé model is ill-conditioned.

All evaluations reuse the same minibatch, so no new prompts, responses or labels are needed; the only extra work is re-scoring the same examples under a perturbed model. The theoretical analysis rests on this "common random numbers" setup: conditioning on the shared examples makes the objective a deterministic function of the scalar step, which allows concentration arguments for the curvature estimates, an approximation guarantee for the model-selected step, and finally a stationarity bound for the whole update sequence. Empirically, the authors compare against AdamW (first-order) and MeZO (zeroth-order) across Qwen-2.5-Math-1.5B, Phi-2, Gemma-2-2B and Llama-3.2-1B on GSM8K, SVAMP, AsDiv, OpenBookQA and a 512/256 train/evaluation subset of MATH, using an RLVR setup with GRPO policy gradient and three seeds per result. The learning rate is fixed to the best value for the first-order baseline from {10⁻⁶, 5×10⁻⁶, 10⁻⁵}, with β then chosen from {3, 5, 10}.

Why This Matters

Impact on research. The paper reframes step-size selection as a local one-dimensional modeling problem rather than a scheduling or tuning problem, and gives it a convergence analysis in which the price of imperfect step selection appears explicitly. It also unifies a family of step rules — Taylor, Padé and the earlier GeN-style quadratic rule — under one fixed-budget abstraction, which makes their assumptions and failure modes comparable for the first time.

Real-world applications.

  • Fine-tuning open-weight reasoning models where retraining runs or large learning-rate sweeps are expensive.
  • Reinforcement-learning-from-verifiable-rewards (RLVR) pipelines, which is the setup used in the experiments, where rollout cost dominates and step instability is a known pain point.
  • Low-memory fine-tuning regimes, since the extra cost is forward-only and comparable to the forward-pass structure of zeroth-order methods like MeZO.
  • Classical scientific or engineering optimization problems with noisy objectives, as illustrated by the nonlinear least squares, Beale and Rosenbrock experiments.

Industry relevance. Hyperparameter tuning of learning rates is a recurring cost in production fine-tuning pipelines, and unstable runs waste GPU hours. A wrapper that adds roughly 5.2% wall-clock overhead and 11.0%–19.1% peak memory while removing some sensitivity to the initial learning rate is directly relevant to teams training on constrained hardware. The code is publicly available at https://github.com/nizswan/Zeroth-First-Order-Framework.

Future Directions

  • Adaptive rather than fixed search radii. The search interval is governed by β, and the ablation shows β and the learning rate interact in a way that can become unstable at large α. An open question is whether β can be set adaptively online instead of chosen from a small candidate set.

  • Model selection without task-specific tuning. The preferred local model varies by objective, and Padé models require a pole-rejection safeguard. Automatically choosing among the Taylor and Padé instances per iteration, or per layer, is a natural extension that the paper does not resolve.

  • Generalization beyond the tested base optimizers and objectives. The framework is presented as compatible with any first-order optimizer and any fixed-batch surrogate, but the reported LLM experiments use AdamW with a GRPO policy-gradient surrogate. Whether the calibration transfers to other optimizers mentioned in the paper, such as Muon, or to other surrogate objectives is not established by the results presented.

  • Scaling the theory and the experiments. The guarantees are stated for a fixed direction and bounded interval, and the empirical study covers models up to the sizes listed (Qwen-2.5-Math-1.5B, Phi-2, Gemma-2-2B, Llama-3.2-1B). Whether the overhead and stability properties hold at substantially larger model and batch scales, and whether the stationary-neighborhood bound can be tightened, remain open.

Target Audience

Researchers and practitioners working on optimization for large-scale neural networks, especially those fine-tuning LLMs or running RLVR-style training pipelines where learning-rate sensitivity and GPU cost matter. It will also be useful to readers interested in hybrid first-order/zeroth-order methods, since it places several existing step-selection rules within a single framework and supplies concentration and stationarity guarantees. Readers without an optimization background can still follow the framework description, the algorithmic procedure and the empirical tables; the convergence analysis requires more mathematical comfort.

Authors’ abstract

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.

Read the original paper