ZFO: Decoupling Direction and Step-Size for Efficient LLM Fine-Tuning
arXiv preprint proposes a lightweight Zero-and-First-Order framework that uses a first-order optimizer for direction and only two extra objective evaluations for…

- ZFO combines trusted first-order direction with zeroth-order evaluations limited to a one-dimensional subspace, reducing cost compared to full line search.
- Theoretical guarantees cover reliable finite-difference curvature estimates, near-optimal local step selection, and convergence to a stationary neighborhood.
- Empirical results on language models and datasets show frequent improvements over fixed-step first-order baselines, varying by objective and chosen local model.
Core Idea
Step-size selection remains a central challenge in large-scale neural network optimization. Conservative steps slow convergence while aggressive ones risk instability.
The proposed ZFO framework decouples direction from step-size. A trusted first-order optimizer determines the search direction while zeroth-order evaluations are performed only along this one-dimensional subspace to decide movement distance.
This produces an adaptive, curvature-aware step within a bounded search interval at lower computational cost than a traditional line search.
Method and Theory
ZFO builds a local model of the objective function using the current gradient information together with two additional objective evaluations.
The preprint establishes that shared-sample evaluations yield reliable finite-difference curvature estimates and that the induced local model selects a near-optimal step along the interval.
Convergence analysis shows that ZFO reaches a neighborhood of a stationary point under standard assumptions.
Experimental Results
Tests across multiple language models and datasets indicate that ZFO often improves both optimization trajectory and final performance relative to fixed-step first-order baselines.
The scale of gains and the best-performing local model depend on the specific objective function being minimized.
Implementation code is publicly released to enable reproduction and further study.
Status and Limitations
This work is an arXiv preprint dated October 1 2026 based solely on the supplied abstract and metadata.
Full experimental details, complete results, and independent validation are not yet available in the captured evidence.
The approach offers a promising direction for making fine-tuning of open-weight models more efficient, though practical impact will depend on outcomes in the complete paper.


