← Back to explorer

Fine-Tuning Language Models with Just Forward Passes

Type
paper
Venue
NeurIPS 2023
Year
2023
Source
x
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

MeZO adapts zeroth-order (gradient-free) SGD to fine-tune language models using only forward passes, operating in-place with inference-level memory. Supports full-parameter and parameter-efficient tuning (LoRA, prefix tuning) and non-differentiable objectives.

Keywords

zeroth-order · fine-tuning · memory-efficient · lora · neurips

Topics

zeroth-order, fine-tuning, memory-efficient, lora, neurips

Research notes

  • Discovery: NOTE: source_url duplicated from the hallucination-benchmark item above by assignment; the intended discovery source for this record is the 'Fine-Tuning Language Models with Just Forward Passes' (MeZO) discussion posted in #random-papers on 2026-09-28. Canonical paper located via arXiv search: 2305.17333v3; code: https://github.com/rxia0716/mezo.
  • Method: Memory-efficient zeroth-order SGD: estimates gradients via forward-pass perturbations (simultaneous perturbation stochastic approximation style) applied in-place, avoiding activation storage for backprop.
  • Key findings: Reported tuning a 30B model on a single A100 80GB (vs 2.7B with backprop), up to 12x lower memory, up to 2x lower GPU-hours than backprop-based tuning, and compatibility with LoRA/prefix tuning.
  • Limitations: Zeroth-order methods typically need more steps / careful hyperparameters vs. backprop; reported gains are from the paper's own experiments. Latest arXiv revision 2024-01-11 (v3).
  • The zeroth-order theme of this batch: MeZO (2023, fine-tuning) vs. EGGROLL (pretraining-scale ES) vs. the unverified @industriaalist pretraining claim.