Fine-Tuning Language Models with Just Forward Passes
- Type
- paper
- Venue
- NeurIPS 2023
- Year
- 2023
- Source
- x
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
MeZO adapts zeroth-order (gradient-free) SGD to fine-tune language models using only forward passes, operating in-place with inference-level memory. Supports full-parameter and parameter-efficient tuning (LoRA, prefix tuning) and non-differentiable objectives.
Keywords
zeroth-order · fine-tuning · memory-efficient · lora · neurips
Topics
zeroth-order, fine-tuning, memory-efficient, lora, neurips
Research notes
- Discovery: NOTE: source_url duplicated from the hallucination-benchmark item above by assignment; the intended discovery source for this record is the 'Fine-Tuning Language Models with Just Forward Passes' (MeZO) discussion posted in #random-papers on 2026-09-28. Canonical paper located via arXiv search: 2305.17333v3; code: https://github.com/rxia0716/mezo.
- Method: Memory-efficient zeroth-order SGD: estimates gradients via forward-pass perturbations (simultaneous perturbation stochastic approximation style) applied in-place, avoiding activation storage for backprop.
- Key findings: Reported tuning a 30B model on a single A100 80GB (vs 2.7B with backprop), up to 12x lower memory, up to 2x lower GPU-hours than backprop-based tuning, and compatibility with LoRA/prefix tuning.
- Limitations: Zeroth-order methods typically need more steps / careful hyperparameters vs. backprop; reported gains are from the paper's own experiments. Latest arXiv revision 2024-01-11 (v3).
- The zeroth-order theme of this batch: MeZO (2023, fine-tuning) vs. EGGROLL (pretraining-scale ES) vs. the unverified @industriaalist pretraining claim.