← Back to explorer

LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters

Type
other
Venue
arXiv / Jagiellonian University / EPFL

Summary

Freezes truncated-SVD factors of W and trains only R∈R^{r×r} between them, so parameter count is independent of hidden size (≈2n/r fewer than LoRA). GPT-3 rank-16 personalization: 96GB vs LoRA 144TB for 1M adapters. RoBERTa-large GLUE: rank 25 / 60K params avg 88.69 vs LoRA 800K 87.82 and VeRA 61K 87.83. LLaMA2-7B commonsense 80.5 vs LoRA 77.6 at 3.67M vs 56M params; LLaMA3-8B 85.3 vs 80.8. Mistral-7B GSM8K 70.35 / MATH 20.96 at 3.67M vs LoRA 168M 67.70 / 19.68. Accepted at ECAI 2025. Code https://github.com/MohammadrezaBanaei/LoRA-XS.

Keywords

lora-xs · peft · svd · vera · glue · llama · ecai · jagiellonian · epfl

Topics

parameter-efficient fine-tuning, LoRA, SVD

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI, cs.CL). License not stated on abs/HTML at check. Equal contrib Bałazy/Banaei; correspondence klaudia.balazy@doctoral.uj.edu.pl. Jagiellonian / EPFL. Code https://github.com/MohammadrezaBanaei/LoRA-XS (56 stars at check). HF paper page 4 upvotes; githubRepo linked; no linked models/datasets. Discord posted abs. Uses public GLUE/GSM8K/MATH/commonsense/MetaMathQA; no new corpus, so no datasets_local row. License field left blank per catalog convention.