← Back to explorer

The Universal Weight Subspace Hypothesis

Type
paper
Venue
alphaXiv (arXiv:2512.05117)
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

The Universal Weight Subspace Hypothesis proposes that neural networks of the same architecture share an architecture-specific low-dimensional 'universal subspace' of weight directions, so new tasks can be stored or trained as coefficients on a frozen shared basis. Experiments across GPT-2, ViT, LLaMA-8B, Flan-T5, and 500 Mistral LoRA adapters show rapid spectral decay and large storage savings, but the evidence is limited to within-architecture comparisons with small evaluation sets.

Keywords

weight-sharing · subspace · model-merging · lora · compression

Topics

weight-sharing, subspace, model-merging, lora, compression

Research notes

  • Method: Collect corresponding weight matrices or LoRA submatrices from same-architecture models, stack, subtract feature-wise mean, then PCA/SVD (or order-1/2 HOSVD) to keep leading per-layer directions; freeze the basis and represent new tasks by coefficients on it.
  • Key findings: Weight variation across 200 GPT-2, 500 ViT, 50 LLaMA-8B, 8 Flan-T5 models concentrates in a small number of layer-wise directions (rapid spectral decay); Held-out ViTs projected onto a 16-component shared subspace: 87.8±1.5% accuracy vs 91.3±2.1% full training (reported as 'no significant drop', no significance test); 500 Mistral-7B LoRA adapters: leading ≤16 directions capture most variance; 19× memory efficiency vs saving all LoRAs; Estimated up to 100× less memory to store 500 ViTs in the subspace (storage estimate only)
  • Limitations: Subspaces are architecture-specific — authors state they lack a method to compare across architectures, leaving the 'ideal' universal hypothesis unverified. Small OOD eval set (4–5 models), ± values undefined, no significance tests; storage estimate is not a training-time/energy saving.
  • Authors not listed on the alphaXiv summary page; verify from arXiv before citing. First/last layers excluded from full-weight ViT analysis.