The Universal Weight Subspace Hypothesis
- Type
- paper
- Venue
- alphaXiv (arXiv:2512.05117)
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
The Universal Weight Subspace Hypothesis proposes that neural networks of the same architecture share an architecture-specific low-dimensional 'universal subspace' of weight directions, so new tasks can be stored or trained as coefficients on a frozen shared basis. Experiments across GPT-2, ViT, LLaMA-8B, Flan-T5, and 500 Mistral LoRA adapters show rapid spectral decay and large storage savings, but the evidence is limited to within-architecture comparisons with small evaluation sets.
Keywords
weight-sharing · subspace · model-merging · lora · compression
Topics
weight-sharing, subspace, model-merging, lora, compression
Research notes
- Method: Collect corresponding weight matrices or LoRA submatrices from same-architecture models, stack, subtract feature-wise mean, then PCA/SVD (or order-1/2 HOSVD) to keep leading per-layer directions; freeze the basis and represent new tasks by coefficients on it.
- Key findings: Weight variation across 200 GPT-2, 500 ViT, 50 LLaMA-8B, 8 Flan-T5 models concentrates in a small number of layer-wise directions (rapid spectral decay); Held-out ViTs projected onto a 16-component shared subspace: 87.8±1.5% accuracy vs 91.3±2.1% full training (reported as 'no significant drop', no significance test); 500 Mistral-7B LoRA adapters: leading ≤16 directions capture most variance; 19× memory efficiency vs saving all LoRAs; Estimated up to 100× less memory to store 500 ViTs in the subspace (storage estimate only)
- Limitations: Subspaces are architecture-specific — authors state they lack a method to compare across architectures, leaving the 'ideal' universal hypothesis unverified. Small OOD eval set (4–5 models), ± values undefined, no significance tests; storage estimate is not a training-time/energy saving.
- Authors not listed on the alphaXiv summary page; verify from arXiv before citing. First/last layers excluded from full-weight ViT analysis.