← Back to explorer

Sparse Weight Decomposition for Efficient Circuit Extraction

Type
repo
Venue
Veri Safe (GitHub)
Year
2026
Source
github
Access
free
Language
en
Added
2026-08-14T19:50:00Z
Verified
2026-08-14T19:50:00Z

Summary

Veri-Safe SWD reference implementation. Factorizes pretrained dense linear maps (GPT-2 Small mlp.c_proj, Qwen2.5 0.5B/1.5B/3B down_proj, Qwen3.5-27B, full GPT-2 MLP and all 48 body projections) via Double Sparse Factorization so replacement CE stays near dense (e.g. GPT-2 L8 c_proj s=0.5: CE delta 0.000889 on 16,384 tokens; Qwen3.5-27B s=0.5: −0.000427). Pipeline: activation-Gram → DSF → fixed-support recovery → CE/KL/recon, task-margin attribution, mean ablation, sufficiency/necessity frontiers vs Transcoder/VPD baselines. Factor-only checkpoints on HF veri-safe/SWD (~1.16 GB usedStorage) plus SWD-Qwen2.5-3B (Qwen Research License). Discord posted the HF Spaces blog veri-safe/SWD-Blog. Apache-2.0. No FineWeb text redistributed.

Keywords

swd · sparse-weight-decomposition · interpretability · circuits · veri-safe · blog

Topics

mechanistic interpretability, sparse factorization, circuit analysis

Research notes

  • Primary: GitHub README (Apache-2.0). Discord posted https://huggingface.co/spaces/veri-safe/SWD-Blog (Space was a loading stub at fetch). Checkpoints https://huggingface.co/veri-safe/SWD and veri-safe/SWD-Qwen2.5-3B. Calibration uses FineWeb-Edu locally; not a new hosted corpus, so no datasets_local row. Individual author names not on the README; org Veri Safe used rather than guess.