Sparser, Faster, Lighter Transformer Language Models
- Type
- other
- Venue
- arXiv / Sakana AI / NVIDIA
Summary
TwELL (tile-wise ELLPACK) materializes ReLU-gated FFN sparsity in the matmul epilogue; fused inference kernel does up+down in one launch; hybrid ELL/dense format stores activations for training despite heavy per-token nnz skew. Mild L1 on gated ReLU FFNs reaches >99% sparsity with negligible task drop through L1=3e-5 (1.5B). At L1=2e-5, 0.5B–2B chinchilla FineWeb models match dense accuracy while forward throughput rises 17.0→20.5% and training −1.5→+21.9% with 19–28% lower peak memory (8×H100). Gains grow with scale as mean active neurons fall (39→24). Code https://github.com/SakanaAI/sparser-faster-llms.
Keywords
sparsity · twell · relu · l1 · cuda · sakana · nvidia · inference · training-kernels
Topics
sparsity, LLM inference, CUDA kernels
Research notes
- Primary: arxiv abs (cs.LG; also cs.CL). License not stated on abs/HTML at check. Cetin/Peluchetti/Castillo core contributors; Cetin corresponding (edo@sakana.ai), Castillo NVIDIA (ecastillo@nvidia.com). Code claimed https://github.com/SakanaAI/sparser-faster-llms (GitHub API rate-limited at check; no star count recorded). HF paper page 2 upvotes; no linked models/datasets. Discord posted abs. Trains on public FineWeb; no new corpus, so no datasets_local row. License field left blank per catalog convention.