Liger Kernel: Efficient Triton Kernels for LLM Training
- Type
- other
- Venue
- arXiv / LinkedIn
Summary
Open-source Triton kernels (RMSNorm, LayerNorm, RoPE, SwiGLU/GeGLU, CrossEntropy, FusedLinearCrossEntropy) using fusion and input chunking. Reports ~20% higher training throughput and ~60% less GPU memory vs HuggingFace implementations. 4×A100 Alpaca seq-512: Llama 3-8B +42.8% throughput / −54.8% memory at batch 64; Qwen2 +25.5% / −56.8% at batch 48. Also Gemma, Mistral, Phi-3. AutoLigerKernelForCausalLM plus HF Trainer/TRL/Axolotl/LLaMA-Factory hooks. Code https://github.com/linkedin/Liger-Kernel.
Keywords
liger · triton · fused-kernels · rmsnorm · rope · swiglu · flce · linkedin · llm-training
Topics
Triton kernels, LLM training efficiency, operator fusion
Research notes
- Primary: arxiv abs (cs.LG; also cs.AI, cs.CL, cs.DC). License CC BY 4.0 on HTML at check. Comment: 17 pages, 12 figures. All authors LinkedIn. Code https://github.com/linkedin/Liger-Kernel (HF githubRepo; 6,567 stars at check; BSD-2-Clause). HF paper page 3 upvotes; unofficial linked models not copied into hf_* fields. Discord posted abs; a follow-up message linked examples/huggingface/training.py. Kernel library, not a new corpus, so no datasets_local row.