← Back to explorer

KTransformers

Type
repo
Venue
kvcache-ai (GitHub)
Year
2026
Source
github
Access
free
Added
2026-08-14T18:45:00Z
Verified
2026-08-14T18:45:00Z

Summary

kvcache-ai/ktransformers is a CPU-GPU hybrid inference/SFT framework (MADSys Tsinghua / Approaching.AI / 9#AISoft). Current tree exposes kt-kernel serving (AMX/AVX INT4/INT8, NUMA-aware MoE, hot-GPU/cold-CPU experts, SGLang integration) and LLaMA-Factory SFT/DPO for ultra-large MoE on limited VRAM. README: DeepSeek-R1/V3 on a single 24GB GPU plus large DRAM; SFT DeepSeek-V3 ~80GB / 3.7 it/s on 4x RTX 4090, claimed 6-12x vs ZeRO-Offload in their MoE SFT benches. SOSP 2025 paper: AMX kernels, async CPU-GPU scheduling, Expert Deferral (CPU util <75% to ~100%, up to 1.45x extra throughput, <=0.5% avg accuracy drop). Apache-2.0. Docs site kvcache-ai.github.io/ktransformers.

Keywords

ktransformers · moe · inference · amx · cpu-gpu · sglang · llama-factory · deepseek · sosp-2025

Topics

LLM inference, MoE, CPU-GPU hybrid serving

Research notes

  • Primary: GitHub README + API (Apache-2.0, Python, 19235 stars / 1523 forks at check). Docs https://kvcache-ai.github.io/ktransformers/. SOSP 2025 paper DOI 10.1145/3731569.3764843; PDF https://madsys.cs.tsinghua.edu.cn/publication/ktransformers-unleashing-the-full-potential-of-cpu/gpu-hybrid-inference-for-moe-models/SOSP25-chen.pdf. Original integrated framework archived under archive/. Discord posted the repo. Inference/SFT framework, not a new hosted corpus, so no datasets_local row.