Shard: getting to 10x KV cache compression
- Type
- other
- Venue
- krishgarg.com
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:20:00Z
- Verified
- 2026-08-14T19:20:00Z
Summary
Engineering writeup of Shard, a drop-in HuggingFace Cache. Reimplementing TurboQuant stalled around 4-6x; authors split K and V: undo RoPE then PCA+int4 on K (rank 192/1024, DP bit allocation with a 4x drop penalty) and Hadamard + 4-d vector quantization on V. Attention scores run on int4 PCA coeffs via a relative-Delta RoPE identity; V Hadamard is applied after the weighted sum. 4 FP16 sink tokens + 64-token recency window. Decode stream uses data-oblivious Hadamard+Lloyd-Max (8-bit path 750/750 FP16 match at 150 tokens). Llama-3.1-8B: 10x at 8K / 11.2x at 32K, NIAH recall 1.000 across 4K-32K, LongBench-E avg 16.19 vs FP16 16.24, WikiText-2 PPL +0.26%. Decode throughput 0.39-0.49x FP16, so a memory win not a latency win. Code krish1905/shard.
Keywords
shard · kv-cache · turboquant · pca · vector-quantization · llama · inference · blog · x
Topics
KV cache compression, LLM inference
Research notes
- Primary: technical blog. Discord/X https://x.com/krishgarg/status/2059041521576648980 via fxtwitter (also http://krishgarg.com/shard). Code https://github.com/krish1905/shard (MIT). Tweet credits @kirrithan; second author name not given so left off. Compares to TurboQuant arXiv 2504.19874 (ICLR 2026); papers_local id 606 is a different TurboQuant-related GitHub (RyanCodrai/turbovec). Inference method, not a new hosted corpus, so no datasets_local row.