VBR: Variable Bit-Rate KV Cache for llama.cpp
- Type
- other
- Venue
- GitHub
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:35:00Z
- Verified
- 2026-08-14T19:35:00Z
Summary
spiritbuun llama.cpp fork (buun-llama-cpp) ships VBR as the default KV cache. Cache starts FP16 and degrades one (layer, side) tensor at a time down a price-ordered ladder (f16 → turbo8 → turbo4 → turbo3_tcq → turbo2_tcq → turbo1_tcq) using TurboQuant/TCQ codecs, following per-model KLD degrade tables. Default implicit VBR floors at turbo4; explicit -ct vbr opens the full ladder to t1. Budget is leftover VRAM after weights/compute; -c caps context so the budget is spent on quality instead of max length. CUDA-first with ROCm/HIP mirror. Related to TurboQuant (papers_local 606/640) but this is a dynamic controller, not a new codec paper.
Keywords
vbr · kv-cache · llama-cpp · turboquant · quantization · inference · x
Topics
KV cache compression, llama.cpp inference
Research notes
- Primary: GitHub fork. Discord/X https://x.com/spiritbuun/status/2075624663598629284 via fxtwitter (1/15 thread; video demo). Related fork spiritbuun/llama-cpp-turboquant-cuda. Enable with -ct vbr / --vbr-floor / --vbr-vram. Inference kernel/code item, not a hosted corpus, so no datasets_local row. papers_local 640 (Shard) and 606 (TurboVec) are different KV-compression artifacts.