← Back to explorer

VBR: Variable Bit-Rate KV Cache for llama.cpp

Type
other
Venue
GitHub
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:35:00Z
Verified
2026-08-14T19:35:00Z

Summary

spiritbuun llama.cpp fork (buun-llama-cpp) ships VBR as the default KV cache. Cache starts FP16 and degrades one (layer, side) tensor at a time down a price-ordered ladder (f16 → turbo8 → turbo4 → turbo3_tcq → turbo2_tcq → turbo1_tcq) using TurboQuant/TCQ codecs, following per-model KLD degrade tables. Default implicit VBR floors at turbo4; explicit -ct vbr opens the full ladder to t1. Budget is leftover VRAM after weights/compute; -c caps context so the budget is spent on quality instead of max length. CUDA-first with ROCm/HIP mirror. Related to TurboQuant (papers_local 606/640) but this is a dynamic controller, not a new codec paper.

Keywords

vbr · kv-cache · llama-cpp · turboquant · quantization · inference · x

Topics

KV cache compression, llama.cpp inference

Research notes

  • Primary: GitHub fork. Discord/X https://x.com/spiritbuun/status/2075624663598629284 via fxtwitter (1/15 thread; video demo). Related fork spiritbuun/llama-cpp-turboquant-cuda. Enable with -ct vbr / --vbr-floor / --vbr-vram. Inference kernel/code item, not a hosted corpus, so no datasets_local row. papers_local 640 (Shard) and 606 (TurboVec) are different KV-compression artifacts.