Still: Amortized KV Cache Compaction in a Single Forward Pass
- Type
- other
- Venue
- arXiv / Baseten
Summary
Still is a per-layer Perceiver (~50M / ~1% of Qwen3-4B at t=128) that cross-attends the full KV cache in a position-free RoPE frame and writes compact (C_k, C_v) in one forward, leaving base weights frozen. Occupies the speed–quality frontier at 8x–200x compression and 8k–128k context vs H2O/SnapKV/StreamingLLM/Attention Matching/KV-Distill. Matched-training RULER: +8–22 points vs KV-Distill in 16/18 cells. HELMET multi_lexsum recovers 74–95% of full-context gain at 8k–64k. Iterative chunked compaction at fixed ratio 1/c is possible because compaction is a forward pass; 8k-trained compactors collapse at 128k. No official code repo on the abs page. Project https://www.baseten.co/research/still-amortized-kv-cache-compaction-in-a-single-forward-pass/.
Keywords
still · kv-cache · perceiver · compaction · amortized-synthesis · baseten · ruler · helmet
Topics
KV cache, long context, inference
Research notes
- Primary: arxiv abs. Baseten; correspondence charlie.oneill/alex.sandomirsky/harry.partridge/mudith@baseten.co and max.kirkby@baseten.co. Discord posted abs v1. HF has no paper page (API 404). No official code on abs. Training MCQ mix is custom and not released as a standalone public corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so license left blank.