Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
- Type
- paper
- Venue
- arXiv / NVIDIA
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:24:56Z
- Verified
- 2026-08-14T16:24:56Z
Summary
Shows within-family matched-KV pairs (shared KV head count and head dim) have substantial linear structure: on Qwen3 14B->32B one source layer explains 56%/32% of target key/value variance, 79%/65% with multiple layers. A closed-form per-head ridge mapper with top-k cross-layer source selection and RoPE-stripped content-space keys retains 73-98% of standalone-prefill accuracy on four of six pairs (Qwen3, Llama 3.1, Ministral 3) and is 2.7-25x faster than re-prefill; two Ministral pairs collapse and a nonlinear MLP recovers up to +37 pp HellaSwag. Attention-output cosine predicts retention better than reconstruction R^2.
Keywords
kv-cache · prefill · ridge-regression · rope · serving · qwen3 · llama · ministral · nvidia
Topics
LLM serving, KV cache, inference
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG). NVIDIA. No official code on abs; HF paper page 1 upvote with two unofficial linked models. Discord posted abs link.