← Back to explorer

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

Type
paper
Venue
arXiv / NVIDIA
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:24:56Z
Verified
2026-08-14T16:24:56Z

Summary

Shows within-family matched-KV pairs (shared KV head count and head dim) have substantial linear structure: on Qwen3 14B->32B one source layer explains 56%/32% of target key/value variance, 79%/65% with multiple layers. A closed-form per-head ridge mapper with top-k cross-layer source selection and RoPE-stripped content-space keys retains 73-98% of standalone-prefill accuracy on four of six pairs (Qwen3, Llama 3.1, Ministral 3) and is 2.7-25x faster than re-prefill; two Ministral pairs collapse and a nonlinear MLP recovers up to +37 pp HellaSwag. Attention-output cosine predicts retention better than reconstruction R^2.

Keywords

kv-cache · prefill · ridge-regression · rope · serving · qwen3 · llama · ministral · nvidia

Topics

LLM serving, KV cache, inference

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG). NVIDIA. No official code on abs; HF paper page 1 upvote with two unofficial linked models. Discord posted abs link.