Do Transformers Need Three Projections? Systematic Study of QKV Variants
- Type
- other
- Venue
- arXiv / BrainChip
Summary
Evaluates Q=K≠V, Q≠K=V, and Q=K=V (plus 2D-PE + variants for non-causal tasks) against standard QKV. On language modeling, Q≠K=V cuts KV cache 50% at +3.1% val PPL (300M) and +2.48% (1.2B) vs QKV; Q=K≠V has no cache benefit, Q=K=V degrades ~25%. Combining Q≠K=V with GQA-4/MQA reaches 87.5%/96.9% cache reduction. K and V projections are highly similar (cosine 0.73) while Q stays distinct, so K=V preserves directional QK^T. Code https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections.
Keywords
qkv · projection-sharing · kv-cache · gqa · mqa · brainchip · icml · slimpajama
Topics
transformer architecture, attention, KV cache
Research notes
- Primary: arxiv abs. ICML 2026. BrainChip Inc.; correspondence akayyam@brainchip.com (Kayyam previously published as Ali Borji). Code MIT https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections (46 stars at check). HF paper page 2 upvotes; no linked models/datasets. Discord posted abs. Trains on SlimPajama and standard vision sets; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.