How Local Mixing Encodes Relative Position in Global NoPE Attention
- Type
- paper
- Venue
- arXiv:2609.38109 (cs.CL), submitted 29 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Explains how hybrid models with local mixing layers (sliding window attention, gated linear attention) and global NoPE attention implicitly encode relative position: local layers induce a recency bias in the residual stream that propagates to and is selected by the global attention logits. Supported by theory and experiments on randomly initialized and trained networks; recency bias persists across long sequences unlike causal-mask-only NoPE, suggesting indefinite length extrapolation.
Keywords
NoPE · positional encoding · sliding window attention · recency bias · hybrid models · length extrapolation · Kimi Delta Attention
Topics
positional encoding, NoPE, hybrid attention, length extrapolation
Research notes
- Discovery: @ZyphraAI X post 2026-10-01 (https://x.com/ZyphraAI/status/2105697705255420318)
- Blog: https://zyphra.com/our-work/learning-position-in-hybrid-nope-models
- Theory + experiments: SWA/gated-linear local layers induce a recency bias in the residual stream that propagates to global NoPE attention logits, explaining how hybrid NoPE models encode relative position; smaller windows predicted text better with less training compute; same bias appears in Kimi Delta Attention; suggests path to indefinite length extrapolation
- License: not stated on arXiv abstract page