← Back to explorer

How Local Mixing Encodes Relative Position in Global NoPE Attention

Type
paper
Venue
arXiv:2609.38109 (cs.CL), submitted 29 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-01
Verified
2026-10-01

Summary

Explains how hybrid models with local mixing layers (sliding window attention, gated linear attention) and global NoPE attention implicitly encode relative position: local layers induce a recency bias in the residual stream that propagates to and is selected by the global attention logits. Supported by theory and experiments on randomly initialized and trained networks; recency bias persists across long sequences unlike causal-mask-only NoPE, suggesting indefinite length extrapolation.

Keywords

NoPE · positional encoding · sliding window attention · recency bias · hybrid models · length extrapolation · Kimi Delta Attention

Topics

positional encoding, NoPE, hybrid attention, length extrapolation

Research notes

  • Discovery: @ZyphraAI X post 2026-10-01 (https://x.com/ZyphraAI/status/2105697705255420318)
  • Blog: https://zyphra.com/our-work/learning-position-in-hybrid-nope-models
  • Theory + experiments: SWA/gated-linear local layers induce a recency bias in the residual stream that propagates to global NoPE attention logits, explaining how hybrid NoPE models encode relative position; smaller windows predicted text better with less training compute; same bias appears in Kimi Delta Attention; suggests path to indefinite length extrapolation
  • License: not stated on arXiv abstract page