← Back to explorer

Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings

Type
paper
Venue
arXiv / Sakana AI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:54:04Z
Verified
2026-08-14T16:54:04Z

Summary

Argues explicit PEs speed pretraining (NoPE has slower positional-bias gradients) but RoPE-scaling (YaRN/NTK/PI) must compress low frequencies, shifting semantic heads so zero-shot long context matches a cropped-window baseline on NIAH. DroPE trains with RoPE, drops PEs, then recalibrates at the original context. From-scratch 0.5B on 16B FineWeb: last 2B tokens without PE matches full-RoPE perplexity and beats YaRN/NTK/ALiBi/NoPE on RULER NIAH at 2x. SmolLM-360M (600B pretrain): 30-120B recalibration recovers in-context benches; LongBench avg 30.52 vs YaRN 19.94, NIAH 74.92 vs 48.25. SmolLM-1.7B (20B rec, 2 percent of pretrain) and Llama-2-7B (20B, 0.5 percent) also beat YaRN/NTK. NIAH at 8x: DroPE 52.20 vs YaRN 12.18 / LongRoPE2 16.45.

Keywords

long-context · positional-embeddings · sakana

Topics

long context, positional embeddings, transformers

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.CL). Code at github.com/SakanaAI/DroPE (220 stars at check). Gelberg Oxford / Sakana. HF paper page 5 upvotes. Recalibration uses FineWeb/FineWeb-Edu and SmolLM-corpus, not a new dataset; no datasets_local row. Discord posted abs link.