Towards Understanding Self-Pretraining for Sequence Classification
- Type
- other
- Venue
- arXiv / Università Campus Bio-Medico di Roma / MPI-IS / ELLIS Institute Tübingen
Summary
Replicates Amos et al. 2024 (ICLR Outstanding Paper): SPT on LRA lifts ListOps/CIFAR10/PathFinder/Retrieval/Text vs from-scratch. Ablations: gains after 1–10 SPT epochs, even in 1-layer models and after swapping the pretraining dataset, pointing to an optimization bottleneck rather than hierarchical features. Freezing random Attention barely hurts from-scratch accuracy; hybrid inits show W_Q/W_K carry the SPT benefit. A 1-layer toy task shows SPT turns additive sinusoidal PEs into a proximity-biased QK map that fine-tuning builds on. Theory: at uniform attention, mean-pooled label loss has zero derivative along some score directions that masked reconstruction can see. No official code on the abs page.
Keywords
self-pretraining · spt · lra · attention · qk · mpi-is · amos2024
Topics
transformers, self-pretraining, attention, Long-Range Arena
Research notes
- Primary: arxiv abs (cs.LG). Campus Bio-Medico di Roma / Umeå / MPI-IS / ELLIS Institute Tübingen; correspondence omarcoser10@gmail.com. Discord posted abs. HF has no paper page (API 404). No official code on abs (experiments reuse Amos et al. 2024). ArXiv license widget not visible in converted abs HTML, so license left blank. Uses public LRA/UCR; no new corpus, so no datasets_local row.