← Back to explorer

Towards Understanding Self-Pretraining for Sequence Classification

Type
other
Venue
arXiv / Università Campus Bio-Medico di Roma / MPI-IS / ELLIS Institute Tübingen

Summary

Replicates Amos et al. 2024 (ICLR Outstanding Paper): SPT on LRA lifts ListOps/CIFAR10/PathFinder/Retrieval/Text vs from-scratch. Ablations: gains after 1–10 SPT epochs, even in 1-layer models and after swapping the pretraining dataset, pointing to an optimization bottleneck rather than hierarchical features. Freezing random Attention barely hurts from-scratch accuracy; hybrid inits show W_Q/W_K carry the SPT benefit. A 1-layer toy task shows SPT turns additive sinusoidal PEs into a proximity-biased QK map that fine-tuning builds on. Theory: at uniform attention, mean-pooled label loss has zero derivative along some score directions that masked reconstruction can see. No official code on the abs page.

Keywords

self-pretraining · spt · lra · attention · qk · mpi-is · amos2024

Topics

transformers, self-pretraining, attention, Long-Range Arena

Research notes

  • Primary: arxiv abs (cs.LG). Campus Bio-Medico di Roma / Umeå / MPI-IS / ELLIS Institute Tübingen; correspondence omarcoser10@gmail.com. Discord posted abs. HF has no paper page (API 404). No official code on abs (experiments reuse Amos et al. 2024). ArXiv license widget not visible in converted abs HTML, so license left blank. Uses public LRA/UCR; no new corpus, so no datasets_local row.