TCNCA: Temporal Convolution Network with Chunked Attention for Scalable Sequence Processing
- Type
- other
- Venue
- arXiv / IBM Research / ETH Zurich
Summary
Swaps MEGA’s FFT-parallel EMA (O(L log L)) for a TCN of dilated-conv residual blocks (one conv per block, O(L)) followed by fixed-window chunked attention. EnWik8 BPC 1.01 vs MEGA 1.02 / Transformer-XL 1.06 at 39M params, with 1.37×/1.24× faster train forward/backward vs MEGA. Dilated conv vs parallel EMA is up to 7.07×/2.86× faster at seq 131k on V100. LRA average 85.5 vs MEGA-chunk 85.6 with 1.28× inference. A simplified TCNCA matches MEGA on associative recall. No official code on abs.
Keywords
tcnca · tcn · mega · chunked-attention · enwik8 · lra · ibm · eth-zurich · long-sequence
Topics
long sequence models, temporal convolution, chunked attention
Research notes
- Primary: arxiv abs (cs.LG; also cs.CV). License CC BY 4.0 on HTML at check. Work at IBM Research – Zurich / ETH Zurich. Correspondence Abbas Rahimi abr@zurich.ibm.com. No official code on abs. HF paper page 4 upvotes; no githubRepo; no linked models/datasets. Discord posted HTML. Uses public EnWik8/LRA rather than a new hosted corpus, so no datasets_local row.