← Back to explorer

TCNCA: Temporal Convolution Network with Chunked Attention for Scalable Sequence Processing

Type
other
Venue
arXiv / IBM Research / ETH Zurich

Summary

Swaps MEGA’s FFT-parallel EMA (O(L log L)) for a TCN of dilated-conv residual blocks (one conv per block, O(L)) followed by fixed-window chunked attention. EnWik8 BPC 1.01 vs MEGA 1.02 / Transformer-XL 1.06 at 39M params, with 1.37×/1.24× faster train forward/backward vs MEGA. Dilated conv vs parallel EMA is up to 7.07×/2.86× faster at seq 131k on V100. LRA average 85.5 vs MEGA-chunk 85.6 with 1.28× inference. A simplified TCNCA matches MEGA on associative recall. No official code on abs.

Keywords

tcnca · tcn · mega · chunked-attention · enwik8 · lra · ibm · eth-zurich · long-sequence

Topics

long sequence models, temporal convolution, chunked attention

Research notes

  • Primary: arxiv abs (cs.LG; also cs.CV). License CC BY 4.0 on HTML at check. Work at IBM Research – Zurich / ETH Zurich. Correspondence Abbas Rahimi abr@zurich.ibm.com. No official code on abs. HF paper page 4 upvotes; no githubRepo; no linked models/datasets. Discord posted HTML. Uses public EnWik8/LRA rather than a new hosted corpus, so no datasets_local row.