← Back to explorer

Charformer: Fast Character Transformers via Gradient-based Subword Tokenization

Type
other
Venue
arXiv / Google Research / DeepMind

Summary

GBST scores candidate byte n-gram blocks (max width 4) position-wise and mixes them, then mean-pools (ds=2–3). Charformer is a T5-style encoder-decoder on the 256-byte vocab. CharformerSBase (134M, deep encoder): English GLUE avg 83.6 vs T5-Base 84.3 / Byte-T5 81.5; Civil Comments 83.0 vs T5 81.2; IMDb 94.4 vs T5 94.2. Multilingual long-PT: TyDiQA-GoldP 81.2/71.3 vs mT5-Base 80.8/70.0, 28% faster with ~3× fewer params. ICLR 2022 camera-ready. Implementation https://github.com/google-research/google-research/tree/master/charformer.

Keywords

charformer · gbst · byte-level · tokenization · iclr · google · deepmind · byt5 · canine

Topics

tokenization, character-level Transformers, byte-level LMs

Research notes

  • Primary: arxiv abs (cs.CL; also cs.AI, cs.LG). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: ICLR 2022 Camera Ready. Equal contrib Tay/Tran. Google Research / DeepMind. Correspondence yitay@google.com, vqtran@google.com. Implementation in google-research/google-research/charformer (monorepo; no standalone official repo on abs). HF paper page 0 upvotes; unofficial linked models orkungedik/charformer-turkish-base and etri-lirs/gbst-kebyt5-* not copied into hf_* fields. Discord posted abs v3. Trains on C4/mC4/GLUE rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.