T-FREE: Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
- Type
- other
- Venue
- arXiv / Aleph Alpha / TU Darmstadt / hessian.AI / DFKI
Summary
T-FREE splits on whitespace/digits/specials, represents each word as n×m hashed character triplets into a v-entry embedding (best v=8k, m=10 at 1B), and trains a multi-label BCE head. Cuts embedding+head params >85% vs a 64k Unigram baseline (1B: 0.84B vs 1.07B params) at competitive 18-benchmark scores. Duplicate-token rate 0% vs 15–35% for BPE/Unigram. Fertility closer to 1.0 across EN/DE/RU/VI/AR than English-centric tokenizers. 3B English→German continual pretrain gains ~5 pp on German HellaSwag/ARC while the Unigram baseline barely moves. Code https://github.com/Aleph-Alpha/trigrams.
Keywords
t-free · tokenizer-free · trigrams · embeddings · aleph-alpha · multilingual · fertility · hash-embeddings
Topics
tokenization, embeddings, multilingual LLMs
Research notes
- Primary: arxiv abs (cs.CL; also cs.AI, cs.LG). License CC BY-NC-SA 4.0 on HTML at check. Deiseroth/Brack/Schramowski/Kersting Aleph Alpha @ IPAI / TU Darmstadt / hessian.AI / DFKI; Weinbach Aleph Alpha. Code on HTML https://github.com/Aleph-Alpha/trigrams (HF githubRepo; now Aleph-Alpha-Research/trigrams, 60 stars at check; GitHub license NOASSERTION). HF paper page 11 upvotes; unofficial linked model krystv/neurolex-creative-name-generator not copied into hf_* fields. Discord posted PDF. Trains on public SlimPajama/Occiglot Fineweb rather than a new hosted corpus, so no datasets_local row.