Tokenization: A Survey for Modern NLP
- Type
- paper
- Venue
- alphaXiv, 2026-09-30 (#12 trending on alphaXiv)
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
The most comprehensive survey of tokenization for modern language models, assembled by 32 tokenizer researchers over ~8 months. Argues tokenization is a wildly understudied area of language modeling despite its effects across all of NLP. Covers every aspect: algorithms, evaluations, multilinguality, encodings, theory, trade-offs and pitfalls of tokenizer choice, and what tokenizers could be replaced with (latent or visual tokenization). Adjacent topics: constrained generation, token healing, tokenizer security concerns. Includes a beginner-friendly nod to Karpathy's "Let's build the GPT Tokenizer" video. Every section comes with questions for prospective researchers and practitioners; readers are invited to report missing papers and join a tokenizer-research Discord.
Keywords
tokenization · NLP · multilingual · survey · evaluation
Topics
tokenization, NLP, multilingual, survey
Research notes
- Discovery: Marco Cognetta (@marco_computers, verified) 2026-09-30 thread: https://x.com/marco_computers/status/2105328448051028117?s=20 (main post plus 6 self-replies)
- alphaXiv: https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp
- Named authors on alphaXiv: Marco Cognetta, Christopher Akiki, Pawan Sasanka Ammanamanchi, Catherine Arnett, Thomas Bauwens, Pavel Chizhov, Pieter Delobelle, Konstantin Dobler, plus 24 more (32 total).
- Connects to the collection's tokenizer and NLP entries.