← Back to explorer

Tokenization: A Survey for Modern NLP

Type
paper
Venue
alphaXiv, 2026-09-30 (#12 trending on alphaXiv)
Year
2026
Source
paper
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

The most comprehensive survey of tokenization for modern language models, assembled by 32 tokenizer researchers over ~8 months. Argues tokenization is a wildly understudied area of language modeling despite its effects across all of NLP. Covers every aspect: algorithms, evaluations, multilinguality, encodings, theory, trade-offs and pitfalls of tokenizer choice, and what tokenizers could be replaced with (latent or visual tokenization). Adjacent topics: constrained generation, token healing, tokenizer security concerns. Includes a beginner-friendly nod to Karpathy's "Let's build the GPT Tokenizer" video. Every section comes with questions for prospective researchers and practitioners; readers are invited to report missing papers and join a tokenizer-research Discord.

Keywords

tokenization · NLP · multilingual · survey · evaluation

Topics

tokenization, NLP, multilingual, survey

Research notes

  • Discovery: Marco Cognetta (@marco_computers, verified) 2026-09-30 thread: https://x.com/marco_computers/status/2105328448051028117?s=20 (main post plus 6 self-replies)
  • alphaXiv: https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp
  • Named authors on alphaXiv: Marco Cognetta, Christopher Akiki, Pawan Sasanka Ammanamanchi, Catherine Arnett, Thomas Bauwens, Pavel Chizhov, Pieter Delobelle, Konstantin Dobler, plus 24 more (32 total).
  • Connects to the collection's tokenizer and NLP entries.