SMOL (Set for Maximal Overall Leverage)
- Type
- dataset
- Venue
- HuggingFace (google)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- 237 languages (221 low-resource languages including Afar, Abkhaz, Abaza, and many others; pivot languages English and Russian)
- Added
- 2026-07-17T20:18:03.632667+00:00
- Verified
- 2026-07-17T20:18:03.632667+00:00
Summary
SMOL (Set for Maximal Overall Leverage) is a Google dataset collection of professional and volunteer translations into 221 low-resource languages, designed for training translation models and increasing NLP representation of under-resourced languages. It contains four resources: SmolDoc (document-level translations into 130 language pairs / 129 unique languages), SmolSent (sentence-level translations into 114 language pairs / 116 unique languages), GATITOS (token-level translations into 181 language pairs / 183 unique languages), and SmolDoc-factuality-annotations (factuality annotations and rationales for 661 documents). The dataset was updated in April 2026 with additional volunteer translations, expanded professional translations, and MediSMOL medical-domain translations. It is described in arXiv papers 2502.12301 and 2303.15265.
Keywords
translation low-resource-languages multilingual gatitos smoldoc smolsent google nlp
Topics
NLP / Translation
Research notes
- Includes both professional and volunteer translations. The April 2026 update added MediSMOL (medical domain translations for ~30 languages) and expanded many languages from tier E to tier C (5x size increase). One subset (gatitos__en_gv) is marked with an error indicator.