← Back to explorer

SMOL (Set for Maximal Overall Leverage)

Type
dataset
Venue
HuggingFace (google)
Year
2026
Source
huggingface
Access
free
Language
237 languages (221 low-resource languages including Afar, Abkhaz, Abaza, and many others; pivot languages English and Russian)
Added
2026-07-17T20:18:03.632667+00:00
Verified
2026-07-17T20:18:03.632667+00:00

Summary

SMOL (Set for Maximal Overall Leverage) is a Google dataset collection of professional and volunteer translations into 221 low-resource languages, designed for training translation models and increasing NLP representation of under-resourced languages. It contains four resources: SmolDoc (document-level translations into 130 language pairs / 129 unique languages), SmolSent (sentence-level translations into 114 language pairs / 116 unique languages), GATITOS (token-level translations into 181 language pairs / 183 unique languages), and SmolDoc-factuality-annotations (factuality annotations and rationales for 661 documents). The dataset was updated in April 2026 with additional volunteer translations, expanded professional translations, and MediSMOL medical-domain translations. It is described in arXiv papers 2502.12301 and 2303.15265.

Keywords

translation low-resource-languages multilingual gatitos smoldoc smolsent google nlp

Topics

NLP / Translation

Research notes

  • Includes both professional and volunteer translations. The April 2026 update added MediSMOL (medical domain translations for ~30 languages) and expanded many languages from tier E to tier C (5x size increase). One subset (gatitos__en_gv) is marked with an error indicator.