SmolLM3: smol, multilingual, long-context reasoner
- Type
- model
- Venue
- Hugging Face (HuggingFaceTB)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:56:35Z
- Verified
- 2026-08-14T18:56:35Z
Summary
HuggingFaceTB 3B decoder (Llama-style, tied embeddings, GQA 4 groups, NoPE every 4th layer, intra-document masking) trained 11.2T tokens on 384 H100s over 24 days. Three-stage web/code/math mix, then 100B long-context (4k to 64k, YaRN to 128k) and 140B reasoning mid-train. SFT 1.8B tokens dual think/no_think, APO alignment, MergeKit soup. Outperforms Llama-3.2-3B and Qwen2.5-3B; competitive with 4B models. Think mode AIME 2025 36.7 percent vs 9.3 percent no-think. Apache-2.0. Languages EN/FR/ES/DE/IT/PT.
Keywords
smollm3 · huggingfacetb · small-lm · reasoning · long-context · multilingual · blog
Topics
small LMs, multilingual, long context, reasoning
Research notes
- Primary: HF blog (public technical report; no arXiv). Instruct HuggingFaceTB/SmolLM3-3B; base SmolLM3-3B-Base. Code https://github.com/huggingface/smollm. SmolTalk2 already papers_export id 57. Cite bakouch2025smollm3. Weights/recipe item, not a new hosted corpus, so no datasets_local row.