← Back to explorer

SmolLM3: smol, multilingual, long-context reasoner

Type
model
Venue
Hugging Face (HuggingFaceTB)
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T18:56:35Z
Verified
2026-08-14T18:56:35Z

Summary

HuggingFaceTB 3B decoder (Llama-style, tied embeddings, GQA 4 groups, NoPE every 4th layer, intra-document masking) trained 11.2T tokens on 384 H100s over 24 days. Three-stage web/code/math mix, then 100B long-context (4k to 64k, YaRN to 128k) and 140B reasoning mid-train. SFT 1.8B tokens dual think/no_think, APO alignment, MergeKit soup. Outperforms Llama-3.2-3B and Qwen2.5-3B; competitive with 4B models. Think mode AIME 2025 36.7 percent vs 9.3 percent no-think. Apache-2.0. Languages EN/FR/ES/DE/IT/PT.

Keywords

smollm3 · huggingfacetb · small-lm · reasoning · long-context · multilingual · blog

Topics

small LMs, multilingual, long context, reasoning

Research notes

  • Primary: HF blog (public technical report; no arXiv). Instruct HuggingFaceTB/SmolLM3-3B; base SmolLM3-3B-Base. Code https://github.com/huggingface/smollm. SmolTalk2 already papers_export id 57. Cite bakouch2025smollm3. Weights/recipe item, not a new hosted corpus, so no datasets_local row.