nanowhale
- Type
- repo
- Venue
- Hugging Face (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:35:00Z
- Verified
- 2026-08-14T18:35:00Z
Summary
huggingface/nanowhale ports DeepSeek-V4 (MLA, 4+1 MoE with top-2, Hyper-Connections hc_mult=4 / Sinkhorn, MTP) into HF Trainer/SFTTrainer at hidden 320 / 8 layers / vocab 129,280 / ctx 2048 (~110M params: 41M embed + 69M non-embed). Pretrain 5K steps on FineWeb-Edu (~2.6B tokens, final loss ~5.3, 1×H100) then 3K-step SmolTalk SFT. Checkpoints on HF as cmpatino/nanowhale-100m-base and cmpatino/nanowhale-100m. README flags bf16 NaNs from Hyper-Connections at this scale (use fp32) and a from_pretrained re-init quirk. README license MIT; GitHub API has no SPDX (no LICENSE file). Companion HF base card is Apache-2.0.
Keywords
nanowhale · deepseek-v4 · moe · mla · hyper-connections · huggingface · small-lm
Topics
DeepSeek-V4, small LMs, MoE, MLA
Research notes
- Primary: GitHub README + API (README MIT, API license null / no LICENSE file; Python; 386 stars / 45 forks at check). Weights live at cmpatino/nanowhale-100m-base (HF card Apache-2.0, 4 likes at check) and cmpatino/nanowhale-100m; not copied into hf_* because the leftover is the GitHub training repo. Discord posted the repo. Training uses public FineWeb-Edu / SmolTalk, so no datasets_local row.