← Back to explorer

nanowhale

Type
repo
Venue
Hugging Face (GitHub)
Year
2026
Source
github
Access
free
Added
2026-08-14T18:35:00Z
Verified
2026-08-14T18:35:00Z

Summary

huggingface/nanowhale ports DeepSeek-V4 (MLA, 4+1 MoE with top-2, Hyper-Connections hc_mult=4 / Sinkhorn, MTP) into HF Trainer/SFTTrainer at hidden 320 / 8 layers / vocab 129,280 / ctx 2048 (~110M params: 41M embed + 69M non-embed). Pretrain 5K steps on FineWeb-Edu (~2.6B tokens, final loss ~5.3, 1×H100) then 3K-step SmolTalk SFT. Checkpoints on HF as cmpatino/nanowhale-100m-base and cmpatino/nanowhale-100m. README flags bf16 NaNs from Hyper-Connections at this scale (use fp32) and a from_pretrained re-init quirk. README license MIT; GitHub API has no SPDX (no LICENSE file). Companion HF base card is Apache-2.0.

Keywords

nanowhale · deepseek-v4 · moe · mla · hyper-connections · huggingface · small-lm

Topics

DeepSeek-V4, small LMs, MoE, MLA

Research notes

  • Primary: GitHub README + API (README MIT, API license null / no LICENSE file; Python; 386 stars / 45 forks at check). Weights live at cmpatino/nanowhale-100m-base (HF card Apache-2.0, 4 likes at check) and cmpatino/nanowhale-100m; not copied into hf_* because the leftover is the GitHub training repo. Discord posted the repo. Training uses public FineWeb-Edu / SmolTalk, so no datasets_local row.