Maple-Preview
- Type
- model
- Venue
- DeepGrove
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:32:00Z
- Verified
- 2026-08-14T18:32:00Z
Summary
Maple-Preview is a 20B-A1B reasoning MoE (24 layers, 256 experts / 8 active, 3:1 SWA-512:GA attention) with ternary weights and custom MapleForCausalLM code. Card: 131,072 context, 5.31 GB packed checkpoint, 218 tok/s on an M4 Mac mini (5–16× vs Gemma 4 / Qwen3.5 / gpt-oss in their figure). Preview is reasoning-focused with only small-scale RL; they flag weaker agentic evals before a full release. Transformers path needs Triton/FlashAttention CUDA; Apple Silicon numbers use a separate runtime. MIT. HF hosts 9 BF16 safetensor shards (~20.21B params).
Keywords
maple · deepgrove · ternary · moe · on-device · reasoning · huggingface
Topics
ternary LLMs, mixture-of-experts, on-device reasoning
Research notes
- Primary: HF model card (MIT, transformers, pipeline_tag text-generation, downloads=5617 likes=358 checked 2026-08-14). No paper URL on the card. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Unofficial spaces ProCreations/maple-webgpu and ajsbsd/maple-preview not copied into hf_* fields. Card 5.31 GB packed size vs ~40 GB BF16 shards on the repo.