← Back to explorer

Maple-Preview

Type
model
Venue
DeepGrove
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T18:32:00Z
Verified
2026-08-14T18:32:00Z

Summary

Maple-Preview is a 20B-A1B reasoning MoE (24 layers, 256 experts / 8 active, 3:1 SWA-512:GA attention) with ternary weights and custom MapleForCausalLM code. Card: 131,072 context, 5.31 GB packed checkpoint, 218 tok/s on an M4 Mac mini (5–16× vs Gemma 4 / Qwen3.5 / gpt-oss in their figure). Preview is reasoning-focused with only small-scale RL; they flag weaker agentic evals before a full release. Transformers path needs Triton/FlashAttention CUDA; Apple Silicon numbers use a separate runtime. MIT. HF hosts 9 BF16 safetensor shards (~20.21B params).

Keywords

maple · deepgrove · ternary · moe · on-device · reasoning · huggingface

Topics

ternary LLMs, mixture-of-experts, on-device reasoning

Research notes

  • Primary: HF model card (MIT, transformers, pipeline_tag text-generation, downloads=5617 likes=358 checked 2026-08-14). No paper URL on the card. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Unofficial spaces ProCreations/maple-webgpu and ajsbsd/maple-preview not copied into hf_* fields. Card 5.31 GB packed size vs ~40 GB BF16 shards on the repo.