Motif-3-Beta
- Type
- other
- Venue
- Motif Technologies (Hugging Face)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en, ko
- Added
- 2026-08-11T21:52:01+00:00
- Verified
- 2026-08-11T21:52:01+00:00
Summary
Preview / beta checkpoint of Motif-3, a large-scale sparse Mixture-of-Experts causal language model that Motif Technologies describes as a fully in-house, proprietary design rather than a re-parameterisation of an existing open-source architecture. The card lists ~314B total parameters with ~13B active per token, 53 layers, hidden size 4096, 384 routed experts with top-8 routing plus 1 shared expert, a natively long 262,144-token (256K) context window, a 220,160-token vocabulary and bfloat16 weights. Custom components include Grouped Differential Latent Attention (GDLA), Grouped PolyNorm activation applied per expert, a modified mHC and a one-layer Multi-Token Prediction head enabling self-speculative decoding. Usage guidance covers vLLM serving (with a Motif reasoning parser and tool-call parser) and HF .generate; the repository ships custom modelling code (trust_remote_code). The card states this is an intermediate checkpoint and that the final Motif-3 release is still to come.
Keywords
hf-model llm moe mixture-of-experts text-generation long-context multilingual korean english preview transformers safetensors custom-code
Topics
Language Modeling
Research notes
- likes=216; downloads last month=5,138 (checked 2026-08-11). Weights are openly downloadable with no access request, but the licence is non-commercial, so it is open-weights rather than fully open-source. Only benchmark on the card is an Artificial Analysis Intelligence Index (AAII) of 44, linked out rather than tabulated; no paper, no eval table, no training-data disclosure. Card explicitly flags this as an intermediate preview checkpoint, not the final release, so figures may change. Hub shows 6 quantized derivatives, 1 Space and a 4-item "Motif 3" collection. Cataloged as item_type=other because the sheet has no model category; hf_kind=model.