← Back to explorer

Motif-3-Beta

Type
other
Venue
Motif Technologies (Hugging Face)
Year
2026
Source
huggingface
Access
free
Language
en, ko
Added
2026-08-11T21:52:01+00:00
Verified
2026-08-11T21:52:01+00:00

Summary

Preview / beta checkpoint of Motif-3, a large-scale sparse Mixture-of-Experts causal language model that Motif Technologies describes as a fully in-house, proprietary design rather than a re-parameterisation of an existing open-source architecture. The card lists ~314B total parameters with ~13B active per token, 53 layers, hidden size 4096, 384 routed experts with top-8 routing plus 1 shared expert, a natively long 262,144-token (256K) context window, a 220,160-token vocabulary and bfloat16 weights. Custom components include Grouped Differential Latent Attention (GDLA), Grouped PolyNorm activation applied per expert, a modified mHC and a one-layer Multi-Token Prediction head enabling self-speculative decoding. Usage guidance covers vLLM serving (with a Motif reasoning parser and tool-call parser) and HF .generate; the repository ships custom modelling code (trust_remote_code). The card states this is an intermediate checkpoint and that the final Motif-3 release is still to come.

Keywords

hf-model llm moe mixture-of-experts text-generation long-context multilingual korean english preview transformers safetensors custom-code

Topics

Language Modeling

Research notes

  • likes=216; downloads last month=5,138 (checked 2026-08-11). Weights are openly downloadable with no access request, but the licence is non-commercial, so it is open-weights rather than fully open-source. Only benchmark on the card is an Artificial Analysis Intelligence Index (AAII) of 44, linked out rather than tabulated; no paper, no eval table, no training-data disclosure. Card explicitly flags this as an intermediate preview checkpoint, not the final release, so figures may change. Hub shows 6 quantized derivatives, 1 Space and a 4-item "Motif 3" collection. Cataloged as item_type=other because the sheet has no model category; hf_kind=model.