Inkling: open-weight controllable-effort multimodal MoE
- Type
- other
- Venue
- Thinking Machines Lab
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:35:00Z
- Verified
- 2026-08-14T19:35:00Z
Summary
Inkling (released 2026-07-15): 66-layer decoder-only MoE, 975B total / 41B active (6-of-256 experts + 2 shared), hybrid local/global attention, native text/image/audio in, text out, 1M context. Sibling Inkling-Small 276B/12B. Apache-2.0 weights on HF. Distinctive RL: Reward = task reward − λ × (# reasoning tokens), with λ and effort-instruction varied across rollouts so the model learns to spend thought rather than maximize it; inference effort is a continuous 0–0.99 knob. Model-card highlights at effort=0.99 include HLE text 29.7% / tools 46.0%, AIME 2026 97.1%, GPQA Diamond 87.2%, SWE-bench Verified 77.6%. Discord post is Vipul Gupta’s note of that RL trick, not a TML thread.
Keywords
inkling · thinking-machines · moe · reasoning-effort · open-weights · rl · x
Topics
open-weight LLMs, reasoning-effort RL, multimodal MoE
Research notes
- Primary: Inkling model card. Discord/X https://x.com/vipul_1011/status/2077865552698298688 via fxtwitter (third-party note of the λ reasoning-token penalty). Product https://thinkingmachines.ai/inkling/. Weights HF collection thinkingmachines/inkling (Inkling, Inkling-NVFP4, Inkling-Small, Inkling-Small-NVFP4). tinker-cookbook scripts under thinking-machines-lab/tinker-cookbook. Open weights, not a new hosted corpus, so no datasets_local row.